822 B
822 B
license, datasets, language, pipeline_tag, library_name
| license | datasets | language | pipeline_tag | library_name | ||
|---|---|---|---|---|---|---|
| cc0-1.0 |
|
|
text-generation | transformers |
Kite
🎉 You are looking at Kite 6.2, which uses a different and cleaner dataset!
Kite is a small, trained, 18 million parameter language model.
Training
It was trained on a tokenized and truncated version of the "multi style" subset of openbmb/Ultra-FineWeb-L3, using 2 epochs, 32 batch size, 1e-3 learning rate, and the pika 4 tokenizer.
Also, evaluation on a tokenized and truncated byunggill/gpt-2-output was done during training.
Limitations
Due to its size, the model is not suitable for production workloads.