Files
kite-6.2-18m/README.md
ModelHub XC 89090973ef 初始化项目,由ModelHub XC社区提供模型
Model: qikp/kite-6.2-18m
Source: Original Platform
2026-08-09 06:37:17 +08:00

822 B

license, datasets, language, pipeline_tag, library_name
license datasets language pipeline_tag library_name
cc0-1.0
openbmb/Ultra-FineWeb-L3
en
text-generation transformers

Kite

🎉 You are looking at Kite 6.2, which uses a different and cleaner dataset!

Kite is a small, trained, 18 million parameter language model.

Training

It was trained on a tokenized and truncated version of the "multi style" subset of openbmb/Ultra-FineWeb-L3, using 2 epochs, 32 batch size, 1e-3 learning rate, and the pika 4 tokenizer.

Also, evaluation on a tokenized and truncated byunggill/gpt-2-output was done during training.

Limitations

Due to its size, the model is not suitable for production workloads.