初始化项目,由ModelHub XC社区提供模型
Model: qikp/kite-6.2-18m Source: Original Platform
This commit is contained in:
25
README.md
Normal file
25
README.md
Normal file
@@ -0,0 +1,25 @@
|
||||
---
|
||||
license: cc0-1.0
|
||||
datasets:
|
||||
- openbmb/Ultra-FineWeb-L3
|
||||
language:
|
||||
- en
|
||||
pipeline_tag: text-generation
|
||||
library_name: transformers
|
||||
---
|
||||
|
||||
# Kite
|
||||
|
||||
🎉 You are looking at Kite 6.2, which uses a different and cleaner dataset!
|
||||
|
||||
Kite is a small, trained, 18 million parameter language model.
|
||||
|
||||
## Training
|
||||
|
||||
It was trained on a tokenized and truncated version of the "multi style" subset of [openbmb/Ultra-FineWeb-L3](https://huggingface.co/datasets/openbmb/Ultra-FineWeb-L3), using 2 epochs, 32 batch size, 1e-3 learning rate, and the [pika 4 tokenizer](https://huggingface.co/qikp/pika-4).
|
||||
|
||||
Also, evaluation on a tokenized and truncated [byunggill/gpt-2-output](https://huggingface.co/datasets/byunggill/gpt-2-output) was done during training.
|
||||
|
||||
## Limitations
|
||||
|
||||
Due to its size, the model is not suitable for production workloads.
|
||||
Reference in New Issue
Block a user