Files
hummingbird-2.5-110m/README.md
ModelHub XC 0450326281 初始化项目,由ModelHub XC社区提供模型
Model: qikp/hummingbird-2.5-110m
Source: Original Platform
2026-09-02 06:52:17 +08:00

44 lines
1.3 KiB
Markdown

---
license: apache-2.0
datasets:
- qikp/aninsthro
- HuggingFaceTB/everyday-conversations-llama3.1-2k
language:
- en
base_model:
- cerebras/Cerebras-GPT-111M
pipeline_tag: text-generation
library_name: transformers
new_version: qikp/hummingbird-2.6-110m
---
# Hummingbird
🎉 You are looking at Hummingbird 2.5, which largely uses Anthropic's dataset instead!
Hummingbird is a Cerebras-GPT derivative trained to be conversational.
## Training
The model was trained using the `paged_adamw_8bit` optimizer, gradient checkpointing, 500 steps, 1 batch size, and 4 gradient accumulation steps.
### Datasets
The training corpus is made up of:
- First 1500 rows of [qikp/aninsthro](https://huggingface.co/datasets/qikp/aninsthro) (a collate of a subset of Anthropic's `hh-rlhf` dataset)
- First 500 rows of [HuggingFaceTB/everyday-conversations-llama3.1-2k](https://huggingface.co/datasets/HuggingFaceTB/everyday-conversations-llama3.1-2k)
The `train` / `train_sft` splits were used.
### Chat template
The Zephyr chat template was used.
## Limitations
The model frequently outputs incorrect information, confirmation with a larger, mature model is advised.
## Benchmark
This model was benchmarked and compared using embeddings. See the results [here](https://codeberg.org/qikp/benchmarks).