Files
kw5-lite-base/README.md
ModelHub XC 28dcb5ee11 初始化项目,由ModelHub XC社区提供模型
Model: regnant-io/kw5-lite-base
Source: Original Platform
2026-09-06 09:12:17 +08:00

183 lines
6.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
language:
- sw
license: apache-2.0
library_name: transformers
tags:
- swahili
- kiswahili
- causal-lm
- base-model
- pretrained
- african-languages
- low-resource
datasets:
- HuggingFaceFW/fineweb-2
pipeline_tag: text-generation
---
# KW5-Lite Base
A 109.5M-parameter Swahili (Kiswahili) language model pretrained from scratch on
1.41B tokens, on a single NVIDIA T4.
Built by [Regnant](https://www.regnant.io/).
**This is a base model.** It does next-token prediction only. It does not follow
instructions and has no chat template. For that, use
[regnant-io/kw5-lite-instruct](https://huggingface.co/regnant-io/kw5-lite-instruct).
## Quick start
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("regnant-io/kw5-lite-base")
tok = AutoTokenizer.from_pretrained("regnant-io/kw5-lite-base")
# Base model: give it a prefix to continue, not an instruction.
# Write <s> into the text — the tokenizer maps it to id 1.
inputs = tok("<s>Tanzania ni nchi", return_tensors="pt", add_special_tokens=False)
out = model.generate(**inputs, max_new_tokens=60)
print(tok.decode(out[0], skip_special_tokens=True))
```
```
Tanzania ni nchi ya amani na utulivu. Ni nchi yenye watu wengi, wenye nguvu
za kiuchumi, wenye uwezo wa kufanya maamuzi magumu kwa wakati mmoja...
```
`add_bos_token` is pinned to `false` so that the same tokenizer settings work
for the instruct model, which renders `<s>` as part of its prompt template.
**Put `<s>` at the start of your text yourself** — every training sequence
began with it.
Write it into the string, as above, rather than concatenating the id onto
`input_ids` afterwards: that leaves `attention_mask` one element shorter than
`input_ids`, and `generate` then fails inside RoPE with
`The size of tensor a (4) must match the size of tensor b (3)`.
Like the instruct model, this one needs a low temperature and a repetition
penalty to stay coherent; the shipped `generation_config.json` defaults to
temperature 0.2 / `repetition_penalty` 1.3.
## Model details
| | |
|---|---|
| Parameters | 109.5M (tied input/output embeddings) |
| Architecture | Llama-compatible decoder-only transformer |
| Layers / hidden / FFN | 12 / 768 / 2048 |
| Attention heads | 12 query, 12 key-value — **standard multi-head attention, not GQA** |
| Normalization | RMSNorm, pre-norm |
| Activation | SwiGLU |
| Position encoding | RoPE, theta 10000 |
| Context length | 2048 |
| Vocabulary | 32,000 SentencePiece BPE |
| Precision | FP32 (trained under FP16 autocast with an FP32 master copy) |
| File size | 438 MB |
> A previous revision of this card claimed "KV heads: 4 (Grouped Query
> Attention)". That was wrong — `num_key_value_heads` is 12 and always has
> been. This model uses standard MHA. Corrected here.
### Tokenizer
32k BPE trained from scratch on the deduplicated Swahili corpus. NFC
normalization applied before training (SentencePiece's own rules are all
NFKC-family, which is lossier), `byte_fallback` so no input can hard-fail, and
digits split.
Special tokens are fixed at low ids: `<unk>`=0, `<s>`=1, `</s>`=2, `<pad>`=3.
Ids 4–6 are `<|system|>`, `<|user|>`, `<|assistant|>`, **reserved but never
trained** — they do not occur in the pretraining corpus, so their embedding
rows are still at initialisation. If you fine-tune and want to use them as
chat-role markers, you must unfreeze `embed_tokens` (or add it to your LoRA
`modules_to_save`), or the model will read them as noise.
## Training
| | |
|---|---|
| Data | [FineWeb-2](https://huggingface.co/datasets/HuggingFaceFW/fineweb-2) `swh_Latn`, cleaned and deduplicated (exact + MinHash LSH) |
| Tokens seen | 1.41B — exactly 2 epochs over the corpus |
| Checkpoint | step 6,150 |
| Context length | 1024 during pretraining, extended to 2048 via NTK-aware RoPE rescaling |
| Precision | FP16 mixed precision with GradScaler (T4/Turing has no bf16 tensor cores) |
| Optimizer | 8-bit AdamW (bitsandbytes), lr 3e-4, weight decay 0.1, grad clip 1.0 |
| Schedule | Warmup-Stable-Decay, stopped at the end of the stable phase |
| Hardware | 1× NVIDIA T4 (16 GB), ~60 GPU-hours across resumable 5-hour sessions |
**On the schedule:** WSD normally ends with an LR decay phase. Training was
stopped at step 6,150 — exactly two epochs — because validation quality began
degrading past that point, so the decay was deliberately not run. Released
checkpoints from later in the run overfit.
## Evaluation
| Task | Metric | KW5-Lite Base | Chance |
|---|---|---|---|
| Belebele-sw (4-way) | Accuracy | 32% | 25% |
| AfriXNLI-sw (3-way) | Accuracy | 32% | 33% |
*n = 50 per task.*
**Interpret these honestly: this model is at or near chance on both.** With
n=50 the standard error is about ±6.6 points, so 32% vs 25% on Belebele is not
a reliable signal, and AfriXNLI is exactly at chance. A 110M model trained on
1.41B tokens is not expected to do multiple-choice reasoning; these numbers
establish a floor, not a capability.
What the model is genuinely good at is fluent, well-formed Swahili
continuation. That is what makes it a useful base to fine-tune, and it is what
the [instruct model](https://huggingface.co/regnant-io/kw5-lite-instruct)
builds on.
## Intended use
Intended as a **starting point for Swahili fine-tuning** — instruction tuning,
domain adaptation, classification heads — where training from scratch is too
expensive and larger multilingual models are too big to serve.
Not intended for direct deployment: no instruction following, no safety tuning,
no factual reliability.
## Limitations
- **Not an instruction model.** It continues text; it does not answer questions.
- **Factual reliability is poor.** Do not use as a knowledge source.
- **Primarily Tanzanian Swahili**, reflecting the corpus distribution.
- **No safety tuning at all.** It will reproduce harmful, biased or explicit
content present in web text.
- Trained on web-scraped data and carries its biases and quality artifacts.
## Files
| File | What it is |
|---|---|
| `model.safetensors` | FP32 weights, tied embeddings, 438 MB |
| `tokenizer.model` | SentencePiece 32k BPE model |
| `tokenizer_config.json` | `add_bos_token=false` — prepend `<s>` yourself |
| `generation_config.json` | Low-temperature defaults that keep this model coherent |
| `training_info.json` | Checkpoint step, tokens seen, benchmark scores |
A previous revision also shipped `pytorch_model.bin` (a stale duplicate of the
weights), duplicate tokenizer files under two names, and a `config.json` whose
`pad_token_id` was 0 while the tokenizer's `<pad>` is 3. All removed or
corrected.
## Citation
```bibtex
@misc{kw5lite2026,
title = {KW5-Lite: a 110M-parameter Swahili language model trained on a single T4},
author = {Regnant},
year = {2026},
url = {https://huggingface.co/regnant-io/kw5-lite-base}
}
```
Apache 2.0.