183 lines
6.7 KiB
Markdown
183 lines
6.7 KiB
Markdown
|
|
---
|
|||
|
|
language:
|
|||
|
|
- sw
|
|||
|
|
license: apache-2.0
|
|||
|
|
library_name: transformers
|
|||
|
|
tags:
|
|||
|
|
- swahili
|
|||
|
|
- kiswahili
|
|||
|
|
- causal-lm
|
|||
|
|
- base-model
|
|||
|
|
- pretrained
|
|||
|
|
- african-languages
|
|||
|
|
- low-resource
|
|||
|
|
datasets:
|
|||
|
|
- HuggingFaceFW/fineweb-2
|
|||
|
|
pipeline_tag: text-generation
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# KW5-Lite Base
|
|||
|
|
|
|||
|
|
A 109.5M-parameter Swahili (Kiswahili) language model pretrained from scratch on
|
|||
|
|
1.41B tokens, on a single NVIDIA T4.
|
|||
|
|
|
|||
|
|
Built by [Regnant](https://www.regnant.io/).
|
|||
|
|
|
|||
|
|
**This is a base model.** It does next-token prediction only. It does not follow
|
|||
|
|
instructions and has no chat template. For that, use
|
|||
|
|
[regnant-io/kw5-lite-instruct](https://huggingface.co/regnant-io/kw5-lite-instruct).
|
|||
|
|
|
|||
|
|
## Quick start
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
from transformers import AutoModelForCausalLM, AutoTokenizer
|
|||
|
|
|
|||
|
|
model = AutoModelForCausalLM.from_pretrained("regnant-io/kw5-lite-base")
|
|||
|
|
tok = AutoTokenizer.from_pretrained("regnant-io/kw5-lite-base")
|
|||
|
|
|
|||
|
|
# Base model: give it a prefix to continue, not an instruction.
|
|||
|
|
# Write <s> into the text — the tokenizer maps it to id 1.
|
|||
|
|
inputs = tok("<s>Tanzania ni nchi", return_tensors="pt", add_special_tokens=False)
|
|||
|
|
|
|||
|
|
out = model.generate(**inputs, max_new_tokens=60)
|
|||
|
|
print(tok.decode(out[0], skip_special_tokens=True))
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
```
|
|||
|
|
Tanzania ni nchi ya amani na utulivu. Ni nchi yenye watu wengi, wenye nguvu
|
|||
|
|
za kiuchumi, wenye uwezo wa kufanya maamuzi magumu kwa wakati mmoja...
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
`add_bos_token` is pinned to `false` so that the same tokenizer settings work
|
|||
|
|
for the instruct model, which renders `<s>` as part of its prompt template.
|
|||
|
|
**Put `<s>` at the start of your text yourself** — every training sequence
|
|||
|
|
began with it.
|
|||
|
|
|
|||
|
|
Write it into the string, as above, rather than concatenating the id onto
|
|||
|
|
`input_ids` afterwards: that leaves `attention_mask` one element shorter than
|
|||
|
|
`input_ids`, and `generate` then fails inside RoPE with
|
|||
|
|
`The size of tensor a (4) must match the size of tensor b (3)`.
|
|||
|
|
|
|||
|
|
Like the instruct model, this one needs a low temperature and a repetition
|
|||
|
|
penalty to stay coherent; the shipped `generation_config.json` defaults to
|
|||
|
|
temperature 0.2 / `repetition_penalty` 1.3.
|
|||
|
|
|
|||
|
|
## Model details
|
|||
|
|
|
|||
|
|
| | |
|
|||
|
|
|---|---|
|
|||
|
|
| Parameters | 109.5M (tied input/output embeddings) |
|
|||
|
|
| Architecture | Llama-compatible decoder-only transformer |
|
|||
|
|
| Layers / hidden / FFN | 12 / 768 / 2048 |
|
|||
|
|
| Attention heads | 12 query, 12 key-value — **standard multi-head attention, not GQA** |
|
|||
|
|
| Normalization | RMSNorm, pre-norm |
|
|||
|
|
| Activation | SwiGLU |
|
|||
|
|
| Position encoding | RoPE, theta 10000 |
|
|||
|
|
| Context length | 2048 |
|
|||
|
|
| Vocabulary | 32,000 SentencePiece BPE |
|
|||
|
|
| Precision | FP32 (trained under FP16 autocast with an FP32 master copy) |
|
|||
|
|
| File size | 438 MB |
|
|||
|
|
|
|||
|
|
> A previous revision of this card claimed "KV heads: 4 (Grouped Query
|
|||
|
|
> Attention)". That was wrong — `num_key_value_heads` is 12 and always has
|
|||
|
|
> been. This model uses standard MHA. Corrected here.
|
|||
|
|
|
|||
|
|
### Tokenizer
|
|||
|
|
|
|||
|
|
32k BPE trained from scratch on the deduplicated Swahili corpus. NFC
|
|||
|
|
normalization applied before training (SentencePiece's own rules are all
|
|||
|
|
NFKC-family, which is lossier), `byte_fallback` so no input can hard-fail, and
|
|||
|
|
digits split.
|
|||
|
|
|
|||
|
|
Special tokens are fixed at low ids: `<unk>`=0, `<s>`=1, `</s>`=2, `<pad>`=3.
|
|||
|
|
|
|||
|
|
Ids 4–6 are `<|system|>`, `<|user|>`, `<|assistant|>`, **reserved but never
|
|||
|
|
trained** — they do not occur in the pretraining corpus, so their embedding
|
|||
|
|
rows are still at initialisation. If you fine-tune and want to use them as
|
|||
|
|
chat-role markers, you must unfreeze `embed_tokens` (or add it to your LoRA
|
|||
|
|
`modules_to_save`), or the model will read them as noise.
|
|||
|
|
|
|||
|
|
## Training
|
|||
|
|
|
|||
|
|
| | |
|
|||
|
|
|---|---|
|
|||
|
|
| Data | [FineWeb-2](https://huggingface.co/datasets/HuggingFaceFW/fineweb-2) `swh_Latn`, cleaned and deduplicated (exact + MinHash LSH) |
|
|||
|
|
| Tokens seen | 1.41B — exactly 2 epochs over the corpus |
|
|||
|
|
| Checkpoint | step 6,150 |
|
|||
|
|
| Context length | 1024 during pretraining, extended to 2048 via NTK-aware RoPE rescaling |
|
|||
|
|
| Precision | FP16 mixed precision with GradScaler (T4/Turing has no bf16 tensor cores) |
|
|||
|
|
| Optimizer | 8-bit AdamW (bitsandbytes), lr 3e-4, weight decay 0.1, grad clip 1.0 |
|
|||
|
|
| Schedule | Warmup-Stable-Decay, stopped at the end of the stable phase |
|
|||
|
|
| Hardware | 1× NVIDIA T4 (16 GB), ~60 GPU-hours across resumable 5-hour sessions |
|
|||
|
|
|
|||
|
|
**On the schedule:** WSD normally ends with an LR decay phase. Training was
|
|||
|
|
stopped at step 6,150 — exactly two epochs — because validation quality began
|
|||
|
|
degrading past that point, so the decay was deliberately not run. Released
|
|||
|
|
checkpoints from later in the run overfit.
|
|||
|
|
|
|||
|
|
## Evaluation
|
|||
|
|
|
|||
|
|
| Task | Metric | KW5-Lite Base | Chance |
|
|||
|
|
|---|---|---|---|
|
|||
|
|
| Belebele-sw (4-way) | Accuracy | 32% | 25% |
|
|||
|
|
| AfriXNLI-sw (3-way) | Accuracy | 32% | 33% |
|
|||
|
|
|
|||
|
|
*n = 50 per task.*
|
|||
|
|
|
|||
|
|
**Interpret these honestly: this model is at or near chance on both.** With
|
|||
|
|
n=50 the standard error is about ±6.6 points, so 32% vs 25% on Belebele is not
|
|||
|
|
a reliable signal, and AfriXNLI is exactly at chance. A 110M model trained on
|
|||
|
|
1.41B tokens is not expected to do multiple-choice reasoning; these numbers
|
|||
|
|
establish a floor, not a capability.
|
|||
|
|
|
|||
|
|
What the model is genuinely good at is fluent, well-formed Swahili
|
|||
|
|
continuation. That is what makes it a useful base to fine-tune, and it is what
|
|||
|
|
the [instruct model](https://huggingface.co/regnant-io/kw5-lite-instruct)
|
|||
|
|
builds on.
|
|||
|
|
|
|||
|
|
## Intended use
|
|||
|
|
|
|||
|
|
Intended as a **starting point for Swahili fine-tuning** — instruction tuning,
|
|||
|
|
domain adaptation, classification heads — where training from scratch is too
|
|||
|
|
expensive and larger multilingual models are too big to serve.
|
|||
|
|
|
|||
|
|
Not intended for direct deployment: no instruction following, no safety tuning,
|
|||
|
|
no factual reliability.
|
|||
|
|
|
|||
|
|
## Limitations
|
|||
|
|
|
|||
|
|
- **Not an instruction model.** It continues text; it does not answer questions.
|
|||
|
|
- **Factual reliability is poor.** Do not use as a knowledge source.
|
|||
|
|
- **Primarily Tanzanian Swahili**, reflecting the corpus distribution.
|
|||
|
|
- **No safety tuning at all.** It will reproduce harmful, biased or explicit
|
|||
|
|
content present in web text.
|
|||
|
|
- Trained on web-scraped data and carries its biases and quality artifacts.
|
|||
|
|
|
|||
|
|
## Files
|
|||
|
|
|
|||
|
|
| File | What it is |
|
|||
|
|
|---|---|
|
|||
|
|
| `model.safetensors` | FP32 weights, tied embeddings, 438 MB |
|
|||
|
|
| `tokenizer.model` | SentencePiece 32k BPE model |
|
|||
|
|
| `tokenizer_config.json` | `add_bos_token=false` — prepend `<s>` yourself |
|
|||
|
|
| `generation_config.json` | Low-temperature defaults that keep this model coherent |
|
|||
|
|
| `training_info.json` | Checkpoint step, tokens seen, benchmark scores |
|
|||
|
|
|
|||
|
|
A previous revision also shipped `pytorch_model.bin` (a stale duplicate of the
|
|||
|
|
weights), duplicate tokenizer files under two names, and a `config.json` whose
|
|||
|
|
`pad_token_id` was 0 while the tokenizer's `<pad>` is 3. All removed or
|
|||
|
|
corrected.
|
|||
|
|
|
|||
|
|
## Citation
|
|||
|
|
|
|||
|
|
```bibtex
|
|||
|
|
@misc{kw5lite2026,
|
|||
|
|
title = {KW5-Lite: a 110M-parameter Swahili language model trained on a single T4},
|
|||
|
|
author = {Regnant},
|
|||
|
|
year = {2026},
|
|||
|
|
url = {https://huggingface.co/regnant-io/kw5-lite-base}
|
|||
|
|
}
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Apache 2.0.
|