初始化项目,由ModelHub XC社区提供模型

Model: regnant-io/kw5-lite-base
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-09-06 09:12:17 +08:00
commit 28dcb5ee11
10 changed files with 255173 additions and 0 deletions

35
.gitattributes vendored Normal file
View File

@@ -0,0 +1,35 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text

182
README.md Normal file
View File

@@ -0,0 +1,182 @@
---
language:
- sw
license: apache-2.0
library_name: transformers
tags:
- swahili
- kiswahili
- causal-lm
- base-model
- pretrained
- african-languages
- low-resource
datasets:
- HuggingFaceFW/fineweb-2
pipeline_tag: text-generation
---
# KW5-Lite Base
A 109.5M-parameter Swahili (Kiswahili) language model pretrained from scratch on
1.41B tokens, on a single NVIDIA T4.
Built by [Regnant](https://www.regnant.io/).
**This is a base model.** It does next-token prediction only. It does not follow
instructions and has no chat template. For that, use
[regnant-io/kw5-lite-instruct](https://huggingface.co/regnant-io/kw5-lite-instruct).
## Quick start
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("regnant-io/kw5-lite-base")
tok = AutoTokenizer.from_pretrained("regnant-io/kw5-lite-base")
# Base model: give it a prefix to continue, not an instruction.
# Write <s> into the text — the tokenizer maps it to id 1.
inputs = tok("<s>Tanzania ni nchi", return_tensors="pt", add_special_tokens=False)
out = model.generate(**inputs, max_new_tokens=60)
print(tok.decode(out[0], skip_special_tokens=True))
```
```
Tanzania ni nchi ya amani na utulivu. Ni nchi yenye watu wengi, wenye nguvu
za kiuchumi, wenye uwezo wa kufanya maamuzi magumu kwa wakati mmoja...
```
`add_bos_token` is pinned to `false` so that the same tokenizer settings work
for the instruct model, which renders `<s>` as part of its prompt template.
**Put `<s>` at the start of your text yourself** — every training sequence
began with it.
Write it into the string, as above, rather than concatenating the id onto
`input_ids` afterwards: that leaves `attention_mask` one element shorter than
`input_ids`, and `generate` then fails inside RoPE with
`The size of tensor a (4) must match the size of tensor b (3)`.
Like the instruct model, this one needs a low temperature and a repetition
penalty to stay coherent; the shipped `generation_config.json` defaults to
temperature 0.2 / `repetition_penalty` 1.3.
## Model details
| | |
|---|---|
| Parameters | 109.5M (tied input/output embeddings) |
| Architecture | Llama-compatible decoder-only transformer |
| Layers / hidden / FFN | 12 / 768 / 2048 |
| Attention heads | 12 query, 12 key-value — **standard multi-head attention, not GQA** |
| Normalization | RMSNorm, pre-norm |
| Activation | SwiGLU |
| Position encoding | RoPE, theta 10000 |
| Context length | 2048 |
| Vocabulary | 32,000 SentencePiece BPE |
| Precision | FP32 (trained under FP16 autocast with an FP32 master copy) |
| File size | 438 MB |
> A previous revision of this card claimed "KV heads: 4 (Grouped Query
> Attention)". That was wrong — `num_key_value_heads` is 12 and always has
> been. This model uses standard MHA. Corrected here.
### Tokenizer
32k BPE trained from scratch on the deduplicated Swahili corpus. NFC
normalization applied before training (SentencePiece's own rules are all
NFKC-family, which is lossier), `byte_fallback` so no input can hard-fail, and
digits split.
Special tokens are fixed at low ids: `<unk>`=0, `<s>`=1, `</s>`=2, `<pad>`=3.
Ids 4–6 are `<|system|>`, `<|user|>`, `<|assistant|>`, **reserved but never
trained** — they do not occur in the pretraining corpus, so their embedding
rows are still at initialisation. If you fine-tune and want to use them as
chat-role markers, you must unfreeze `embed_tokens` (or add it to your LoRA
`modules_to_save`), or the model will read them as noise.
## Training
| | |
|---|---|
| Data | [FineWeb-2](https://huggingface.co/datasets/HuggingFaceFW/fineweb-2) `swh_Latn`, cleaned and deduplicated (exact + MinHash LSH) |
| Tokens seen | 1.41B — exactly 2 epochs over the corpus |
| Checkpoint | step 6,150 |
| Context length | 1024 during pretraining, extended to 2048 via NTK-aware RoPE rescaling |
| Precision | FP16 mixed precision with GradScaler (T4/Turing has no bf16 tensor cores) |
| Optimizer | 8-bit AdamW (bitsandbytes), lr 3e-4, weight decay 0.1, grad clip 1.0 |
| Schedule | Warmup-Stable-Decay, stopped at the end of the stable phase |
| Hardware | 1× NVIDIA T4 (16 GB), ~60 GPU-hours across resumable 5-hour sessions |
**On the schedule:** WSD normally ends with an LR decay phase. Training was
stopped at step 6,150 — exactly two epochs — because validation quality began
degrading past that point, so the decay was deliberately not run. Released
checkpoints from later in the run overfit.
## Evaluation
| Task | Metric | KW5-Lite Base | Chance |
|---|---|---|---|
| Belebele-sw (4-way) | Accuracy | 32% | 25% |
| AfriXNLI-sw (3-way) | Accuracy | 32% | 33% |
*n = 50 per task.*
**Interpret these honestly: this model is at or near chance on both.** With
n=50 the standard error is about ±6.6 points, so 32% vs 25% on Belebele is not
a reliable signal, and AfriXNLI is exactly at chance. A 110M model trained on
1.41B tokens is not expected to do multiple-choice reasoning; these numbers
establish a floor, not a capability.
What the model is genuinely good at is fluent, well-formed Swahili
continuation. That is what makes it a useful base to fine-tune, and it is what
the [instruct model](https://huggingface.co/regnant-io/kw5-lite-instruct)
builds on.
## Intended use
Intended as a **starting point for Swahili fine-tuning** — instruction tuning,
domain adaptation, classification heads — where training from scratch is too
expensive and larger multilingual models are too big to serve.
Not intended for direct deployment: no instruction following, no safety tuning,
no factual reliability.
## Limitations
- **Not an instruction model.** It continues text; it does not answer questions.
- **Factual reliability is poor.** Do not use as a knowledge source.
- **Primarily Tanzanian Swahili**, reflecting the corpus distribution.
- **No safety tuning at all.** It will reproduce harmful, biased or explicit
content present in web text.
- Trained on web-scraped data and carries its biases and quality artifacts.
## Files
| File | What it is |
|---|---|
| `model.safetensors` | FP32 weights, tied embeddings, 438 MB |
| `tokenizer.model` | SentencePiece 32k BPE model |
| `tokenizer_config.json` | `add_bos_token=false` — prepend `<s>` yourself |
| `generation_config.json` | Low-temperature defaults that keep this model coherent |
| `training_info.json` | Checkpoint step, tokens seen, benchmark scores |
A previous revision also shipped `pytorch_model.bin` (a stale duplicate of the
weights), duplicate tokenizer files under two names, and a `config.json` whose
`pad_token_id` was 0 while the tokenizer's `<pad>` is 3. All removed or
corrected.
## Citation
```bibtex
@misc{kw5lite2026,
title = {KW5-Lite: a 110M-parameter Swahili language model trained on a single T4},
author = {Regnant},
year = {2026},
url = {https://huggingface.co/regnant-io/kw5-lite-base}
}
```
Apache 2.0.

32
config.json Normal file
View File

@@ -0,0 +1,32 @@
{
"architectures": [
"LlamaForCausalLM"
],
"attention_bias": false,
"attention_dropout": 0.0,
"bos_token_id": 1,
"dtype": "float32",
"eos_token_id": 2,
"head_dim": 64,
"hidden_act": "silu",
"hidden_size": 768,
"initializer_range": 0.02,
"intermediate_size": 2048,
"max_position_embeddings": 2048,
"mlp_bias": false,
"model_type": "llama",
"num_attention_heads": 12,
"num_hidden_layers": 12,
"num_key_value_heads": 12,
"pad_token_id": 3,
"pretraining_tp": 1,
"rms_norm_eps": 1e-05,
"rope_parameters": {
"rope_theta": 10000.0,
"rope_type": "default"
},
"tie_word_embeddings": true,
"transformers_version": "5.16.1",
"use_cache": true,
"vocab_size": 32000
}

10
generation_config.json Normal file
View File

@@ -0,0 +1,10 @@
{
"bos_token_id": 1,
"eos_token_id": 2,
"pad_token_id": 3,
"do_sample": true,
"temperature": 0.2,
"top_p": 0.9,
"repetition_penalty": 1.3,
"max_new_tokens": 512
}

3
model.safetensors Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:961e6c296d49aeb6c0df43d6a0c9f933193fddfaafc8ff215af3374643e75414
size 438131648

6
special_tokens_map.json Normal file
View File

@@ -0,0 +1,6 @@
{
"bos_token": "<s>",
"eos_token": "</s>",
"unk_token": "<unk>",
"pad_token": "<pad>"
}

254873
tokenizer.json Normal file

File diff suppressed because it is too large Load Diff

3
tokenizer.model Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:696c2b00d0212edfe7b9f934fb13bab07be92d533182f27ad2355c1350161312
size 542002

14
tokenizer_config.json Normal file
View File

@@ -0,0 +1,14 @@
{
"tokenizer_class": "LlamaTokenizer",
"model_max_length": 2048,
"bos_token": "<s>",
"eos_token": "</s>",
"unk_token": "<unk>",
"pad_token": "<pad>",
"add_bos_token": false,
"add_eos_token": false,
"clean_up_tokenization_spaces": false,
"legacy": false,
"use_default_system_prompt": false,
"chat_template": "{%- set ns = namespace(out=bos_token, sys='', first=true) -%}{%- for m in messages -%}{%- if m['role'] == 'system' -%}{%- set ns.sys = m['content'] -%}{%- endif -%}{%- endfor -%}{%- for m in messages -%}{%- if m['role'] == 'user' -%}{%- set ns.out = ns.out + '[INST] ' -%}{%- if ns.first and ns.sys -%}{%- set ns.out = ns.out + '<<SYS>>\n' + ns.sys + '\n<</SYS>>\n\n' -%}{%- endif -%}{%- set ns.first = false -%}{%- set ns.out = ns.out + m['content'] + ' [/INST]' -%}{%- elif m['role'] == 'assistant' -%}{%- set ns.out = ns.out + ' ' + m['content'] + eos_token -%}{%- endif -%}{%- endfor -%}{{- ns.out -}}"
}

15
training_info.json Normal file
View File

@@ -0,0 +1,15 @@
{
"checkpoint_step": 6150,
"tokens_seen": 1410662400,
"epochs": 2.0,
"note": "Stopped at exactly 2 epochs; WSD decay deliberately not run because later checkpoints overfit.",
"benchmark_scores": {
"belebele_sw": 0.32,
"afrixnli_sw": 0.32,
"n_samples": 50,
"chance": {
"belebele_sw": 0.25,
"afrixnli_sw": 0.333
}
}
}