初始化项目,由ModelHub XC社区提供模型
Model: regnant-io/kw5-lite-base Source: Original Platform
This commit is contained in:
35
.gitattributes
vendored
Normal file
35
.gitattributes
vendored
Normal file
@@ -0,0 +1,35 @@
|
||||
*.7z filter=lfs diff=lfs merge=lfs -text
|
||||
*.arrow filter=lfs diff=lfs merge=lfs -text
|
||||
*.bin filter=lfs diff=lfs merge=lfs -text
|
||||
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
||||
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
||||
*.ftz filter=lfs diff=lfs merge=lfs -text
|
||||
*.gz filter=lfs diff=lfs merge=lfs -text
|
||||
*.h5 filter=lfs diff=lfs merge=lfs -text
|
||||
*.joblib filter=lfs diff=lfs merge=lfs -text
|
||||
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
||||
*.model filter=lfs diff=lfs merge=lfs -text
|
||||
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
||||
*.npy filter=lfs diff=lfs merge=lfs -text
|
||||
*.npz filter=lfs diff=lfs merge=lfs -text
|
||||
*.onnx filter=lfs diff=lfs merge=lfs -text
|
||||
*.ot filter=lfs diff=lfs merge=lfs -text
|
||||
*.parquet filter=lfs diff=lfs merge=lfs -text
|
||||
*.pb filter=lfs diff=lfs merge=lfs -text
|
||||
*.pickle filter=lfs diff=lfs merge=lfs -text
|
||||
*.pkl filter=lfs diff=lfs merge=lfs -text
|
||||
*.pt filter=lfs diff=lfs merge=lfs -text
|
||||
*.pth filter=lfs diff=lfs merge=lfs -text
|
||||
*.rar filter=lfs diff=lfs merge=lfs -text
|
||||
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
||||
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar filter=lfs diff=lfs merge=lfs -text
|
||||
*.tflite filter=lfs diff=lfs merge=lfs -text
|
||||
*.tgz filter=lfs diff=lfs merge=lfs -text
|
||||
*.wasm filter=lfs diff=lfs merge=lfs -text
|
||||
*.xz filter=lfs diff=lfs merge=lfs -text
|
||||
*.zip filter=lfs diff=lfs merge=lfs -text
|
||||
*.zst filter=lfs diff=lfs merge=lfs -text
|
||||
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
||||
182
README.md
Normal file
182
README.md
Normal file
@@ -0,0 +1,182 @@
|
||||
---
|
||||
language:
|
||||
- sw
|
||||
license: apache-2.0
|
||||
library_name: transformers
|
||||
tags:
|
||||
- swahili
|
||||
- kiswahili
|
||||
- causal-lm
|
||||
- base-model
|
||||
- pretrained
|
||||
- african-languages
|
||||
- low-resource
|
||||
datasets:
|
||||
- HuggingFaceFW/fineweb-2
|
||||
pipeline_tag: text-generation
|
||||
---
|
||||
|
||||
# KW5-Lite Base
|
||||
|
||||
A 109.5M-parameter Swahili (Kiswahili) language model pretrained from scratch on
|
||||
1.41B tokens, on a single NVIDIA T4.
|
||||
|
||||
Built by [Regnant](https://www.regnant.io/).
|
||||
|
||||
**This is a base model.** It does next-token prediction only. It does not follow
|
||||
instructions and has no chat template. For that, use
|
||||
[regnant-io/kw5-lite-instruct](https://huggingface.co/regnant-io/kw5-lite-instruct).
|
||||
|
||||
## Quick start
|
||||
|
||||
```python
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
|
||||
model = AutoModelForCausalLM.from_pretrained("regnant-io/kw5-lite-base")
|
||||
tok = AutoTokenizer.from_pretrained("regnant-io/kw5-lite-base")
|
||||
|
||||
# Base model: give it a prefix to continue, not an instruction.
|
||||
# Write <s> into the text — the tokenizer maps it to id 1.
|
||||
inputs = tok("<s>Tanzania ni nchi", return_tensors="pt", add_special_tokens=False)
|
||||
|
||||
out = model.generate(**inputs, max_new_tokens=60)
|
||||
print(tok.decode(out[0], skip_special_tokens=True))
|
||||
```
|
||||
|
||||
```
|
||||
Tanzania ni nchi ya amani na utulivu. Ni nchi yenye watu wengi, wenye nguvu
|
||||
za kiuchumi, wenye uwezo wa kufanya maamuzi magumu kwa wakati mmoja...
|
||||
```
|
||||
|
||||
`add_bos_token` is pinned to `false` so that the same tokenizer settings work
|
||||
for the instruct model, which renders `<s>` as part of its prompt template.
|
||||
**Put `<s>` at the start of your text yourself** — every training sequence
|
||||
began with it.
|
||||
|
||||
Write it into the string, as above, rather than concatenating the id onto
|
||||
`input_ids` afterwards: that leaves `attention_mask` one element shorter than
|
||||
`input_ids`, and `generate` then fails inside RoPE with
|
||||
`The size of tensor a (4) must match the size of tensor b (3)`.
|
||||
|
||||
Like the instruct model, this one needs a low temperature and a repetition
|
||||
penalty to stay coherent; the shipped `generation_config.json` defaults to
|
||||
temperature 0.2 / `repetition_penalty` 1.3.
|
||||
|
||||
## Model details
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| Parameters | 109.5M (tied input/output embeddings) |
|
||||
| Architecture | Llama-compatible decoder-only transformer |
|
||||
| Layers / hidden / FFN | 12 / 768 / 2048 |
|
||||
| Attention heads | 12 query, 12 key-value — **standard multi-head attention, not GQA** |
|
||||
| Normalization | RMSNorm, pre-norm |
|
||||
| Activation | SwiGLU |
|
||||
| Position encoding | RoPE, theta 10000 |
|
||||
| Context length | 2048 |
|
||||
| Vocabulary | 32,000 SentencePiece BPE |
|
||||
| Precision | FP32 (trained under FP16 autocast with an FP32 master copy) |
|
||||
| File size | 438 MB |
|
||||
|
||||
> A previous revision of this card claimed "KV heads: 4 (Grouped Query
|
||||
> Attention)". That was wrong — `num_key_value_heads` is 12 and always has
|
||||
> been. This model uses standard MHA. Corrected here.
|
||||
|
||||
### Tokenizer
|
||||
|
||||
32k BPE trained from scratch on the deduplicated Swahili corpus. NFC
|
||||
normalization applied before training (SentencePiece's own rules are all
|
||||
NFKC-family, which is lossier), `byte_fallback` so no input can hard-fail, and
|
||||
digits split.
|
||||
|
||||
Special tokens are fixed at low ids: `<unk>`=0, `<s>`=1, `</s>`=2, `<pad>`=3.
|
||||
|
||||
Ids 4–6 are `<|system|>`, `<|user|>`, `<|assistant|>`, **reserved but never
|
||||
trained** — they do not occur in the pretraining corpus, so their embedding
|
||||
rows are still at initialisation. If you fine-tune and want to use them as
|
||||
chat-role markers, you must unfreeze `embed_tokens` (or add it to your LoRA
|
||||
`modules_to_save`), or the model will read them as noise.
|
||||
|
||||
## Training
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| Data | [FineWeb-2](https://huggingface.co/datasets/HuggingFaceFW/fineweb-2) `swh_Latn`, cleaned and deduplicated (exact + MinHash LSH) |
|
||||
| Tokens seen | 1.41B — exactly 2 epochs over the corpus |
|
||||
| Checkpoint | step 6,150 |
|
||||
| Context length | 1024 during pretraining, extended to 2048 via NTK-aware RoPE rescaling |
|
||||
| Precision | FP16 mixed precision with GradScaler (T4/Turing has no bf16 tensor cores) |
|
||||
| Optimizer | 8-bit AdamW (bitsandbytes), lr 3e-4, weight decay 0.1, grad clip 1.0 |
|
||||
| Schedule | Warmup-Stable-Decay, stopped at the end of the stable phase |
|
||||
| Hardware | 1× NVIDIA T4 (16 GB), ~60 GPU-hours across resumable 5-hour sessions |
|
||||
|
||||
**On the schedule:** WSD normally ends with an LR decay phase. Training was
|
||||
stopped at step 6,150 — exactly two epochs — because validation quality began
|
||||
degrading past that point, so the decay was deliberately not run. Released
|
||||
checkpoints from later in the run overfit.
|
||||
|
||||
## Evaluation
|
||||
|
||||
| Task | Metric | KW5-Lite Base | Chance |
|
||||
|---|---|---|---|
|
||||
| Belebele-sw (4-way) | Accuracy | 32% | 25% |
|
||||
| AfriXNLI-sw (3-way) | Accuracy | 32% | 33% |
|
||||
|
||||
*n = 50 per task.*
|
||||
|
||||
**Interpret these honestly: this model is at or near chance on both.** With
|
||||
n=50 the standard error is about ±6.6 points, so 32% vs 25% on Belebele is not
|
||||
a reliable signal, and AfriXNLI is exactly at chance. A 110M model trained on
|
||||
1.41B tokens is not expected to do multiple-choice reasoning; these numbers
|
||||
establish a floor, not a capability.
|
||||
|
||||
What the model is genuinely good at is fluent, well-formed Swahili
|
||||
continuation. That is what makes it a useful base to fine-tune, and it is what
|
||||
the [instruct model](https://huggingface.co/regnant-io/kw5-lite-instruct)
|
||||
builds on.
|
||||
|
||||
## Intended use
|
||||
|
||||
Intended as a **starting point for Swahili fine-tuning** — instruction tuning,
|
||||
domain adaptation, classification heads — where training from scratch is too
|
||||
expensive and larger multilingual models are too big to serve.
|
||||
|
||||
Not intended for direct deployment: no instruction following, no safety tuning,
|
||||
no factual reliability.
|
||||
|
||||
## Limitations
|
||||
|
||||
- **Not an instruction model.** It continues text; it does not answer questions.
|
||||
- **Factual reliability is poor.** Do not use as a knowledge source.
|
||||
- **Primarily Tanzanian Swahili**, reflecting the corpus distribution.
|
||||
- **No safety tuning at all.** It will reproduce harmful, biased or explicit
|
||||
content present in web text.
|
||||
- Trained on web-scraped data and carries its biases and quality artifacts.
|
||||
|
||||
## Files
|
||||
|
||||
| File | What it is |
|
||||
|---|---|
|
||||
| `model.safetensors` | FP32 weights, tied embeddings, 438 MB |
|
||||
| `tokenizer.model` | SentencePiece 32k BPE model |
|
||||
| `tokenizer_config.json` | `add_bos_token=false` — prepend `<s>` yourself |
|
||||
| `generation_config.json` | Low-temperature defaults that keep this model coherent |
|
||||
| `training_info.json` | Checkpoint step, tokens seen, benchmark scores |
|
||||
|
||||
A previous revision also shipped `pytorch_model.bin` (a stale duplicate of the
|
||||
weights), duplicate tokenizer files under two names, and a `config.json` whose
|
||||
`pad_token_id` was 0 while the tokenizer's `<pad>` is 3. All removed or
|
||||
corrected.
|
||||
|
||||
## Citation
|
||||
|
||||
```bibtex
|
||||
@misc{kw5lite2026,
|
||||
title = {KW5-Lite: a 110M-parameter Swahili language model trained on a single T4},
|
||||
author = {Regnant},
|
||||
year = {2026},
|
||||
url = {https://huggingface.co/regnant-io/kw5-lite-base}
|
||||
}
|
||||
```
|
||||
|
||||
Apache 2.0.
|
||||
32
config.json
Normal file
32
config.json
Normal file
@@ -0,0 +1,32 @@
|
||||
{
|
||||
"architectures": [
|
||||
"LlamaForCausalLM"
|
||||
],
|
||||
"attention_bias": false,
|
||||
"attention_dropout": 0.0,
|
||||
"bos_token_id": 1,
|
||||
"dtype": "float32",
|
||||
"eos_token_id": 2,
|
||||
"head_dim": 64,
|
||||
"hidden_act": "silu",
|
||||
"hidden_size": 768,
|
||||
"initializer_range": 0.02,
|
||||
"intermediate_size": 2048,
|
||||
"max_position_embeddings": 2048,
|
||||
"mlp_bias": false,
|
||||
"model_type": "llama",
|
||||
"num_attention_heads": 12,
|
||||
"num_hidden_layers": 12,
|
||||
"num_key_value_heads": 12,
|
||||
"pad_token_id": 3,
|
||||
"pretraining_tp": 1,
|
||||
"rms_norm_eps": 1e-05,
|
||||
"rope_parameters": {
|
||||
"rope_theta": 10000.0,
|
||||
"rope_type": "default"
|
||||
},
|
||||
"tie_word_embeddings": true,
|
||||
"transformers_version": "5.16.1",
|
||||
"use_cache": true,
|
||||
"vocab_size": 32000
|
||||
}
|
||||
10
generation_config.json
Normal file
10
generation_config.json
Normal file
@@ -0,0 +1,10 @@
|
||||
{
|
||||
"bos_token_id": 1,
|
||||
"eos_token_id": 2,
|
||||
"pad_token_id": 3,
|
||||
"do_sample": true,
|
||||
"temperature": 0.2,
|
||||
"top_p": 0.9,
|
||||
"repetition_penalty": 1.3,
|
||||
"max_new_tokens": 512
|
||||
}
|
||||
3
model.safetensors
Normal file
3
model.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:961e6c296d49aeb6c0df43d6a0c9f933193fddfaafc8ff215af3374643e75414
|
||||
size 438131648
|
||||
6
special_tokens_map.json
Normal file
6
special_tokens_map.json
Normal file
@@ -0,0 +1,6 @@
|
||||
{
|
||||
"bos_token": "<s>",
|
||||
"eos_token": "</s>",
|
||||
"unk_token": "<unk>",
|
||||
"pad_token": "<pad>"
|
||||
}
|
||||
254873
tokenizer.json
Normal file
254873
tokenizer.json
Normal file
File diff suppressed because it is too large
Load Diff
3
tokenizer.model
Normal file
3
tokenizer.model
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:696c2b00d0212edfe7b9f934fb13bab07be92d533182f27ad2355c1350161312
|
||||
size 542002
|
||||
14
tokenizer_config.json
Normal file
14
tokenizer_config.json
Normal file
@@ -0,0 +1,14 @@
|
||||
{
|
||||
"tokenizer_class": "LlamaTokenizer",
|
||||
"model_max_length": 2048,
|
||||
"bos_token": "<s>",
|
||||
"eos_token": "</s>",
|
||||
"unk_token": "<unk>",
|
||||
"pad_token": "<pad>",
|
||||
"add_bos_token": false,
|
||||
"add_eos_token": false,
|
||||
"clean_up_tokenization_spaces": false,
|
||||
"legacy": false,
|
||||
"use_default_system_prompt": false,
|
||||
"chat_template": "{%- set ns = namespace(out=bos_token, sys='', first=true) -%}{%- for m in messages -%}{%- if m['role'] == 'system' -%}{%- set ns.sys = m['content'] -%}{%- endif -%}{%- endfor -%}{%- for m in messages -%}{%- if m['role'] == 'user' -%}{%- set ns.out = ns.out + '[INST] ' -%}{%- if ns.first and ns.sys -%}{%- set ns.out = ns.out + '<<SYS>>\n' + ns.sys + '\n<</SYS>>\n\n' -%}{%- endif -%}{%- set ns.first = false -%}{%- set ns.out = ns.out + m['content'] + ' [/INST]' -%}{%- elif m['role'] == 'assistant' -%}{%- set ns.out = ns.out + ' ' + m['content'] + eos_token -%}{%- endif -%}{%- endfor -%}{{- ns.out -}}"
|
||||
}
|
||||
15
training_info.json
Normal file
15
training_info.json
Normal file
@@ -0,0 +1,15 @@
|
||||
{
|
||||
"checkpoint_step": 6150,
|
||||
"tokens_seen": 1410662400,
|
||||
"epochs": 2.0,
|
||||
"note": "Stopped at exactly 2 epochs; WSD decay deliberately not run because later checkpoints overfit.",
|
||||
"benchmark_scores": {
|
||||
"belebele_sw": 0.32,
|
||||
"afrixnli_sw": 0.32,
|
||||
"n_samples": 50,
|
||||
"chance": {
|
||||
"belebele_sw": 0.25,
|
||||
"afrixnli_sw": 0.333
|
||||
}
|
||||
}
|
||||
}
|
||||
Reference in New Issue
Block a user