初始化项目,由ModelHub XC社区提供模型
Model: regnant-io/kw5-lite-base Source: Original Platform
This commit is contained in:
35
.gitattributes
vendored
Normal file
35
.gitattributes
vendored
Normal file
@@ -0,0 +1,35 @@
|
|||||||
|
*.7z filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.arrow filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.bin filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.ftz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.gz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.h5 filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.joblib filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.model filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.npy filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.npz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.onnx filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.ot filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.parquet filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pb filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pickle filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pkl filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pt filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pth filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.rar filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
||||||
|
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tar filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tflite filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tgz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.wasm filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.xz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.zip filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.zst filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
||||||
182
README.md
Normal file
182
README.md
Normal file
@@ -0,0 +1,182 @@
|
|||||||
|
---
|
||||||
|
language:
|
||||||
|
- sw
|
||||||
|
license: apache-2.0
|
||||||
|
library_name: transformers
|
||||||
|
tags:
|
||||||
|
- swahili
|
||||||
|
- kiswahili
|
||||||
|
- causal-lm
|
||||||
|
- base-model
|
||||||
|
- pretrained
|
||||||
|
- african-languages
|
||||||
|
- low-resource
|
||||||
|
datasets:
|
||||||
|
- HuggingFaceFW/fineweb-2
|
||||||
|
pipeline_tag: text-generation
|
||||||
|
---
|
||||||
|
|
||||||
|
# KW5-Lite Base
|
||||||
|
|
||||||
|
A 109.5M-parameter Swahili (Kiswahili) language model pretrained from scratch on
|
||||||
|
1.41B tokens, on a single NVIDIA T4.
|
||||||
|
|
||||||
|
Built by [Regnant](https://www.regnant.io/).
|
||||||
|
|
||||||
|
**This is a base model.** It does next-token prediction only. It does not follow
|
||||||
|
instructions and has no chat template. For that, use
|
||||||
|
[regnant-io/kw5-lite-instruct](https://huggingface.co/regnant-io/kw5-lite-instruct).
|
||||||
|
|
||||||
|
## Quick start
|
||||||
|
|
||||||
|
```python
|
||||||
|
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||||
|
|
||||||
|
model = AutoModelForCausalLM.from_pretrained("regnant-io/kw5-lite-base")
|
||||||
|
tok = AutoTokenizer.from_pretrained("regnant-io/kw5-lite-base")
|
||||||
|
|
||||||
|
# Base model: give it a prefix to continue, not an instruction.
|
||||||
|
# Write <s> into the text — the tokenizer maps it to id 1.
|
||||||
|
inputs = tok("<s>Tanzania ni nchi", return_tensors="pt", add_special_tokens=False)
|
||||||
|
|
||||||
|
out = model.generate(**inputs, max_new_tokens=60)
|
||||||
|
print(tok.decode(out[0], skip_special_tokens=True))
|
||||||
|
```
|
||||||
|
|
||||||
|
```
|
||||||
|
Tanzania ni nchi ya amani na utulivu. Ni nchi yenye watu wengi, wenye nguvu
|
||||||
|
za kiuchumi, wenye uwezo wa kufanya maamuzi magumu kwa wakati mmoja...
|
||||||
|
```
|
||||||
|
|
||||||
|
`add_bos_token` is pinned to `false` so that the same tokenizer settings work
|
||||||
|
for the instruct model, which renders `<s>` as part of its prompt template.
|
||||||
|
**Put `<s>` at the start of your text yourself** — every training sequence
|
||||||
|
began with it.
|
||||||
|
|
||||||
|
Write it into the string, as above, rather than concatenating the id onto
|
||||||
|
`input_ids` afterwards: that leaves `attention_mask` one element shorter than
|
||||||
|
`input_ids`, and `generate` then fails inside RoPE with
|
||||||
|
`The size of tensor a (4) must match the size of tensor b (3)`.
|
||||||
|
|
||||||
|
Like the instruct model, this one needs a low temperature and a repetition
|
||||||
|
penalty to stay coherent; the shipped `generation_config.json` defaults to
|
||||||
|
temperature 0.2 / `repetition_penalty` 1.3.
|
||||||
|
|
||||||
|
## Model details
|
||||||
|
|
||||||
|
| | |
|
||||||
|
|---|---|
|
||||||
|
| Parameters | 109.5M (tied input/output embeddings) |
|
||||||
|
| Architecture | Llama-compatible decoder-only transformer |
|
||||||
|
| Layers / hidden / FFN | 12 / 768 / 2048 |
|
||||||
|
| Attention heads | 12 query, 12 key-value — **standard multi-head attention, not GQA** |
|
||||||
|
| Normalization | RMSNorm, pre-norm |
|
||||||
|
| Activation | SwiGLU |
|
||||||
|
| Position encoding | RoPE, theta 10000 |
|
||||||
|
| Context length | 2048 |
|
||||||
|
| Vocabulary | 32,000 SentencePiece BPE |
|
||||||
|
| Precision | FP32 (trained under FP16 autocast with an FP32 master copy) |
|
||||||
|
| File size | 438 MB |
|
||||||
|
|
||||||
|
> A previous revision of this card claimed "KV heads: 4 (Grouped Query
|
||||||
|
> Attention)". That was wrong — `num_key_value_heads` is 12 and always has
|
||||||
|
> been. This model uses standard MHA. Corrected here.
|
||||||
|
|
||||||
|
### Tokenizer
|
||||||
|
|
||||||
|
32k BPE trained from scratch on the deduplicated Swahili corpus. NFC
|
||||||
|
normalization applied before training (SentencePiece's own rules are all
|
||||||
|
NFKC-family, which is lossier), `byte_fallback` so no input can hard-fail, and
|
||||||
|
digits split.
|
||||||
|
|
||||||
|
Special tokens are fixed at low ids: `<unk>`=0, `<s>`=1, `</s>`=2, `<pad>`=3.
|
||||||
|
|
||||||
|
Ids 4–6 are `<|system|>`, `<|user|>`, `<|assistant|>`, **reserved but never
|
||||||
|
trained** — they do not occur in the pretraining corpus, so their embedding
|
||||||
|
rows are still at initialisation. If you fine-tune and want to use them as
|
||||||
|
chat-role markers, you must unfreeze `embed_tokens` (or add it to your LoRA
|
||||||
|
`modules_to_save`), or the model will read them as noise.
|
||||||
|
|
||||||
|
## Training
|
||||||
|
|
||||||
|
| | |
|
||||||
|
|---|---|
|
||||||
|
| Data | [FineWeb-2](https://huggingface.co/datasets/HuggingFaceFW/fineweb-2) `swh_Latn`, cleaned and deduplicated (exact + MinHash LSH) |
|
||||||
|
| Tokens seen | 1.41B — exactly 2 epochs over the corpus |
|
||||||
|
| Checkpoint | step 6,150 |
|
||||||
|
| Context length | 1024 during pretraining, extended to 2048 via NTK-aware RoPE rescaling |
|
||||||
|
| Precision | FP16 mixed precision with GradScaler (T4/Turing has no bf16 tensor cores) |
|
||||||
|
| Optimizer | 8-bit AdamW (bitsandbytes), lr 3e-4, weight decay 0.1, grad clip 1.0 |
|
||||||
|
| Schedule | Warmup-Stable-Decay, stopped at the end of the stable phase |
|
||||||
|
| Hardware | 1× NVIDIA T4 (16 GB), ~60 GPU-hours across resumable 5-hour sessions |
|
||||||
|
|
||||||
|
**On the schedule:** WSD normally ends with an LR decay phase. Training was
|
||||||
|
stopped at step 6,150 — exactly two epochs — because validation quality began
|
||||||
|
degrading past that point, so the decay was deliberately not run. Released
|
||||||
|
checkpoints from later in the run overfit.
|
||||||
|
|
||||||
|
## Evaluation
|
||||||
|
|
||||||
|
| Task | Metric | KW5-Lite Base | Chance |
|
||||||
|
|---|---|---|---|
|
||||||
|
| Belebele-sw (4-way) | Accuracy | 32% | 25% |
|
||||||
|
| AfriXNLI-sw (3-way) | Accuracy | 32% | 33% |
|
||||||
|
|
||||||
|
*n = 50 per task.*
|
||||||
|
|
||||||
|
**Interpret these honestly: this model is at or near chance on both.** With
|
||||||
|
n=50 the standard error is about ±6.6 points, so 32% vs 25% on Belebele is not
|
||||||
|
a reliable signal, and AfriXNLI is exactly at chance. A 110M model trained on
|
||||||
|
1.41B tokens is not expected to do multiple-choice reasoning; these numbers
|
||||||
|
establish a floor, not a capability.
|
||||||
|
|
||||||
|
What the model is genuinely good at is fluent, well-formed Swahili
|
||||||
|
continuation. That is what makes it a useful base to fine-tune, and it is what
|
||||||
|
the [instruct model](https://huggingface.co/regnant-io/kw5-lite-instruct)
|
||||||
|
builds on.
|
||||||
|
|
||||||
|
## Intended use
|
||||||
|
|
||||||
|
Intended as a **starting point for Swahili fine-tuning** — instruction tuning,
|
||||||
|
domain adaptation, classification heads — where training from scratch is too
|
||||||
|
expensive and larger multilingual models are too big to serve.
|
||||||
|
|
||||||
|
Not intended for direct deployment: no instruction following, no safety tuning,
|
||||||
|
no factual reliability.
|
||||||
|
|
||||||
|
## Limitations
|
||||||
|
|
||||||
|
- **Not an instruction model.** It continues text; it does not answer questions.
|
||||||
|
- **Factual reliability is poor.** Do not use as a knowledge source.
|
||||||
|
- **Primarily Tanzanian Swahili**, reflecting the corpus distribution.
|
||||||
|
- **No safety tuning at all.** It will reproduce harmful, biased or explicit
|
||||||
|
content present in web text.
|
||||||
|
- Trained on web-scraped data and carries its biases and quality artifacts.
|
||||||
|
|
||||||
|
## Files
|
||||||
|
|
||||||
|
| File | What it is |
|
||||||
|
|---|---|
|
||||||
|
| `model.safetensors` | FP32 weights, tied embeddings, 438 MB |
|
||||||
|
| `tokenizer.model` | SentencePiece 32k BPE model |
|
||||||
|
| `tokenizer_config.json` | `add_bos_token=false` — prepend `<s>` yourself |
|
||||||
|
| `generation_config.json` | Low-temperature defaults that keep this model coherent |
|
||||||
|
| `training_info.json` | Checkpoint step, tokens seen, benchmark scores |
|
||||||
|
|
||||||
|
A previous revision also shipped `pytorch_model.bin` (a stale duplicate of the
|
||||||
|
weights), duplicate tokenizer files under two names, and a `config.json` whose
|
||||||
|
`pad_token_id` was 0 while the tokenizer's `<pad>` is 3. All removed or
|
||||||
|
corrected.
|
||||||
|
|
||||||
|
## Citation
|
||||||
|
|
||||||
|
```bibtex
|
||||||
|
@misc{kw5lite2026,
|
||||||
|
title = {KW5-Lite: a 110M-parameter Swahili language model trained on a single T4},
|
||||||
|
author = {Regnant},
|
||||||
|
year = {2026},
|
||||||
|
url = {https://huggingface.co/regnant-io/kw5-lite-base}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Apache 2.0.
|
||||||
32
config.json
Normal file
32
config.json
Normal file
@@ -0,0 +1,32 @@
|
|||||||
|
{
|
||||||
|
"architectures": [
|
||||||
|
"LlamaForCausalLM"
|
||||||
|
],
|
||||||
|
"attention_bias": false,
|
||||||
|
"attention_dropout": 0.0,
|
||||||
|
"bos_token_id": 1,
|
||||||
|
"dtype": "float32",
|
||||||
|
"eos_token_id": 2,
|
||||||
|
"head_dim": 64,
|
||||||
|
"hidden_act": "silu",
|
||||||
|
"hidden_size": 768,
|
||||||
|
"initializer_range": 0.02,
|
||||||
|
"intermediate_size": 2048,
|
||||||
|
"max_position_embeddings": 2048,
|
||||||
|
"mlp_bias": false,
|
||||||
|
"model_type": "llama",
|
||||||
|
"num_attention_heads": 12,
|
||||||
|
"num_hidden_layers": 12,
|
||||||
|
"num_key_value_heads": 12,
|
||||||
|
"pad_token_id": 3,
|
||||||
|
"pretraining_tp": 1,
|
||||||
|
"rms_norm_eps": 1e-05,
|
||||||
|
"rope_parameters": {
|
||||||
|
"rope_theta": 10000.0,
|
||||||
|
"rope_type": "default"
|
||||||
|
},
|
||||||
|
"tie_word_embeddings": true,
|
||||||
|
"transformers_version": "5.16.1",
|
||||||
|
"use_cache": true,
|
||||||
|
"vocab_size": 32000
|
||||||
|
}
|
||||||
10
generation_config.json
Normal file
10
generation_config.json
Normal file
@@ -0,0 +1,10 @@
|
|||||||
|
{
|
||||||
|
"bos_token_id": 1,
|
||||||
|
"eos_token_id": 2,
|
||||||
|
"pad_token_id": 3,
|
||||||
|
"do_sample": true,
|
||||||
|
"temperature": 0.2,
|
||||||
|
"top_p": 0.9,
|
||||||
|
"repetition_penalty": 1.3,
|
||||||
|
"max_new_tokens": 512
|
||||||
|
}
|
||||||
3
model.safetensors
Normal file
3
model.safetensors
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:961e6c296d49aeb6c0df43d6a0c9f933193fddfaafc8ff215af3374643e75414
|
||||||
|
size 438131648
|
||||||
6
special_tokens_map.json
Normal file
6
special_tokens_map.json
Normal file
@@ -0,0 +1,6 @@
|
|||||||
|
{
|
||||||
|
"bos_token": "<s>",
|
||||||
|
"eos_token": "</s>",
|
||||||
|
"unk_token": "<unk>",
|
||||||
|
"pad_token": "<pad>"
|
||||||
|
}
|
||||||
254873
tokenizer.json
Normal file
254873
tokenizer.json
Normal file
File diff suppressed because it is too large
Load Diff
3
tokenizer.model
Normal file
3
tokenizer.model
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:696c2b00d0212edfe7b9f934fb13bab07be92d533182f27ad2355c1350161312
|
||||||
|
size 542002
|
||||||
14
tokenizer_config.json
Normal file
14
tokenizer_config.json
Normal file
@@ -0,0 +1,14 @@
|
|||||||
|
{
|
||||||
|
"tokenizer_class": "LlamaTokenizer",
|
||||||
|
"model_max_length": 2048,
|
||||||
|
"bos_token": "<s>",
|
||||||
|
"eos_token": "</s>",
|
||||||
|
"unk_token": "<unk>",
|
||||||
|
"pad_token": "<pad>",
|
||||||
|
"add_bos_token": false,
|
||||||
|
"add_eos_token": false,
|
||||||
|
"clean_up_tokenization_spaces": false,
|
||||||
|
"legacy": false,
|
||||||
|
"use_default_system_prompt": false,
|
||||||
|
"chat_template": "{%- set ns = namespace(out=bos_token, sys='', first=true) -%}{%- for m in messages -%}{%- if m['role'] == 'system' -%}{%- set ns.sys = m['content'] -%}{%- endif -%}{%- endfor -%}{%- for m in messages -%}{%- if m['role'] == 'user' -%}{%- set ns.out = ns.out + '[INST] ' -%}{%- if ns.first and ns.sys -%}{%- set ns.out = ns.out + '<<SYS>>\n' + ns.sys + '\n<</SYS>>\n\n' -%}{%- endif -%}{%- set ns.first = false -%}{%- set ns.out = ns.out + m['content'] + ' [/INST]' -%}{%- elif m['role'] == 'assistant' -%}{%- set ns.out = ns.out + ' ' + m['content'] + eos_token -%}{%- endif -%}{%- endfor -%}{{- ns.out -}}"
|
||||||
|
}
|
||||||
15
training_info.json
Normal file
15
training_info.json
Normal file
@@ -0,0 +1,15 @@
|
|||||||
|
{
|
||||||
|
"checkpoint_step": 6150,
|
||||||
|
"tokens_seen": 1410662400,
|
||||||
|
"epochs": 2.0,
|
||||||
|
"note": "Stopped at exactly 2 epochs; WSD decay deliberately not run because later checkpoints overfit.",
|
||||||
|
"benchmark_scores": {
|
||||||
|
"belebele_sw": 0.32,
|
||||||
|
"afrixnli_sw": 0.32,
|
||||||
|
"n_samples": 50,
|
||||||
|
"chance": {
|
||||||
|
"belebele_sw": 0.25,
|
||||||
|
"afrixnli_sw": 0.333
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
Reference in New Issue
Block a user