143 lines
6.0 KiB
Markdown
143 lines
6.0 KiB
Markdown
---
|
|
license: apache-2.0
|
|
datasets:
|
|
- HuggingFaceFW/fineweb
|
|
- bigcode/the-stack-dedup
|
|
- GAIR/lima
|
|
language:
|
|
- en
|
|
pipeline_tag: text-classification
|
|
tags:
|
|
- khazri
|
|
- softyu
|
|
- ai
|
|
- llm
|
|
- azerbaijani-national-llm
|
|
- azerbaijan
|
|
---
|
|
<p align="center">
|
|
<img src="./assets/khazri-2-mini-banner.png" alt="Khazri 2 Mini — compact open-weight language model" width="100%">
|
|
</p>
|
|
|
|
# Khazri 2 Mini — 100M
|
|
|
|
**Khazri 2 Mini** is a compact, open-weight decoder-only language model in the Khazri family. It combines a modern LLaMA-style Transformer with a custom Byte-Level BPE tokenizer and a 2B-token English training corpus.
|
|
|
|
[Hugging Face](https://huggingface.co/Yusiko/khazri-2-mini) · [Khazri](https://khazri.dev) · [Contact](mailto:contact@khazri.dev)
|
|
|
|
## At a glance
|
|
|
|
| Item | Detail |
|
|
| --- | --- |
|
|
| Model | Khazri 2 Mini |
|
|
| Parameters | **100.68M** |
|
|
| Status | Open weights on Hugging Face |
|
|
| Architecture | LLaMA-style, decoder-only Transformer |
|
|
| Training precision | bf16 |
|
|
| Context configured for training | 1,024 tokens |
|
|
| Vocabulary | 32,768 tokens |
|
|
| Tokenizer | Custom Byte-Level BPE |
|
|
| Attention | Grouped-Query Attention: 12 query heads / 4 KV heads |
|
|
|
|
## Architecture
|
|
|
|
| Component | Configuration |
|
|
| --- | --- |
|
|
| Hidden size | 768 |
|
|
| Transformer layers | 12 |
|
|
| Attention heads | 12 |
|
|
| Key/value heads | 4 |
|
|
| MLP intermediate size | 2,048 |
|
|
| Positional encoding | RoPE |
|
|
| Normalization | RMSNorm |
|
|
| MLP activation | SwiGLU / SiLU |
|
|
| Attention backend | FlashAttention-2 where available; PyTorch SDPA fallback |
|
|
| Embeddings | Tied input/output embeddings |
|
|
|
|
## Training data
|
|
|
|
Khazri 2 Mini is trained on a custom, pretokenized **English-only** corpus with a target size of **2,000,000,000 tokens**. The corpus is packed into **1,953,125 sequences** of 1,024 tokens and stored in Arrow shards with source identifiers.
|
|
|
|
The documented token budget is:
|
|
|
|
| Source | Token budget | Share | Role |
|
|
| --- | ---: | ---: | --- |
|
|
| [Cosmopedia](https://huggingface.co/datasets/HuggingFaceTB/cosmopedia) | 850M | 42.5% | General English educational and synthetic-text coverage |
|
|
| [The Stack v2 Dedup](https://huggingface.co/datasets/bigcode/the-stack-v2-dedup) | 450M | 22.5% | Code from Python, JavaScript, TypeScript, Java, C++, C, Go and Rust |
|
|
| [TinyStories](https://huggingface.co/datasets/roneneldan/TinyStories) | 300M | 15.0% | Simple narrative language |
|
|
| [LIMA](https://huggingface.co/datasets/GAIR/lima) | 50M | 2.5% | Instruction and conversation examples |
|
|
| SYNAPSE synthetic instruction data | 350M | 17.5% | Arithmetic, context, abstention, web-needed, identity, symbolic-math and general-assistant routes |
|
|
|
|
The listed values are the documented source-token budget. The release manifest should be used for the final source counts of a particular weight revision.
|
|
|
|
### Data processing and safeguards
|
|
|
|
- Only English text is retained for this corpus. Short records are excluded, very long records are capped, and language/character checks are applied before tokenization.
|
|
- The code portion is limited to the eight languages listed above.
|
|
- The training mix combines general English, narrative, code, instruction and route-aware synthetic material. This preserves general capabilities while teaching specialized SYNAPSE behaviours.
|
|
- Checkpoint evaluation should cover general Q&A, code, arithmetic, context extraction, abstention and current-information requests. A route-specific gain should not be accepted if it degrades general behaviour.
|
|
- Original source datasets remain subject to their own terms and licences. Consult their source pages and the model repository licence before use.
|
|
|
|
## Tokenizer
|
|
|
|
Khazri 2 Mini uses a 32,768-token custom Byte-Level BPE tokenizer. It reserves structural whitespace, chat and SYNAPSE route tokens as single tokens, preserves indentation for code, and uses single-digit splitting to make arithmetic strings more explicit to the model.
|
|
|
|
## Installation
|
|
|
|
~~~bash
|
|
pip install -U torch transformers accelerate safetensors
|
|
~~~
|
|
|
|
## Quick start
|
|
|
|
~~~python
|
|
import torch
|
|
from transformers import AutoModelForCausalLM, AutoTokenizer
|
|
|
|
MODEL_ID = "Yusiko/khazri-2-mini"
|
|
|
|
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
|
|
model = AutoModelForCausalLM.from_pretrained(
|
|
MODEL_ID,
|
|
torch_dtype="auto",
|
|
device_map="auto",
|
|
)
|
|
|
|
prompt = "Write a concise explanation of a small language model."
|
|
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
|
|
|
|
with torch.inference_mode():
|
|
output = model.generate(
|
|
**inputs,
|
|
max_new_tokens=160,
|
|
do_sample=False,
|
|
pad_token_id=tokenizer.eos_token_id,
|
|
)
|
|
|
|
print(tokenizer.decode(output[0], skip_special_tokens=True))
|
|
~~~
|
|
|
|
## Khazri 2 Preview comparison
|
|
|
|
The following table reports the project-provided compact-model comparison for **Khazri 2 Preview**, not Khazri 2 Mini. Higher is better for every listed task.
|
|
|
|
| Model | Parameters | Context extraction | Mixed speed/proxy | Arithmetic | Word problems | Abstention |
|
|
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
|
|
| **Khazri 2 Preview** | ~250M | **100%** | **62%** | **99%** | **99%** | **97.4%** |
|
|
| Gemma 3 | 270M | 100% | 36% | 0% | 0% | 18% |
|
|
| Qwen 2.5 | 0.5B | 89% | 45% | 14% | 28% | 46% |
|
|
| Pythia | 160M | 22% | 10% | 0% | 2% | 1% |
|
|
|
|
These are internal preview results on selected compact-model tests. They are not independently audited and should not be used to make claims about Khazri 2 Mini. Publish prompts, model revisions, scoring rules, hardware and complete evaluation assets with any future benchmark announcement.
|
|
|
|
## Responsible use
|
|
|
|
Khazri 2 Mini can produce incorrect, incomplete or biased outputs. Evaluate it on your own task, verify material claims and keep a human in the loop for consequential decisions. Do not rely on it as the sole basis for legal, medical, financial, hiring, safety or other high-impact decisions.
|
|
|
|
## Roadmap
|
|
|
|
Khazri 2 Mini is part of the second Khazri generation. The next planned stage is **Khazri 3**: a larger parameter scale and stronger results.
|
|
|
|
## Contact
|
|
|
|
For research, integration or partnership inquiries, visit [khazri.dev](https://khazri.dev) or email [contact@khazri.dev](mailto:contact@khazri.dev). |