142 lines
4.1 KiB
Markdown
142 lines
4.1 KiB
Markdown
---
|
|
language:
|
|
- it
|
|
- en
|
|
license: apache-2.0
|
|
tags:
|
|
- italian
|
|
- causal-lm
|
|
- small-language-model
|
|
- trained-from-scratch
|
|
- chatml
|
|
- conversational
|
|
pipeline_tag: text-generation
|
|
---
|
|
|
|
# ⚛ Quark-50M-v2
|
|
|
|
**43.8M parameter Italian-first bilingual language model, trained from scratch by ThingAI.**
|
|
|
|
Quark-50M is an ultra-compact causal language model that speaks fluent Italian. Designed as a proof-of-concept for small, efficient, Italian-centric AI.
|
|
|
|
## Highlights
|
|
|
|
- **43.8M parameters** — runs on any device, even CPU
|
|
- **Italian-first** — trained on 60% Italian data (books, Wikipedia, web)
|
|
- **ChatML format** — `<|im_start|>user`/`<|im_start|>assistant`
|
|
- **Custom tokenizer** — 16k BPE, optimized for Italian (4.15 chars/token)
|
|
- **Trained from scratch** — architecture, tokenizer, and training pipeline all custom
|
|
|
|
## Architecture
|
|
|
|
| Component | Value |
|
|
|-----------|-------|
|
|
| Parameters | 43.8M |
|
|
| Vocabulary | 16,384 (BPE) |
|
|
| Dimensions | 512 |
|
|
| Layers | 12 |
|
|
| Heads | 8 (4 KV heads, GQA) |
|
|
| FFN | 1,408 (SwiGLU) |
|
|
| Context | 2,048 tokens |
|
|
| Normalization | RMSNorm |
|
|
| Position | RoPE |
|
|
| Weight Tying | Yes |
|
|
|
|
## Training
|
|
|
|
**Pretraining:** 5B tokens on a curated mix:
|
|
|
|
| Dataset | Weight | Type |
|
|
|---------|--------|------|
|
|
| PleIAs/Italian-PD | 25% | 171K Italian books (public domain) |
|
|
| FineWeb-2 Italian | 20% | Cleaned, deduplicated web |
|
|
| Wikipedia IT | 15% | Encyclopedia |
|
|
| Cosmopedia | 15% | Synthetic educational |
|
|
| SmolLM-Corpus | 10% | Curated mix |
|
|
| StarCoder Python | 8% | Code |
|
|
| OpenWebMath | 7% | Mathematics |
|
|
|
|
**SFT:** Fine-tuned on [quattro-chiacchiere](https://huggingface.co/datasets/ThingAI/quattro-chiacchiere), a synthetic Italian Q&A dataset generated with [Alembic](https://github.com/skein-labs/Alembic).
|
|
|
|
## Usage
|
|
|
|
```python
|
|
import torch
|
|
from huggingface_hub import hf_hub_download
|
|
from transformers import PreTrainedTokenizerFast
|
|
|
|
# Load
|
|
ckpt_path = hf_hub_download("ThingAI/Quark-50M", "model.pt")
|
|
model_py = hf_hub_download("ThingAI/Quark-50M", "model.py")
|
|
tok_file = hf_hub_download("ThingAI/Quark-50M", "tokenizer.json")
|
|
|
|
# Tokenizer
|
|
tokenizer = PreTrainedTokenizerFast(tokenizer_file=tok_file)
|
|
tokenizer.eos_token = "<|endoftext|>"
|
|
|
|
# Model
|
|
import importlib.util
|
|
spec = importlib.util.spec_from_file_location("model", model_py)
|
|
mod = importlib.util.module_from_spec(spec)
|
|
spec.loader.exec_module(mod)
|
|
|
|
ckpt = torch.load(ckpt_path, map_location="cpu", weights_only=False)
|
|
cfg = mod.ModelConfig(**ckpt["model_cfg"])
|
|
model = mod.Quark(cfg).eval()
|
|
model.load_state_dict(ckpt["model"])
|
|
|
|
# Chat
|
|
prompt = "<|im_start|>user\nQual è la capitale d'Italia?<|im_end|>\n<|im_start|>assistant\n"
|
|
ids = tokenizer.encode(prompt, return_tensors="pt")
|
|
with torch.no_grad():
|
|
for _ in range(100):
|
|
logits = model(ids)[1][:, -1, :].float()
|
|
nxt = logits.argmax(-1, keepdim=True)
|
|
if nxt.item() == tokenizer.convert_tokens_to_ids("<|im_end|>"): break
|
|
ids = torch.cat([ids, nxt], -1)
|
|
print(tokenizer.decode(ids[0], skip_special_tokens=True))
|
|
# La capitale d'Italia è Roma.
|
|
```
|
|
|
|
## Examples
|
|
|
|
```
|
|
Tu: Qual è la capitale d'Italia?
|
|
Quark: La capitale d'Italia è Roma.
|
|
|
|
Tu: Chi sei?
|
|
Quark: Sono Quark, piacere di conoscerti.
|
|
|
|
Tu: Come ti chiami?
|
|
Quark: Mi chiamo Quark, piacere di conoscerti.
|
|
```
|
|
|
|
## Limitations
|
|
|
|
- **43.8M parameters** — cannot perform complex reasoning or long-form generation
|
|
- **Factual accuracy** — may hallucinate facts, especially on niche topics
|
|
- **SFT dataset** — currently limited; more data will improve reliability
|
|
- **No safety training** — not recommended for production without guardrails
|
|
|
|
## Related
|
|
|
|
- [Quark3Tokenizer](https://huggingface.co/ThingAI/Quark3Tokenizer) — the tokenizer
|
|
- [quattro-chiacchiere](https://huggingface.co/datasets/ThingAI/quattro-chiacchiere) — the SFT dataset
|
|
- [Alembic](https://github.com/skein-labs/Alembic) — the dataset distillation tool
|
|
- [Glyph](https://huggingface.co/ThingAI/Glyph) — multi-task text classifier by ThingAI
|
|
|
|
## Citation
|
|
|
|
```bibtex
|
|
@misc{quark50m,
|
|
author = {ThingAI},
|
|
title = {Quark-50M-v2: Italian-First Small Language Model},
|
|
year = {2026},
|
|
url = {https://huggingface.co/ThingAI/Quark-50M}
|
|
}
|
|
```
|
|
|
|
## License
|
|
|
|
Apache 2.0
|