初始化项目,由ModelHub XC社区提供模型
Model: ThingAI/Quark-50m-v2 Source: Original Platform
This commit is contained in:
141
README.md
Normal file
141
README.md
Normal file
@@ -0,0 +1,141 @@
|
||||
---
|
||||
language:
|
||||
- it
|
||||
- en
|
||||
license: apache-2.0
|
||||
tags:
|
||||
- italian
|
||||
- causal-lm
|
||||
- small-language-model
|
||||
- trained-from-scratch
|
||||
- chatml
|
||||
- conversational
|
||||
pipeline_tag: text-generation
|
||||
---
|
||||
|
||||
# ⚛ Quark-50M-v2
|
||||
|
||||
**43.8M parameter Italian-first bilingual language model, trained from scratch by ThingAI.**
|
||||
|
||||
Quark-50M is an ultra-compact causal language model that speaks fluent Italian. Designed as a proof-of-concept for small, efficient, Italian-centric AI.
|
||||
|
||||
## Highlights
|
||||
|
||||
- **43.8M parameters** — runs on any device, even CPU
|
||||
- **Italian-first** — trained on 60% Italian data (books, Wikipedia, web)
|
||||
- **ChatML format** — `<|im_start|>user`/`<|im_start|>assistant`
|
||||
- **Custom tokenizer** — 16k BPE, optimized for Italian (4.15 chars/token)
|
||||
- **Trained from scratch** — architecture, tokenizer, and training pipeline all custom
|
||||
|
||||
## Architecture
|
||||
|
||||
| Component | Value |
|
||||
|-----------|-------|
|
||||
| Parameters | 43.8M |
|
||||
| Vocabulary | 16,384 (BPE) |
|
||||
| Dimensions | 512 |
|
||||
| Layers | 12 |
|
||||
| Heads | 8 (4 KV heads, GQA) |
|
||||
| FFN | 1,408 (SwiGLU) |
|
||||
| Context | 2,048 tokens |
|
||||
| Normalization | RMSNorm |
|
||||
| Position | RoPE |
|
||||
| Weight Tying | Yes |
|
||||
|
||||
## Training
|
||||
|
||||
**Pretraining:** 5B tokens on a curated mix:
|
||||
|
||||
| Dataset | Weight | Type |
|
||||
|---------|--------|------|
|
||||
| PleIAs/Italian-PD | 25% | 171K Italian books (public domain) |
|
||||
| FineWeb-2 Italian | 20% | Cleaned, deduplicated web |
|
||||
| Wikipedia IT | 15% | Encyclopedia |
|
||||
| Cosmopedia | 15% | Synthetic educational |
|
||||
| SmolLM-Corpus | 10% | Curated mix |
|
||||
| StarCoder Python | 8% | Code |
|
||||
| OpenWebMath | 7% | Mathematics |
|
||||
|
||||
**SFT:** Fine-tuned on [quattro-chiacchiere](https://huggingface.co/datasets/ThingAI/quattro-chiacchiere), a synthetic Italian Q&A dataset generated with [Alembic](https://github.com/skein-labs/Alembic).
|
||||
|
||||
## Usage
|
||||
|
||||
```python
|
||||
import torch
|
||||
from huggingface_hub import hf_hub_download
|
||||
from transformers import PreTrainedTokenizerFast
|
||||
|
||||
# Load
|
||||
ckpt_path = hf_hub_download("ThingAI/Quark-50M", "model.pt")
|
||||
model_py = hf_hub_download("ThingAI/Quark-50M", "model.py")
|
||||
tok_file = hf_hub_download("ThingAI/Quark-50M", "tokenizer.json")
|
||||
|
||||
# Tokenizer
|
||||
tokenizer = PreTrainedTokenizerFast(tokenizer_file=tok_file)
|
||||
tokenizer.eos_token = "<|endoftext|>"
|
||||
|
||||
# Model
|
||||
import importlib.util
|
||||
spec = importlib.util.spec_from_file_location("model", model_py)
|
||||
mod = importlib.util.module_from_spec(spec)
|
||||
spec.loader.exec_module(mod)
|
||||
|
||||
ckpt = torch.load(ckpt_path, map_location="cpu", weights_only=False)
|
||||
cfg = mod.ModelConfig(**ckpt["model_cfg"])
|
||||
model = mod.Quark(cfg).eval()
|
||||
model.load_state_dict(ckpt["model"])
|
||||
|
||||
# Chat
|
||||
prompt = "<|im_start|>user\nQual è la capitale d'Italia?<|im_end|>\n<|im_start|>assistant\n"
|
||||
ids = tokenizer.encode(prompt, return_tensors="pt")
|
||||
with torch.no_grad():
|
||||
for _ in range(100):
|
||||
logits = model(ids)[1][:, -1, :].float()
|
||||
nxt = logits.argmax(-1, keepdim=True)
|
||||
if nxt.item() == tokenizer.convert_tokens_to_ids("<|im_end|>"): break
|
||||
ids = torch.cat([ids, nxt], -1)
|
||||
print(tokenizer.decode(ids[0], skip_special_tokens=True))
|
||||
# La capitale d'Italia è Roma.
|
||||
```
|
||||
|
||||
## Examples
|
||||
|
||||
```
|
||||
Tu: Qual è la capitale d'Italia?
|
||||
Quark: La capitale d'Italia è Roma.
|
||||
|
||||
Tu: Chi sei?
|
||||
Quark: Sono Quark, piacere di conoscerti.
|
||||
|
||||
Tu: Come ti chiami?
|
||||
Quark: Mi chiamo Quark, piacere di conoscerti.
|
||||
```
|
||||
|
||||
## Limitations
|
||||
|
||||
- **43.8M parameters** — cannot perform complex reasoning or long-form generation
|
||||
- **Factual accuracy** — may hallucinate facts, especially on niche topics
|
||||
- **SFT dataset** — currently limited; more data will improve reliability
|
||||
- **No safety training** — not recommended for production without guardrails
|
||||
|
||||
## Related
|
||||
|
||||
- [Quark3Tokenizer](https://huggingface.co/ThingAI/Quark3Tokenizer) — the tokenizer
|
||||
- [quattro-chiacchiere](https://huggingface.co/datasets/ThingAI/quattro-chiacchiere) — the SFT dataset
|
||||
- [Alembic](https://github.com/skein-labs/Alembic) — the dataset distillation tool
|
||||
- [Glyph](https://huggingface.co/ThingAI/Glyph) — multi-task text classifier by ThingAI
|
||||
|
||||
## Citation
|
||||
|
||||
```bibtex
|
||||
@misc{quark50m,
|
||||
author = {ThingAI},
|
||||
title = {Quark-50M-v2: Italian-First Small Language Model},
|
||||
year = {2026},
|
||||
url = {https://huggingface.co/ThingAI/Quark-50M}
|
||||
}
|
||||
```
|
||||
|
||||
## License
|
||||
|
||||
Apache 2.0
|
||||
Reference in New Issue
Block a user