Files
dante-2b-ita-instruct/README.md
ModelHub XC ce8c59b2a2 初始化项目,由ModelHub XC社区提供模型
Model: SerFabio89/dante-2b-ita-instruct
Source: Original Platform
2026-05-30 22:23:31 +08:00

232 lines
9.0 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
license: apache-2.0
language:
- it
- en
tags:
- italian
- bilingual
- from-scratch
- decoder-only
- llama-style
- gqa
- causal-lm
- text-generation
- sft
- instruct
library_name: transformers
pipeline_tag: text-generation
model-index:
- name: Dante-2B
results: []
---
# Dante-2B
**A 2.1B parameter bilingual Italian/English language model, trained entirely from scratch.**
Dante-2B is a decoder-only transformer built by a single developer on 2× NVIDIA H200 NVL GPUs. Everything is native — the tokenizer, the architecture, the training pipeline — no fine-tune of an existing English model, no retrofitted multilingual vocabulary. Italian is a first-class citizen from byte zero.
| | |
|---|---|
| **Parameters** | 2.1B (all active, no MoE) |
| **Architecture** | Llama-style decoder-only transformer |
| **Languages** | Italian, English |
| **Tokenizer** | Custom 64K BPE, Italian-native |
| **Context** | 4,096 tokens |
| **Training** | 120B tokens across 3 phases |
| **Hardware** | 2× NVIDIA H200 NVL |
| **License** | Apache 2.0 |
## Why Dante-2B?
Most "Italian" language models are English-first architectures with Italian bolted on during fine-tuning. Their tokenizers fragment Italian text into too many subwords, wasting context window and degrading fluency. Dante-2B takes a different path: the tokenizer was trained on a balanced Italian/English corpus from the start, so Italian apostrophe contractions (`l'`, `dell'`, `un'`), accented vowels, and common morphological patterns are single tokens — not sequences of three.
This is not a research lab effort with 512 GPUs. It's a from-scratch build on commodity hardware, with every engineering decision documented — including the failures.
## Architecture
Dante-2B uses a pure Llama-style architecture optimized for training speed on limited hardware.
| Component | Detail |
|---|---|
| Type | Decoder-only dense transformer |
| `d_model` | 2,560 |
| Layers | 28 |
| Attention | Grouped Query Attention (GQA) — 20 query heads, 4 KV heads (5:1) |
| `d_head` | 128 (optimal for Flash Attention) |
| FFN | SwiGLU, `d_ff` = 6,912 |
| Normalization | RMSNorm (ε = 1e-6) |
| Position encoding | RoPE (θ = 10,000) |
| Vocab size | 64,000 |
| Weight tying | Embedding ↔ LM head |
| Dropout | 0.0 |
## Tokenizer
A custom 64K BPE tokenizer trained on a character-balanced subset of the corpus (45% Italian, 45% English, 10% code).
Key properties:
- Italian accented vowels (à, è, é, ì, ò, ù) are always single tokens
- Italian apostrophe contractions (`l'intelligenza`, `dell'algoritmo`) are handled as single tokens
- Custom pre-tokenization regex preserves Italian morphology
- 36 special tokens including ChatML markers, thinking tokens, tool-use tokens, and expert routing tokens (future-proofing)
Fertility (tokens per word): ~1.35 Italian, ~1.20 English. For comparison, LLaMA's tokenizer scores ~1.85 on Italian.
## Training
### Phase 1 — Base Pretraining
90B tokens at sequence length 2,048 over ~57,200 steps.
- Optimizer: AdamW (β₁=0.9, β₂=0.95, weight decay 0.1)
- LR schedule: cosine, 3e-4 → 3e-5, 2,000-step warmup
- Batch size: micro_batch 24 × grad_accum 16 × 2 GPUs = 1,572,864 tokens/step
- Infrastructure: DeepSpeed ZeRO-2, FP8 (torchao), torch.compile (reduce-overhead), Flash Attention 2
- Throughput: ~88–89K tokens/sec, 28% MFU
- Final loss: ~1.85–1.93
### Phase 2 — Continual Pretraining (Context Extension)
30B tokens at sequence length 4,096 over ~28,600 steps.
- LR schedule: cosine, 6e-5 → 6e-6, 500-step warmup
- Purpose: extend context window from 2,048 → 4,096 while consolidating learned representations
- Final loss: ~1.75–1.80 (expected plateau — gains manifest in downstream performance)
- Inference test: coherent Italian at 125.7 tok/s
### Phase 3 — Supervised Fine-Tuning (SFT)
540K conversations (approximately 40% Italian), 1 epoch in ~3.5 hours.
- Micro_batch 12, grad_accum 10, ~89K tok/s, 33.4% MFU
- Eval loss: 1.23 (declining toward ~1.15)
- Strategy: single-epoch SFT to avoid overfitting
- Loss masking: only assistant turns contribute to loss; user turns and headers are masked with -100
**SFT Datasets:**
- HuggingFaceH4/no_robots
- OpenAssistant/oasst1 (via Guanaco)
- HuggingFaceH4/ultrachat_200k
- efederici/capybara-claude-15k-ita
- anakin87/fine-instructions-ita-70k
- Mattimax/Camoscio-ITA
- mchl-labs/stambecco_data_it
- teelinsan/camoscio
- DeepMount00/Sonnet-3.5-ITA-INSTRUCT
- DeepMount00/italian_conversations
## Pretraining Data
The pretraining corpus is a bilingual mix (approximately 45% Italian, 45% English, 10% code by character count):
**Italian:** FineWeb-2 IT, FineWiki IT, Italian Public Domain (PleIAs), Gazzetta Ufficiale, EuroParl IT, FinePDFs IT, Clean mC4 IT
**English:** FineWeb-Edu (100B sample), FineWiki EN
**Code:** StarCoderData
## Chat Format
Dante-2B uses a ChatML-style template:
```
<|begin_of_text|><|start_header|>user<|end_header|>
Come funziona la fotosintesi?<|eot|>
<|start_header|>assistant<|end_header|>
La fotosintesi è il processo attraverso cui le piante convertono...
```
## Usage
```python
import torch
from tokenizers import Tokenizer
from model_architecture import DanteConfig, DanteModel
# Load model
config = DanteConfig.load("config.json")
model = DanteModel(config)
model.load_state_dict(torch.load("model.pt", map_location="cpu"), strict=False)
model.eval().cuda()
# Load tokenizer
tokenizer = Tokenizer.from_file("tokenizer.json")
# Chat
prompt = "<|begin_of_text|><|start_header|>user<|end_header|>\nCiao, come stai?<|eot|>\n<|start_header|>assistant<|end_header|>\n"
input_ids = torch.tensor([tokenizer.encode(prompt).ids], device="cuda")
with torch.no_grad():
# Use KV-cache for efficient generation
kv_cache = None
for _ in range(256):
out = model(input_ids if kv_cache is None else input_ids[:, -1:],
kv_cache=kv_cache, use_cache=True)
kv_cache = out["kv_cache"]
next_id = out["logits"][:, -1].argmax(dim=-1, keepdim=True)
input_ids = torch.cat([input_ids, next_id], dim=-1)
token = tokenizer.decode([next_id.item()])
if token in ["<|eot|>", "<|end_of_text|>"]:
break
print(token, end="", flush=True)
```
> **Tip:** use `repetition_penalty=1.15` for longer generations to avoid repetition loops.
## Limitations
- **Scale:** At 2.1B parameters, Dante-2B is not competing with 7B+ models on complex reasoning. It is a capable small model for Italian-language tasks, on-device deployment, and research.
- **Knowledge cutoff:** The pretraining data has a knowledge cutoff determined by the source datasets (primarily 2023–2024 web crawls).
- **Benchmarks:** Formal evaluation on ITA-Bench and standard benchmarks is in progress and will be added here.
- **Safety:** This model has not undergone RLHF or safety-specific alignment. Use responsibly.
## Training Stack
| Component | Version / Detail |
|---|---|
| Framework | PyTorch 2.11 + CUDA 12.8 |
| Distributed | DeepSpeed ZeRO-2 |
| Mixed precision | FP8 via torchao |
| Compilation | torch.compile (reduce-overhead) |
| Attention | Flash Attention 2 |
| Hardware | 2× NVIDIA H200 NVL (141GB each) |
| Peak FLOPS | 1,671 TFLOPS per GPU (NVL spec) |
## Repository Structure
```
├── model_architecture.py # DanteConfig + DanteModel (full architecture)
├── config.json # Model configuration
├── tokenizer.json # 64K BPE tokenizer
├── model.pt # Model weights
└── README.md # This file
```
The full training codebase — from corpus download to SFT — is available on GitHub: **[link to be added]**
## Citation
If you use Dante-2B in your research or applications, please cite:
```bibtex
@misc{angeletti2025dante2b,
title={Dante-2B: A Bilingual Italian/English Language Model Trained from Scratch},
author={Fabio Angeletti},
year={2025},
url={https://huggingface.co/leafsrls/dante-2b}
}
```
## About
Dante-2B is built and maintained by **Fabio Angeletti**, PhD in Computer Engineering, Adjunct Professor at LUISS and LUISS Business School, and Founder & CEO of [LEAF srls](https://leafsrls.com).
**LEAF** is a specialized AI company based in Italy, operating at the intersection of Artificial Intelligence, Machine Learning, IoT, and Cybersecurity. We deliver secure, efficient, and transformative digital strategies — from AI model development and system integration to low-power wireless monitoring and edge AI deployment. Our mission is to empower businesses to navigate and succeed in the digital economy through advanced technological innovation.
For collaborations, commercial licensing, or inquiries: **info@leafsrls.com**
## License
This model is released under the [Apache License 2.0](https://www.apache.org/licenses/LICENSE-2.0). You are free to use, modify, and distribute it for any purpose, including commercial use.