初始化项目,由ModelHub XC社区提供模型
Model: SerFabio89/dante-2b-ita-instruct Source: Original Platform
This commit is contained in:
35
.gitattributes
vendored
Normal file
35
.gitattributes
vendored
Normal file
@@ -0,0 +1,35 @@
|
|||||||
|
*.7z filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.arrow filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.bin filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.ftz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.gz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.h5 filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.joblib filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.model filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.npy filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.npz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.onnx filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.ot filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.parquet filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pb filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pickle filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pkl filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pt filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pth filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.rar filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
||||||
|
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tar filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tflite filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tgz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.wasm filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.xz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.zip filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.zst filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
||||||
232
README.md
Normal file
232
README.md
Normal file
@@ -0,0 +1,232 @@
|
|||||||
|
---
|
||||||
|
license: apache-2.0
|
||||||
|
language:
|
||||||
|
- it
|
||||||
|
- en
|
||||||
|
tags:
|
||||||
|
- italian
|
||||||
|
- bilingual
|
||||||
|
- from-scratch
|
||||||
|
- decoder-only
|
||||||
|
- llama-style
|
||||||
|
- gqa
|
||||||
|
- causal-lm
|
||||||
|
- text-generation
|
||||||
|
- sft
|
||||||
|
- instruct
|
||||||
|
library_name: transformers
|
||||||
|
pipeline_tag: text-generation
|
||||||
|
model-index:
|
||||||
|
- name: Dante-2B
|
||||||
|
results: []
|
||||||
|
---
|
||||||
|
|
||||||
|
# Dante-2B
|
||||||
|
|
||||||
|
**A 2.1B parameter bilingual Italian/English language model, trained entirely from scratch.**
|
||||||
|
|
||||||
|
Dante-2B is a decoder-only transformer built by a single developer on 2× NVIDIA H200 NVL GPUs. Everything is native — the tokenizer, the architecture, the training pipeline — no fine-tune of an existing English model, no retrofitted multilingual vocabulary. Italian is a first-class citizen from byte zero.
|
||||||
|
|
||||||
|
| | |
|
||||||
|
|---|---|
|
||||||
|
| **Parameters** | 2.1B (all active, no MoE) |
|
||||||
|
| **Architecture** | Llama-style decoder-only transformer |
|
||||||
|
| **Languages** | Italian, English |
|
||||||
|
| **Tokenizer** | Custom 64K BPE, Italian-native |
|
||||||
|
| **Context** | 4,096 tokens |
|
||||||
|
| **Training** | 120B tokens across 3 phases |
|
||||||
|
| **Hardware** | 2× NVIDIA H200 NVL |
|
||||||
|
| **License** | Apache 2.0 |
|
||||||
|
|
||||||
|
## Why Dante-2B?
|
||||||
|
|
||||||
|
Most "Italian" language models are English-first architectures with Italian bolted on during fine-tuning. Their tokenizers fragment Italian text into too many subwords, wasting context window and degrading fluency. Dante-2B takes a different path: the tokenizer was trained on a balanced Italian/English corpus from the start, so Italian apostrophe contractions (`l'`, `dell'`, `un'`), accented vowels, and common morphological patterns are single tokens — not sequences of three.
|
||||||
|
|
||||||
|
This is not a research lab effort with 512 GPUs. It's a from-scratch build on commodity hardware, with every engineering decision documented — including the failures.
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
Dante-2B uses a pure Llama-style architecture optimized for training speed on limited hardware.
|
||||||
|
|
||||||
|
| Component | Detail |
|
||||||
|
|---|---|
|
||||||
|
| Type | Decoder-only dense transformer |
|
||||||
|
| `d_model` | 2,560 |
|
||||||
|
| Layers | 28 |
|
||||||
|
| Attention | Grouped Query Attention (GQA) — 20 query heads, 4 KV heads (5:1) |
|
||||||
|
| `d_head` | 128 (optimal for Flash Attention) |
|
||||||
|
| FFN | SwiGLU, `d_ff` = 6,912 |
|
||||||
|
| Normalization | RMSNorm (ε = 1e-6) |
|
||||||
|
| Position encoding | RoPE (θ = 10,000) |
|
||||||
|
| Vocab size | 64,000 |
|
||||||
|
| Weight tying | Embedding ↔ LM head |
|
||||||
|
| Dropout | 0.0 |
|
||||||
|
|
||||||
|
## Tokenizer
|
||||||
|
|
||||||
|
A custom 64K BPE tokenizer trained on a character-balanced subset of the corpus (45% Italian, 45% English, 10% code).
|
||||||
|
|
||||||
|
Key properties:
|
||||||
|
- Italian accented vowels (à, è, é, ì, ò, ù) are always single tokens
|
||||||
|
- Italian apostrophe contractions (`l'intelligenza`, `dell'algoritmo`) are handled as single tokens
|
||||||
|
- Custom pre-tokenization regex preserves Italian morphology
|
||||||
|
- 36 special tokens including ChatML markers, thinking tokens, tool-use tokens, and expert routing tokens (future-proofing)
|
||||||
|
|
||||||
|
Fertility (tokens per word): ~1.35 Italian, ~1.20 English. For comparison, LLaMA's tokenizer scores ~1.85 on Italian.
|
||||||
|
|
||||||
|
## Training
|
||||||
|
|
||||||
|
### Phase 1 — Base Pretraining
|
||||||
|
|
||||||
|
90B tokens at sequence length 2,048 over ~57,200 steps.
|
||||||
|
|
||||||
|
- Optimizer: AdamW (β₁=0.9, β₂=0.95, weight decay 0.1)
|
||||||
|
- LR schedule: cosine, 3e-4 → 3e-5, 2,000-step warmup
|
||||||
|
- Batch size: micro_batch 24 × grad_accum 16 × 2 GPUs = 1,572,864 tokens/step
|
||||||
|
- Infrastructure: DeepSpeed ZeRO-2, FP8 (torchao), torch.compile (reduce-overhead), Flash Attention 2
|
||||||
|
- Throughput: ~88–89K tokens/sec, 28% MFU
|
||||||
|
- Final loss: ~1.85–1.93
|
||||||
|
|
||||||
|
### Phase 2 — Continual Pretraining (Context Extension)
|
||||||
|
|
||||||
|
30B tokens at sequence length 4,096 over ~28,600 steps.
|
||||||
|
|
||||||
|
- LR schedule: cosine, 6e-5 → 6e-6, 500-step warmup
|
||||||
|
- Purpose: extend context window from 2,048 → 4,096 while consolidating learned representations
|
||||||
|
- Final loss: ~1.75–1.80 (expected plateau — gains manifest in downstream performance)
|
||||||
|
- Inference test: coherent Italian at 125.7 tok/s
|
||||||
|
|
||||||
|
### Phase 3 — Supervised Fine-Tuning (SFT)
|
||||||
|
|
||||||
|
540K conversations (approximately 40% Italian), 1 epoch in ~3.5 hours.
|
||||||
|
|
||||||
|
- Micro_batch 12, grad_accum 10, ~89K tok/s, 33.4% MFU
|
||||||
|
- Eval loss: 1.23 (declining toward ~1.15)
|
||||||
|
- Strategy: single-epoch SFT to avoid overfitting
|
||||||
|
- Loss masking: only assistant turns contribute to loss; user turns and headers are masked with -100
|
||||||
|
|
||||||
|
**SFT Datasets:**
|
||||||
|
- HuggingFaceH4/no_robots
|
||||||
|
- OpenAssistant/oasst1 (via Guanaco)
|
||||||
|
- HuggingFaceH4/ultrachat_200k
|
||||||
|
- efederici/capybara-claude-15k-ita
|
||||||
|
- anakin87/fine-instructions-ita-70k
|
||||||
|
- Mattimax/Camoscio-ITA
|
||||||
|
- mchl-labs/stambecco_data_it
|
||||||
|
- teelinsan/camoscio
|
||||||
|
- DeepMount00/Sonnet-3.5-ITA-INSTRUCT
|
||||||
|
- DeepMount00/italian_conversations
|
||||||
|
|
||||||
|
## Pretraining Data
|
||||||
|
|
||||||
|
The pretraining corpus is a bilingual mix (approximately 45% Italian, 45% English, 10% code by character count):
|
||||||
|
|
||||||
|
**Italian:** FineWeb-2 IT, FineWiki IT, Italian Public Domain (PleIAs), Gazzetta Ufficiale, EuroParl IT, FinePDFs IT, Clean mC4 IT
|
||||||
|
|
||||||
|
**English:** FineWeb-Edu (100B sample), FineWiki EN
|
||||||
|
|
||||||
|
**Code:** StarCoderData
|
||||||
|
|
||||||
|
## Chat Format
|
||||||
|
|
||||||
|
Dante-2B uses a ChatML-style template:
|
||||||
|
|
||||||
|
```
|
||||||
|
<|begin_of_text|><|start_header|>user<|end_header|>
|
||||||
|
Come funziona la fotosintesi?<|eot|>
|
||||||
|
<|start_header|>assistant<|end_header|>
|
||||||
|
La fotosintesi è il processo attraverso cui le piante convertono...
|
||||||
|
```
|
||||||
|
|
||||||
|
## Usage
|
||||||
|
|
||||||
|
```python
|
||||||
|
import torch
|
||||||
|
from tokenizers import Tokenizer
|
||||||
|
from model_architecture import DanteConfig, DanteModel
|
||||||
|
|
||||||
|
# Load model
|
||||||
|
config = DanteConfig.load("config.json")
|
||||||
|
model = DanteModel(config)
|
||||||
|
model.load_state_dict(torch.load("model.pt", map_location="cpu"), strict=False)
|
||||||
|
model.eval().cuda()
|
||||||
|
|
||||||
|
# Load tokenizer
|
||||||
|
tokenizer = Tokenizer.from_file("tokenizer.json")
|
||||||
|
|
||||||
|
# Chat
|
||||||
|
prompt = "<|begin_of_text|><|start_header|>user<|end_header|>\nCiao, come stai?<|eot|>\n<|start_header|>assistant<|end_header|>\n"
|
||||||
|
input_ids = torch.tensor([tokenizer.encode(prompt).ids], device="cuda")
|
||||||
|
|
||||||
|
with torch.no_grad():
|
||||||
|
# Use KV-cache for efficient generation
|
||||||
|
kv_cache = None
|
||||||
|
for _ in range(256):
|
||||||
|
out = model(input_ids if kv_cache is None else input_ids[:, -1:],
|
||||||
|
kv_cache=kv_cache, use_cache=True)
|
||||||
|
kv_cache = out["kv_cache"]
|
||||||
|
next_id = out["logits"][:, -1].argmax(dim=-1, keepdim=True)
|
||||||
|
input_ids = torch.cat([input_ids, next_id], dim=-1)
|
||||||
|
token = tokenizer.decode([next_id.item()])
|
||||||
|
if token in ["<|eot|>", "<|end_of_text|>"]:
|
||||||
|
break
|
||||||
|
print(token, end="", flush=True)
|
||||||
|
```
|
||||||
|
|
||||||
|
> **Tip:** use `repetition_penalty=1.15` for longer generations to avoid repetition loops.
|
||||||
|
|
||||||
|
## Limitations
|
||||||
|
|
||||||
|
- **Scale:** At 2.1B parameters, Dante-2B is not competing with 7B+ models on complex reasoning. It is a capable small model for Italian-language tasks, on-device deployment, and research.
|
||||||
|
- **Knowledge cutoff:** The pretraining data has a knowledge cutoff determined by the source datasets (primarily 2023–2024 web crawls).
|
||||||
|
- **Benchmarks:** Formal evaluation on ITA-Bench and standard benchmarks is in progress and will be added here.
|
||||||
|
- **Safety:** This model has not undergone RLHF or safety-specific alignment. Use responsibly.
|
||||||
|
|
||||||
|
## Training Stack
|
||||||
|
|
||||||
|
| Component | Version / Detail |
|
||||||
|
|---|---|
|
||||||
|
| Framework | PyTorch 2.11 + CUDA 12.8 |
|
||||||
|
| Distributed | DeepSpeed ZeRO-2 |
|
||||||
|
| Mixed precision | FP8 via torchao |
|
||||||
|
| Compilation | torch.compile (reduce-overhead) |
|
||||||
|
| Attention | Flash Attention 2 |
|
||||||
|
| Hardware | 2× NVIDIA H200 NVL (141GB each) |
|
||||||
|
| Peak FLOPS | 1,671 TFLOPS per GPU (NVL spec) |
|
||||||
|
|
||||||
|
## Repository Structure
|
||||||
|
|
||||||
|
```
|
||||||
|
├── model_architecture.py # DanteConfig + DanteModel (full architecture)
|
||||||
|
├── config.json # Model configuration
|
||||||
|
├── tokenizer.json # 64K BPE tokenizer
|
||||||
|
├── model.pt # Model weights
|
||||||
|
└── README.md # This file
|
||||||
|
```
|
||||||
|
|
||||||
|
The full training codebase — from corpus download to SFT — is available on GitHub: **[link to be added]**
|
||||||
|
|
||||||
|
## Citation
|
||||||
|
|
||||||
|
If you use Dante-2B in your research or applications, please cite:
|
||||||
|
|
||||||
|
```bibtex
|
||||||
|
@misc{angeletti2025dante2b,
|
||||||
|
title={Dante-2B: A Bilingual Italian/English Language Model Trained from Scratch},
|
||||||
|
author={Fabio Angeletti},
|
||||||
|
year={2025},
|
||||||
|
url={https://huggingface.co/leafsrls/dante-2b}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## About
|
||||||
|
|
||||||
|
Dante-2B is built and maintained by **Fabio Angeletti**, PhD in Computer Engineering, Adjunct Professor at LUISS and LUISS Business School, and Founder & CEO of [LEAF srls](https://leafsrls.com).
|
||||||
|
|
||||||
|
**LEAF** is a specialized AI company based in Italy, operating at the intersection of Artificial Intelligence, Machine Learning, IoT, and Cybersecurity. We deliver secure, efficient, and transformative digital strategies — from AI model development and system integration to low-power wireless monitoring and edge AI deployment. Our mission is to empower businesses to navigate and succeed in the digital economy through advanced technological innovation.
|
||||||
|
|
||||||
|
For collaborations, commercial licensing, or inquiries: **info@leafsrls.com**
|
||||||
|
|
||||||
|
## License
|
||||||
|
|
||||||
|
This model is released under the [Apache License 2.0](https://www.apache.org/licenses/LICENSE-2.0). You are free to use, modify, and distribute it for any purpose, including commercial use.
|
||||||
32
config.json
Normal file
32
config.json
Normal file
@@ -0,0 +1,32 @@
|
|||||||
|
{
|
||||||
|
"architectures": [
|
||||||
|
"LlamaForCausalLM"
|
||||||
|
],
|
||||||
|
"attention_bias": false,
|
||||||
|
"attention_dropout": 0.0,
|
||||||
|
"bos_token_id": 0,
|
||||||
|
"dtype": "bfloat16",
|
||||||
|
"eos_token_id": 1,
|
||||||
|
"head_dim": 128,
|
||||||
|
"hidden_act": "silu",
|
||||||
|
"hidden_size": 2560,
|
||||||
|
"initializer_range": 0.02,
|
||||||
|
"intermediate_size": 6912,
|
||||||
|
"max_position_embeddings": 4096,
|
||||||
|
"mlp_bias": false,
|
||||||
|
"model_type": "llama",
|
||||||
|
"num_attention_heads": 20,
|
||||||
|
"num_hidden_layers": 28,
|
||||||
|
"num_key_value_heads": 4,
|
||||||
|
"pad_token_id": 2,
|
||||||
|
"pretraining_tp": 1,
|
||||||
|
"rms_norm_eps": 1e-06,
|
||||||
|
"rope_parameters": {
|
||||||
|
"rope_theta": 10000.0,
|
||||||
|
"rope_type": "default"
|
||||||
|
},
|
||||||
|
"tie_word_embeddings": true,
|
||||||
|
"transformers_version": "5.5.0",
|
||||||
|
"use_cache": true,
|
||||||
|
"vocab_size": 64000
|
||||||
|
}
|
||||||
17
dante_export_manifest.json
Normal file
17
dante_export_manifest.json
Normal file
@@ -0,0 +1,17 @@
|
|||||||
|
{
|
||||||
|
"source_checkpoint_dir": "/dir1/llm_corpus/checkpoints/dante_2b_instruct_v2/final_hf",
|
||||||
|
"project_dir": "/home/angeletti/project_slm",
|
||||||
|
"output_dir": "/dir1/llm_corpus/checkpoints/dante_2b_instruct_v2/final_transformers",
|
||||||
|
"source_weight_file": "model.bin",
|
||||||
|
"target_architecture": "LlamaForCausalLM",
|
||||||
|
"safe_serialization": true,
|
||||||
|
"dtype": "bfloat16",
|
||||||
|
"num_layers": 28,
|
||||||
|
"hidden_size": 2560,
|
||||||
|
"num_attention_heads": 20,
|
||||||
|
"num_key_value_heads": 4,
|
||||||
|
"intermediate_size": 6912,
|
||||||
|
"max_position_embeddings": 4096,
|
||||||
|
"rope_theta": 10000.0,
|
||||||
|
"tokenizer_model_max_length": 4096
|
||||||
|
}
|
||||||
8
generation_config.json
Normal file
8
generation_config.json
Normal file
@@ -0,0 +1,8 @@
|
|||||||
|
{
|
||||||
|
"bos_token_id": 0,
|
||||||
|
"do_sample": false,
|
||||||
|
"eos_token_id": 1,
|
||||||
|
"pad_token_id": 2,
|
||||||
|
"transformers_version": "5.5.0",
|
||||||
|
"use_cache": true
|
||||||
|
}
|
||||||
3
model.safetensors
Normal file
3
model.safetensors
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:f5956922446015327823f84774079a030d243fa0a90334df976478116fce2933
|
||||||
|
size 4181517984
|
||||||
40
special_tokens_map.json
Normal file
40
special_tokens_map.json
Normal file
@@ -0,0 +1,40 @@
|
|||||||
|
{
|
||||||
|
"bos_token": "<|begin_of_text|>",
|
||||||
|
"eos_token": "<|end_of_text|>",
|
||||||
|
"pad_token": "<|pad|>",
|
||||||
|
"unk_token": "<|unk|>",
|
||||||
|
"sep_token": "<|sep|>",
|
||||||
|
"mask_token": "<|mask|>",
|
||||||
|
"additional_special_tokens": [
|
||||||
|
"<|start_header|>",
|
||||||
|
"<|end_header|>",
|
||||||
|
"<|eot|>",
|
||||||
|
"<|system|>",
|
||||||
|
"<|user|>",
|
||||||
|
"<|assistant|>",
|
||||||
|
"<think>",
|
||||||
|
"</think>",
|
||||||
|
"<|code|>",
|
||||||
|
"<|/code|>",
|
||||||
|
"<|tool_call|>",
|
||||||
|
"<|tool_result|>",
|
||||||
|
"<|/tool_call|>",
|
||||||
|
"<|/tool_result|>",
|
||||||
|
"<|expert_0|>",
|
||||||
|
"<|expert_1|>",
|
||||||
|
"<|expert_2|>",
|
||||||
|
"<|expert_3|>",
|
||||||
|
"<|expert_4|>",
|
||||||
|
"<|expert_5|>",
|
||||||
|
"<|expert_6|>",
|
||||||
|
"<|expert_7|>",
|
||||||
|
"<|expert_8|>",
|
||||||
|
"<|expert_9|>",
|
||||||
|
"<|expert_10|>",
|
||||||
|
"<|expert_11|>",
|
||||||
|
"<|expert_12|>",
|
||||||
|
"<|expert_13|>",
|
||||||
|
"<|expert_14|>",
|
||||||
|
"<|expert_15|>"
|
||||||
|
]
|
||||||
|
}
|
||||||
316327
tokenizer.json
Normal file
316327
tokenizer.json
Normal file
File diff suppressed because it is too large
Load Diff
333
tokenizer_config.json
Normal file
333
tokenizer_config.json
Normal file
@@ -0,0 +1,333 @@
|
|||||||
|
{
|
||||||
|
"added_tokens_decoder": {
|
||||||
|
"0": {
|
||||||
|
"content": "<|begin_of_text|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"1": {
|
||||||
|
"content": "<|end_of_text|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"2": {
|
||||||
|
"content": "<|pad|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"3": {
|
||||||
|
"content": "<|unk|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"4": {
|
||||||
|
"content": "<|sep|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"5": {
|
||||||
|
"content": "<|mask|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"6": {
|
||||||
|
"content": "<|start_header|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"7": {
|
||||||
|
"content": "<|end_header|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"8": {
|
||||||
|
"content": "<|eot|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"9": {
|
||||||
|
"content": "<|system|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"10": {
|
||||||
|
"content": "<|user|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"11": {
|
||||||
|
"content": "<|assistant|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"12": {
|
||||||
|
"content": "<think>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"13": {
|
||||||
|
"content": "</think>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"14": {
|
||||||
|
"content": "<|code|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"15": {
|
||||||
|
"content": "<|/code|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"16": {
|
||||||
|
"content": "<|tool_call|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"17": {
|
||||||
|
"content": "<|tool_result|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"18": {
|
||||||
|
"content": "<|/tool_call|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"19": {
|
||||||
|
"content": "<|/tool_result|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"20": {
|
||||||
|
"content": "<|expert_0|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"21": {
|
||||||
|
"content": "<|expert_1|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"22": {
|
||||||
|
"content": "<|expert_2|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"23": {
|
||||||
|
"content": "<|expert_3|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"24": {
|
||||||
|
"content": "<|expert_4|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"25": {
|
||||||
|
"content": "<|expert_5|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"26": {
|
||||||
|
"content": "<|expert_6|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"27": {
|
||||||
|
"content": "<|expert_7|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"28": {
|
||||||
|
"content": "<|expert_8|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"29": {
|
||||||
|
"content": "<|expert_9|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"30": {
|
||||||
|
"content": "<|expert_10|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"31": {
|
||||||
|
"content": "<|expert_11|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"32": {
|
||||||
|
"content": "<|expert_12|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"33": {
|
||||||
|
"content": "<|expert_13|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"34": {
|
||||||
|
"content": "<|expert_14|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
},
|
||||||
|
"35": {
|
||||||
|
"content": "<|expert_15|>",
|
||||||
|
"lstrip": false,
|
||||||
|
"normalized": false,
|
||||||
|
"rstrip": false,
|
||||||
|
"single_word": true,
|
||||||
|
"special": true
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"additional_special_tokens": [
|
||||||
|
"<|start_header|>",
|
||||||
|
"<|end_header|>",
|
||||||
|
"<|eot|>",
|
||||||
|
"<|system|>",
|
||||||
|
"<|user|>",
|
||||||
|
"<|assistant|>",
|
||||||
|
"<think>",
|
||||||
|
"</think>",
|
||||||
|
"<|code|>",
|
||||||
|
"<|/code|>",
|
||||||
|
"<|tool_call|>",
|
||||||
|
"<|tool_result|>",
|
||||||
|
"<|/tool_call|>",
|
||||||
|
"<|/tool_result|>",
|
||||||
|
"<|expert_0|>",
|
||||||
|
"<|expert_1|>",
|
||||||
|
"<|expert_2|>",
|
||||||
|
"<|expert_3|>",
|
||||||
|
"<|expert_4|>",
|
||||||
|
"<|expert_5|>",
|
||||||
|
"<|expert_6|>",
|
||||||
|
"<|expert_7|>",
|
||||||
|
"<|expert_8|>",
|
||||||
|
"<|expert_9|>",
|
||||||
|
"<|expert_10|>",
|
||||||
|
"<|expert_11|>",
|
||||||
|
"<|expert_12|>",
|
||||||
|
"<|expert_13|>",
|
||||||
|
"<|expert_14|>",
|
||||||
|
"<|expert_15|>"
|
||||||
|
],
|
||||||
|
"bos_token": "<|begin_of_text|>",
|
||||||
|
"eos_token": "<|end_of_text|>",
|
||||||
|
"pad_token": "<|pad|>",
|
||||||
|
"unk_token": "<|unk|>",
|
||||||
|
"sep_token": "<|sep|>",
|
||||||
|
"mask_token": "<|mask|>",
|
||||||
|
"clean_up_tokenization_spaces": false,
|
||||||
|
"model_max_length": 4096,
|
||||||
|
"tokenizer_class": "PreTrainedTokenizerFast"
|
||||||
|
}
|
||||||
Reference in New Issue
Block a user