初始化项目,由ModelHub XC社区提供模型
Model: SerFabio89/dante-2b-ita-instruct Source: Original Platform
This commit is contained in:
35
.gitattributes
vendored
Normal file
35
.gitattributes
vendored
Normal file
@@ -0,0 +1,35 @@
|
||||
*.7z filter=lfs diff=lfs merge=lfs -text
|
||||
*.arrow filter=lfs diff=lfs merge=lfs -text
|
||||
*.bin filter=lfs diff=lfs merge=lfs -text
|
||||
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
||||
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
||||
*.ftz filter=lfs diff=lfs merge=lfs -text
|
||||
*.gz filter=lfs diff=lfs merge=lfs -text
|
||||
*.h5 filter=lfs diff=lfs merge=lfs -text
|
||||
*.joblib filter=lfs diff=lfs merge=lfs -text
|
||||
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
||||
*.model filter=lfs diff=lfs merge=lfs -text
|
||||
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
||||
*.npy filter=lfs diff=lfs merge=lfs -text
|
||||
*.npz filter=lfs diff=lfs merge=lfs -text
|
||||
*.onnx filter=lfs diff=lfs merge=lfs -text
|
||||
*.ot filter=lfs diff=lfs merge=lfs -text
|
||||
*.parquet filter=lfs diff=lfs merge=lfs -text
|
||||
*.pb filter=lfs diff=lfs merge=lfs -text
|
||||
*.pickle filter=lfs diff=lfs merge=lfs -text
|
||||
*.pkl filter=lfs diff=lfs merge=lfs -text
|
||||
*.pt filter=lfs diff=lfs merge=lfs -text
|
||||
*.pth filter=lfs diff=lfs merge=lfs -text
|
||||
*.rar filter=lfs diff=lfs merge=lfs -text
|
||||
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
||||
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar filter=lfs diff=lfs merge=lfs -text
|
||||
*.tflite filter=lfs diff=lfs merge=lfs -text
|
||||
*.tgz filter=lfs diff=lfs merge=lfs -text
|
||||
*.wasm filter=lfs diff=lfs merge=lfs -text
|
||||
*.xz filter=lfs diff=lfs merge=lfs -text
|
||||
*.zip filter=lfs diff=lfs merge=lfs -text
|
||||
*.zst filter=lfs diff=lfs merge=lfs -text
|
||||
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
||||
232
README.md
Normal file
232
README.md
Normal file
@@ -0,0 +1,232 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
language:
|
||||
- it
|
||||
- en
|
||||
tags:
|
||||
- italian
|
||||
- bilingual
|
||||
- from-scratch
|
||||
- decoder-only
|
||||
- llama-style
|
||||
- gqa
|
||||
- causal-lm
|
||||
- text-generation
|
||||
- sft
|
||||
- instruct
|
||||
library_name: transformers
|
||||
pipeline_tag: text-generation
|
||||
model-index:
|
||||
- name: Dante-2B
|
||||
results: []
|
||||
---
|
||||
|
||||
# Dante-2B
|
||||
|
||||
**A 2.1B parameter bilingual Italian/English language model, trained entirely from scratch.**
|
||||
|
||||
Dante-2B is a decoder-only transformer built by a single developer on 2× NVIDIA H200 NVL GPUs. Everything is native — the tokenizer, the architecture, the training pipeline — no fine-tune of an existing English model, no retrofitted multilingual vocabulary. Italian is a first-class citizen from byte zero.
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| **Parameters** | 2.1B (all active, no MoE) |
|
||||
| **Architecture** | Llama-style decoder-only transformer |
|
||||
| **Languages** | Italian, English |
|
||||
| **Tokenizer** | Custom 64K BPE, Italian-native |
|
||||
| **Context** | 4,096 tokens |
|
||||
| **Training** | 120B tokens across 3 phases |
|
||||
| **Hardware** | 2× NVIDIA H200 NVL |
|
||||
| **License** | Apache 2.0 |
|
||||
|
||||
## Why Dante-2B?
|
||||
|
||||
Most "Italian" language models are English-first architectures with Italian bolted on during fine-tuning. Their tokenizers fragment Italian text into too many subwords, wasting context window and degrading fluency. Dante-2B takes a different path: the tokenizer was trained on a balanced Italian/English corpus from the start, so Italian apostrophe contractions (`l'`, `dell'`, `un'`), accented vowels, and common morphological patterns are single tokens — not sequences of three.
|
||||
|
||||
This is not a research lab effort with 512 GPUs. It's a from-scratch build on commodity hardware, with every engineering decision documented — including the failures.
|
||||
|
||||
## Architecture
|
||||
|
||||
Dante-2B uses a pure Llama-style architecture optimized for training speed on limited hardware.
|
||||
|
||||
| Component | Detail |
|
||||
|---|---|
|
||||
| Type | Decoder-only dense transformer |
|
||||
| `d_model` | 2,560 |
|
||||
| Layers | 28 |
|
||||
| Attention | Grouped Query Attention (GQA) — 20 query heads, 4 KV heads (5:1) |
|
||||
| `d_head` | 128 (optimal for Flash Attention) |
|
||||
| FFN | SwiGLU, `d_ff` = 6,912 |
|
||||
| Normalization | RMSNorm (ε = 1e-6) |
|
||||
| Position encoding | RoPE (θ = 10,000) |
|
||||
| Vocab size | 64,000 |
|
||||
| Weight tying | Embedding ↔ LM head |
|
||||
| Dropout | 0.0 |
|
||||
|
||||
## Tokenizer
|
||||
|
||||
A custom 64K BPE tokenizer trained on a character-balanced subset of the corpus (45% Italian, 45% English, 10% code).
|
||||
|
||||
Key properties:
|
||||
- Italian accented vowels (à, è, é, ì, ò, ù) are always single tokens
|
||||
- Italian apostrophe contractions (`l'intelligenza`, `dell'algoritmo`) are handled as single tokens
|
||||
- Custom pre-tokenization regex preserves Italian morphology
|
||||
- 36 special tokens including ChatML markers, thinking tokens, tool-use tokens, and expert routing tokens (future-proofing)
|
||||
|
||||
Fertility (tokens per word): ~1.35 Italian, ~1.20 English. For comparison, LLaMA's tokenizer scores ~1.85 on Italian.
|
||||
|
||||
## Training
|
||||
|
||||
### Phase 1 — Base Pretraining
|
||||
|
||||
90B tokens at sequence length 2,048 over ~57,200 steps.
|
||||
|
||||
- Optimizer: AdamW (β₁=0.9, β₂=0.95, weight decay 0.1)
|
||||
- LR schedule: cosine, 3e-4 → 3e-5, 2,000-step warmup
|
||||
- Batch size: micro_batch 24 × grad_accum 16 × 2 GPUs = 1,572,864 tokens/step
|
||||
- Infrastructure: DeepSpeed ZeRO-2, FP8 (torchao), torch.compile (reduce-overhead), Flash Attention 2
|
||||
- Throughput: ~88–89K tokens/sec, 28% MFU
|
||||
- Final loss: ~1.85–1.93
|
||||
|
||||
### Phase 2 — Continual Pretraining (Context Extension)
|
||||
|
||||
30B tokens at sequence length 4,096 over ~28,600 steps.
|
||||
|
||||
- LR schedule: cosine, 6e-5 → 6e-6, 500-step warmup
|
||||
- Purpose: extend context window from 2,048 → 4,096 while consolidating learned representations
|
||||
- Final loss: ~1.75–1.80 (expected plateau — gains manifest in downstream performance)
|
||||
- Inference test: coherent Italian at 125.7 tok/s
|
||||
|
||||
### Phase 3 — Supervised Fine-Tuning (SFT)
|
||||
|
||||
540K conversations (approximately 40% Italian), 1 epoch in ~3.5 hours.
|
||||
|
||||
- Micro_batch 12, grad_accum 10, ~89K tok/s, 33.4% MFU
|
||||
- Eval loss: 1.23 (declining toward ~1.15)
|
||||
- Strategy: single-epoch SFT to avoid overfitting
|
||||
- Loss masking: only assistant turns contribute to loss; user turns and headers are masked with -100
|
||||
|
||||
**SFT Datasets:**
|
||||
- HuggingFaceH4/no_robots
|
||||
- OpenAssistant/oasst1 (via Guanaco)
|
||||
- HuggingFaceH4/ultrachat_200k
|
||||
- efederici/capybara-claude-15k-ita
|
||||
- anakin87/fine-instructions-ita-70k
|
||||
- Mattimax/Camoscio-ITA
|
||||
- mchl-labs/stambecco_data_it
|
||||
- teelinsan/camoscio
|
||||
- DeepMount00/Sonnet-3.5-ITA-INSTRUCT
|
||||
- DeepMount00/italian_conversations
|
||||
|
||||
## Pretraining Data
|
||||
|
||||
The pretraining corpus is a bilingual mix (approximately 45% Italian, 45% English, 10% code by character count):
|
||||
|
||||
**Italian:** FineWeb-2 IT, FineWiki IT, Italian Public Domain (PleIAs), Gazzetta Ufficiale, EuroParl IT, FinePDFs IT, Clean mC4 IT
|
||||
|
||||
**English:** FineWeb-Edu (100B sample), FineWiki EN
|
||||
|
||||
**Code:** StarCoderData
|
||||
|
||||
## Chat Format
|
||||
|
||||
Dante-2B uses a ChatML-style template:
|
||||
|
||||
```
|
||||
<|begin_of_text|><|start_header|>user<|end_header|>
|
||||
Come funziona la fotosintesi?<|eot|>
|
||||
<|start_header|>assistant<|end_header|>
|
||||
La fotosintesi è il processo attraverso cui le piante convertono...
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
```python
|
||||
import torch
|
||||
from tokenizers import Tokenizer
|
||||
from model_architecture import DanteConfig, DanteModel
|
||||
|
||||
# Load model
|
||||
config = DanteConfig.load("config.json")
|
||||
model = DanteModel(config)
|
||||
model.load_state_dict(torch.load("model.pt", map_location="cpu"), strict=False)
|
||||
model.eval().cuda()
|
||||
|
||||
# Load tokenizer
|
||||
tokenizer = Tokenizer.from_file("tokenizer.json")
|
||||
|
||||
# Chat
|
||||
prompt = "<|begin_of_text|><|start_header|>user<|end_header|>\nCiao, come stai?<|eot|>\n<|start_header|>assistant<|end_header|>\n"
|
||||
input_ids = torch.tensor([tokenizer.encode(prompt).ids], device="cuda")
|
||||
|
||||
with torch.no_grad():
|
||||
# Use KV-cache for efficient generation
|
||||
kv_cache = None
|
||||
for _ in range(256):
|
||||
out = model(input_ids if kv_cache is None else input_ids[:, -1:],
|
||||
kv_cache=kv_cache, use_cache=True)
|
||||
kv_cache = out["kv_cache"]
|
||||
next_id = out["logits"][:, -1].argmax(dim=-1, keepdim=True)
|
||||
input_ids = torch.cat([input_ids, next_id], dim=-1)
|
||||
token = tokenizer.decode([next_id.item()])
|
||||
if token in ["<|eot|>", "<|end_of_text|>"]:
|
||||
break
|
||||
print(token, end="", flush=True)
|
||||
```
|
||||
|
||||
> **Tip:** use `repetition_penalty=1.15` for longer generations to avoid repetition loops.
|
||||
|
||||
## Limitations
|
||||
|
||||
- **Scale:** At 2.1B parameters, Dante-2B is not competing with 7B+ models on complex reasoning. It is a capable small model for Italian-language tasks, on-device deployment, and research.
|
||||
- **Knowledge cutoff:** The pretraining data has a knowledge cutoff determined by the source datasets (primarily 2023–2024 web crawls).
|
||||
- **Benchmarks:** Formal evaluation on ITA-Bench and standard benchmarks is in progress and will be added here.
|
||||
- **Safety:** This model has not undergone RLHF or safety-specific alignment. Use responsibly.
|
||||
|
||||
## Training Stack
|
||||
|
||||
| Component | Version / Detail |
|
||||
|---|---|
|
||||
| Framework | PyTorch 2.11 + CUDA 12.8 |
|
||||
| Distributed | DeepSpeed ZeRO-2 |
|
||||
| Mixed precision | FP8 via torchao |
|
||||
| Compilation | torch.compile (reduce-overhead) |
|
||||
| Attention | Flash Attention 2 |
|
||||
| Hardware | 2× NVIDIA H200 NVL (141GB each) |
|
||||
| Peak FLOPS | 1,671 TFLOPS per GPU (NVL spec) |
|
||||
|
||||
## Repository Structure
|
||||
|
||||
```
|
||||
├── model_architecture.py # DanteConfig + DanteModel (full architecture)
|
||||
├── config.json # Model configuration
|
||||
├── tokenizer.json # 64K BPE tokenizer
|
||||
├── model.pt # Model weights
|
||||
└── README.md # This file
|
||||
```
|
||||
|
||||
The full training codebase — from corpus download to SFT — is available on GitHub: **[link to be added]**
|
||||
|
||||
## Citation
|
||||
|
||||
If you use Dante-2B in your research or applications, please cite:
|
||||
|
||||
```bibtex
|
||||
@misc{angeletti2025dante2b,
|
||||
title={Dante-2B: A Bilingual Italian/English Language Model Trained from Scratch},
|
||||
author={Fabio Angeletti},
|
||||
year={2025},
|
||||
url={https://huggingface.co/leafsrls/dante-2b}
|
||||
}
|
||||
```
|
||||
|
||||
## About
|
||||
|
||||
Dante-2B is built and maintained by **Fabio Angeletti**, PhD in Computer Engineering, Adjunct Professor at LUISS and LUISS Business School, and Founder & CEO of [LEAF srls](https://leafsrls.com).
|
||||
|
||||
**LEAF** is a specialized AI company based in Italy, operating at the intersection of Artificial Intelligence, Machine Learning, IoT, and Cybersecurity. We deliver secure, efficient, and transformative digital strategies — from AI model development and system integration to low-power wireless monitoring and edge AI deployment. Our mission is to empower businesses to navigate and succeed in the digital economy through advanced technological innovation.
|
||||
|
||||
For collaborations, commercial licensing, or inquiries: **info@leafsrls.com**
|
||||
|
||||
## License
|
||||
|
||||
This model is released under the [Apache License 2.0](https://www.apache.org/licenses/LICENSE-2.0). You are free to use, modify, and distribute it for any purpose, including commercial use.
|
||||
32
config.json
Normal file
32
config.json
Normal file
@@ -0,0 +1,32 @@
|
||||
{
|
||||
"architectures": [
|
||||
"LlamaForCausalLM"
|
||||
],
|
||||
"attention_bias": false,
|
||||
"attention_dropout": 0.0,
|
||||
"bos_token_id": 0,
|
||||
"dtype": "bfloat16",
|
||||
"eos_token_id": 1,
|
||||
"head_dim": 128,
|
||||
"hidden_act": "silu",
|
||||
"hidden_size": 2560,
|
||||
"initializer_range": 0.02,
|
||||
"intermediate_size": 6912,
|
||||
"max_position_embeddings": 4096,
|
||||
"mlp_bias": false,
|
||||
"model_type": "llama",
|
||||
"num_attention_heads": 20,
|
||||
"num_hidden_layers": 28,
|
||||
"num_key_value_heads": 4,
|
||||
"pad_token_id": 2,
|
||||
"pretraining_tp": 1,
|
||||
"rms_norm_eps": 1e-06,
|
||||
"rope_parameters": {
|
||||
"rope_theta": 10000.0,
|
||||
"rope_type": "default"
|
||||
},
|
||||
"tie_word_embeddings": true,
|
||||
"transformers_version": "5.5.0",
|
||||
"use_cache": true,
|
||||
"vocab_size": 64000
|
||||
}
|
||||
17
dante_export_manifest.json
Normal file
17
dante_export_manifest.json
Normal file
@@ -0,0 +1,17 @@
|
||||
{
|
||||
"source_checkpoint_dir": "/dir1/llm_corpus/checkpoints/dante_2b_instruct_v2/final_hf",
|
||||
"project_dir": "/home/angeletti/project_slm",
|
||||
"output_dir": "/dir1/llm_corpus/checkpoints/dante_2b_instruct_v2/final_transformers",
|
||||
"source_weight_file": "model.bin",
|
||||
"target_architecture": "LlamaForCausalLM",
|
||||
"safe_serialization": true,
|
||||
"dtype": "bfloat16",
|
||||
"num_layers": 28,
|
||||
"hidden_size": 2560,
|
||||
"num_attention_heads": 20,
|
||||
"num_key_value_heads": 4,
|
||||
"intermediate_size": 6912,
|
||||
"max_position_embeddings": 4096,
|
||||
"rope_theta": 10000.0,
|
||||
"tokenizer_model_max_length": 4096
|
||||
}
|
||||
8
generation_config.json
Normal file
8
generation_config.json
Normal file
@@ -0,0 +1,8 @@
|
||||
{
|
||||
"bos_token_id": 0,
|
||||
"do_sample": false,
|
||||
"eos_token_id": 1,
|
||||
"pad_token_id": 2,
|
||||
"transformers_version": "5.5.0",
|
||||
"use_cache": true
|
||||
}
|
||||
3
model.safetensors
Normal file
3
model.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:f5956922446015327823f84774079a030d243fa0a90334df976478116fce2933
|
||||
size 4181517984
|
||||
40
special_tokens_map.json
Normal file
40
special_tokens_map.json
Normal file
@@ -0,0 +1,40 @@
|
||||
{
|
||||
"bos_token": "<|begin_of_text|>",
|
||||
"eos_token": "<|end_of_text|>",
|
||||
"pad_token": "<|pad|>",
|
||||
"unk_token": "<|unk|>",
|
||||
"sep_token": "<|sep|>",
|
||||
"mask_token": "<|mask|>",
|
||||
"additional_special_tokens": [
|
||||
"<|start_header|>",
|
||||
"<|end_header|>",
|
||||
"<|eot|>",
|
||||
"<|system|>",
|
||||
"<|user|>",
|
||||
"<|assistant|>",
|
||||
"<think>",
|
||||
"</think>",
|
||||
"<|code|>",
|
||||
"<|/code|>",
|
||||
"<|tool_call|>",
|
||||
"<|tool_result|>",
|
||||
"<|/tool_call|>",
|
||||
"<|/tool_result|>",
|
||||
"<|expert_0|>",
|
||||
"<|expert_1|>",
|
||||
"<|expert_2|>",
|
||||
"<|expert_3|>",
|
||||
"<|expert_4|>",
|
||||
"<|expert_5|>",
|
||||
"<|expert_6|>",
|
||||
"<|expert_7|>",
|
||||
"<|expert_8|>",
|
||||
"<|expert_9|>",
|
||||
"<|expert_10|>",
|
||||
"<|expert_11|>",
|
||||
"<|expert_12|>",
|
||||
"<|expert_13|>",
|
||||
"<|expert_14|>",
|
||||
"<|expert_15|>"
|
||||
]
|
||||
}
|
||||
316327
tokenizer.json
Normal file
316327
tokenizer.json
Normal file
File diff suppressed because it is too large
Load Diff
333
tokenizer_config.json
Normal file
333
tokenizer_config.json
Normal file
@@ -0,0 +1,333 @@
|
||||
{
|
||||
"added_tokens_decoder": {
|
||||
"0": {
|
||||
"content": "<|begin_of_text|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"1": {
|
||||
"content": "<|end_of_text|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"2": {
|
||||
"content": "<|pad|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"3": {
|
||||
"content": "<|unk|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"4": {
|
||||
"content": "<|sep|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"5": {
|
||||
"content": "<|mask|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"6": {
|
||||
"content": "<|start_header|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"7": {
|
||||
"content": "<|end_header|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"8": {
|
||||
"content": "<|eot|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"9": {
|
||||
"content": "<|system|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"10": {
|
||||
"content": "<|user|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"11": {
|
||||
"content": "<|assistant|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"12": {
|
||||
"content": "<think>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"13": {
|
||||
"content": "</think>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"14": {
|
||||
"content": "<|code|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"15": {
|
||||
"content": "<|/code|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"16": {
|
||||
"content": "<|tool_call|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"17": {
|
||||
"content": "<|tool_result|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"18": {
|
||||
"content": "<|/tool_call|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"19": {
|
||||
"content": "<|/tool_result|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"20": {
|
||||
"content": "<|expert_0|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"21": {
|
||||
"content": "<|expert_1|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"22": {
|
||||
"content": "<|expert_2|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"23": {
|
||||
"content": "<|expert_3|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"24": {
|
||||
"content": "<|expert_4|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"25": {
|
||||
"content": "<|expert_5|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"26": {
|
||||
"content": "<|expert_6|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"27": {
|
||||
"content": "<|expert_7|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"28": {
|
||||
"content": "<|expert_8|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"29": {
|
||||
"content": "<|expert_9|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"30": {
|
||||
"content": "<|expert_10|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"31": {
|
||||
"content": "<|expert_11|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"32": {
|
||||
"content": "<|expert_12|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"33": {
|
||||
"content": "<|expert_13|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"34": {
|
||||
"content": "<|expert_14|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
},
|
||||
"35": {
|
||||
"content": "<|expert_15|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": true,
|
||||
"special": true
|
||||
}
|
||||
},
|
||||
"additional_special_tokens": [
|
||||
"<|start_header|>",
|
||||
"<|end_header|>",
|
||||
"<|eot|>",
|
||||
"<|system|>",
|
||||
"<|user|>",
|
||||
"<|assistant|>",
|
||||
"<think>",
|
||||
"</think>",
|
||||
"<|code|>",
|
||||
"<|/code|>",
|
||||
"<|tool_call|>",
|
||||
"<|tool_result|>",
|
||||
"<|/tool_call|>",
|
||||
"<|/tool_result|>",
|
||||
"<|expert_0|>",
|
||||
"<|expert_1|>",
|
||||
"<|expert_2|>",
|
||||
"<|expert_3|>",
|
||||
"<|expert_4|>",
|
||||
"<|expert_5|>",
|
||||
"<|expert_6|>",
|
||||
"<|expert_7|>",
|
||||
"<|expert_8|>",
|
||||
"<|expert_9|>",
|
||||
"<|expert_10|>",
|
||||
"<|expert_11|>",
|
||||
"<|expert_12|>",
|
||||
"<|expert_13|>",
|
||||
"<|expert_14|>",
|
||||
"<|expert_15|>"
|
||||
],
|
||||
"bos_token": "<|begin_of_text|>",
|
||||
"eos_token": "<|end_of_text|>",
|
||||
"pad_token": "<|pad|>",
|
||||
"unk_token": "<|unk|>",
|
||||
"sep_token": "<|sep|>",
|
||||
"mask_token": "<|mask|>",
|
||||
"clean_up_tokenization_spaces": false,
|
||||
"model_max_length": 4096,
|
||||
"tokenizer_class": "PreTrainedTokenizerFast"
|
||||
}
|
||||
Reference in New Issue
Block a user