Files
ModelHub XC bc49f1e127 初始化项目,由ModelHub XC社区提供模型
Model: fgoose180/zagreus-0.4B-xmoons-improved
Source: Original Platform
2026-09-12 04:28:16 +08:00

243 lines
8.5 KiB
Markdown

---
library_name: transformers
pipeline_tag: text-generation
language:
- it
license: llama3
base_model:
- mii-llm/zagreus-0.4B-ita
datasets:
- efederici/pinocchio
- FinancialSupport/italic_sft
- FinancialSupport/italic_sft_ext
- FinancialSupport/quiz_militare
tags:
- llama
- italian
- small-language-model
- multiple-choice
- supervised-fine-tuning
- post-training
- italic
---
# Zagreus 0.4B xmoons improved
Zagreus 0.4B xmoons improved is an Italian multiple-choice model fine-tuned
from [`mii-llm/zagreus-0.4B-ita`](https://huggingface.co/mii-llm/zagreus-0.4B-ita).
It is the selected checkpoint from the reproducible H100 run that reached
**44.35% accuracy** on the official 10,000-question ITALIC benchmark in fast,
five-shot mode through vLLM.
The model was developed for the
[`mii-llm/Post-Training-Challenge`](https://huggingface.co/spaces/mii-llm/Post-Training-Challenge).
The model is a compact Llama-style causal language model with 437,760,960
parameters. Its tokenizer and Llama 3 chat template come from
[`swap-uniba/LLaMAntino-3-ANITA-8B-Inst-DPO-ITA`](https://huggingface.co/swap-uniba/LLaMAntino-3-ANITA-8B-Inst-DPO-ITA).
The repository is a self-contained Transformers checkpoint and uses
`<|eot_id|>` as both the end-of-turn and generation stop token.
## Intended use
The model is intended for:
- research on compact Italian language models;
- Italian multiple-choice question answering;
- participation in the MII Post-Training Challenge;
- reproduction and analysis of the ITALIC post-training experiment.
It is not designed as a general-purpose assistant, a factual authority, or a
component for medical, legal, financial, or other high-stakes decisions.
## Usage
Install PyTorch and Transformers, then load the checkpoint from the repository
root:
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "fgoose180/zagreus-0.4B-xmoons-improved"
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.bfloat16 if device == "cuda" else torch.float32
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
dtype=dtype,
).to(device)
messages = [
{"role": "system", "content": "Sei un assistente utile."},
{
"role": "user",
"content": """Rispondi alla seguente domanda a scelta multipla sull'argomento 'data_geography'. La tua risposta deve essere nel seguente formato: 'LETTERA' (senza virgolette) dove LETTERA è una tra ABCD. Scrivi solo la lettera corrispondente alla tua risposta senza spiegazioni.
Qual è la capitale d'Italia?
A) Roma
B) Milano
C) Torino
D) Napoli
Risposta:""",
},
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
return_tensors="pt",
return_dict=True,
).to(device)
with torch.inference_mode():
output = model.generate(
**inputs,
max_new_tokens=8,
do_sample=False,
eos_token_id=tokenizer.eos_token_id,
pad_token_id=tokenizer.pad_token_id,
)
answer = tokenizer.decode(
output[0, inputs["input_ids"].shape[1] :],
skip_special_tokens=True,
).strip()
print(answer)
```
The reported benchmark result uses the five fixed official ITALIC
demonstrations. Zero-shot use, different prompts, sampling, or another chat
template should not be expected to reproduce the reported score.
## Model architecture
| Property | Value |
|---|---|
| Architecture | `LlamaForCausalLM` |
| Model type | Llama-style decoder-only transformer |
| Parameters | 437,760,960 |
| Hidden size | 960 |
| Layers | 32 |
| Attention heads / KV heads | 15 / 5 |
| Vocabulary size | 128,256 |
| Maximum trained context | 2,048 tokens during SFT |
| Checkpoint serialization | SafeTensors |
## Training
### Data
The final training stream was built online from pinned revisions of four public
datasets:
| Dataset | Main contribution |
|---|---|
| `efederici/pinocchio` | Italian language, culture, and general-knowledge MCQs |
| `FinancialSupport/italic_sft` | ITALIC-style Italian MCQs |
| `FinancialSupport/italic_sft_ext` | Extended ITALIC-style MCQs |
| `FinancialSupport/quiz_militare` | Italian civic and general-knowledge MCQs |
Rows were normalized, exactly deduplicated, split before upsampling, and
decontaminated against the official ITALIC test and five demonstrations using
character TF-IDF cosine and word-shingle MinHash/LSH. The final artifacts were:
| Artifact | Rows | Unique questions |
|---|---:|---:|
| Clean training stream | 108,288 | 67,654 |
| Clean validation holdout | 6,745 | 6,745 |
The training stream contains 53,591 language draws and 54,697 culture draws.
The benchmark labels were not used for checkpoint selection.
### Objective and hyperparameters
The model was trained for one epoch with completion-only supervised loss. Each
target was rendered with five distinct training-pool demonstrations and
independently permuted answer options. The official five ITALIC demonstrations
were reserved for validation and final evaluation.
| Parameter | Value |
|---|---:|
| Hardware | 1x NVIDIA H100 |
| Optimizer steps | 6,768 |
| Batch size | 16 |
| Initial learning rate | `3e-4` |
| Schedule | 50-step warmup, cosine decay |
| Weight decay | `0.0` |
| Maximum sequence length | 2,048 |
| Forward precision | BF16 autocast |
| Parameters and optimizer updates | FP32 |
| Gradient clipping | `1.0` |
| Random seed | 0 |
| Few-shot probability / count | `1.0 / 5` |
| Option permutation probability | `1.0` |
The detached Modal pipeline took approximately 84 minutes end to end. Carbon
emissions were not measured, so no emissions estimate is reported.
## Evaluation
The selected checkpoint was chosen at step 6,768 using only a clean validation
holdout and the criterion
`0.5 * language_accuracy + 0.5 * culture_accuracy`.
| Evaluation | Questions | Accuracy | Balanced accuracy | Unparsed |
|---|---:|---:|---:|---:|
| Validation holdout | 6,745 | 55.97% | 61.38% | 0 |
| Official ITALIC fast five-shot, vLLM 0.26 | 10,000 | **44.35%** | not reported | 0 |
The official evaluation used greedy decoding, the official fast answer
extractor, a pinned ITALIC harness, and the official five demonstrations. The
result contained 4,435 correct answers out of 10,000.
## Reproducibility
The training and evaluation implementation is available in
[`mattiacurri/zagreus-italic-challenge-xmoons`](https://github.com/mattiacurri/zagreus-italic-challenge-xmoons).
The exact local source snapshot used for this release is identified by Git
commit `afc6b7f09a1e107f38ae04358ece7ed85f6be7a3`.
Selected artifact identities:
```text
d679f3e79221aea95c64ef0be177530b02dce25d4745b0247cbae7b677f1ceab training pool
749dbcdea8b244c3609acecddf618d9780559f501da18b09f650dfe10f2fb881 validation pool
0d534fd8eceec72b4fc3179d77afe7d6291766f5831bc4c8b69b9c45463c1656 selection.json
2ac2f6e412ea1e8f67ce6d8395c1c9845270101590d84de1be748e5919833c93 official result JSON
```
The source repository contains the full experiment and reproducibility reports.
## Limitations and risks
- The model is optimized for Italian multiple-choice prompts and often emits
only an answer letter. It is not a broadly instruction-tuned chat model.
- The 44.35% figure is one benchmark result under one exact prompt and backend
configuration; it does not measure general Italian language competence.
- Exact and fuzzy decontamination reduce known overlap but cannot prove the
absence of semantic contamination or benchmark-distribution overfitting.
- The public training sources may contain factual errors, social biases,
stereotypes, or outdated information that can be inherited by the model.
- Outputs outside the trained answer format may be unreliable.
## License and attribution
The base model weights are published under Apache-2.0. The included tokenizer
and chat template come from a Llama 3 derivative and remain subject to the
published Llama 3 terms; the metadata therefore uses the more restrictive
`llama3` license identifier. Training datasets and benchmark assets retain
their own terms. Some `FinancialSupport` dataset cards do not declare a license,
so users should verify those terms before redistribution or commercial use.
## Acknowledgements
- [`firekern/zagreus-italic-challenge-xmoons`](https://github.com/firekern/zagreus-italic-challenge-xmoons), the upstream repository
- [`mii-llm/zagreus-0.4B-ita`](https://huggingface.co/mii-llm/zagreus-0.4B-ita)
- [`swap-uniba/LLaMAntino-3-ANITA-8B-Inst-DPO-ITA`](https://huggingface.co/swap-uniba/LLaMAntino-3-ANITA-8B-Inst-DPO-ITA)
- [`mii-llm/Post-Training-Challenge`](https://huggingface.co/spaces/mii-llm/Post-Training-Challenge)
- [ITALIC](https://italicbench.it/)