Files
zagreus_0.4_competition/README.md

59 lines
2.9 KiB
Markdown
Raw Permalink Normal View History

---
base_model: mii-llm/nesso-0.4B-agentic
language:
- it
pipeline_tag: text-generation
tags:
- distillation
- on-policy-distillation
- italic
- italian
- mcqa
---
# zagreus_0.4_competition
**nesso-0.4B-agentic improved on the [ITALIC](https://github.com/Crisp-Unimib/ITALIC) benchmark by on-policy distillation from Coloss/nesso-3B.**
Official ITALIC harness (`run_eval.py`, full 10,000 questions, 5-shot):
| Model | fast | slow (CoT) |
|---|---:|---:|
| mii-llm/nesso-0.4B-agentic (baseline) | 33.1% | 0.0%* |
| **this model** | **37.2%** | **33.2%** |
| Coloss/nesso-3B (teacher) | 50.7% | 2.5%* |
\* baseline/teacher answer with a bare letter, which the official slow-mode extractor cannot parse; this model answers parseably in both modes.
## Method
On-policy distillation ([tinker-cookbook recipe](https://github.com/thinking-machines-lab/tinker-cookbook/blob/main/tinker_cookbook/recipes/distillation/on_policy_distillation.py), reimplemented in [mii-llm/palingenesis](https://github.com/mii-llm/palingenesis) branch `odp`): every step the student samples completions with its current weights and the loss is the full-distribution reverse KL to the teacher over exactly those tokens. Student (ChatML) and teacher (Llama-3 template) share the Llama-3 base vocabulary; prompts are rendered per-model and the student's `<|im_end|>` mass is merged into the teacher's `<|eot_id|>` slot, so stopping behavior is distilled too.
The winning run (`opd_v3`, checkpoint step 550 of 600):
- pool of 49,643 ITALIC-format MCQA prompts (pinocchio-raw + MMLU-ita sources, deduped against the ITALIC 10k test set by normalized-question hash), **filtered to rows the teacher answers correctly** (`pgs distill-score` — pure KL would distill the teacher's errors too)
- Italian-language (vocabulary/grammar/comprehension) rows upweighted ×4
- 80% of prompts rendered with ITALIC's official 5 shots, `max_new_tokens 8` (terse supervision)
- 600 steps, batch 32 prompts, lr 1e-5 cosine, full reverse KL, ~40 min on one A100 80GB
Training/eval code: [giux78/zagreus_0.4_competition](https://github.com/giux78/zagreus_0.4_competition) (experiment) + [mii-llm/palingenesis](https://github.com/mii-llm/palingenesis) `odp` branch (library, `pgs distill`).
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("giux78/zagreus_0.4_competition")
model = AutoModelForCausalLM.from_pretrained("giux78/zagreus_0.4_competition", dtype="bfloat16")
messages = [
{"role": "system", "content": "Sei un assistente utile."},
{"role": "user", "content": "Rispondi alla seguente domanda a scelta multipla..."},
]
ids = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt")
out = model.generate(ids, max_new_tokens=8, do_sample=False)
print(tok.decode(out[0, ids.shape[1]:], skip_special_tokens=True))
```
For ITALIC-style evaluation use `max_new_tokens=8` and greedy decoding (the model answers with the bare option letter).