--- language: - it base_model: - mii-llm/zagreus-0.4B-ita datasets: - antoniogr7/italic-lexical-elicitation pipeline_tag: text-generation library_name: transformers tags: - italian - italic - elicitation - multiple-choice --- # zagreus-0.4B-ita-elicit An **elicitation** fine-tune of [`mii-llm/zagreus-0.4B-ita`](https://huggingface.co/mii-llm/zagreus-0.4B-ita) for the [ITALIC](https://github.com/Crisp-Unimib/ITALIC) benchmark. It does **not** inject new knowledge (the model is only 0.4B); it surfaces the lexical/semantic knowledge the base model *already* holds and routes it onto the answer-letter channel, while removing the base model's heavy answer-letter bias. **No ITALIC data is used for training** — the benchmark is only read to guard against leakage. Net result: **ITALIC 0.280 → 0.3126** (fast, zero-shot, greedy, n=10,000). ## Results ITALIC, **fast / zero-shot**, greedy, n=10,000 (0 unparsed). Base = `mii-llm/zagreus-0.4B-ita`. | | base | **this model** | Δ | |---|---|---|---| | **overall** | 0.280 | **0.3126** | **+3.2** | | language capability (macro) | 0.2705 | 0.3133 | +4.3 | | culture and commonsense (macro) | 0.2868 | 0.3121 | +2.5 | | `lexicon` | 0.2584 | **0.3841** | **+12.6** | | `synonyms_and_antonyms` | 0.2915 | 0.3275 | +3.6 | The intended target (`lexicon`) moves the most. The culture categories rise modestly too — not from injected facts but from repairing the shared letter-emission head; per-fact culture remains a wall for a 0.4B, so those gains are small. ## Method (brief) Full fine-tune with an elicitation objective, per rendered option ordering: 1. **`CE(gold)`** — commit probability mass to the correct answer letter (peaking anchor, the primary driver). 2. **`λ·KL(q ‖ p)`** — distill a **PMI-cloze teacher** computed over the option *text* (`q`) onto the answer-letter logits (`p`), down-weighted (`λ=0.3`) to avoid flattening. 3. **format term** — keep probability mass on the option-letter tokens. Options are re-permuted every epoch so the model must bind meaning to the letter it lands on, not to a position. Training data is a synthetic Italian lexical/semantic MCQ pool built from **Italian Wiktionary** (kaikki.org): [`antoniogr7/italic-lexical-elicitation`](https://huggingface.co/datasets/antoniogr7/italic-lexical-elicitation). Trained and evaluated in **bfloat16**. ## Usage **Serve/evaluate zero-shot.** The model was trained and measured 0-shot; few-shot demos shift the emission distribution it was tuned for and regress the score. The ITALIC chat template is baked into the tokenizer. ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer name = "antoniogr7/zagreus-0.4B-ita-elicit" tok = AutoTokenizer.from_pretrained(name) model = AutoModelForCausalLM.from_pretrained(name, dtype=torch.bfloat16).cuda().eval() question = "Qual è il sinonimo di «celere»?" options = {"A": "lento", "B": "rapido", "C": "grande", "D": "scuro"} body = question + "\n\n" + "\n".join(f"{k}) {v}" for k, v in options.items()) + "\n\nRisposta:\n" msgs = [{"role": "system", "content": "Sei un assistente utile."}, {"role": "user", "content": body}] ids = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt").cuda() out = model.generate(ids, max_new_tokens=2, do_sample=False) print(tok.decode(out[0, ids.shape[1]:], skip_special_tokens=True).strip()) # -> "B" ``` For faithful benchmark scoring use the official ITALIC harness (vLLM + `run_eval.py`, `fast: true`, `few_shot_file: null`). ## Limitations - **0.4B model.** Only the lexical/semantic slice of ITALIC (~19.5%) is addressable by elicitation; culture is per-fact knowledge (a wall) and grammar is a capacity wall. The realistic ceiling of this approach is ~0.31–0.33 overall. - **Zero-shot only** (see Usage). - Italian only. ## License & attribution The model inherits the license of the base model `mii-llm/zagreus-0.4B-ita`. The training dataset is derived from Italian Wiktionary via kaikki.org and released under **CC BY-SA 4.0** (attribution: Wiktionary contributors; extraction by wiktextract / kaikki.org).