189 lines
9.8 KiB
Markdown
189 lines
9.8 KiB
Markdown
|
|
---
|
|||
|
|
license: apache-2.0
|
|||
|
|
base_model: Qwen/Qwen2.5-14B-Instruct-AWQ
|
|||
|
|
tags:
|
|||
|
|
- iol-ai-2026
|
|||
|
|
- linguistics
|
|||
|
|
- reasoning
|
|||
|
|
language:
|
|||
|
|
- en
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# IOL-AI 2026 — Qwen2.5-14B-Instruct-AWQ
|
|||
|
|
|
|||
|
|
Submission for the [IOL-AI 2026 Linguistics Olympiad Challenge](https://iolai.org).
|
|||
|
|
|
|||
|
|
The weights are an unmodified copy of
|
|||
|
|
[`Qwen/Qwen2.5-14B-Instruct-AWQ`](https://huggingface.co/Qwen/Qwen2.5-14B-Instruct-AWQ)
|
|||
|
|
(Apache-2.0, redistributable), shipped in-repo because the evaluation sandbox
|
|||
|
|
has no internet access. **All of the work is in `script.py`.**
|
|||
|
|
|
|||
|
|
## Approach
|
|||
|
|
|
|||
|
|
The eval budget is 30 minutes on a 16 GB T4 for a test set of only ~90
|
|||
|
|
sub-questions, so compute per problem is abundant while *reliability* is
|
|||
|
|
scarce. The script is built around that asymmetry.
|
|||
|
|
|
|||
|
|
**1. Alignment first.** Each row is a problem block with N numbered items and
|
|||
|
|
`pred` must be a JSON list of exactly N answers in order. A single missing line
|
|||
|
|
shifts every later answer and zeroes the whole block on both exact-match and
|
|||
|
|
chrF. `detect_n_items` recovers N from the query — handling numbered lines,
|
|||
|
|
`(1)` blank markers, stated ranges, lettered items, unnumbered one-per-line
|
|||
|
|
lists, and the `match_letters` shape whose items live in the shared context.
|
|||
|
|
Measured on the 160 public Linguini problems it puts **98.4% of items in
|
|||
|
|
correctly-sized blocks**. Model output is then force-fitted to N, preferring
|
|||
|
|
the model's own numbering when it supplies it.
|
|||
|
|
|
|||
|
|
**2. Never emit an empty answer.** The final score is a geometric mean of exact
|
|||
|
|
match and chrF, so a blank scores zero on both and is strictly worse than a
|
|||
|
|
wrong guess. Every path ends in a non-empty string.
|
|||
|
|
|
|||
|
|
**2b. Answer style: a hypothesis that was tested and rejected.** Gold answers do
|
|||
|
|
follow the conventions of whatever language the answer is in (measured over the
|
|||
|
|
920 public Linguini answers, into-English golds that are full sentences are 99%
|
|||
|
|
capitalised, while the 157 that are bare clauses are only 10% capitalised, and
|
|||
|
|
the style matches the problem's own glosses in 36/36 measurable cases). Encoding
|
|||
|
|
that as prompt guidance nevertheless *lowered* exact match on the hidden set
|
|||
|
|
twice (0.0250 -> 0.0218 -> 0.0000). It is therefore not in the shipped script.
|
|||
|
|
The lesson recorded here for anyone rerunning this: a correct statistical
|
|||
|
|
description of the gold format did not translate into a better prompt.
|
|||
|
|
|
|||
|
|
**3. Monotone improvement under a hard deadline.** A complete, correctly-shaped
|
|||
|
|
`submission.csv` is written *before the model is loaded*, then overwritten after
|
|||
|
|
every improvement: greedy pass → each self-consistency pass → explanations.
|
|||
|
|
A crash or a timeout leaves the best result reached so far on disk rather than
|
|||
|
|
nothing. The script tracks its own remaining budget and stops adding passes
|
|||
|
|
when one more would not fit.
|
|||
|
|
|
|||
|
|
**4. Greedy-anchored voting.** After the greedy pass, sampled passes (T=0.5)
|
|||
|
|
run while budget remains, but the greedy answer is the default and sampled
|
|||
|
|
answers may only displace it when at least two of them agree on the same
|
|||
|
|
normalised form *and* that form outpolls the greedy one.
|
|||
|
|
|
|||
|
|
The asymmetry is empirical. A symmetric version — majority, else "most central
|
|||
|
|
by chrF" — was measurably worse than not voting at all: with only a handful of
|
|||
|
|
samples the centrality fallback is ill-defined (with two candidates pairwise
|
|||
|
|
chrF is symmetric, so it degenerated into preferring the shorter string) and it
|
|||
|
|
swapped the greedy answer for a sampled one about half the time. On the mock
|
|||
|
|
set that cost 4x exact match (EM 0.044 -> 0.011). Anchoring makes the procedure
|
|||
|
|
monotone: it can only fire on genuine agreement. chrF is implemented inline so
|
|||
|
|
the script carries no dependency the sandbox might lack.
|
|||
|
|
|
|||
|
|
**5. `match_letters` as an assignment problem.** Free-form generation answers
|
|||
|
|
this task type with the identity permutation (A, B, C, ...), which is a *valid*
|
|||
|
|
permutation, so duplicate-repair never fires and it scores ~0. `solve_matching`
|
|||
|
|
instead scores every (item, option) pair from the next-token distribution and
|
|||
|
|
takes the optimal one-to-one assignment, enforcing the bijection exactly.
|
|||
|
|
Duplicate-repair is retained only as a fallback for when that solver declines.
|
|||
|
|
|
|||
|
|
## Human Evaluation Challenge
|
|||
|
|
|
|||
|
|
`submission.csv` includes an `explanation` column: a short, human-readable
|
|||
|
|
statement of the rules behind each answer (not a raw reasoning trace),
|
|||
|
|
generated after the answers are fixed.
|
|||
|
|
|
|||
|
|
## Reproducing
|
|||
|
|
|
|||
|
|
```bash
|
|||
|
|
python script.py # reads /tmp/data/test.csv, writes submission.csv
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Environment knobs (all optional, defaults match the platform):
|
|||
|
|
`IOL_TEST_CSV`, `IOL_OUT_CSV`, `IOL_MODEL`, `IOL_TIME_LIMIT`, `IOL_BATCH`,
|
|||
|
|
`IOL_EXPLAIN`.
|
|||
|
|
|
|||
|
|
|
|||
|
|
## Revision history (measured on the hidden set, not guessed)
|
|||
|
|
|
|||
|
|
| submission | change | score | chrF | exact match |
|
|||
|
|
|---|---|---|---|---|
|
|||
|
|
| 1 | symmetric self-consistency vote | 0.0686 | 0.1882 | 0.0250 |
|
|||
|
|
| 2 | greedy-anchored voting (vote no longer fires) | **0.0712** | 0.2029 | 0.0250 |
|
|||
|
|
| 3 | + "capitalise English answers" style rule | 0.0679 | 0.2117 | 0.0218 |
|
|||
|
|
| 4 | + mirror-gloss-style, answer normalisation, equation hints | 0.0000 | 0.1527 | 0.0000 |
|
|||
|
|
| 5 | revert to 2, plus the match_letters assignment solver | 0.0712 | 0.2029 | 0.0250 |
|
|||
|
|
| 6 | + `repetition_penalty=1.0` (the model ships 1.05) | 0.0830 | 0.2067 | 0.0333 |
|
|||
|
|
| 8 (v8) | **baseline replication + `repetition_penalty=1.0`** | **0.2245** | 0.3150 | 0.1600 |
|
|||
|
|
| 10 (v10) | v8 but batch=4 (left-padded batching) | 0.1940 | 0.2818 | 0.1336 |
|
|||
|
|
| 11 (v11) | v8 + beam search `num_beams=4` on the answer pass | 0.1964 | 0.2888 | 0.1336 |
|
|||
|
|
| 12 (v12) | v8 + assignment solver for `match_letters` only | 0.1624 | 0.2526 | 0.1044 |
|
|||
|
|
| 13 (v13) | v8 + a one-shot worked exemplar as chat turns | 0.0920 | 0.2259 | 0.0375 |
|
|||
|
|
|
|||
|
|
Every layer of prompt/post-processing cleverness measurably *hurt*. Submission 5
|
|||
|
|
therefore reverts to the configuration of submission 2 and adds exactly one
|
|||
|
|
change, motivated by a specific measured failure:
|
|||
|
|
|
|||
|
|
**`match_letters` was being answered with the identity permutation.** Replaying
|
|||
|
|
seven parser variants over saved raw generations gave exact match 0.0000 for all
|
|||
|
|
seven, which exonerates the parser — the model simply was not solving the task,
|
|||
|
|
emitting the option labels in order (A, B, C, ...). Because the identity is a
|
|||
|
|
valid permutation, `repair_bijection` never fired. `solve_matching` replaces
|
|||
|
|
free-form generation for this task type: it scores every (item, option) pair
|
|||
|
|
from the next-token distribution and takes the optimal one-to-one assignment,
|
|||
|
|
so the bijection constraint is enforced exactly rather than hoped for.
|
|||
|
|
|
|||
|
|
|
|||
|
|
## The silent decoding bug
|
|||
|
|
|
|||
|
|
`Qwen/Qwen2.5-14B-Instruct-AWQ` ships `generation_config.json` containing
|
|||
|
|
`repetition_penalty: 1.05`. Greedy decoding ignores `temperature`, `top_p` and
|
|||
|
|
`top_k` — and transformers emits a warning for each of those — but a repetition
|
|||
|
|
penalty **is** applied under greedy decoding, with no warning at all.
|
|||
|
|
|
|||
|
|
That matters here specifically: 34% of the 920 public gold answers repeat some
|
|||
|
|
letter three or more times, because these languages are agglutinative and the
|
|||
|
|
answers look like `ɨmpʼuhurʼu` and `ɨŋɡɨrʼɨ`. A 5% penalty on repeated tokens
|
|||
|
|
biases the model away from exactly the strings the task requires. The script now
|
|||
|
|
passes `repetition_penalty=1.0` explicitly.
|
|||
|
|
|
|||
|
|
NFC normalisation of answers was considered and rejected: 98.15% of public golds
|
|||
|
|
are already NFC, but 13 of them are NFD-and-not-NFC, so forcing NFC would break
|
|||
|
|
those for an unmeasured gain.
|
|||
|
|
|
|||
|
|
|
|||
|
|
## v8 — faithful baseline replication
|
|||
|
|
|
|||
|
|
The organizers' reference script reaches exact match **0.0729** on the hidden set
|
|||
|
|
with these exact weights. Our best is 0.0333. Before adding anything further we
|
|||
|
|
need to know whether that number is reproducible by us at all, so v8 replicates
|
|||
|
|
their script literally — trivial system prompt, no chain-of-thought, **batch 1 (no padding at all)**,
|
|||
|
|
naive line split, and **no forcing to N answers** — changing exactly one thing:
|
|||
|
|
`repetition_penalty=1.0`. Generation is EOS-limited rather than cap-limited:
|
|||
|
|
without chain-of-thought the model emits a few short answer lines and stops.
|
|||
|
|
|
|||
|
|
**Result: 0.2245 (chrF 0.3150, exact match 0.1600) — first place of 43 teams.**
|
|||
|
|
|
|||
|
|
That is 2.7x our best engineered pipeline (0.0830) and +83% on the organizers'
|
|||
|
|
own baseline (0.1227), the entire delta over their number being
|
|||
|
|
`repetition_penalty=1.0`.
|
|||
|
|
|
|||
|
|
The lesson is uncomfortable and worth recording plainly: every layer we added on
|
|||
|
|
top of the reference structure — chain-of-thought, an `ANSWERS:` block, answer
|
|||
|
|
style rules, output normalisation, forcing exactly N answers — reduced exact
|
|||
|
|
match. Submissions 2 -> 3 -> 4 fell 0.0712 -> 0.0679 -> 0.0000 as more
|
|||
|
|
engineering went in. The winning move was deleting all of it and fixing one
|
|||
|
|
decoding flag.
|
|||
|
|
|
|||
|
|
## Final-day probes (all negative, all single changes on v8)
|
|||
|
|
|
|||
|
|
Three orthogonal, individually-motivated improvements were each tested as a
|
|||
|
|
minimal diff on the frozen v8 script, one variable at a time:
|
|||
|
|
|
|||
|
|
* **v11 — beam search** (`num_beams=4`, batch 1, greedy fallback on OOM/low
|
|||
|
|
budget): 0.2245 → 0.1964. Exact match fell to 0.1336 — the same value as
|
|||
|
|
left-padded batch=4 — suggesting a shared fp16-numerics mechanism in
|
|||
|
|
multi-sequence forward passes rather than anything about search.
|
|||
|
|
* **v12 — assignment solver for `match_letters`**: 0.2245 → 0.1624, and
|
|||
|
|
explanation coverage fell to 50% from the solver's extra forward passes.
|
|||
|
|
The hidden set evidently does not reward bare option letters where
|
|||
|
|
free-form text had been earning chrF credit.
|
|||
|
|
* **v13 — one-shot worked exemplar** (a public-Linguini Kayapo problem as a
|
|||
|
|
genuine user→assistant exchange): 0.2245 → 0.0920. The demonstration
|
|||
|
|
derailed the model far more than any instruction-style prompt addition.
|
|||
|
|
|
|||
|
|
With those, every direction adjacent to v8 has been measured: chain-of-thought,
|
|||
|
|
answer-style rules, output normalisation, N-forcing, batching, beam search,
|
|||
|
|
constrained decoding, few-shot. All reduced the score. The shipped
|
|||
|
|
configuration — the organizers' minimal structure plus `repetition_penalty=1.0`
|
|||
|
|
— is a sharp local optimum, and `script.py` on `main` is exactly that config.
|