189 lines
9.8 KiB
Markdown
189 lines
9.8 KiB
Markdown
---
|
||
license: apache-2.0
|
||
base_model: Qwen/Qwen2.5-14B-Instruct-AWQ
|
||
tags:
|
||
- iol-ai-2026
|
||
- linguistics
|
||
- reasoning
|
||
language:
|
||
- en
|
||
---
|
||
|
||
# IOL-AI 2026 — Qwen2.5-14B-Instruct-AWQ
|
||
|
||
Submission for the [IOL-AI 2026 Linguistics Olympiad Challenge](https://iolai.org).
|
||
|
||
The weights are an unmodified copy of
|
||
[`Qwen/Qwen2.5-14B-Instruct-AWQ`](https://huggingface.co/Qwen/Qwen2.5-14B-Instruct-AWQ)
|
||
(Apache-2.0, redistributable), shipped in-repo because the evaluation sandbox
|
||
has no internet access. **All of the work is in `script.py`.**
|
||
|
||
## Approach
|
||
|
||
The eval budget is 30 minutes on a 16 GB T4 for a test set of only ~90
|
||
sub-questions, so compute per problem is abundant while *reliability* is
|
||
scarce. The script is built around that asymmetry.
|
||
|
||
**1. Alignment first.** Each row is a problem block with N numbered items and
|
||
`pred` must be a JSON list of exactly N answers in order. A single missing line
|
||
shifts every later answer and zeroes the whole block on both exact-match and
|
||
chrF. `detect_n_items` recovers N from the query — handling numbered lines,
|
||
`(1)` blank markers, stated ranges, lettered items, unnumbered one-per-line
|
||
lists, and the `match_letters` shape whose items live in the shared context.
|
||
Measured on the 160 public Linguini problems it puts **98.4% of items in
|
||
correctly-sized blocks**. Model output is then force-fitted to N, preferring
|
||
the model's own numbering when it supplies it.
|
||
|
||
**2. Never emit an empty answer.** The final score is a geometric mean of exact
|
||
match and chrF, so a blank scores zero on both and is strictly worse than a
|
||
wrong guess. Every path ends in a non-empty string.
|
||
|
||
**2b. Answer style: a hypothesis that was tested and rejected.** Gold answers do
|
||
follow the conventions of whatever language the answer is in (measured over the
|
||
920 public Linguini answers, into-English golds that are full sentences are 99%
|
||
capitalised, while the 157 that are bare clauses are only 10% capitalised, and
|
||
the style matches the problem's own glosses in 36/36 measurable cases). Encoding
|
||
that as prompt guidance nevertheless *lowered* exact match on the hidden set
|
||
twice (0.0250 -> 0.0218 -> 0.0000). It is therefore not in the shipped script.
|
||
The lesson recorded here for anyone rerunning this: a correct statistical
|
||
description of the gold format did not translate into a better prompt.
|
||
|
||
**3. Monotone improvement under a hard deadline.** A complete, correctly-shaped
|
||
`submission.csv` is written *before the model is loaded*, then overwritten after
|
||
every improvement: greedy pass → each self-consistency pass → explanations.
|
||
A crash or a timeout leaves the best result reached so far on disk rather than
|
||
nothing. The script tracks its own remaining budget and stops adding passes
|
||
when one more would not fit.
|
||
|
||
**4. Greedy-anchored voting.** After the greedy pass, sampled passes (T=0.5)
|
||
run while budget remains, but the greedy answer is the default and sampled
|
||
answers may only displace it when at least two of them agree on the same
|
||
normalised form *and* that form outpolls the greedy one.
|
||
|
||
The asymmetry is empirical. A symmetric version — majority, else "most central
|
||
by chrF" — was measurably worse than not voting at all: with only a handful of
|
||
samples the centrality fallback is ill-defined (with two candidates pairwise
|
||
chrF is symmetric, so it degenerated into preferring the shorter string) and it
|
||
swapped the greedy answer for a sampled one about half the time. On the mock
|
||
set that cost 4x exact match (EM 0.044 -> 0.011). Anchoring makes the procedure
|
||
monotone: it can only fire on genuine agreement. chrF is implemented inline so
|
||
the script carries no dependency the sandbox might lack.
|
||
|
||
**5. `match_letters` as an assignment problem.** Free-form generation answers
|
||
this task type with the identity permutation (A, B, C, ...), which is a *valid*
|
||
permutation, so duplicate-repair never fires and it scores ~0. `solve_matching`
|
||
instead scores every (item, option) pair from the next-token distribution and
|
||
takes the optimal one-to-one assignment, enforcing the bijection exactly.
|
||
Duplicate-repair is retained only as a fallback for when that solver declines.
|
||
|
||
## Human Evaluation Challenge
|
||
|
||
`submission.csv` includes an `explanation` column: a short, human-readable
|
||
statement of the rules behind each answer (not a raw reasoning trace),
|
||
generated after the answers are fixed.
|
||
|
||
## Reproducing
|
||
|
||
```bash
|
||
python script.py # reads /tmp/data/test.csv, writes submission.csv
|
||
```
|
||
|
||
Environment knobs (all optional, defaults match the platform):
|
||
`IOL_TEST_CSV`, `IOL_OUT_CSV`, `IOL_MODEL`, `IOL_TIME_LIMIT`, `IOL_BATCH`,
|
||
`IOL_EXPLAIN`.
|
||
|
||
|
||
## Revision history (measured on the hidden set, not guessed)
|
||
|
||
| submission | change | score | chrF | exact match |
|
||
|---|---|---|---|---|
|
||
| 1 | symmetric self-consistency vote | 0.0686 | 0.1882 | 0.0250 |
|
||
| 2 | greedy-anchored voting (vote no longer fires) | **0.0712** | 0.2029 | 0.0250 |
|
||
| 3 | + "capitalise English answers" style rule | 0.0679 | 0.2117 | 0.0218 |
|
||
| 4 | + mirror-gloss-style, answer normalisation, equation hints | 0.0000 | 0.1527 | 0.0000 |
|
||
| 5 | revert to 2, plus the match_letters assignment solver | 0.0712 | 0.2029 | 0.0250 |
|
||
| 6 | + `repetition_penalty=1.0` (the model ships 1.05) | 0.0830 | 0.2067 | 0.0333 |
|
||
| 8 (v8) | **baseline replication + `repetition_penalty=1.0`** | **0.2245** | 0.3150 | 0.1600 |
|
||
| 10 (v10) | v8 but batch=4 (left-padded batching) | 0.1940 | 0.2818 | 0.1336 |
|
||
| 11 (v11) | v8 + beam search `num_beams=4` on the answer pass | 0.1964 | 0.2888 | 0.1336 |
|
||
| 12 (v12) | v8 + assignment solver for `match_letters` only | 0.1624 | 0.2526 | 0.1044 |
|
||
| 13 (v13) | v8 + a one-shot worked exemplar as chat turns | 0.0920 | 0.2259 | 0.0375 |
|
||
|
||
Every layer of prompt/post-processing cleverness measurably *hurt*. Submission 5
|
||
therefore reverts to the configuration of submission 2 and adds exactly one
|
||
change, motivated by a specific measured failure:
|
||
|
||
**`match_letters` was being answered with the identity permutation.** Replaying
|
||
seven parser variants over saved raw generations gave exact match 0.0000 for all
|
||
seven, which exonerates the parser — the model simply was not solving the task,
|
||
emitting the option labels in order (A, B, C, ...). Because the identity is a
|
||
valid permutation, `repair_bijection` never fired. `solve_matching` replaces
|
||
free-form generation for this task type: it scores every (item, option) pair
|
||
from the next-token distribution and takes the optimal one-to-one assignment,
|
||
so the bijection constraint is enforced exactly rather than hoped for.
|
||
|
||
|
||
## The silent decoding bug
|
||
|
||
`Qwen/Qwen2.5-14B-Instruct-AWQ` ships `generation_config.json` containing
|
||
`repetition_penalty: 1.05`. Greedy decoding ignores `temperature`, `top_p` and
|
||
`top_k` — and transformers emits a warning for each of those — but a repetition
|
||
penalty **is** applied under greedy decoding, with no warning at all.
|
||
|
||
That matters here specifically: 34% of the 920 public gold answers repeat some
|
||
letter three or more times, because these languages are agglutinative and the
|
||
answers look like `ɨmpʼuhurʼu` and `ɨŋɡɨrʼɨ`. A 5% penalty on repeated tokens
|
||
biases the model away from exactly the strings the task requires. The script now
|
||
passes `repetition_penalty=1.0` explicitly.
|
||
|
||
NFC normalisation of answers was considered and rejected: 98.15% of public golds
|
||
are already NFC, but 13 of them are NFD-and-not-NFC, so forcing NFC would break
|
||
those for an unmeasured gain.
|
||
|
||
|
||
## v8 — faithful baseline replication
|
||
|
||
The organizers' reference script reaches exact match **0.0729** on the hidden set
|
||
with these exact weights. Our best is 0.0333. Before adding anything further we
|
||
need to know whether that number is reproducible by us at all, so v8 replicates
|
||
their script literally — trivial system prompt, no chain-of-thought, **batch 1 (no padding at all)**,
|
||
naive line split, and **no forcing to N answers** — changing exactly one thing:
|
||
`repetition_penalty=1.0`. Generation is EOS-limited rather than cap-limited:
|
||
without chain-of-thought the model emits a few short answer lines and stops.
|
||
|
||
**Result: 0.2245 (chrF 0.3150, exact match 0.1600) — first place of 43 teams.**
|
||
|
||
That is 2.7x our best engineered pipeline (0.0830) and +83% on the organizers'
|
||
own baseline (0.1227), the entire delta over their number being
|
||
`repetition_penalty=1.0`.
|
||
|
||
The lesson is uncomfortable and worth recording plainly: every layer we added on
|
||
top of the reference structure — chain-of-thought, an `ANSWERS:` block, answer
|
||
style rules, output normalisation, forcing exactly N answers — reduced exact
|
||
match. Submissions 2 -> 3 -> 4 fell 0.0712 -> 0.0679 -> 0.0000 as more
|
||
engineering went in. The winning move was deleting all of it and fixing one
|
||
decoding flag.
|
||
|
||
## Final-day probes (all negative, all single changes on v8)
|
||
|
||
Three orthogonal, individually-motivated improvements were each tested as a
|
||
minimal diff on the frozen v8 script, one variable at a time:
|
||
|
||
* **v11 — beam search** (`num_beams=4`, batch 1, greedy fallback on OOM/low
|
||
budget): 0.2245 → 0.1964. Exact match fell to 0.1336 — the same value as
|
||
left-padded batch=4 — suggesting a shared fp16-numerics mechanism in
|
||
multi-sequence forward passes rather than anything about search.
|
||
* **v12 — assignment solver for `match_letters`**: 0.2245 → 0.1624, and
|
||
explanation coverage fell to 50% from the solver's extra forward passes.
|
||
The hidden set evidently does not reward bare option letters where
|
||
free-form text had been earning chrF credit.
|
||
* **v13 — one-shot worked exemplar** (a public-Linguini Kayapo problem as a
|
||
genuine user→assistant exchange): 0.2245 → 0.0920. The demonstration
|
||
derailed the model far more than any instruction-style prompt addition.
|
||
|
||
With those, every direction adjacent to v8 has been measured: chain-of-thought,
|
||
answer-style rules, output normalisation, N-forcing, batching, beam search,
|
||
constrained decoding, few-shot. All reduced the score. The shipped
|
||
configuration — the organizers' minimal structure plus `repetition_penalty=1.0`
|
||
— is a sharp local optimum, and `script.py` on `main` is exactly that config.
|