Files
iolai-2026-qwen25-14b-awq/README.md
ModelHub XC 3532add6de 初始化项目,由ModelHub XC社区提供模型
Model: arvindcr4/iolai-2026-qwen25-14b-awq
Source: Original Platform
2026-09-17 18:40:18 +08:00

189 lines
9.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
license: apache-2.0
base_model: Qwen/Qwen2.5-14B-Instruct-AWQ
tags:
- iol-ai-2026
- linguistics
- reasoning
language:
- en
---
# IOL-AI 2026 — Qwen2.5-14B-Instruct-AWQ
Submission for the [IOL-AI 2026 Linguistics Olympiad Challenge](https://iolai.org).
The weights are an unmodified copy of
[`Qwen/Qwen2.5-14B-Instruct-AWQ`](https://huggingface.co/Qwen/Qwen2.5-14B-Instruct-AWQ)
(Apache-2.0, redistributable), shipped in-repo because the evaluation sandbox
has no internet access. **All of the work is in `script.py`.**
## Approach
The eval budget is 30 minutes on a 16 GB T4 for a test set of only ~90
sub-questions, so compute per problem is abundant while *reliability* is
scarce. The script is built around that asymmetry.
**1. Alignment first.** Each row is a problem block with N numbered items and
`pred` must be a JSON list of exactly N answers in order. A single missing line
shifts every later answer and zeroes the whole block on both exact-match and
chrF. `detect_n_items` recovers N from the query — handling numbered lines,
`(1)` blank markers, stated ranges, lettered items, unnumbered one-per-line
lists, and the `match_letters` shape whose items live in the shared context.
Measured on the 160 public Linguini problems it puts **98.4% of items in
correctly-sized blocks**. Model output is then force-fitted to N, preferring
the model's own numbering when it supplies it.
**2. Never emit an empty answer.** The final score is a geometric mean of exact
match and chrF, so a blank scores zero on both and is strictly worse than a
wrong guess. Every path ends in a non-empty string.
**2b. Answer style: a hypothesis that was tested and rejected.** Gold answers do
follow the conventions of whatever language the answer is in (measured over the
920 public Linguini answers, into-English golds that are full sentences are 99%
capitalised, while the 157 that are bare clauses are only 10% capitalised, and
the style matches the problem's own glosses in 36/36 measurable cases). Encoding
that as prompt guidance nevertheless *lowered* exact match on the hidden set
twice (0.0250 -> 0.0218 -> 0.0000). It is therefore not in the shipped script.
The lesson recorded here for anyone rerunning this: a correct statistical
description of the gold format did not translate into a better prompt.
**3. Monotone improvement under a hard deadline.** A complete, correctly-shaped
`submission.csv` is written *before the model is loaded*, then overwritten after
every improvement: greedy pass → each self-consistency pass → explanations.
A crash or a timeout leaves the best result reached so far on disk rather than
nothing. The script tracks its own remaining budget and stops adding passes
when one more would not fit.
**4. Greedy-anchored voting.** After the greedy pass, sampled passes (T=0.5)
run while budget remains, but the greedy answer is the default and sampled
answers may only displace it when at least two of them agree on the same
normalised form *and* that form outpolls the greedy one.
The asymmetry is empirical. A symmetric version — majority, else "most central
by chrF" — was measurably worse than not voting at all: with only a handful of
samples the centrality fallback is ill-defined (with two candidates pairwise
chrF is symmetric, so it degenerated into preferring the shorter string) and it
swapped the greedy answer for a sampled one about half the time. On the mock
set that cost 4x exact match (EM 0.044 -> 0.011). Anchoring makes the procedure
monotone: it can only fire on genuine agreement. chrF is implemented inline so
the script carries no dependency the sandbox might lack.
**5. `match_letters` as an assignment problem.** Free-form generation answers
this task type with the identity permutation (A, B, C, ...), which is a *valid*
permutation, so duplicate-repair never fires and it scores ~0. `solve_matching`
instead scores every (item, option) pair from the next-token distribution and
takes the optimal one-to-one assignment, enforcing the bijection exactly.
Duplicate-repair is retained only as a fallback for when that solver declines.
## Human Evaluation Challenge
`submission.csv` includes an `explanation` column: a short, human-readable
statement of the rules behind each answer (not a raw reasoning trace),
generated after the answers are fixed.
## Reproducing
```bash
python script.py # reads /tmp/data/test.csv, writes submission.csv
```
Environment knobs (all optional, defaults match the platform):
`IOL_TEST_CSV`, `IOL_OUT_CSV`, `IOL_MODEL`, `IOL_TIME_LIMIT`, `IOL_BATCH`,
`IOL_EXPLAIN`.
## Revision history (measured on the hidden set, not guessed)
| submission | change | score | chrF | exact match |
|---|---|---|---|---|
| 1 | symmetric self-consistency vote | 0.0686 | 0.1882 | 0.0250 |
| 2 | greedy-anchored voting (vote no longer fires) | **0.0712** | 0.2029 | 0.0250 |
| 3 | + "capitalise English answers" style rule | 0.0679 | 0.2117 | 0.0218 |
| 4 | + mirror-gloss-style, answer normalisation, equation hints | 0.0000 | 0.1527 | 0.0000 |
| 5 | revert to 2, plus the match_letters assignment solver | 0.0712 | 0.2029 | 0.0250 |
| 6 | + `repetition_penalty=1.0` (the model ships 1.05) | 0.0830 | 0.2067 | 0.0333 |
| 8 (v8) | **baseline replication + `repetition_penalty=1.0`** | **0.2245** | 0.3150 | 0.1600 |
| 10 (v10) | v8 but batch=4 (left-padded batching) | 0.1940 | 0.2818 | 0.1336 |
| 11 (v11) | v8 + beam search `num_beams=4` on the answer pass | 0.1964 | 0.2888 | 0.1336 |
| 12 (v12) | v8 + assignment solver for `match_letters` only | 0.1624 | 0.2526 | 0.1044 |
| 13 (v13) | v8 + a one-shot worked exemplar as chat turns | 0.0920 | 0.2259 | 0.0375 |
Every layer of prompt/post-processing cleverness measurably *hurt*. Submission 5
therefore reverts to the configuration of submission 2 and adds exactly one
change, motivated by a specific measured failure:
**`match_letters` was being answered with the identity permutation.** Replaying
seven parser variants over saved raw generations gave exact match 0.0000 for all
seven, which exonerates the parser — the model simply was not solving the task,
emitting the option labels in order (A, B, C, ...). Because the identity is a
valid permutation, `repair_bijection` never fired. `solve_matching` replaces
free-form generation for this task type: it scores every (item, option) pair
from the next-token distribution and takes the optimal one-to-one assignment,
so the bijection constraint is enforced exactly rather than hoped for.
## The silent decoding bug
`Qwen/Qwen2.5-14B-Instruct-AWQ` ships `generation_config.json` containing
`repetition_penalty: 1.05`. Greedy decoding ignores `temperature`, `top_p` and
`top_k` — and transformers emits a warning for each of those — but a repetition
penalty **is** applied under greedy decoding, with no warning at all.
That matters here specifically: 34% of the 920 public gold answers repeat some
letter three or more times, because these languages are agglutinative and the
answers look like `ɨmpʼuhurʼu` and `ɨŋɡɨrʼɨ`. A 5% penalty on repeated tokens
biases the model away from exactly the strings the task requires. The script now
passes `repetition_penalty=1.0` explicitly.
NFC normalisation of answers was considered and rejected: 98.15% of public golds
are already NFC, but 13 of them are NFD-and-not-NFC, so forcing NFC would break
those for an unmeasured gain.
## v8 — faithful baseline replication
The organizers' reference script reaches exact match **0.0729** on the hidden set
with these exact weights. Our best is 0.0333. Before adding anything further we
need to know whether that number is reproducible by us at all, so v8 replicates
their script literally — trivial system prompt, no chain-of-thought, **batch 1 (no padding at all)**,
naive line split, and **no forcing to N answers** — changing exactly one thing:
`repetition_penalty=1.0`. Generation is EOS-limited rather than cap-limited:
without chain-of-thought the model emits a few short answer lines and stops.
**Result: 0.2245 (chrF 0.3150, exact match 0.1600) — first place of 43 teams.**
That is 2.7x our best engineered pipeline (0.0830) and +83% on the organizers'
own baseline (0.1227), the entire delta over their number being
`repetition_penalty=1.0`.
The lesson is uncomfortable and worth recording plainly: every layer we added on
top of the reference structure — chain-of-thought, an `ANSWERS:` block, answer
style rules, output normalisation, forcing exactly N answers — reduced exact
match. Submissions 2 -> 3 -> 4 fell 0.0712 -> 0.0679 -> 0.0000 as more
engineering went in. The winning move was deleting all of it and fixing one
decoding flag.
## Final-day probes (all negative, all single changes on v8)
Three orthogonal, individually-motivated improvements were each tested as a
minimal diff on the frozen v8 script, one variable at a time:
* **v11 — beam search** (`num_beams=4`, batch 1, greedy fallback on OOM/low
budget): 0.2245 → 0.1964. Exact match fell to 0.1336 — the same value as
left-padded batch=4 — suggesting a shared fp16-numerics mechanism in
multi-sequence forward passes rather than anything about search.
* **v12 — assignment solver for `match_letters`**: 0.2245 → 0.1624, and
explanation coverage fell to 50% from the solver's extra forward passes.
The hidden set evidently does not reward bare option letters where
free-form text had been earning chrF credit.
* **v13 — one-shot worked exemplar** (a public-Linguini Kayapo problem as a
genuine user→assistant exchange): 0.2245 → 0.0920. The demonstration
derailed the model far more than any instruction-style prompt addition.
With those, every direction adjacent to v8 has been measured: chain-of-thought,
answer-style rules, output normalisation, N-forcing, batching, beam search,
constrained decoding, few-shot. All reduced the score. The shipped
configuration — the organizers' minimal structure plus `repetition_penalty=1.0`
— is a sharp local optimum, and `script.py` on `main` is exactly that config.