初始化项目,由ModelHub XC社区提供模型
Model: arvindcr4/iolai-2026-qwen25-14b-awq Source: Original Platform
This commit is contained in:
188
README.md
Normal file
188
README.md
Normal file
@@ -0,0 +1,188 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
base_model: Qwen/Qwen2.5-14B-Instruct-AWQ
|
||||
tags:
|
||||
- iol-ai-2026
|
||||
- linguistics
|
||||
- reasoning
|
||||
language:
|
||||
- en
|
||||
---
|
||||
|
||||
# IOL-AI 2026 — Qwen2.5-14B-Instruct-AWQ
|
||||
|
||||
Submission for the [IOL-AI 2026 Linguistics Olympiad Challenge](https://iolai.org).
|
||||
|
||||
The weights are an unmodified copy of
|
||||
[`Qwen/Qwen2.5-14B-Instruct-AWQ`](https://huggingface.co/Qwen/Qwen2.5-14B-Instruct-AWQ)
|
||||
(Apache-2.0, redistributable), shipped in-repo because the evaluation sandbox
|
||||
has no internet access. **All of the work is in `script.py`.**
|
||||
|
||||
## Approach
|
||||
|
||||
The eval budget is 30 minutes on a 16 GB T4 for a test set of only ~90
|
||||
sub-questions, so compute per problem is abundant while *reliability* is
|
||||
scarce. The script is built around that asymmetry.
|
||||
|
||||
**1. Alignment first.** Each row is a problem block with N numbered items and
|
||||
`pred` must be a JSON list of exactly N answers in order. A single missing line
|
||||
shifts every later answer and zeroes the whole block on both exact-match and
|
||||
chrF. `detect_n_items` recovers N from the query — handling numbered lines,
|
||||
`(1)` blank markers, stated ranges, lettered items, unnumbered one-per-line
|
||||
lists, and the `match_letters` shape whose items live in the shared context.
|
||||
Measured on the 160 public Linguini problems it puts **98.4% of items in
|
||||
correctly-sized blocks**. Model output is then force-fitted to N, preferring
|
||||
the model's own numbering when it supplies it.
|
||||
|
||||
**2. Never emit an empty answer.** The final score is a geometric mean of exact
|
||||
match and chrF, so a blank scores zero on both and is strictly worse than a
|
||||
wrong guess. Every path ends in a non-empty string.
|
||||
|
||||
**2b. Answer style: a hypothesis that was tested and rejected.** Gold answers do
|
||||
follow the conventions of whatever language the answer is in (measured over the
|
||||
920 public Linguini answers, into-English golds that are full sentences are 99%
|
||||
capitalised, while the 157 that are bare clauses are only 10% capitalised, and
|
||||
the style matches the problem's own glosses in 36/36 measurable cases). Encoding
|
||||
that as prompt guidance nevertheless *lowered* exact match on the hidden set
|
||||
twice (0.0250 -> 0.0218 -> 0.0000). It is therefore not in the shipped script.
|
||||
The lesson recorded here for anyone rerunning this: a correct statistical
|
||||
description of the gold format did not translate into a better prompt.
|
||||
|
||||
**3. Monotone improvement under a hard deadline.** A complete, correctly-shaped
|
||||
`submission.csv` is written *before the model is loaded*, then overwritten after
|
||||
every improvement: greedy pass → each self-consistency pass → explanations.
|
||||
A crash or a timeout leaves the best result reached so far on disk rather than
|
||||
nothing. The script tracks its own remaining budget and stops adding passes
|
||||
when one more would not fit.
|
||||
|
||||
**4. Greedy-anchored voting.** After the greedy pass, sampled passes (T=0.5)
|
||||
run while budget remains, but the greedy answer is the default and sampled
|
||||
answers may only displace it when at least two of them agree on the same
|
||||
normalised form *and* that form outpolls the greedy one.
|
||||
|
||||
The asymmetry is empirical. A symmetric version — majority, else "most central
|
||||
by chrF" — was measurably worse than not voting at all: with only a handful of
|
||||
samples the centrality fallback is ill-defined (with two candidates pairwise
|
||||
chrF is symmetric, so it degenerated into preferring the shorter string) and it
|
||||
swapped the greedy answer for a sampled one about half the time. On the mock
|
||||
set that cost 4x exact match (EM 0.044 -> 0.011). Anchoring makes the procedure
|
||||
monotone: it can only fire on genuine agreement. chrF is implemented inline so
|
||||
the script carries no dependency the sandbox might lack.
|
||||
|
||||
**5. `match_letters` as an assignment problem.** Free-form generation answers
|
||||
this task type with the identity permutation (A, B, C, ...), which is a *valid*
|
||||
permutation, so duplicate-repair never fires and it scores ~0. `solve_matching`
|
||||
instead scores every (item, option) pair from the next-token distribution and
|
||||
takes the optimal one-to-one assignment, enforcing the bijection exactly.
|
||||
Duplicate-repair is retained only as a fallback for when that solver declines.
|
||||
|
||||
## Human Evaluation Challenge
|
||||
|
||||
`submission.csv` includes an `explanation` column: a short, human-readable
|
||||
statement of the rules behind each answer (not a raw reasoning trace),
|
||||
generated after the answers are fixed.
|
||||
|
||||
## Reproducing
|
||||
|
||||
```bash
|
||||
python script.py # reads /tmp/data/test.csv, writes submission.csv
|
||||
```
|
||||
|
||||
Environment knobs (all optional, defaults match the platform):
|
||||
`IOL_TEST_CSV`, `IOL_OUT_CSV`, `IOL_MODEL`, `IOL_TIME_LIMIT`, `IOL_BATCH`,
|
||||
`IOL_EXPLAIN`.
|
||||
|
||||
|
||||
## Revision history (measured on the hidden set, not guessed)
|
||||
|
||||
| submission | change | score | chrF | exact match |
|
||||
|---|---|---|---|---|
|
||||
| 1 | symmetric self-consistency vote | 0.0686 | 0.1882 | 0.0250 |
|
||||
| 2 | greedy-anchored voting (vote no longer fires) | **0.0712** | 0.2029 | 0.0250 |
|
||||
| 3 | + "capitalise English answers" style rule | 0.0679 | 0.2117 | 0.0218 |
|
||||
| 4 | + mirror-gloss-style, answer normalisation, equation hints | 0.0000 | 0.1527 | 0.0000 |
|
||||
| 5 | revert to 2, plus the match_letters assignment solver | 0.0712 | 0.2029 | 0.0250 |
|
||||
| 6 | + `repetition_penalty=1.0` (the model ships 1.05) | 0.0830 | 0.2067 | 0.0333 |
|
||||
| 8 (v8) | **baseline replication + `repetition_penalty=1.0`** | **0.2245** | 0.3150 | 0.1600 |
|
||||
| 10 (v10) | v8 but batch=4 (left-padded batching) | 0.1940 | 0.2818 | 0.1336 |
|
||||
| 11 (v11) | v8 + beam search `num_beams=4` on the answer pass | 0.1964 | 0.2888 | 0.1336 |
|
||||
| 12 (v12) | v8 + assignment solver for `match_letters` only | 0.1624 | 0.2526 | 0.1044 |
|
||||
| 13 (v13) | v8 + a one-shot worked exemplar as chat turns | 0.0920 | 0.2259 | 0.0375 |
|
||||
|
||||
Every layer of prompt/post-processing cleverness measurably *hurt*. Submission 5
|
||||
therefore reverts to the configuration of submission 2 and adds exactly one
|
||||
change, motivated by a specific measured failure:
|
||||
|
||||
**`match_letters` was being answered with the identity permutation.** Replaying
|
||||
seven parser variants over saved raw generations gave exact match 0.0000 for all
|
||||
seven, which exonerates the parser — the model simply was not solving the task,
|
||||
emitting the option labels in order (A, B, C, ...). Because the identity is a
|
||||
valid permutation, `repair_bijection` never fired. `solve_matching` replaces
|
||||
free-form generation for this task type: it scores every (item, option) pair
|
||||
from the next-token distribution and takes the optimal one-to-one assignment,
|
||||
so the bijection constraint is enforced exactly rather than hoped for.
|
||||
|
||||
|
||||
## The silent decoding bug
|
||||
|
||||
`Qwen/Qwen2.5-14B-Instruct-AWQ` ships `generation_config.json` containing
|
||||
`repetition_penalty: 1.05`. Greedy decoding ignores `temperature`, `top_p` and
|
||||
`top_k` — and transformers emits a warning for each of those — but a repetition
|
||||
penalty **is** applied under greedy decoding, with no warning at all.
|
||||
|
||||
That matters here specifically: 34% of the 920 public gold answers repeat some
|
||||
letter three or more times, because these languages are agglutinative and the
|
||||
answers look like `ɨmpʼuhurʼu` and `ɨŋɡɨrʼɨ`. A 5% penalty on repeated tokens
|
||||
biases the model away from exactly the strings the task requires. The script now
|
||||
passes `repetition_penalty=1.0` explicitly.
|
||||
|
||||
NFC normalisation of answers was considered and rejected: 98.15% of public golds
|
||||
are already NFC, but 13 of them are NFD-and-not-NFC, so forcing NFC would break
|
||||
those for an unmeasured gain.
|
||||
|
||||
|
||||
## v8 — faithful baseline replication
|
||||
|
||||
The organizers' reference script reaches exact match **0.0729** on the hidden set
|
||||
with these exact weights. Our best is 0.0333. Before adding anything further we
|
||||
need to know whether that number is reproducible by us at all, so v8 replicates
|
||||
their script literally — trivial system prompt, no chain-of-thought, **batch 1 (no padding at all)**,
|
||||
naive line split, and **no forcing to N answers** — changing exactly one thing:
|
||||
`repetition_penalty=1.0`. Generation is EOS-limited rather than cap-limited:
|
||||
without chain-of-thought the model emits a few short answer lines and stops.
|
||||
|
||||
**Result: 0.2245 (chrF 0.3150, exact match 0.1600) — first place of 43 teams.**
|
||||
|
||||
That is 2.7x our best engineered pipeline (0.0830) and +83% on the organizers'
|
||||
own baseline (0.1227), the entire delta over their number being
|
||||
`repetition_penalty=1.0`.
|
||||
|
||||
The lesson is uncomfortable and worth recording plainly: every layer we added on
|
||||
top of the reference structure — chain-of-thought, an `ANSWERS:` block, answer
|
||||
style rules, output normalisation, forcing exactly N answers — reduced exact
|
||||
match. Submissions 2 -> 3 -> 4 fell 0.0712 -> 0.0679 -> 0.0000 as more
|
||||
engineering went in. The winning move was deleting all of it and fixing one
|
||||
decoding flag.
|
||||
|
||||
## Final-day probes (all negative, all single changes on v8)
|
||||
|
||||
Three orthogonal, individually-motivated improvements were each tested as a
|
||||
minimal diff on the frozen v8 script, one variable at a time:
|
||||
|
||||
* **v11 — beam search** (`num_beams=4`, batch 1, greedy fallback on OOM/low
|
||||
budget): 0.2245 → 0.1964. Exact match fell to 0.1336 — the same value as
|
||||
left-padded batch=4 — suggesting a shared fp16-numerics mechanism in
|
||||
multi-sequence forward passes rather than anything about search.
|
||||
* **v12 — assignment solver for `match_letters`**: 0.2245 → 0.1624, and
|
||||
explanation coverage fell to 50% from the solver's extra forward passes.
|
||||
The hidden set evidently does not reward bare option letters where
|
||||
free-form text had been earning chrF credit.
|
||||
* **v13 — one-shot worked exemplar** (a public-Linguini Kayapo problem as a
|
||||
genuine user→assistant exchange): 0.2245 → 0.0920. The demonstration
|
||||
derailed the model far more than any instruction-style prompt addition.
|
||||
|
||||
With those, every direction adjacent to v8 has been measured: chain-of-thought,
|
||||
answer-style rules, output normalisation, N-forcing, batching, beam search,
|
||||
constrained decoding, few-shot. All reduced the score. The shipped
|
||||
configuration — the organizers' minimal structure plus `repetition_penalty=1.0`
|
||||
— is a sharp local optimum, and `script.py` on `main` is exactly that config.
|
||||
Reference in New Issue
Block a user