初始化项目,由ModelHub XC社区提供模型

Model: arvindcr4/iolai-2026-qwen25-14b-awq
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-09-17 18:40:18 +08:00
commit 3532add6de
14 changed files with 457415 additions and 0 deletions

35
.gitattributes vendored Normal file
View File

@@ -0,0 +1,35 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text

202
LICENSE Normal file
View File

@@ -0,0 +1,202 @@
Apache License
Version 2.0, January 2004
http://www.apache.org/licenses/
TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
1. Definitions.
"License" shall mean the terms and conditions for use, reproduction,
and distribution as defined by Sections 1 through 9 of this document.
"Licensor" shall mean the copyright owner or entity authorized by
the copyright owner that is granting the License.
"Legal Entity" shall mean the union of the acting entity and all
other entities that control, are controlled by, or are under common
control with that entity. For the purposes of this definition,
"control" means (i) the power, direct or indirect, to cause the
direction or management of such entity, whether by contract or
otherwise, or (ii) ownership of fifty percent (50%) or more of the
outstanding shares, or (iii) beneficial ownership of such entity.
"You" (or "Your") shall mean an individual or Legal Entity
exercising permissions granted by this License.
"Source" form shall mean the preferred form for making modifications,
including but not limited to software source code, documentation
source, and configuration files.
"Object" form shall mean any form resulting from mechanical
transformation or translation of a Source form, including but
not limited to compiled object code, generated documentation,
and conversions to other media types.
"Work" shall mean the work of authorship, whether in Source or
Object form, made available under the License, as indicated by a
copyright notice that is included in or attached to the work
(an example is provided in the Appendix below).
"Derivative Works" shall mean any work, whether in Source or Object
form, that is based on (or derived from) the Work and for which the
editorial revisions, annotations, elaborations, or other modifications
represent, as a whole, an original work of authorship. For the purposes
of this License, Derivative Works shall not include works that remain
separable from, or merely link (or bind by name) to the interfaces of,
the Work and Derivative Works thereof.
"Contribution" shall mean any work of authorship, including
the original version of the Work and any modifications or additions
to that Work or Derivative Works thereof, that is intentionally
submitted to Licensor for inclusion in the Work by the copyright owner
or by an individual or Legal Entity authorized to submit on behalf of
the copyright owner. For the purposes of this definition, "submitted"
means any form of electronic, verbal, or written communication sent
to the Licensor or its representatives, including but not limited to
communication on electronic mailing lists, source code control systems,
and issue tracking systems that are managed by, or on behalf of, the
Licensor for the purpose of discussing and improving the Work, but
excluding communication that is conspicuously marked or otherwise
designated in writing by the copyright owner as "Not a Contribution."
"Contributor" shall mean Licensor and any individual or Legal Entity
on behalf of whom a Contribution has been received by Licensor and
subsequently incorporated within the Work.
2. Grant of Copyright License. Subject to the terms and conditions of
this License, each Contributor hereby grants to You a perpetual,
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
copyright license to reproduce, prepare Derivative Works of,
publicly display, publicly perform, sublicense, and distribute the
Work and such Derivative Works in Source or Object form.
3. Grant of Patent License. Subject to the terms and conditions of
this License, each Contributor hereby grants to You a perpetual,
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
(except as stated in this section) patent license to make, have made,
use, offer to sell, sell, import, and otherwise transfer the Work,
where such license applies only to those patent claims licensable
by such Contributor that are necessarily infringed by their
Contribution(s) alone or by combination of their Contribution(s)
with the Work to which such Contribution(s) was submitted. If You
institute patent litigation against any entity (including a
cross-claim or counterclaim in a lawsuit) alleging that the Work
or a Contribution incorporated within the Work constitutes direct
or contributory patent infringement, then any patent licenses
granted to You under this License for that Work shall terminate
as of the date such litigation is filed.
4. Redistribution. You may reproduce and distribute copies of the
Work or Derivative Works thereof in any medium, with or without
modifications, and in Source or Object form, provided that You
meet the following conditions:
(a) You must give any other recipients of the Work or
Derivative Works a copy of this License; and
(b) You must cause any modified files to carry prominent notices
stating that You changed the files; and
(c) You must retain, in the Source form of any Derivative Works
that You distribute, all copyright, patent, trademark, and
attribution notices from the Source form of the Work,
excluding those notices that do not pertain to any part of
the Derivative Works; and
(d) If the Work includes a "NOTICE" text file as part of its
distribution, then any Derivative Works that You distribute must
include a readable copy of the attribution notices contained
within such NOTICE file, excluding those notices that do not
pertain to any part of the Derivative Works, in at least one
of the following places: within a NOTICE text file distributed
as part of the Derivative Works; within the Source form or
documentation, if provided along with the Derivative Works; or,
within a display generated by the Derivative Works, if and
wherever such third-party notices normally appear. The contents
of the NOTICE file are for informational purposes only and
do not modify the License. You may add Your own attribution
notices within Derivative Works that You distribute, alongside
or as an addendum to the NOTICE text from the Work, provided
that such additional attribution notices cannot be construed
as modifying the License.
You may add Your own copyright statement to Your modifications and
may provide additional or different license terms and conditions
for use, reproduction, or distribution of Your modifications, or
for any such Derivative Works as a whole, provided Your use,
reproduction, and distribution of the Work otherwise complies with
the conditions stated in this License.
5. Submission of Contributions. Unless You explicitly state otherwise,
any Contribution intentionally submitted for inclusion in the Work
by You to the Licensor shall be under the terms and conditions of
this License, without any additional terms or conditions.
Notwithstanding the above, nothing herein shall supersede or modify
the terms of any separate license agreement you may have executed
with Licensor regarding such Contributions.
6. Trademarks. This License does not grant permission to use the trade
names, trademarks, service marks, or product names of the Licensor,
except as required for reasonable and customary use in describing the
origin of the Work and reproducing the content of the NOTICE file.
7. Disclaimer of Warranty. Unless required by applicable law or
agreed to in writing, Licensor provides the Work (and each
Contributor provides its Contributions) on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
implied, including, without limitation, any warranties or conditions
of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
PARTICULAR PURPOSE. You are solely responsible for determining the
appropriateness of using or redistributing the Work and assume any
risks associated with Your exercise of permissions under this License.
8. Limitation of Liability. In no event and under no legal theory,
whether in tort (including negligence), contract, or otherwise,
unless required by applicable law (such as deliberate and grossly
negligent acts) or agreed to in writing, shall any Contributor be
liable to You for damages, including any direct, indirect, special,
incidental, or consequential damages of any character arising as a
result of this License or out of the use or inability to use the
Work (including but not limited to damages for loss of goodwill,
work stoppage, computer failure or malfunction, or any and all
other commercial damages or losses), even if such Contributor
has been advised of the possibility of such damages.
9. Accepting Warranty or Additional Liability. While redistributing
the Work or Derivative Works thereof, You may choose to offer,
and charge a fee for, acceptance of support, warranty, indemnity,
or other liability obligations and/or rights consistent with this
License. However, in accepting such obligations, You may act only
on Your own behalf and on Your sole responsibility, not on behalf
of any other Contributor, and only if You agree to indemnify,
defend, and hold each Contributor harmless for any liability
incurred by, or claims asserted against, such Contributor by reason
of your accepting any such warranty or additional liability.
END OF TERMS AND CONDITIONS
APPENDIX: How to apply the Apache License to your work.
To apply the Apache License to your work, attach the following
boilerplate notice, with the fields enclosed by brackets "[]"
replaced with your own identifying information. (Don't include
the brackets!) The text should be enclosed in the appropriate
comment syntax for the file format. We also recommend that a
file or class name and description of purpose be included on the
same "printed page" as the copyright notice for easier
identification within third-party archives.
Copyright 2024 Alibaba Cloud
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.

188
README.md Normal file
View File

@@ -0,0 +1,188 @@
---
license: apache-2.0
base_model: Qwen/Qwen2.5-14B-Instruct-AWQ
tags:
- iol-ai-2026
- linguistics
- reasoning
language:
- en
---
# IOL-AI 2026 — Qwen2.5-14B-Instruct-AWQ
Submission for the [IOL-AI 2026 Linguistics Olympiad Challenge](https://iolai.org).
The weights are an unmodified copy of
[`Qwen/Qwen2.5-14B-Instruct-AWQ`](https://huggingface.co/Qwen/Qwen2.5-14B-Instruct-AWQ)
(Apache-2.0, redistributable), shipped in-repo because the evaluation sandbox
has no internet access. **All of the work is in `script.py`.**
## Approach
The eval budget is 30 minutes on a 16 GB T4 for a test set of only ~90
sub-questions, so compute per problem is abundant while *reliability* is
scarce. The script is built around that asymmetry.
**1. Alignment first.** Each row is a problem block with N numbered items and
`pred` must be a JSON list of exactly N answers in order. A single missing line
shifts every later answer and zeroes the whole block on both exact-match and
chrF. `detect_n_items` recovers N from the query — handling numbered lines,
`(1)` blank markers, stated ranges, lettered items, unnumbered one-per-line
lists, and the `match_letters` shape whose items live in the shared context.
Measured on the 160 public Linguini problems it puts **98.4% of items in
correctly-sized blocks**. Model output is then force-fitted to N, preferring
the model's own numbering when it supplies it.
**2. Never emit an empty answer.** The final score is a geometric mean of exact
match and chrF, so a blank scores zero on both and is strictly worse than a
wrong guess. Every path ends in a non-empty string.
**2b. Answer style: a hypothesis that was tested and rejected.** Gold answers do
follow the conventions of whatever language the answer is in (measured over the
920 public Linguini answers, into-English golds that are full sentences are 99%
capitalised, while the 157 that are bare clauses are only 10% capitalised, and
the style matches the problem's own glosses in 36/36 measurable cases). Encoding
that as prompt guidance nevertheless *lowered* exact match on the hidden set
twice (0.0250 -> 0.0218 -> 0.0000). It is therefore not in the shipped script.
The lesson recorded here for anyone rerunning this: a correct statistical
description of the gold format did not translate into a better prompt.
**3. Monotone improvement under a hard deadline.** A complete, correctly-shaped
`submission.csv` is written *before the model is loaded*, then overwritten after
every improvement: greedy pass → each self-consistency pass → explanations.
A crash or a timeout leaves the best result reached so far on disk rather than
nothing. The script tracks its own remaining budget and stops adding passes
when one more would not fit.
**4. Greedy-anchored voting.** After the greedy pass, sampled passes (T=0.5)
run while budget remains, but the greedy answer is the default and sampled
answers may only displace it when at least two of them agree on the same
normalised form *and* that form outpolls the greedy one.
The asymmetry is empirical. A symmetric version — majority, else "most central
by chrF" — was measurably worse than not voting at all: with only a handful of
samples the centrality fallback is ill-defined (with two candidates pairwise
chrF is symmetric, so it degenerated into preferring the shorter string) and it
swapped the greedy answer for a sampled one about half the time. On the mock
set that cost 4x exact match (EM 0.044 -> 0.011). Anchoring makes the procedure
monotone: it can only fire on genuine agreement. chrF is implemented inline so
the script carries no dependency the sandbox might lack.
**5. `match_letters` as an assignment problem.** Free-form generation answers
this task type with the identity permutation (A, B, C, ...), which is a *valid*
permutation, so duplicate-repair never fires and it scores ~0. `solve_matching`
instead scores every (item, option) pair from the next-token distribution and
takes the optimal one-to-one assignment, enforcing the bijection exactly.
Duplicate-repair is retained only as a fallback for when that solver declines.
## Human Evaluation Challenge
`submission.csv` includes an `explanation` column: a short, human-readable
statement of the rules behind each answer (not a raw reasoning trace),
generated after the answers are fixed.
## Reproducing
```bash
python script.py # reads /tmp/data/test.csv, writes submission.csv
```
Environment knobs (all optional, defaults match the platform):
`IOL_TEST_CSV`, `IOL_OUT_CSV`, `IOL_MODEL`, `IOL_TIME_LIMIT`, `IOL_BATCH`,
`IOL_EXPLAIN`.
## Revision history (measured on the hidden set, not guessed)
| submission | change | score | chrF | exact match |
|---|---|---|---|---|
| 1 | symmetric self-consistency vote | 0.0686 | 0.1882 | 0.0250 |
| 2 | greedy-anchored voting (vote no longer fires) | **0.0712** | 0.2029 | 0.0250 |
| 3 | + "capitalise English answers" style rule | 0.0679 | 0.2117 | 0.0218 |
| 4 | + mirror-gloss-style, answer normalisation, equation hints | 0.0000 | 0.1527 | 0.0000 |
| 5 | revert to 2, plus the match_letters assignment solver | 0.0712 | 0.2029 | 0.0250 |
| 6 | + `repetition_penalty=1.0` (the model ships 1.05) | 0.0830 | 0.2067 | 0.0333 |
| 8 (v8) | **baseline replication + `repetition_penalty=1.0`** | **0.2245** | 0.3150 | 0.1600 |
| 10 (v10) | v8 but batch=4 (left-padded batching) | 0.1940 | 0.2818 | 0.1336 |
| 11 (v11) | v8 + beam search `num_beams=4` on the answer pass | 0.1964 | 0.2888 | 0.1336 |
| 12 (v12) | v8 + assignment solver for `match_letters` only | 0.1624 | 0.2526 | 0.1044 |
| 13 (v13) | v8 + a one-shot worked exemplar as chat turns | 0.0920 | 0.2259 | 0.0375 |
Every layer of prompt/post-processing cleverness measurably *hurt*. Submission 5
therefore reverts to the configuration of submission 2 and adds exactly one
change, motivated by a specific measured failure:
**`match_letters` was being answered with the identity permutation.** Replaying
seven parser variants over saved raw generations gave exact match 0.0000 for all
seven, which exonerates the parser — the model simply was not solving the task,
emitting the option labels in order (A, B, C, ...). Because the identity is a
valid permutation, `repair_bijection` never fired. `solve_matching` replaces
free-form generation for this task type: it scores every (item, option) pair
from the next-token distribution and takes the optimal one-to-one assignment,
so the bijection constraint is enforced exactly rather than hoped for.
## The silent decoding bug
`Qwen/Qwen2.5-14B-Instruct-AWQ` ships `generation_config.json` containing
`repetition_penalty: 1.05`. Greedy decoding ignores `temperature`, `top_p` and
`top_k` — and transformers emits a warning for each of those — but a repetition
penalty **is** applied under greedy decoding, with no warning at all.
That matters here specifically: 34% of the 920 public gold answers repeat some
letter three or more times, because these languages are agglutinative and the
answers look like `ɨmpʼuhurʼu` and `ɨŋɡɨrʼɨ`. A 5% penalty on repeated tokens
biases the model away from exactly the strings the task requires. The script now
passes `repetition_penalty=1.0` explicitly.
NFC normalisation of answers was considered and rejected: 98.15% of public golds
are already NFC, but 13 of them are NFD-and-not-NFC, so forcing NFC would break
those for an unmeasured gain.
## v8 — faithful baseline replication
The organizers' reference script reaches exact match **0.0729** on the hidden set
with these exact weights. Our best is 0.0333. Before adding anything further we
need to know whether that number is reproducible by us at all, so v8 replicates
their script literally — trivial system prompt, no chain-of-thought, **batch 1 (no padding at all)**,
naive line split, and **no forcing to N answers** — changing exactly one thing:
`repetition_penalty=1.0`. Generation is EOS-limited rather than cap-limited:
without chain-of-thought the model emits a few short answer lines and stops.
**Result: 0.2245 (chrF 0.3150, exact match 0.1600) — first place of 43 teams.**
That is 2.7x our best engineered pipeline (0.0830) and +83% on the organizers'
own baseline (0.1227), the entire delta over their number being
`repetition_penalty=1.0`.
The lesson is uncomfortable and worth recording plainly: every layer we added on
top of the reference structure — chain-of-thought, an `ANSWERS:` block, answer
style rules, output normalisation, forcing exactly N answers — reduced exact
match. Submissions 2 -> 3 -> 4 fell 0.0712 -> 0.0679 -> 0.0000 as more
engineering went in. The winning move was deleting all of it and fixing one
decoding flag.
## Final-day probes (all negative, all single changes on v8)
Three orthogonal, individually-motivated improvements were each tested as a
minimal diff on the frozen v8 script, one variable at a time:
* **v11 — beam search** (`num_beams=4`, batch 1, greedy fallback on OOM/low
budget): 0.2245 → 0.1964. Exact match fell to 0.1336 — the same value as
left-padded batch=4 — suggesting a shared fp16-numerics mechanism in
multi-sequence forward passes rather than anything about search.
* **v12 — assignment solver for `match_letters`**: 0.2245 → 0.1624, and
explanation coverage fell to 50% from the solver's extra forward passes.
The hidden set evidently does not reward bare option letters where
free-form text had been earning chrF credit.
* **v13 — one-shot worked exemplar** (a public-Linguini Kayapo problem as a
genuine user→assistant exchange): 0.2245 → 0.0920. The demonstration
derailed the model far more than any instruction-style prompt addition.
With those, every direction adjacent to v8 has been measured: chain-of-thought,
answer-style rules, output normalisation, N-forcing, batching, beam search,
constrained decoding, few-shot. All reduced the score. The shipped
configuration — the organizers' minimal structure plus `repetition_penalty=1.0`
— is a sharp local optimum, and `script.py` on `main` is exactly that config.

35
config.json Normal file
View File

@@ -0,0 +1,35 @@
{
"architectures": [
"Qwen2ForCausalLM"
],
"attention_dropout": 0.0,
"bos_token_id": 151643,
"eos_token_id": 151645,
"hidden_act": "silu",
"hidden_size": 5120,
"initializer_range": 0.02,
"intermediate_size": 13824,
"max_position_embeddings": 32768,
"max_window_layers": 70,
"model_type": "qwen2",
"num_attention_heads": 40,
"num_hidden_layers": 48,
"num_key_value_heads": 8,
"quantization_config": {
"bits": 4,
"group_size": 128,
"modules_to_not_convert": null,
"quant_method": "awq",
"version": "gemm",
"zero_point": true
},
"rms_norm_eps": 1e-06,
"rope_theta": 1000000.0,
"sliding_window": 131072,
"tie_word_embeddings": false,
"torch_dtype": "float16",
"transformers_version": "4.41.1",
"use_cache": true,
"use_sliding_window": false,
"vocab_size": 152064
}

14
generation_config.json Normal file
View File

@@ -0,0 +1,14 @@
{
"bos_token_id": 151643,
"do_sample": true,
"eos_token_id": [
151645,
151643
],
"pad_token_id": 151643,
"repetition_penalty": 1.05,
"temperature": 0.7,
"top_k": 20,
"top_p": 0.8,
"transformers_version": "4.41.1"
}

151387
merges.txt Normal file

File diff suppressed because it is too large Load Diff

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4e874dc3b1febb5b22fe74a8793066ae430d90cdbc51765d0a4eb44a82a1fbbd
size 3988804408

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b3b25da74cc854cdc726956f8152f1dda8519c7bb7d4724ac12c1312463e61c8
size 3968309440

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:25b97bf28033ed4560387293ce76bd3dc55a22882fea005a10c137b6478dae85
size 2023056736

1258
model.safetensors.index.json Normal file

File diff suppressed because it is too large Load Diff

797
script.py Normal file
View File

@@ -0,0 +1,797 @@
#!/usr/bin/env python
"""IOL-AI 2026 submission -- International Linguistics Olympiad solver.
Design notes (the eval sandbox is unforgiving, so these matter):
* HARD 30-MINUTE LIMIT. A killed process means no score at all, so the script
is structured as a monotonically-improving pipeline: it writes a complete,
correctly-shaped submission.csv *before* the model is even loaded, then
overwrites it after every improvement. Any crash or timeout leaves the best
result reached so far on disk.
* ALIGNMENT IS EVERYTHING. Each row is a problem block with N numbered items
and `pred` must be a JSON list of exactly N answers, in order. One missing
line shifts every later answer and zeroes the whole block on both metrics.
So N is detected from the query and the model output is force-fitted to it.
* NEVER EMIT AN EMPTY STRING. The final score is a geometric mean of exact
match and chrF, so an empty answer scores zero on both. A wrong guess is
strictly better than a blank.
* Environment is transformers 4.44.1 / torch 2.4.0 / autoawq on a 16GB T4
(fp16 only, no bf16, no flash-attn), with no internet.
"""
import os
import re
import json
import time
import unicodedata
from collections import Counter, defaultdict
T0 = time.time()
# The platform allows 30 minutes. Reserve a margin for model load overhead we
# can't predict and for the final write; being 60s early costs a little
# accuracy, being 1s late costs the entire submission.
TIME_LIMIT = float(os.environ.get("IOL_TIME_LIMIT", "1800"))
SAFETY = float(os.environ.get("IOL_SAFETY", "150"))
DEADLINE = T0 + TIME_LIMIT - SAFETY
TEST_CSV = os.environ.get("IOL_TEST_CSV", "/tmp/data/test.csv")
OUT_CSV = os.environ.get("IOL_OUT_CSV", "submission.csv")
MODEL_ID = os.environ.get("IOL_MODEL", ".")
WANT_EXPLANATION = os.environ.get("IOL_EXPLAIN", "1") == "1"
MAX_NEW = int(os.environ.get("IOL_MAXNEW", "900")) # reasoning budget/item
MAX_SAMPLES = int(os.environ.get("IOL_MAXSAMPLES", "8")) # self-consistency cap
# BASELINE REPLICATION MODE. The organizers' reference script reaches exact match
# 0.0729 on the hidden set with THESE EXACT WEIGHTS; our best is 0.0333. Before
# adding anything else we need to know whether that number is reproducible by us
# at all. This mode replicates their script literally -- trivial prompt, no CoT,
# 512 tokens, batch 1 (no padding at all), naive line split, NO forcing to N --
# and changes exactly one thing: repetition_penalty=1.0, our one proven fix.
BASELINE_MODE = os.environ.get("IOL_BASELINE", "1") == "1" # v8: ON by default
# Lower than the usual 0.7: samples only earn a vote by agreeing with each
# other, so keeping them near the greedy mode makes agreement meaningful.
SAMPLE_TEMP = float(os.environ.get("IOL_TEMP", "0.5"))
os.environ.setdefault("HF_HUB_OFFLINE", "1")
os.environ.setdefault("TRANSFORMERS_OFFLINE", "1")
os.environ.setdefault("TOKENIZERS_PARALLELISM", "false")
# Reduce allocator fragmentation: at batch 4 the T4 has only ~2GB spare.
os.environ.setdefault("PYTORCH_CUDA_ALLOC_CONF", "expandable_segments:True")
def log(msg):
print(f"[{time.time() - T0:7.1f}s] {msg}", flush=True)
def left():
return DEADLINE - time.time()
# ===========================================================================
# Item-count detection (validated: 98.4% of Linguini items land in
# correctly-sized blocks)
# ===========================================================================
_LINE_NUM = re.compile(r"^[ \t]*(\d{1,3})[.)\]]", re.M)
_PAREN_NUM = re.compile(r"\((\d{1,3})\)")
_RANGE = re.compile(r"\(?(\d{1,3})\s*(?:[-–—]|to)\s*(\d{1,3})\)?")
_LINE_LETTER = re.compile(r"^[ \t]*([A-Z])[.)\]]\s", re.M)
_PAREN_LETTER = re.compile(r"\(([A-Z])\)")
def detect_n_items(query, task_type="", context=""):
"""How many numbered sub-items this problem asks for. Never < 1."""
q = query or ""
line_nums = [int(m) for m in _LINE_NUM.findall(q)]
paren_nums = [int(m) for m in _PAREN_NUM.findall(q)]
range_n = 0
for a, b in _RANGE.findall(q):
a, b = int(a), int(b)
if 0 < b - a < 60:
range_n = max(range_n, b - a + 1)
cand = max(len(set(line_nums)), len(set(paren_nums)))
if range_n and cand and range_n != cand:
# A stated range ("items 1-4") can disagree with the markers actually
# present; the markers are what we have to answer, so they win.
return cand
cand = max(cand,
len(set(_LINE_LETTER.findall(q))),
len(set(_PAREN_LETTER.findall(q))))
n = max(range_n, cand)
if n > 1:
return n
# Unnumbered "Translate into X:" followed by one item per line.
lines = [l.strip() for l in q.splitlines() if l.strip()]
if len(lines) > 1:
head = lines[0]
body = lines[1:] if head.endswith((":", ".")) else lines
if body:
return len(body)
# Bare instruction ("Determine the correct correspondences."): items are in
# the shared context (this is the match_letters shape).
if context:
c_nums = len(set(int(m) for m in _LINE_NUM.findall(context)))
if c_nums > 1:
return c_nums
c_lets = len(set(_LINE_LETTER.findall(context)))
if c_lets > 1:
return c_lets
return max(n, 1)
# ===========================================================================
# Output parsing / repair
# ===========================================================================
_STRIP_PREFIX = re.compile(r"^\s*(?:\(?\d{1,3}\)?[.):\]]\s*|[-*•]\s+)")
_FENCE = re.compile(r"^```[a-zA-Z]*\s*$")
_CHATTY = re.compile(
r"^\s*(?:here (?:are|is)\b|answers?\s*:?\s*$|explanation\b|note\b|okay\b|"
r"solution\b|reasoning\b|analysis\b|translations?\s*:?\s*$|the answers?\b|"
r"let me\b|first,|so,|therefore\b|thus\b)",
re.I,
)
def clean_line(s):
s = s.strip()
s = _STRIP_PREFIX.sub("", s)
s = s.strip().strip("`").strip()
if len(s) >= 2 and s[0] == s[-1] and s[0] in "\"'“”":
s = s[1:-1].strip()
# "word | gloss" answer lines: keep the side being asked for is ambiguous,
# so keep the whole line -- chrF still gives partial credit.
return s.strip()
def extract_item_sources(query, n):
"""The source text of each numbered item, used as a last-resort fallback.
A blank scores zero on both metrics; echoing the item's own source string is
strictly better, and on transcription / fill-the-blank tasks the source and
the target share a lot of characters, so it collects real chrF credit.
"""
q = query or ""
out = []
for ln in q.splitlines():
s = ln.strip()
if not s:
continue
m = re.match(r"^\(?(\d{1,3})\)?[.):\]]\s*(.+)$", s)
if m:
out.append(m.group(2).strip())
if not out:
lines = [l.strip() for l in q.splitlines() if l.strip()]
if len(lines) > 1 and lines[0].endswith((":", ".")):
out = lines[1:]
# "form | gloss" items: the left side is the thing being asked about.
out = [o.split("|")[0].strip() if "|" in o else o for o in out]
out = [o for o in out if o]
while len(out) < n:
out.append(out[-1] if out else "?")
return out[:n]
def parse_answers(text, n, fallback=None):
"""Raw model output -> exactly n non-empty answers."""
if not text:
return list(fallback[:n]) if fallback else ["?"] * n
# Prefer the explicit final block the prompt asks for.
m = None
for m2 in re.finditer(r"(?:^|\n)\s*(?:final\s+)?answers?\s*:\s*\n?", text, re.I):
m = m2
body = text[m.end():] if m else text
numbered, raw = [], []
for ln in body.splitlines():
if _FENCE.match(ln):
continue
mm = re.match(r"^\s*\(?(\d{1,3})\)?[.):\]]\s*(.+)$", ln.strip())
if mm:
val = clean_line(mm.group(2))
if val and not _CHATTY.match(val):
numbered.append((int(mm.group(1)), val))
c = clean_line(ln)
if c and not _CHATTY.match(c):
raw.append(c)
# If the model numbered its answers, trust those labels for placement.
if len(numbered) >= n:
by_label = {}
for lab, val in numbered:
by_label[lab] = val # last write wins (models restate)
labs = sorted(by_label)
if len(labs) >= n:
return [by_label[l] for l in labs[:n]]
return fit_to_n(raw, n, fallback)
def fit_to_n(items, n, fallback=None):
items = [i for i in items if i and i.strip()]
if len(items) > n:
# Take the LAST n. The prompt asks for reasoning first and the answers
# last, so when there is no ANSWERS: marker to slice on, the tail is the
# answer block and the head is reasoning prose.
items = items[-n:]
while len(items) < n:
if fallback and len(items) < len(fallback):
items.append(fallback[len(items)])
else:
items.append(items[-1] if items else "?")
return items[:n]
def norm(s):
s = unicodedata.normalize("NFC", (s or "").strip().lower())
s = re.sub(r"\s+", " ", s)
return s.strip(" .!?;:,")
# ===========================================================================
# chrF (inline, dependency-free) -- used only to pick the most "central"
# candidate when self-consistency voting has no majority. sacrebleu is not
# guaranteed to be importable inside the sandbox.
# ===========================================================================
def _ngrams(s, k):
s = re.sub(r"\s+", "", s)
return Counter(s[i:i + k] for i in range(len(s) - k + 1)) if len(s) >= k else Counter()
def chrf_sim(hyp, ref, order=6, beta=2.0):
if not hyp or not ref:
return 0.0
ps, rs = [], []
for k in range(1, order + 1):
h, r = _ngrams(hyp, k), _ngrams(ref, k)
if not h or not r:
continue
overlap = sum((h & r).values())
ps.append(overlap / max(1, sum(h.values())))
rs.append(overlap / max(1, sum(r.values())))
if not ps:
return 0.0
p, r = sum(ps) / len(ps), sum(rs) / len(rs)
if p + r == 0:
return 0.0
b2 = beta * beta
return (1 + b2) * p * r / (b2 * p + r)
def vote(cands, anchor=None):
"""Pick one answer for an item, given the greedy answer plus samples.
`anchor` is the greedy (temperature-0) answer and is the default. Sampled
answers may only displace it when at least two of them agree on the same
normalised form AND that form has strictly more support than the anchor's.
This asymmetry is empirically necessary, not decorative. An earlier version
treated all candidates equally and fell back to "most central by chrF" when
no majority existed. With only a handful of samples that fallback is
ill-defined -- with two candidates the pairwise chrF is symmetric, so it
degenerated to picking the shorter string -- and it replaced the greedy
answer with a temperature-0.7 sample about half the time. Measured on the
mock set that cost 4x exact match (EM 0.044 -> 0.011). Anchoring makes
voting monotone: it can only fire on genuine agreement.
"""
cands = [c for c in cands if c and c.strip()]
if anchor is None:
anchor = cands[0] if cands else "?"
if len(cands) < 3:
return anchor
groups = defaultdict(list)
for c in cands:
groups[norm(c)].append(c)
anchor_support = len(groups.get(norm(anchor), []))
best_key, best_n = None, 0
for k, v in groups.items():
if len(v) > best_n:
best_key, best_n = k, len(v)
if best_key is not None and best_n >= 2 and best_n > anchor_support:
return Counter(groups[best_key]).most_common(1)[0][0]
return anchor
_OPT_LINE = re.compile(r"^[ \t]*([A-Za-z])[.)]\s+(.+)$", re.M)
_ITEM_LINE = re.compile(r"^[ \t]*(\d{1,3})[.)]\s+(.+)$", re.M)
def parse_matching_block(context):
"""For match_letters: the numbered items and the lettered options."""
items = [(int(a), b.strip()) for a, b in _ITEM_LINE.findall(context or "")]
opts = [(a, b.strip()) for a, b in _OPT_LINE.findall(context or "")]
seen = set()
items = [x for x in items if not (x[0] in seen or seen.add(x[0]))]
seen = set()
opts = [x for x in opts if not (x[0] in seen or seen.add(x[0]))]
return items, opts
def best_assignment(score):
"""Max-weight one-to-one assignment. scipy if present, else greedy+swaps."""
n, m = len(score), len(score[0])
try:
from scipy.optimize import linear_sum_assignment
import numpy as _np
r, c = linear_sum_assignment(-_np.array(score))
return list(c)
except Exception:
pass
used, out = set(), [0] * n
order = sorted(range(n), key=lambda i: -(max(score[i]) - sorted(score[i])[-2]
if m > 1 else 0))
for i in order:
j = max((j for j in range(m) if j not in used),
key=lambda j: score[i][j], default=0)
used.add(j)
out[i] = j
for _ in range(4): # local 2-swaps
improved = False
for a in range(n):
for b in range(a + 1, n):
cur = score[a][out[a]] + score[b][out[b]]
alt = score[a][out[b]] + score[b][out[a]]
if alt > cur + 1e-9:
out[a], out[b] = out[b], out[a]
improved = True
if not improved:
break
return out
def repair_bijection(answers):
"""match_letters answers are usually a permutation of the option letters.
When every answer is a single letter and there are as many items as
distinct letters available, duplicates are certainly wrong. Reassign the
duplicated slots to the unused letters. Strictly guarded so it is a no-op
on anything that isn't this shape.
"""
if len(answers) < 3:
return answers
if not all(re.fullmatch(r"[A-Z]", a or "") for a in answers):
return answers
n = len(answers)
universe = [chr(ord("A") + i) for i in range(n)]
if len(set(answers)) == n:
return answers
unused = [l for l in universe if l not in set(answers)]
if not unused:
return answers
seen, out = set(), []
for a in answers:
if a in seen and unused:
out.append(unused.pop(0))
else:
seen.add(a)
out.append(a)
return out
# ===========================================================================
# Prompting
# ===========================================================================
SYSTEM = (
"You are a gold medallist at the International Linguistics Olympiad.\n"
"Each problem gives data from a language you have never seen. Everything "
"you need is in the problem itself; no outside knowledge is required or "
"allowed.\n"
"Method: line up the given examples, segment the words, identify the "
"recurring morphemes and the rules that order them, check your rules "
"against EVERY example, then apply them to the items asked for.\n"
"Be concise while reasoning. Then output a final block that begins with a "
"line containing exactly ANSWERS: followed by one answer per line, in the "
"order asked, with no numbering, no commentary and no blank lines.\n"
"Give your best guess for every item. Never leave one blank."
)
# Exact match is half the score, so the answer's *form* matters as much as its
# content. test.csv states the task type, so say precisely what a well-formed
# answer looks like. Unknown/absent types simply get no hint.
TASK_HINTS = {
"translation": "Each answer is the translation alone -- no source text, no "
"gloss, no notes, no quotation marks.",
"match_letters": "Each answer is a single capital letter identifying the "
"match for that numbered item. Every letter is used "
"exactly once, so no letter may repeat.",
"fill_blanks": "Each answer is only the missing form that belongs in that "
"blank -- not the whole line, not the gloss.",
"text_to_num": "Each answer is written in digits only (e.g. 111).",
"num_to_text": "Each answer is the number written out in the problem "
"language, words only.",
}
def build_prompt(row, n):
hint = TASK_HINTS.get((row.get("task_type") or "").strip().lower(), "")
return (
f"{row['context'].strip()}\n\n{row['query'].strip()}\n\n"
f"There are exactly {n} item{'s' if n != 1 else ''} to answer."
+ (f" {hint}" if hint else "") +
f"\nAfter your reasoning, write ANSWERS: on its own line and then exactly "
f"{n} line{'s' if n != 1 else ''}, one answer per item, in order."
)
EXPLAIN_SYSTEM = (
"You explain International Linguistics Olympiad solutions to a human judge. "
"Given a problem and the answers produced, state the key rules of the "
"language that justify them: the relevant morphemes, word order and any "
"sound changes. Be specific and concise (2-4 sentences or a few short "
"bullets). Do not restate the reasoning as a stream of thought."
)
def build_explain_prompt(row, answers):
return (
f"{row['context'].strip()}\n\n{row['query'].strip()}\n\n"
f"Answers given:\n" + "\n".join(f"- {a}" for a in answers) +
"\n\nBriefly explain the linguistic rules behind these answers."
)
# ===========================================================================
# Main
# ===========================================================================
def dev_score(preds):
"""Offline diagnostic: score against a gold file when IOL_GOLD is set.
Never runs on the platform (the answers are hidden, so the variable is
unset there); it exists so one benchmark run reveals the whole learning
curve -- greedy, then after each self-consistency pass -- instead of a
single final number.
"""
gold_path = os.environ.get("IOL_GOLD")
if not gold_path or not os.path.exists(gold_path):
return
try:
import ast
import pandas as pd
g = pd.read_csv(gold_path, dtype=str)
ems, cfs = [], []
for _, r in g.iterrows():
gold = ast.literal_eval(r["answer"])
p = preds.get(str(r["id"]), [])
p = list(p)[:len(gold)] + [""] * max(0, len(gold) - len(p))
for gi, pi in zip(gold, p):
alts = gi if isinstance(gi, (list, tuple)) else [gi]
alts = [str(a) for a in alts]
ems.append(1.0 if any(pi.strip() == a.strip() for a in alts) else 0.0)
cfs.append(max(chrf_sim(pi, a) for a in alts))
em = sum(ems) / max(1, len(ems))
cf = sum(cfs) / max(1, len(cfs))
log(f" [dev] EM={em:.4f} chrF~={cf:.4f} score~={(em * cf) ** 0.5:.4f} "
f"over {len(ems)} items")
except Exception as e:
log(f" [dev] scoring failed: {type(e).__name__}: {e}")
def write_submission(path, ids, preds, explanations=None):
import pandas as pd
rows = []
for i in ids:
rec = {"id": i, "pred": json.dumps(preds[i], ensure_ascii=False)}
if explanations is not None:
rec["explanation"] = explanations.get(i, "")
rows.append(rec)
pd.DataFrame(rows).to_csv(path, index=False)
def main():
import pandas as pd
df = pd.read_csv(TEST_CSV, dtype=str).fillna("")
ids = [str(x) for x in df["id"].tolist()]
ns = [detect_n_items(r.get("query", ""), r.get("task_type", ""), r.get("context", ""))
for _, r in df.iterrows()]
total_items = sum(ns)
log(f"loaded {len(df)} problems, {total_items} items "
f"(min={min(ns)} max={max(ns)} mean={total_items / len(ns):.1f})")
srcs = {i: extract_item_sources(r.get("query", ""), n)
for i, (_, r), n in zip(ids, df.iterrows(), ns)}
# --- 1. Baseline submission on disk before anything can go wrong --------
preds = {i: list(srcs[i]) for i in ids}
explanations = {i: "" for i in ids} if WANT_EXPLANATION else None
write_submission(OUT_CSV, ids, preds, explanations)
log(f"wrote placeholder {OUT_CSV} ({len(ids)} rows)")
# --- 2. Load model -----------------------------------------------------
import torch
from transformers import (AutoTokenizer, AutoModelForCausalLM,
StoppingCriteria, StoppingCriteriaList)
class Deadline(StoppingCriteria):
"""Abort generation on wall-clock, checked every token.
Without this the budget is only checked between batches, so a batch
started near the limit runs past it and the platform kills the process.
"""
def __init__(self, stop_at):
self.stop_at = stop_at
def __call__(self, input_ids, scores, **kw):
return time.time() > self.stop_at
log("loading tokenizer/model ...")
tok = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True)
if tok.pad_token is None:
tok.pad_token = tok.eos_token
tok.padding_side = "left"
# Pin every layer to the GPU. device_map="auto" is free to spill layers to
# CPU when it thinks VRAM is tight, and a couple of offloaded layers make
# generation ~100x slower without any error -- the worst kind of failure
# here. Falling back to "auto" only if the explicit placement fails.
def _load(dev_map):
# transformers 4.44 (the sandbox) wants torch_dtype=; 5.x renamed it to
# dtype=. Accept either so the same file runs in both.
try:
return AutoModelForCausalLM.from_pretrained(
MODEL_ID, torch_dtype=torch.float16, device_map=dev_map,
trust_remote_code=True).eval()
except TypeError:
return AutoModelForCausalLM.from_pretrained(
MODEL_ID, dtype=torch.float16, device_map=dev_map,
trust_remote_code=True).eval()
try:
model = _load({"": 0} if torch.cuda.is_available() else "auto")
except Exception as e:
log(f"pinned load failed ({type(e).__name__}: {e}); falling back to auto")
model = _load("auto")
devs = set(str(p.device) for p in model.parameters())
log(f"model ready on {sorted(devs)} ({left():.0f}s of budget left)")
if any(d.startswith("cpu") or d == "meta" for d in devs):
log("WARNING: part of the model is off-GPU; generation will be very slow")
if torch.cuda.is_available():
log(f" VRAM allocated {torch.cuda.memory_allocated()/1e9:.2f} GB / "
f"{torch.cuda.get_device_properties(0).total_memory/1e9:.1f} GB")
prompts = []
for (_, r), n in zip(df.iterrows(), ns):
if BASELINE_MODE:
msgs = [{"role": "system", "content":
"You solve International Linguistics Olympiad problems. "
"Answer every numbered item. Put each answer on its own line, "
"in order, with no numbering and no extra text."},
{"role": "user", "content":
f"{r['context'].strip()}\n\n{r['query'].strip()}"}]
else:
msgs = [{"role": "system", "content": SYSTEM},
{"role": "user", "content": build_prompt(r, n)}]
prompts.append(tok.apply_chat_template(msgs, tokenize=False,
add_generation_prompt=True))
batch_size = 1 if BASELINE_MODE else int(os.environ.get("IOL_BATCH", "4"))
def generate(texts, max_new, sample, temp=0.7):
"""Batched generation with OOM backoff. Returns list of strings."""
nonlocal batch_size
out = [""] * len(texts)
order = sorted(range(len(texts)), key=lambda i: len(texts[i]))
i = 0
while i < len(order):
if left() < 25:
log(" out of time inside generate(); returning partial")
break
idx = order[i:i + batch_size]
chunk = [texts[j] for j in idx]
try:
enc = tok(chunk, return_tensors="pt", padding=True,
truncation=True, max_length=6144).to(model.device)
# repetition_penalty=1.0 EXPLICITLY. Qwen2.5-14B-Instruct-AWQ
# ships generation_config.json with repetition_penalty=1.05,
# and unlike temperature/top_p/top_k (which greedy ignores, and
# which transformers warns about) a repetition penalty IS
# applied under greedy decoding -- silently, with no warning.
# 34% of the public gold answers repeat a letter 3+ times
# (agglutinative morphology like 'ɨmpʼuhurʼu'), so a 5% penalty
# pushes the model off exactly the strings we need.
kw = dict(max_new_tokens=max_new, pad_token_id=tok.pad_token_id,
repetition_penalty=1.0,
stopping_criteria=StoppingCriteriaList(
[Deadline(DEADLINE - 10)]))
if sample:
kw.update(do_sample=True, temperature=temp, top_p=0.95)
else:
kw.update(do_sample=False)
with torch.no_grad():
o = model.generate(**enc, **kw)
for k, j in enumerate(idx):
out[j] = tok.decode(o[k][enc["input_ids"].shape[1]:],
skip_special_tokens=True)
i += batch_size
except torch.cuda.OutOfMemoryError:
torch.cuda.empty_cache()
if batch_size == 1:
log(" OOM at batch=1; skipping this item")
i += 1
else:
batch_size = max(1, batch_size // 2)
log(f" OOM -> batch_size={batch_size}")
except Exception as e: # never die mid-run
log(f" generate error: {type(e).__name__}: {e}")
i += batch_size
return out
def solve_matching(row, n):
"""Score every (item, option) pair and take the best one-to-one assignment.
Free-form generation fails badly here: measured on the benchmark the
model just emits the option labels in order (A, B, C, ... == the
identity permutation), which is a *valid* permutation so no repair
fires, and it scores ~0. Asking for one letter at a time and reading
the next-token distribution turns the task into an assignment problem
the model is actually good at, and the one-to-one constraint is then
enforced exactly rather than hoped for.
"""
items, opts = parse_matching_block(row.get("context", ""))
if len(items) < 3 or len(opts) < 3 or len(items) != n:
return None
letters = [o[0] for o in opts]
# token id for each option letter, bare and space-prefixed
cand_ids = []
for L in letters:
ids = set()
for form in (L, " " + L):
t = tok.encode(form, add_special_tokens=False)
if t:
ids.add(t[0])
cand_ids.append(sorted(ids))
ctx = row["context"].strip()
prompts_m = []
for num, itext in items:
msgs = [
{"role": "system", "content":
"You match items to their correct counterparts in a "
"linguistics problem. Reply with one option letter only."},
{"role": "user", "content":
f"{ctx}\n\nWhich lettered option corresponds to item {num} "
f"({itext})? Reply with the option letter only."},
]
prompts_m.append(tok.apply_chat_template(
msgs, tokenize=False, add_generation_prompt=True))
score = []
bs = 4
for s0 in range(0, len(prompts_m), bs):
if left() < 30:
return None
chunk = prompts_m[s0:s0 + bs]
enc = tok(chunk, return_tensors="pt", padding=True,
truncation=True, max_length=6144).to(model.device)
with torch.no_grad():
logits = model(**enc).logits[:, -1, :].float()
logprobs = torch.log_softmax(logits, dim=-1)
for b in range(len(chunk)):
score.append([max(logprobs[b, i].item() for i in ids)
for ids in cand_ids])
col = best_assignment(score)
return [letters[c] for c in col]
# --- 3. Pass 1: greedy, guarantees a full answer set --------------------
# Size the reasoning budget to the actual problem count. Measured on the
# eval hardware (T4, 14B AWQ, batch 4) throughput is ~32 tok/s, so the whole
# 30 minutes buys only ~50k generated tokens. With ~16 problem blocks that
# affords full-length reasoning; if the platform instead ships one row per
# sub-question (~90 rows) a fixed 900-token budget would not even finish a
# single pass. Spend at most ~40% of what's left on pass 1.
TOK_PER_S = float(os.environ.get("IOL_TOKS", "30"))
adaptive = int(0.40 * max(1.0, left()) * TOK_PER_S / max(1, len(df)))
max_new = max(192, min(MAX_NEW, adaptive))
log(f"reasoning budget: {max_new} new tokens/problem "
f"(adaptive={adaptive}, cap={MAX_NEW}, {len(df)} problems)")
t = time.time()
texts = generate(prompts, max_new=max_new, sample=False)
pass1_cost = time.time() - t
samples = {i: [] for i in ids}
n_matched = 0
for (i, n, txt), (_, row) in zip(zip(ids, ns, texts), df.iterrows()):
if BASELINE_MODE:
# literally the organizers' parse: every non-empty stripped line,
# however many there are. No cleaning, no fallback, no forcing.
preds[i] = [ln.strip() for ln in (txt or "").splitlines() if ln.strip()]
samples[i].append(preds[i])
continue
a = repair_bijection(parse_answers(txt, n, srcs[i]))
# match_letters: free-form generation emits the identity permutation
# (A, B, C, ...) and scores ~0, so solve it as an assignment instead.
if (row.get("task_type") or "").strip().lower() == "match_letters":
try:
mm_ = solve_matching(row, n)
if mm_ and len(mm_) == n:
a = mm_
n_matched += 1
except Exception as e:
log(f" matching solver failed on {i}: {type(e).__name__}: {e}")
preds[i] = a
samples[i].append(a)
if n_matched:
log(f"assignment solver used on {n_matched} match_letters problem(s)")
write_submission(OUT_CSV, ids, preds, explanations)
# How often did reasoning run past the token budget before the model got to
# its ANSWERS: block? Those problems fall back to salvaged lines, so a high
# count means max_new is too small rather than the model being wrong.
no_block = sum(1 for txt in texts
if not re.search(r"answers?\s*:", txt or "", re.I))
empty = sum(1 for txt in texts if not (txt or "").strip())
log(f"pass 1 (greedy) done in {pass1_cost:.0f}s -> submission written "
f"({no_block}/{len(texts)} without an ANSWERS: block, {empty} empty)")
dev_score(preds)
# --- 4. Self-consistency passes while budget allows ---------------------
reserve = 0.0
if WANT_EXPLANATION:
reserve = min(300.0, 0.25 * pass1_cost + 60) # explanations are short
n_extra = 0
while left() - reserve > pass1_cost * 1.25 and n_extra < MAX_SAMPLES:
n_extra += 1
log(f"self-consistency pass {n_extra} ({left():.0f}s left)")
texts = generate(prompts, max_new=max_new, sample=True, temp=SAMPLE_TEMP)
for i, n, txt in zip(ids, ns, texts):
if txt:
samples[i].append(repair_bijection(parse_answers(txt, n, srcs[i])))
for i, n in zip(ids, ns):
# samples[i][0] is the greedy pass; it anchors every item.
if len(samples[i]) >= 3:
greedy = samples[i][0]
preds[i] = repair_bijection(
[vote([s[k] for s in samples[i] if k < len(s)],
anchor=greedy[k] if k < len(greedy) else None)
for k in range(n)])
write_submission(OUT_CSV, ids, preds, explanations)
log(f" voted over {n_extra + 1} samples (greedy-anchored) -> written")
dev_score(preds)
# --- 5. Explanations for the jury track ---------------------------------
if WANT_EXPLANATION and left() > 60:
log(f"generating explanations ({left():.0f}s left)")
ex_prompts = []
for (_, r), i in zip(df.iterrows(), ids):
msgs = [{"role": "system", "content": EXPLAIN_SYSTEM},
{"role": "user", "content": build_explain_prompt(r, preds[i])}]
ex_prompts.append(tok.apply_chat_template(
msgs, tokenize=False, add_generation_prompt=True))
ex = generate(ex_prompts, max_new=200, sample=False)
for i, e in zip(ids, ex):
e = re.sub(r"\s+", " ", (e or "").strip())
if e:
explanations[i] = e[:1200]
write_submission(OUT_CSV, ids, preds, explanations)
log("explanations written")
# --- 6. Final integrity check ------------------------------------------
bad = [i for i, n in zip(ids, ns) if len(preds[i]) != n or any(
not str(x).strip() for x in preds[i])]
if bad:
log(f"repairing {len(bad)} malformed rows")
for i, n in zip(ids, ns):
preds[i] = fit_to_n([x for x in preds[i] if str(x).strip()], n, srcs[i])
write_submission(OUT_CSV, ids, preds, explanations)
log(f"DONE. {len(ids)} rows, {sum(len(v) for v in preds.values())} answers, "
f"{time.time() - T0:.0f}s elapsed")
if __name__ == "__main__":
main()

303282
tokenizer.json Normal file

File diff suppressed because it is too large Load Diff

207
tokenizer_config.json Normal file
View File

@@ -0,0 +1,207 @@
{
"add_bos_token": false,
"add_prefix_space": false,
"added_tokens_decoder": {
"151643": {
"content": "<|endoftext|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151644": {
"content": "<|im_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151645": {
"content": "<|im_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151646": {
"content": "<|object_ref_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151647": {
"content": "<|object_ref_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151648": {
"content": "<|box_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151649": {
"content": "<|box_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151650": {
"content": "<|quad_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151651": {
"content": "<|quad_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151652": {
"content": "<|vision_start|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151653": {
"content": "<|vision_end|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151654": {
"content": "<|vision_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151655": {
"content": "<|image_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151656": {
"content": "<|video_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": true
},
"151657": {
"content": "<tool_call>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151658": {
"content": "</tool_call>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151659": {
"content": "<|fim_prefix|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151660": {
"content": "<|fim_middle|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151661": {
"content": "<|fim_suffix|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151662": {
"content": "<|fim_pad|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151663": {
"content": "<|repo_name|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
},
"151664": {
"content": "<|file_sep|>",
"lstrip": false,
"normalized": false,
"rstrip": false,
"single_word": false,
"special": false
}
},
"additional_special_tokens": [
"<|im_start|>",
"<|im_end|>",
"<|object_ref_start|>",
"<|object_ref_end|>",
"<|box_start|>",
"<|box_end|>",
"<|quad_start|>",
"<|quad_end|>",
"<|vision_start|>",
"<|vision_end|>",
"<|vision_pad|>",
"<|image_pad|>",
"<|video_pad|>"
],
"bos_token": null,
"chat_template": "{%- if tools %}\n {{- '<|im_start|>system\\n' }}\n {%- if messages[0]['role'] == 'system' %}\n {{- messages[0]['content'] }}\n {%- else %}\n {{- 'You are Qwen, created by Alibaba Cloud. You are a helpful assistant.' }}\n {%- endif %}\n {{- \"\\n\\n# Tools\\n\\nYou may call one or more functions to assist with the user query.\\n\\nYou are provided with function signatures within <tools></tools> XML tags:\\n<tools>\" }}\n {%- for tool in tools %}\n {{- \"\\n\" }}\n {{- tool | tojson }}\n {%- endfor %}\n {{- \"\\n</tools>\\n\\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\\n<tool_call>\\n{\\\"name\\\": <function-name>, \\\"arguments\\\": <args-json-object>}\\n</tool_call><|im_end|>\\n\" }}\n{%- else %}\n {%- if messages[0]['role'] == 'system' %}\n {{- '<|im_start|>system\\n' + messages[0]['content'] + '<|im_end|>\\n' }}\n {%- else %}\n {{- '<|im_start|>system\\nYou are Qwen, created by Alibaba Cloud. You are a helpful assistant.<|im_end|>\\n' }}\n {%- endif %}\n{%- endif %}\n{%- for message in messages %}\n {%- if (message.role == \"user\") or (message.role == \"system\" and not loop.first) or (message.role == \"assistant\" and not message.tool_calls) %}\n {{- '<|im_start|>' + message.role + '\\n' + message.content + '<|im_end|>' + '\\n' }}\n {%- elif message.role == \"assistant\" %}\n {{- '<|im_start|>' + message.role }}\n {%- if message.content %}\n {{- '\\n' + message.content }}\n {%- endif %}\n {%- for tool_call in message.tool_calls %}\n {%- if tool_call.function is defined %}\n {%- set tool_call = tool_call.function %}\n {%- endif %}\n {{- '\\n<tool_call>\\n{\"name\": \"' }}\n {{- tool_call.name }}\n {{- '\", \"arguments\": ' }}\n {{- tool_call.arguments | tojson }}\n {{- '}\\n</tool_call>' }}\n {%- endfor %}\n {{- '<|im_end|>\\n' }}\n {%- elif message.role == \"tool\" %}\n {%- if (loop.index0 == 0) or (messages[loop.index0 - 1].role != \"tool\") %}\n {{- '<|im_start|>user' }}\n {%- endif %}\n {{- '\\n<tool_response>\\n' }}\n {{- message.content }}\n {{- '\\n</tool_response>' }}\n {%- if loop.last or (messages[loop.index0 + 1].role != \"tool\") %}\n {{- '<|im_end|>\\n' }}\n {%- endif %}\n {%- endif %}\n{%- endfor %}\n{%- if add_generation_prompt %}\n {{- '<|im_start|>assistant\\n' }}\n{%- endif %}\n",
"clean_up_tokenization_spaces": false,
"eos_token": "<|im_end|>",
"errors": "replace",
"model_max_length": 131072,
"pad_token": "<|endoftext|>",
"split_special_tokens": false,
"tokenizer_class": "Qwen2Tokenizer",
"unk_token": null
}

1
vocab.json Normal file

File diff suppressed because one or more lines are too long