Files
iolai26-solve/fable.md
ModelHub XC 5b016c1af1 初始化项目,由ModelHub XC社区提供模型
Model: rpant/iolai26-solve
Source: Original Platform
2026-07-28 09:36:12 +08:00

219 lines
12 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# fable.md — research-direction notes from implementation
Notes written while building the IOL-AI solver seed (see AGENTS.md for the plan,
README.md for repo state). Each item is something the implementation surfaced
that the research plan should absorb.
**Update (real-data phase):** items 4, 8, 9 are now RESOLVED — the official
eval notebook (rita-berrada/iolai-2026-workshop) fixed the semantics, and
facebook/linguini gave 160 real dev puzzles. New findings start at §11.
## 11. Real-data baseline resets expectations (and the plan's emphasis)
Symbolic-only on all 160 Linguini puzzles: **EM 2.4%, chrF 20, in ~9 s.**
(The synthetic dev set scored 1.0 — constructed puzzles flatter substitution
methods.) Real translation items are compositional, real paradigms involve
metathesis/harmony/infixation, real numeral systems have morphophonology that
breaks token-level CSP. Consequences:
- The LLM direct path (notebook format) is the main score carrier for now;
the symbolic stack's near-term value is (a) exact hits where structure is
clean, (b) the chrF floor, (c) **verified confidences that route the LLM
budget** (157/160 puzzles flagged low-confidence — correctly).
- The DSL/CEGIS path is the ceiling-raiser, not the floor: its job is to
convert LLM linguistic insight into verified exact hits. Ablate it against
direct-LLM on the T4 before investing in finetuning.
## 12. Official semantics differ from my assumptions in scoring-relevant ways
- EM is `strip().lower()` ONLY — final punctuation and internal spacing are
significant. Answers must carry the gold's punctuation conventions; a
format-inducer pass (§3) is now demonstrably worth EM points.
- Gold items can be a LIST OF ALTERNATIVES (14/160 rows); scorer takes max.
- chrF is sacrebleu with effective-order smoothing; my reimplementation now
matches to 0.0 delta over 500 fuzz cases (tests/test_scorer.py).
## 13. Linguini formats: what the parser must survive (now does, 158/160)
Pipe tables 2-5 columns with header rows; numbered example sentences whose
numbering the query CONTINUES (item "17." refers to nothing in the query);
bare-line items with no numbering; (k)-blank markers in any column, both
directions in one puzzle, including damaged rows ("(5) to tie" merged cell);
items living in the CONTEXT while the query is instruction-only
("Determine the correct correspondences", "Fill in the blanks (114)");
task-language strings that end in ':' (vowel length) — never use trailing
colon alone to detect instructions. Two rows remain unparseable-by-count: one
has misaligned gold (7 answers, 6 items), one enumerates payloads inline in
prose. Position-aligned scoring makes over/under-parsing cost only the
misaligned tail — worth a guard that pads rather than truncates.
## 14. match_letters needs morpheme-level CSP, not surface matching
16 puzzles / 223 items (25% of all items) are unordered form↔meaning
matching with usually ZERO attested pairs. Surface similarity carries no
signal (5.4% EM ≈ barely above random). The real structure: recurring
morphemes across forms must map consistently to recurring words across
meanings (Zuni doko:ko ↔ 'chicken', mo:chikwa ↔ 'peach'). That is a small
constraint-satisfaction / bilingual-lexicon-induction problem over sets —
highly verifiable (a candidate assignment implies a consistent lexicon or it
doesn't) and a perfect fit for the propose-verify architecture: LLM proposes
morpheme↔word hypotheses, a solver checks global consistency, Hungarian
finishes. My soft quadratic-assignment attempt (power iteration over
similarity graphs) was neutral; the discrete version is the right next try.
## 15. Throughput reality on the T4 (from the notebook)
The baseline runs ~1 min/problem at 1536 new tokens on the 14B AWQ — a full
160-puzzle set would need ~2.7 h sequential. The 30-minute budget therefore
REQUIRES the two-pass design: symbolic first (~9 s), then batched direct-LLM
on the low-confidence subset, weakest-first, with a hard budget cutoff
(script.py implements this). Concrete knobs to tune on Colab: batch size
(KV-cache limited on 16 GB with a 14B model — try 2-4), MAX_NEW_TOKENS
(1536 baseline; shorter for translation-only puzzles), and whether a 7B
model with bigger batches beats the 14B with tiny batches at fixed wall
clock. That ablation needs the actual T4.
## 16a. v1 migration to single-shot-with-tools (done 2026-07-21)
The architecture inversion recommended by the landscape survey is
implemented: the tools now feed the answering model (scaffold in the prompt:
segmentation, alignment, verified numeral values, labeled symbolic
candidates), verified symbolic output overrides the model, unverified output
is demoted to hints, and CEGIS left the runtime path. Explanations ride in
the same generation (`EXPLANATION:` after `FINAL ANSWERS:`) — the
human-judged track costs zero extra passes. Open empirical questions for the
T4: does the scaffold help a 14B AWQ model the way it helped GPT-4-class
models (the segmentation evidence says gains concentrate above ~15% baseline
— a 14B may be below it); and the deferred CoT-vs-IO prompt A/B (§ "Probing
LLMs", arXiv:2502.00817).
## 16b. First live submission: the T4 constraint is prefill logits, not KV
The first real submission OOMed: on the sandbox's transformers, `generate`
computes LM-head logits over EVERY prompt position and casts to float32 —
batch 6 × ~2.9k tokens × 152k vocab × 4B = the exact 9.82 GiB failed
allocation. KV cache (what batch sizing usually optimizes) is negligible
under GQA; **total prompt tokens per batch** is the real T4 memory bound at
~0.9 MB/token. Fixes: token-budget batch packing (3,500), OOM → retry
singly → abstain, `logits_to_keep=1` where supported. The deeper lesson:
the crash killed the script before submission.csv existed — "never empty"
must hold at the FILE level, not the item level. submission.csv is now
checkpoint-written from ~15 s in and atomically replaced as results improve.
## 16. The sandbox runs OLDER transformers than Colab
The notebook comments that `apply_chat_template(..., return_dict=True)`
returns a dict on Colab but a bare tensor in the sandbox. llm.py sidesteps
the whole class of drift by using `tokenize=False` + a normal tokenizer
call, which is stable across versions. Keep any future transformers usage on
the oldest-common-API path; there is no way to pin the sandbox's version.
## 1. The symbolic template translator may be the real workhorse, not the DSL
The plan positions LLM→DSL synthesis as the core translation engine. But the
minimal-pair **substitution translator** (`solver/template.py`: find the
attested sentence closest to the query, swap the differing tokens through
alignment links) solved every synthetic translation item exactly, with zero
LLM tokens, once three refinements landed:
- **competition / explaining-away in alignment** (`align.py`): demote a target
token already strongly claimed by another source token. Tiny corpora make
co-occurrence ties pervasive; this one change flipped several wrong answers.
- **morphological back-off**: single-word glosses (`moko = dog`) project into
inflected forms containing them (`namoko`).
- **affix-matched substitution**: when replacing a form, prefer the candidate
sharing an affix with the form replaced (kupu:nakupu :: moko:namoko) — a
proportional analogy inside the template frame.
Implication: budget the DSL/CEGIS path for the residual — items where the
query is far (bag-distance > ~len/2) from every attested sentence, long
compositional sentences, and phenomena the template can't reach (reordering,
harmony). Measure on real dev data what fraction that residual actually is;
it directly sets the LLM compute budget and the finetuning priority.
## 2. Verifier honesty requires two scoring regimes (found a real trap)
Memorizing predictors (template translator, fallback) score EM=1.0 on attested
pairs by construction, so scoring every candidate with plain `evaluate()` lets
memorization beat a generalizing grammar on ties — with MDL penalty it beats
it *always*. The fix in `router._pick_direction_solver`: fixed programs
(grammars) are scored directly (they must reproduce the data), while
fit-from-data predictors are scored **leave-one-out**. Any future candidate
added to the selection pool must declare which regime it belongs to. This also
matters for RL later: the reward must be LOO-style for anything with data
access, or the policy learns to memorize.
## 3. Answer-format induction is a first-class subproblem the plan ignores
The synthetic num_to_text case exposed it: for 23, `hun ox` (20+3) and
`kan hun ox` (1×20+3) are both arithmetically valid; the gold followed the
attested *style* (explicit unit multiplier). The fix was a style model scoring
candidates by consistency with attested phrasing. This generalizes: whenever
multiple surface answers are semantically correct, the scorer only rewards the
one matching the dataset's convention (capitalization, articles, hyphenation
of morph boundaries, multiplier explicitness). Recommend: an explicit
"format inducer" pass that learns per-puzzle answer conventions from context
examples, applied as the last stage of every task type. Cheap, and it converts
chrF-close answers into EM hits — which the geometric mean doubly rewards.
## 4. Context parsing is still the top empirical risk
Everything downstream consumes `extract_pairs`. The current parser
(separator voting + per-line fallback) handles the formats I could construct,
but real Linguini CSVs were not available in this environment. **Before any
model work, pull the actual Linguini/PuzzLing dev CSVs and fuzz the parser
against every context in them**; count rows where zero pairs are extracted —
each such row is guaranteed near-zero score. A cheap LLM repair path (ask the
model to emit the pair list as JSON when the symbolic parser yields < 2 pairs)
would cap this risk at one batched call per malformed puzzle.
## 5. Direction detection deserves attested-data validation, not just wording
`detect_direction` keys off phrases like "into English". Real queries may
say "What does X mean?", "Give the form meaning ...", or nothing explicit.
A robust cross-check: try both directions and see which side of the attested
pairs the query string is *script-similar* to (character n-gram overlap with
task-language vs work-language material). If the query looks like the unknown
language, it's an analysis item. Cheap and language-agnostic.
## 6. CEGIS should get verifier-guided repair hints, not just failures
The refine prompt currently shows failing (input, expected, got) triples. The
symbolic layer knows more: which morpheme boundary the output diverged at,
which affix went unused, whether the failure is pure reordering. Feeding a
one-line diagnosis per failure ("output differs only in word order",
"expected form contains attested morph 'na-' that your grammar never
attaches") should cut CEGIS rounds worth an ablation axis in phase 2/3.
## 7. Self-consistency has a free implementation via round-trip
The interpreter runs both directions, so `analyze(generate(x)) == x` is a
zero-cost consistency check usable as (a) a tie-breaker in grammar selection
(already anticipated in AGENTS.md), and (b) a confidence signal for budget
allocation spend refinement rounds on items whose round-trip fails.
## 8. Points weighting has no data column — resolve early
The scorer supports per-item weights, but the Linguini schema
(id/context/query/...) carries no points column. If official IOL point values
weight the leaderboard metric, they must be embedded somewhere (query text?
separate mapping by id?). Resolve this with a probe submission early it
changes budget allocation (high-point items deserve the CEGIS rounds).
## 9. Local scorer vs sacrebleu chrF: verify parity once, in the eval image
`eval/scorer.py` implements chrF2 (n6, β=2, whitespace stripped) by hand to
stay dependency-free. sacrebleu differs in epsilon smoothing on zero-match
orders. Before trusting local ablations, run both on a few hundred string
pairs in the submission container and confirm the delta is < 1e-3; otherwise
model selection could silently optimize the wrong metric.
## 10. Prompt library format drifted from the plan (deliberately)
AGENTS.md says YAML prompts; the repo uses plain `.md` templates with
`str.format` placeholders (`prompts/*.md`). Rationale: zero dependencies, no
YAML-escaping pain with multiline linguistic data. If config-driven prompt
variants are needed for best-of-N diversity (the greedy-decoding constraint
means diversity must come from prompts, not sampling temperature see
`synth.synthesize_best_of_n`), add a small variants list per template rather
than reintroducing YAML.