Files
iolai26-solve/fable.md

219 lines
12 KiB
Markdown
Raw Normal View History

# fable.md — research-direction notes from implementation
Notes written while building the IOL-AI solver seed (see AGENTS.md for the plan,
README.md for repo state). Each item is something the implementation surfaced
that the research plan should absorb.
**Update (real-data phase):** items 4, 8, 9 are now RESOLVED — the official
eval notebook (rita-berrada/iolai-2026-workshop) fixed the semantics, and
facebook/linguini gave 160 real dev puzzles. New findings start at §11.
## 11. Real-data baseline resets expectations (and the plan's emphasis)
Symbolic-only on all 160 Linguini puzzles: **EM 2.4%, chrF 20, in ~9 s.**
(The synthetic dev set scored 1.0 — constructed puzzles flatter substitution
methods.) Real translation items are compositional, real paradigms involve
metathesis/harmony/infixation, real numeral systems have morphophonology that
breaks token-level CSP. Consequences:
- The LLM direct path (notebook format) is the main score carrier for now;
the symbolic stack's near-term value is (a) exact hits where structure is
clean, (b) the chrF floor, (c) **verified confidences that route the LLM
budget** (157/160 puzzles flagged low-confidence — correctly).
- The DSL/CEGIS path is the ceiling-raiser, not the floor: its job is to
convert LLM linguistic insight into verified exact hits. Ablate it against
direct-LLM on the T4 before investing in finetuning.
## 12. Official semantics differ from my assumptions in scoring-relevant ways
- EM is `strip().lower()` ONLY — final punctuation and internal spacing are
significant. Answers must carry the gold's punctuation conventions; a
format-inducer pass (§3) is now demonstrably worth EM points.
- Gold items can be a LIST OF ALTERNATIVES (14/160 rows); scorer takes max.
- chrF is sacrebleu with effective-order smoothing; my reimplementation now
matches to 0.0 delta over 500 fuzz cases (tests/test_scorer.py).
## 13. Linguini formats: what the parser must survive (now does, 158/160)
Pipe tables 2-5 columns with header rows; numbered example sentences whose
numbering the query CONTINUES (item "17." refers to nothing in the query);
bare-line items with no numbering; (k)-blank markers in any column, both
directions in one puzzle, including damaged rows ("(5) to tie" merged cell);
items living in the CONTEXT while the query is instruction-only
("Determine the correct correspondences", "Fill in the blanks (114)");
task-language strings that end in ':' (vowel length) — never use trailing
colon alone to detect instructions. Two rows remain unparseable-by-count: one
has misaligned gold (7 answers, 6 items), one enumerates payloads inline in
prose. Position-aligned scoring makes over/under-parsing cost only the
misaligned tail — worth a guard that pads rather than truncates.
## 14. match_letters needs morpheme-level CSP, not surface matching
16 puzzles / 223 items (25% of all items) are unordered form↔meaning
matching with usually ZERO attested pairs. Surface similarity carries no
signal (5.4% EM ≈ barely above random). The real structure: recurring
morphemes across forms must map consistently to recurring words across
meanings (Zuni doko:ko ↔ 'chicken', mo:chikwa ↔ 'peach'). That is a small
constraint-satisfaction / bilingual-lexicon-induction problem over sets —
highly verifiable (a candidate assignment implies a consistent lexicon or it
doesn't) and a perfect fit for the propose-verify architecture: LLM proposes
morpheme↔word hypotheses, a solver checks global consistency, Hungarian
finishes. My soft quadratic-assignment attempt (power iteration over
similarity graphs) was neutral; the discrete version is the right next try.
## 15. Throughput reality on the T4 (from the notebook)
The baseline runs ~1 min/problem at 1536 new tokens on the 14B AWQ — a full
160-puzzle set would need ~2.7 h sequential. The 30-minute budget therefore
REQUIRES the two-pass design: symbolic first (~9 s), then batched direct-LLM
on the low-confidence subset, weakest-first, with a hard budget cutoff
(script.py implements this). Concrete knobs to tune on Colab: batch size
(KV-cache limited on 16 GB with a 14B model — try 2-4), MAX_NEW_TOKENS
(1536 baseline; shorter for translation-only puzzles), and whether a 7B
model with bigger batches beats the 14B with tiny batches at fixed wall
clock. That ablation needs the actual T4.
## 16a. v1 migration to single-shot-with-tools (done 2026-07-21)
The architecture inversion recommended by the landscape survey is
implemented: the tools now feed the answering model (scaffold in the prompt:
segmentation, alignment, verified numeral values, labeled symbolic
candidates), verified symbolic output overrides the model, unverified output
is demoted to hints, and CEGIS left the runtime path. Explanations ride in
the same generation (`EXPLANATION:` after `FINAL ANSWERS:`) — the
human-judged track costs zero extra passes. Open empirical questions for the
T4: does the scaffold help a 14B AWQ model the way it helped GPT-4-class
models (the segmentation evidence says gains concentrate above ~15% baseline
— a 14B may be below it); and the deferred CoT-vs-IO prompt A/B (§ "Probing
LLMs", arXiv:2502.00817).
## 16b. First live submission: the T4 constraint is prefill logits, not KV
The first real submission OOMed: on the sandbox's transformers, `generate`
computes LM-head logits over EVERY prompt position and casts to float32 —
batch 6 × ~2.9k tokens × 152k vocab × 4B = the exact 9.82 GiB failed
allocation. KV cache (what batch sizing usually optimizes) is negligible
under GQA; **total prompt tokens per batch** is the real T4 memory bound at
~0.9 MB/token. Fixes: token-budget batch packing (3,500), OOM → retry
singly → abstain, `logits_to_keep=1` where supported. The deeper lesson:
the crash killed the script before submission.csv existed — "never empty"
must hold at the FILE level, not the item level. submission.csv is now
checkpoint-written from ~15 s in and atomically replaced as results improve.
## 16. The sandbox runs OLDER transformers than Colab
The notebook comments that `apply_chat_template(..., return_dict=True)`
returns a dict on Colab but a bare tensor in the sandbox. llm.py sidesteps
the whole class of drift by using `tokenize=False` + a normal tokenizer
call, which is stable across versions. Keep any future transformers usage on
the oldest-common-API path; there is no way to pin the sandbox's version.
## 1. The symbolic template translator may be the real workhorse, not the DSL
The plan positions LLM→DSL synthesis as the core translation engine. But the
minimal-pair **substitution translator** (`solver/template.py`: find the
attested sentence closest to the query, swap the differing tokens through
alignment links) solved every synthetic translation item exactly, with zero
LLM tokens, once three refinements landed:
- **competition / explaining-away in alignment** (`align.py`): demote a target
token already strongly claimed by another source token. Tiny corpora make
co-occurrence ties pervasive; this one change flipped several wrong answers.
- **morphological back-off**: single-word glosses (`moko = dog`) project into
inflected forms containing them (`namoko`).
- **affix-matched substitution**: when replacing a form, prefer the candidate
sharing an affix with the form replaced (kupu:nakupu :: moko:namoko) — a
proportional analogy inside the template frame.
Implication: budget the DSL/CEGIS path for the residual — items where the
query is far (bag-distance > ~len/2) from every attested sentence, long
compositional sentences, and phenomena the template can't reach (reordering,
harmony). Measure on real dev data what fraction that residual actually is;
it directly sets the LLM compute budget and the finetuning priority.
## 2. Verifier honesty requires two scoring regimes (found a real trap)
Memorizing predictors (template translator, fallback) score EM=1.0 on attested
pairs by construction, so scoring every candidate with plain `evaluate()` lets
memorization beat a generalizing grammar on ties — with MDL penalty it beats
it *always*. The fix in `router._pick_direction_solver`: fixed programs
(grammars) are scored directly (they must reproduce the data), while
fit-from-data predictors are scored **leave-one-out**. Any future candidate
added to the selection pool must declare which regime it belongs to. This also
matters for RL later: the reward must be LOO-style for anything with data
access, or the policy learns to memorize.
## 3. Answer-format induction is a first-class subproblem the plan ignores
The synthetic num_to_text case exposed it: for 23, `hun ox` (20+3) and
`kan hun ox` (1×20+3) are both arithmetically valid; the gold followed the
attested *style* (explicit unit multiplier). The fix was a style model scoring
candidates by consistency with attested phrasing. This generalizes: whenever
multiple surface answers are semantically correct, the scorer only rewards the
one matching the dataset's convention (capitalization, articles, hyphenation
of morph boundaries, multiplier explicitness). Recommend: an explicit
"format inducer" pass that learns per-puzzle answer conventions from context
examples, applied as the last stage of every task type. Cheap, and it converts
chrF-close answers into EM hits — which the geometric mean doubly rewards.
## 4. Context parsing is still the top empirical risk
Everything downstream consumes `extract_pairs`. The current parser
(separator voting + per-line fallback) handles the formats I could construct,
but real Linguini CSVs were not available in this environment. **Before any
model work, pull the actual Linguini/PuzzLing dev CSVs and fuzz the parser
against every context in them**; count rows where zero pairs are extracted —
each such row is guaranteed near-zero score. A cheap LLM repair path (ask the
model to emit the pair list as JSON when the symbolic parser yields < 2 pairs)
would cap this risk at one batched call per malformed puzzle.
## 5. Direction detection deserves attested-data validation, not just wording
`detect_direction` keys off phrases like "into English". Real queries may
say "What does X mean?", "Give the form meaning ...", or nothing explicit.
A robust cross-check: try both directions and see which side of the attested
pairs the query string is *script-similar* to (character n-gram overlap with
task-language vs work-language material). If the query looks like the unknown
language, it's an analysis item. Cheap and language-agnostic.
## 6. CEGIS should get verifier-guided repair hints, not just failures
The refine prompt currently shows failing (input, expected, got) triples. The
symbolic layer knows more: which morpheme boundary the output diverged at,
which affix went unused, whether the failure is pure reordering. Feeding a
one-line diagnosis per failure ("output differs only in word order",
"expected form contains attested morph 'na-' that your grammar never
attaches") should cut CEGIS rounds — worth an ablation axis in phase 2/3.
## 7. Self-consistency has a free implementation via round-trip
The interpreter runs both directions, so `analyze(generate(x)) == x` is a
zero-cost consistency check usable as (a) a tie-breaker in grammar selection
(already anticipated in AGENTS.md), and (b) a confidence signal for budget
allocation — spend refinement rounds on items whose round-trip fails.
## 8. Points weighting has no data column — resolve early
The scorer supports per-item weights, but the Linguini schema
(id/context/query/...) carries no points column. If official IOL point values
weight the leaderboard metric, they must be embedded somewhere (query text?
separate mapping by id?). Resolve this with a probe submission early — it
changes budget allocation (high-point items deserve the CEGIS rounds).
## 9. Local scorer vs sacrebleu chrF: verify parity once, in the eval image
`eval/scorer.py` implements chrF2 (n≤6, β=2, whitespace stripped) by hand to
stay dependency-free. sacrebleu differs in epsilon smoothing on zero-match
orders. Before trusting local ablations, run both on a few hundred string
pairs in the submission container and confirm the delta is < 1e-3; otherwise
model selection could silently optimize the wrong metric.
## 10. Prompt library format drifted from the plan (deliberately)
AGENTS.md says YAML prompts; the repo uses plain `.md` templates with
`str.format` placeholders (`prompts/*.md`). Rationale: zero dependencies, no
YAML-escaping pain with multiline linguistic data. If config-driven prompt
variants are needed for best-of-N diversity (the greedy-decoding constraint
means diversity must come from prompts, not sampling temperature — see
`synth.synthesize_best_of_n`), add a small variants list per template rather
than reintroducing YAML.