219 lines
12 KiB
Markdown
219 lines
12 KiB
Markdown
# fable.md — research-direction notes from implementation
|
||
|
||
Notes written while building the IOL-AI solver seed (see AGENTS.md for the plan,
|
||
README.md for repo state). Each item is something the implementation surfaced
|
||
that the research plan should absorb.
|
||
|
||
**Update (real-data phase):** items 4, 8, 9 are now RESOLVED — the official
|
||
eval notebook (rita-berrada/iolai-2026-workshop) fixed the semantics, and
|
||
facebook/linguini gave 160 real dev puzzles. New findings start at §11.
|
||
|
||
## 11. Real-data baseline resets expectations (and the plan's emphasis)
|
||
|
||
Symbolic-only on all 160 Linguini puzzles: **EM 2.4%, chrF 20, in ~9 s.**
|
||
(The synthetic dev set scored 1.0 — constructed puzzles flatter substitution
|
||
methods.) Real translation items are compositional, real paradigms involve
|
||
metathesis/harmony/infixation, real numeral systems have morphophonology that
|
||
breaks token-level CSP. Consequences:
|
||
- The LLM direct path (notebook format) is the main score carrier for now;
|
||
the symbolic stack's near-term value is (a) exact hits where structure is
|
||
clean, (b) the chrF floor, (c) **verified confidences that route the LLM
|
||
budget** (157/160 puzzles flagged low-confidence — correctly).
|
||
- The DSL/CEGIS path is the ceiling-raiser, not the floor: its job is to
|
||
convert LLM linguistic insight into verified exact hits. Ablate it against
|
||
direct-LLM on the T4 before investing in finetuning.
|
||
|
||
## 12. Official semantics differ from my assumptions in scoring-relevant ways
|
||
|
||
- EM is `strip().lower()` ONLY — final punctuation and internal spacing are
|
||
significant. Answers must carry the gold's punctuation conventions; a
|
||
format-inducer pass (§3) is now demonstrably worth EM points.
|
||
- Gold items can be a LIST OF ALTERNATIVES (14/160 rows); scorer takes max.
|
||
- chrF is sacrebleu with effective-order smoothing; my reimplementation now
|
||
matches to 0.0 delta over 500 fuzz cases (tests/test_scorer.py).
|
||
|
||
## 13. Linguini formats: what the parser must survive (now does, 158/160)
|
||
|
||
Pipe tables 2-5 columns with header rows; numbered example sentences whose
|
||
numbering the query CONTINUES (item "17." refers to nothing in the query);
|
||
bare-line items with no numbering; (k)-blank markers in any column, both
|
||
directions in one puzzle, including damaged rows ("(5) to tie" merged cell);
|
||
items living in the CONTEXT while the query is instruction-only
|
||
("Determine the correct correspondences", "Fill in the blanks (1–14)");
|
||
task-language strings that end in ':' (vowel length) — never use trailing
|
||
colon alone to detect instructions. Two rows remain unparseable-by-count: one
|
||
has misaligned gold (7 answers, 6 items), one enumerates payloads inline in
|
||
prose. Position-aligned scoring makes over/under-parsing cost only the
|
||
misaligned tail — worth a guard that pads rather than truncates.
|
||
|
||
## 14. match_letters needs morpheme-level CSP, not surface matching
|
||
|
||
16 puzzles / 223 items (25% of all items) are unordered form↔meaning
|
||
matching with usually ZERO attested pairs. Surface similarity carries no
|
||
signal (5.4% EM ≈ barely above random). The real structure: recurring
|
||
morphemes across forms must map consistently to recurring words across
|
||
meanings (Zuni doko:ko ↔ 'chicken', mo:chikwa ↔ 'peach'). That is a small
|
||
constraint-satisfaction / bilingual-lexicon-induction problem over sets —
|
||
highly verifiable (a candidate assignment implies a consistent lexicon or it
|
||
doesn't) and a perfect fit for the propose-verify architecture: LLM proposes
|
||
morpheme↔word hypotheses, a solver checks global consistency, Hungarian
|
||
finishes. My soft quadratic-assignment attempt (power iteration over
|
||
similarity graphs) was neutral; the discrete version is the right next try.
|
||
|
||
## 15. Throughput reality on the T4 (from the notebook)
|
||
|
||
The baseline runs ~1 min/problem at 1536 new tokens on the 14B AWQ — a full
|
||
160-puzzle set would need ~2.7 h sequential. The 30-minute budget therefore
|
||
REQUIRES the two-pass design: symbolic first (~9 s), then batched direct-LLM
|
||
on the low-confidence subset, weakest-first, with a hard budget cutoff
|
||
(script.py implements this). Concrete knobs to tune on Colab: batch size
|
||
(KV-cache limited on 16 GB with a 14B model — try 2-4), MAX_NEW_TOKENS
|
||
(1536 baseline; shorter for translation-only puzzles), and whether a 7B
|
||
model with bigger batches beats the 14B with tiny batches at fixed wall
|
||
clock. That ablation needs the actual T4.
|
||
|
||
## 16a. v1 migration to single-shot-with-tools (done 2026-07-21)
|
||
|
||
The architecture inversion recommended by the landscape survey is
|
||
implemented: the tools now feed the answering model (scaffold in the prompt:
|
||
segmentation, alignment, verified numeral values, labeled symbolic
|
||
candidates), verified symbolic output overrides the model, unverified output
|
||
is demoted to hints, and CEGIS left the runtime path. Explanations ride in
|
||
the same generation (`EXPLANATION:` after `FINAL ANSWERS:`) — the
|
||
human-judged track costs zero extra passes. Open empirical questions for the
|
||
T4: does the scaffold help a 14B AWQ model the way it helped GPT-4-class
|
||
models (the segmentation evidence says gains concentrate above ~15% baseline
|
||
— a 14B may be below it); and the deferred CoT-vs-IO prompt A/B (§ "Probing
|
||
LLMs", arXiv:2502.00817).
|
||
|
||
## 16b. First live submission: the T4 constraint is prefill logits, not KV
|
||
|
||
The first real submission OOMed: on the sandbox's transformers, `generate`
|
||
computes LM-head logits over EVERY prompt position and casts to float32 —
|
||
batch 6 × ~2.9k tokens × 152k vocab × 4B = the exact 9.82 GiB failed
|
||
allocation. KV cache (what batch sizing usually optimizes) is negligible
|
||
under GQA; **total prompt tokens per batch** is the real T4 memory bound at
|
||
~0.9 MB/token. Fixes: token-budget batch packing (3,500), OOM → retry
|
||
singly → abstain, `logits_to_keep=1` where supported. The deeper lesson:
|
||
the crash killed the script before submission.csv existed — "never empty"
|
||
must hold at the FILE level, not the item level. submission.csv is now
|
||
checkpoint-written from ~15 s in and atomically replaced as results improve.
|
||
|
||
## 16. The sandbox runs OLDER transformers than Colab
|
||
|
||
The notebook comments that `apply_chat_template(..., return_dict=True)`
|
||
returns a dict on Colab but a bare tensor in the sandbox. llm.py sidesteps
|
||
the whole class of drift by using `tokenize=False` + a normal tokenizer
|
||
call, which is stable across versions. Keep any future transformers usage on
|
||
the oldest-common-API path; there is no way to pin the sandbox's version.
|
||
|
||
## 1. The symbolic template translator may be the real workhorse, not the DSL
|
||
|
||
The plan positions LLM→DSL synthesis as the core translation engine. But the
|
||
minimal-pair **substitution translator** (`solver/template.py`: find the
|
||
attested sentence closest to the query, swap the differing tokens through
|
||
alignment links) solved every synthetic translation item exactly, with zero
|
||
LLM tokens, once three refinements landed:
|
||
|
||
- **competition / explaining-away in alignment** (`align.py`): demote a target
|
||
token already strongly claimed by another source token. Tiny corpora make
|
||
co-occurrence ties pervasive; this one change flipped several wrong answers.
|
||
- **morphological back-off**: single-word glosses (`moko = dog`) project into
|
||
inflected forms containing them (`namoko`).
|
||
- **affix-matched substitution**: when replacing a form, prefer the candidate
|
||
sharing an affix with the form replaced (kupu:nakupu :: moko:namoko) — a
|
||
proportional analogy inside the template frame.
|
||
|
||
Implication: budget the DSL/CEGIS path for the residual — items where the
|
||
query is far (bag-distance > ~len/2) from every attested sentence, long
|
||
compositional sentences, and phenomena the template can't reach (reordering,
|
||
harmony). Measure on real dev data what fraction that residual actually is;
|
||
it directly sets the LLM compute budget and the finetuning priority.
|
||
|
||
## 2. Verifier honesty requires two scoring regimes (found a real trap)
|
||
|
||
Memorizing predictors (template translator, fallback) score EM=1.0 on attested
|
||
pairs by construction, so scoring every candidate with plain `evaluate()` lets
|
||
memorization beat a generalizing grammar on ties — with MDL penalty it beats
|
||
it *always*. The fix in `router._pick_direction_solver`: fixed programs
|
||
(grammars) are scored directly (they must reproduce the data), while
|
||
fit-from-data predictors are scored **leave-one-out**. Any future candidate
|
||
added to the selection pool must declare which regime it belongs to. This also
|
||
matters for RL later: the reward must be LOO-style for anything with data
|
||
access, or the policy learns to memorize.
|
||
|
||
## 3. Answer-format induction is a first-class subproblem the plan ignores
|
||
|
||
The synthetic num_to_text case exposed it: for 23, `hun ox` (20+3) and
|
||
`kan hun ox` (1×20+3) are both arithmetically valid; the gold followed the
|
||
attested *style* (explicit unit multiplier). The fix was a style model scoring
|
||
candidates by consistency with attested phrasing. This generalizes: whenever
|
||
multiple surface answers are semantically correct, the scorer only rewards the
|
||
one matching the dataset's convention (capitalization, articles, hyphenation
|
||
of morph boundaries, multiplier explicitness). Recommend: an explicit
|
||
"format inducer" pass that learns per-puzzle answer conventions from context
|
||
examples, applied as the last stage of every task type. Cheap, and it converts
|
||
chrF-close answers into EM hits — which the geometric mean doubly rewards.
|
||
|
||
## 4. Context parsing is still the top empirical risk
|
||
|
||
Everything downstream consumes `extract_pairs`. The current parser
|
||
(separator voting + per-line fallback) handles the formats I could construct,
|
||
but real Linguini CSVs were not available in this environment. **Before any
|
||
model work, pull the actual Linguini/PuzzLing dev CSVs and fuzz the parser
|
||
against every context in them**; count rows where zero pairs are extracted —
|
||
each such row is guaranteed near-zero score. A cheap LLM repair path (ask the
|
||
model to emit the pair list as JSON when the symbolic parser yields < 2 pairs)
|
||
would cap this risk at one batched call per malformed puzzle.
|
||
|
||
## 5. Direction detection deserves attested-data validation, not just wording
|
||
|
||
`detect_direction` keys off phrases like "into English". Real queries may
|
||
say "What does X mean?", "Give the form meaning ...", or nothing explicit.
|
||
A robust cross-check: try both directions and see which side of the attested
|
||
pairs the query string is *script-similar* to (character n-gram overlap with
|
||
task-language vs work-language material). If the query looks like the unknown
|
||
language, it's an analysis item. Cheap and language-agnostic.
|
||
|
||
## 6. CEGIS should get verifier-guided repair hints, not just failures
|
||
|
||
The refine prompt currently shows failing (input, expected, got) triples. The
|
||
symbolic layer knows more: which morpheme boundary the output diverged at,
|
||
which affix went unused, whether the failure is pure reordering. Feeding a
|
||
one-line diagnosis per failure ("output differs only in word order",
|
||
"expected form contains attested morph 'na-' that your grammar never
|
||
attaches") should cut CEGIS rounds — worth an ablation axis in phase 2/3.
|
||
|
||
## 7. Self-consistency has a free implementation via round-trip
|
||
|
||
The interpreter runs both directions, so `analyze(generate(x)) == x` is a
|
||
zero-cost consistency check usable as (a) a tie-breaker in grammar selection
|
||
(already anticipated in AGENTS.md), and (b) a confidence signal for budget
|
||
allocation — spend refinement rounds on items whose round-trip fails.
|
||
|
||
## 8. Points weighting has no data column — resolve early
|
||
|
||
The scorer supports per-item weights, but the Linguini schema
|
||
(id/context/query/...) carries no points column. If official IOL point values
|
||
weight the leaderboard metric, they must be embedded somewhere (query text?
|
||
separate mapping by id?). Resolve this with a probe submission early — it
|
||
changes budget allocation (high-point items deserve the CEGIS rounds).
|
||
|
||
## 9. Local scorer vs sacrebleu chrF: verify parity once, in the eval image
|
||
|
||
`eval/scorer.py` implements chrF2 (n≤6, β=2, whitespace stripped) by hand to
|
||
stay dependency-free. sacrebleu differs in epsilon smoothing on zero-match
|
||
orders. Before trusting local ablations, run both on a few hundred string
|
||
pairs in the submission container and confirm the delta is < 1e-3; otherwise
|
||
model selection could silently optimize the wrong metric.
|
||
|
||
## 10. Prompt library format drifted from the plan (deliberately)
|
||
|
||
AGENTS.md says YAML prompts; the repo uses plain `.md` templates with
|
||
`str.format` placeholders (`prompts/*.md`). Rationale: zero dependencies, no
|
||
YAML-escaping pain with multiline linguistic data. If config-driven prompt
|
||
variants are needed for best-of-N diversity (the greedy-decoding constraint
|
||
means diversity must come from prompts, not sampling temperature — see
|
||
`synth.synthesize_best_of_n`), add a small variants list per template rather
|
||
than reintroducing YAML.
|