初始化项目,由ModelHub XC社区提供模型
Model: rpant/iolai26-solve Source: Original Platform
This commit is contained in:
218
fable.md
Normal file
218
fable.md
Normal file
@@ -0,0 +1,218 @@
|
||||
# fable.md — research-direction notes from implementation
|
||||
|
||||
Notes written while building the IOL-AI solver seed (see AGENTS.md for the plan,
|
||||
README.md for repo state). Each item is something the implementation surfaced
|
||||
that the research plan should absorb.
|
||||
|
||||
**Update (real-data phase):** items 4, 8, 9 are now RESOLVED — the official
|
||||
eval notebook (rita-berrada/iolai-2026-workshop) fixed the semantics, and
|
||||
facebook/linguini gave 160 real dev puzzles. New findings start at §11.
|
||||
|
||||
## 11. Real-data baseline resets expectations (and the plan's emphasis)
|
||||
|
||||
Symbolic-only on all 160 Linguini puzzles: **EM 2.4%, chrF 20, in ~9 s.**
|
||||
(The synthetic dev set scored 1.0 — constructed puzzles flatter substitution
|
||||
methods.) Real translation items are compositional, real paradigms involve
|
||||
metathesis/harmony/infixation, real numeral systems have morphophonology that
|
||||
breaks token-level CSP. Consequences:
|
||||
- The LLM direct path (notebook format) is the main score carrier for now;
|
||||
the symbolic stack's near-term value is (a) exact hits where structure is
|
||||
clean, (b) the chrF floor, (c) **verified confidences that route the LLM
|
||||
budget** (157/160 puzzles flagged low-confidence — correctly).
|
||||
- The DSL/CEGIS path is the ceiling-raiser, not the floor: its job is to
|
||||
convert LLM linguistic insight into verified exact hits. Ablate it against
|
||||
direct-LLM on the T4 before investing in finetuning.
|
||||
|
||||
## 12. Official semantics differ from my assumptions in scoring-relevant ways
|
||||
|
||||
- EM is `strip().lower()` ONLY — final punctuation and internal spacing are
|
||||
significant. Answers must carry the gold's punctuation conventions; a
|
||||
format-inducer pass (§3) is now demonstrably worth EM points.
|
||||
- Gold items can be a LIST OF ALTERNATIVES (14/160 rows); scorer takes max.
|
||||
- chrF is sacrebleu with effective-order smoothing; my reimplementation now
|
||||
matches to 0.0 delta over 500 fuzz cases (tests/test_scorer.py).
|
||||
|
||||
## 13. Linguini formats: what the parser must survive (now does, 158/160)
|
||||
|
||||
Pipe tables 2-5 columns with header rows; numbered example sentences whose
|
||||
numbering the query CONTINUES (item "17." refers to nothing in the query);
|
||||
bare-line items with no numbering; (k)-blank markers in any column, both
|
||||
directions in one puzzle, including damaged rows ("(5) to tie" merged cell);
|
||||
items living in the CONTEXT while the query is instruction-only
|
||||
("Determine the correct correspondences", "Fill in the blanks (1–14)");
|
||||
task-language strings that end in ':' (vowel length) — never use trailing
|
||||
colon alone to detect instructions. Two rows remain unparseable-by-count: one
|
||||
has misaligned gold (7 answers, 6 items), one enumerates payloads inline in
|
||||
prose. Position-aligned scoring makes over/under-parsing cost only the
|
||||
misaligned tail — worth a guard that pads rather than truncates.
|
||||
|
||||
## 14. match_letters needs morpheme-level CSP, not surface matching
|
||||
|
||||
16 puzzles / 223 items (25% of all items) are unordered form↔meaning
|
||||
matching with usually ZERO attested pairs. Surface similarity carries no
|
||||
signal (5.4% EM ≈ barely above random). The real structure: recurring
|
||||
morphemes across forms must map consistently to recurring words across
|
||||
meanings (Zuni doko:ko ↔ 'chicken', mo:chikwa ↔ 'peach'). That is a small
|
||||
constraint-satisfaction / bilingual-lexicon-induction problem over sets —
|
||||
highly verifiable (a candidate assignment implies a consistent lexicon or it
|
||||
doesn't) and a perfect fit for the propose-verify architecture: LLM proposes
|
||||
morpheme↔word hypotheses, a solver checks global consistency, Hungarian
|
||||
finishes. My soft quadratic-assignment attempt (power iteration over
|
||||
similarity graphs) was neutral; the discrete version is the right next try.
|
||||
|
||||
## 15. Throughput reality on the T4 (from the notebook)
|
||||
|
||||
The baseline runs ~1 min/problem at 1536 new tokens on the 14B AWQ — a full
|
||||
160-puzzle set would need ~2.7 h sequential. The 30-minute budget therefore
|
||||
REQUIRES the two-pass design: symbolic first (~9 s), then batched direct-LLM
|
||||
on the low-confidence subset, weakest-first, with a hard budget cutoff
|
||||
(script.py implements this). Concrete knobs to tune on Colab: batch size
|
||||
(KV-cache limited on 16 GB with a 14B model — try 2-4), MAX_NEW_TOKENS
|
||||
(1536 baseline; shorter for translation-only puzzles), and whether a 7B
|
||||
model with bigger batches beats the 14B with tiny batches at fixed wall
|
||||
clock. That ablation needs the actual T4.
|
||||
|
||||
## 16a. v1 migration to single-shot-with-tools (done 2026-07-21)
|
||||
|
||||
The architecture inversion recommended by the landscape survey is
|
||||
implemented: the tools now feed the answering model (scaffold in the prompt:
|
||||
segmentation, alignment, verified numeral values, labeled symbolic
|
||||
candidates), verified symbolic output overrides the model, unverified output
|
||||
is demoted to hints, and CEGIS left the runtime path. Explanations ride in
|
||||
the same generation (`EXPLANATION:` after `FINAL ANSWERS:`) — the
|
||||
human-judged track costs zero extra passes. Open empirical questions for the
|
||||
T4: does the scaffold help a 14B AWQ model the way it helped GPT-4-class
|
||||
models (the segmentation evidence says gains concentrate above ~15% baseline
|
||||
— a 14B may be below it); and the deferred CoT-vs-IO prompt A/B (§ "Probing
|
||||
LLMs", arXiv:2502.00817).
|
||||
|
||||
## 16b. First live submission: the T4 constraint is prefill logits, not KV
|
||||
|
||||
The first real submission OOMed: on the sandbox's transformers, `generate`
|
||||
computes LM-head logits over EVERY prompt position and casts to float32 —
|
||||
batch 6 × ~2.9k tokens × 152k vocab × 4B = the exact 9.82 GiB failed
|
||||
allocation. KV cache (what batch sizing usually optimizes) is negligible
|
||||
under GQA; **total prompt tokens per batch** is the real T4 memory bound at
|
||||
~0.9 MB/token. Fixes: token-budget batch packing (3,500), OOM → retry
|
||||
singly → abstain, `logits_to_keep=1` where supported. The deeper lesson:
|
||||
the crash killed the script before submission.csv existed — "never empty"
|
||||
must hold at the FILE level, not the item level. submission.csv is now
|
||||
checkpoint-written from ~15 s in and atomically replaced as results improve.
|
||||
|
||||
## 16. The sandbox runs OLDER transformers than Colab
|
||||
|
||||
The notebook comments that `apply_chat_template(..., return_dict=True)`
|
||||
returns a dict on Colab but a bare tensor in the sandbox. llm.py sidesteps
|
||||
the whole class of drift by using `tokenize=False` + a normal tokenizer
|
||||
call, which is stable across versions. Keep any future transformers usage on
|
||||
the oldest-common-API path; there is no way to pin the sandbox's version.
|
||||
|
||||
## 1. The symbolic template translator may be the real workhorse, not the DSL
|
||||
|
||||
The plan positions LLM→DSL synthesis as the core translation engine. But the
|
||||
minimal-pair **substitution translator** (`solver/template.py`: find the
|
||||
attested sentence closest to the query, swap the differing tokens through
|
||||
alignment links) solved every synthetic translation item exactly, with zero
|
||||
LLM tokens, once three refinements landed:
|
||||
|
||||
- **competition / explaining-away in alignment** (`align.py`): demote a target
|
||||
token already strongly claimed by another source token. Tiny corpora make
|
||||
co-occurrence ties pervasive; this one change flipped several wrong answers.
|
||||
- **morphological back-off**: single-word glosses (`moko = dog`) project into
|
||||
inflected forms containing them (`namoko`).
|
||||
- **affix-matched substitution**: when replacing a form, prefer the candidate
|
||||
sharing an affix with the form replaced (kupu:nakupu :: moko:namoko) — a
|
||||
proportional analogy inside the template frame.
|
||||
|
||||
Implication: budget the DSL/CEGIS path for the residual — items where the
|
||||
query is far (bag-distance > ~len/2) from every attested sentence, long
|
||||
compositional sentences, and phenomena the template can't reach (reordering,
|
||||
harmony). Measure on real dev data what fraction that residual actually is;
|
||||
it directly sets the LLM compute budget and the finetuning priority.
|
||||
|
||||
## 2. Verifier honesty requires two scoring regimes (found a real trap)
|
||||
|
||||
Memorizing predictors (template translator, fallback) score EM=1.0 on attested
|
||||
pairs by construction, so scoring every candidate with plain `evaluate()` lets
|
||||
memorization beat a generalizing grammar on ties — with MDL penalty it beats
|
||||
it *always*. The fix in `router._pick_direction_solver`: fixed programs
|
||||
(grammars) are scored directly (they must reproduce the data), while
|
||||
fit-from-data predictors are scored **leave-one-out**. Any future candidate
|
||||
added to the selection pool must declare which regime it belongs to. This also
|
||||
matters for RL later: the reward must be LOO-style for anything with data
|
||||
access, or the policy learns to memorize.
|
||||
|
||||
## 3. Answer-format induction is a first-class subproblem the plan ignores
|
||||
|
||||
The synthetic num_to_text case exposed it: for 23, `hun ox` (20+3) and
|
||||
`kan hun ox` (1×20+3) are both arithmetically valid; the gold followed the
|
||||
attested *style* (explicit unit multiplier). The fix was a style model scoring
|
||||
candidates by consistency with attested phrasing. This generalizes: whenever
|
||||
multiple surface answers are semantically correct, the scorer only rewards the
|
||||
one matching the dataset's convention (capitalization, articles, hyphenation
|
||||
of morph boundaries, multiplier explicitness). Recommend: an explicit
|
||||
"format inducer" pass that learns per-puzzle answer conventions from context
|
||||
examples, applied as the last stage of every task type. Cheap, and it converts
|
||||
chrF-close answers into EM hits — which the geometric mean doubly rewards.
|
||||
|
||||
## 4. Context parsing is still the top empirical risk
|
||||
|
||||
Everything downstream consumes `extract_pairs`. The current parser
|
||||
(separator voting + per-line fallback) handles the formats I could construct,
|
||||
but real Linguini CSVs were not available in this environment. **Before any
|
||||
model work, pull the actual Linguini/PuzzLing dev CSVs and fuzz the parser
|
||||
against every context in them**; count rows where zero pairs are extracted —
|
||||
each such row is guaranteed near-zero score. A cheap LLM repair path (ask the
|
||||
model to emit the pair list as JSON when the symbolic parser yields < 2 pairs)
|
||||
would cap this risk at one batched call per malformed puzzle.
|
||||
|
||||
## 5. Direction detection deserves attested-data validation, not just wording
|
||||
|
||||
`detect_direction` keys off phrases like "into English". Real queries may
|
||||
say "What does X mean?", "Give the form meaning ...", or nothing explicit.
|
||||
A robust cross-check: try both directions and see which side of the attested
|
||||
pairs the query string is *script-similar* to (character n-gram overlap with
|
||||
task-language vs work-language material). If the query looks like the unknown
|
||||
language, it's an analysis item. Cheap and language-agnostic.
|
||||
|
||||
## 6. CEGIS should get verifier-guided repair hints, not just failures
|
||||
|
||||
The refine prompt currently shows failing (input, expected, got) triples. The
|
||||
symbolic layer knows more: which morpheme boundary the output diverged at,
|
||||
which affix went unused, whether the failure is pure reordering. Feeding a
|
||||
one-line diagnosis per failure ("output differs only in word order",
|
||||
"expected form contains attested morph 'na-' that your grammar never
|
||||
attaches") should cut CEGIS rounds — worth an ablation axis in phase 2/3.
|
||||
|
||||
## 7. Self-consistency has a free implementation via round-trip
|
||||
|
||||
The interpreter runs both directions, so `analyze(generate(x)) == x` is a
|
||||
zero-cost consistency check usable as (a) a tie-breaker in grammar selection
|
||||
(already anticipated in AGENTS.md), and (b) a confidence signal for budget
|
||||
allocation — spend refinement rounds on items whose round-trip fails.
|
||||
|
||||
## 8. Points weighting has no data column — resolve early
|
||||
|
||||
The scorer supports per-item weights, but the Linguini schema
|
||||
(id/context/query/...) carries no points column. If official IOL point values
|
||||
weight the leaderboard metric, they must be embedded somewhere (query text?
|
||||
separate mapping by id?). Resolve this with a probe submission early — it
|
||||
changes budget allocation (high-point items deserve the CEGIS rounds).
|
||||
|
||||
## 9. Local scorer vs sacrebleu chrF: verify parity once, in the eval image
|
||||
|
||||
`eval/scorer.py` implements chrF2 (n≤6, β=2, whitespace stripped) by hand to
|
||||
stay dependency-free. sacrebleu differs in epsilon smoothing on zero-match
|
||||
orders. Before trusting local ablations, run both on a few hundred string
|
||||
pairs in the submission container and confirm the delta is < 1e-3; otherwise
|
||||
model selection could silently optimize the wrong metric.
|
||||
|
||||
## 10. Prompt library format drifted from the plan (deliberately)
|
||||
|
||||
AGENTS.md says YAML prompts; the repo uses plain `.md` templates with
|
||||
`str.format` placeholders (`prompts/*.md`). Rationale: zero dependencies, no
|
||||
YAML-escaping pain with multiline linguistic data. If config-driven prompt
|
||||
variants are needed for best-of-N diversity (the greedy-decoding constraint
|
||||
means diversity must come from prompts, not sampling temperature — see
|
||||
`synth.synthesize_best_of_n`), add a small variants list per template rather
|
||||
than reintroducing YAML.
|
||||
Reference in New Issue
Block a user