--- license: other license_name: weft-research license_link: https://github.com/wallfacers/data-weave pipeline_tag: text-generation library_name: transformers base_model: Qwen/Qwen2.5-Coder-0.5B-Instruct datasets: - wallfacers/weft-script-lineage-synth language: - en tags: - research-artifact - negative-result - memorization - domain-shift - data-lineage - etl - lora model-index: - name: weft-lineage-extractor-0.5b results: - task: type: table-level-lineage-extraction name: ETL table-level data-lineage extraction dataset: type: synthetic-etl name: synthetic held-out (structural-form isolated) metrics: - type: precision value: 0.994 name: Table precision (synthetic held-out) - task: type: table-level-lineage-extraction name: ETL table-level data-lineage extraction dataset: type: real-github-etl name: real GitHub ETL (human gold) metrics: - type: precision value: 0.243 name: Table precision (real, out-of-distribution) --- # weft-lineage-extractor-0.5b — smallest scale point of a NEGATIVE RESULT > ## ⚠️ RESEARCH ARTIFACT — the 0.5B point of a *synthetic-only* training study. Not a production tool. > > ### ✅ Resolved by real data: use **[weft-lineage-extractor-3b](https://huggingface.co/wallfacers/weft-lineage-extractor-3b)** (real corpus, real precision **0.64**). Full study: **[weft-lineage-extractor-1.5b](https://huggingface.co/wallfacers/weft-lineage-extractor-1.5b)**. The **0.5B** point of a study showing that **synthetic-only training** induces a **verbatim memorization leak** in small models for ETL table-lineage extraction. It is the smallest scale point and shows the **worst** leak: near-perfect synthetic precision (0.994) collapses to **0.243** on real GitHub scripts, with **37.4%** of hallucinations being table names recited verbatim from the synthetic training pool. Same recipe as the [1.5B main model](https://huggingface.co/wallfacers/weft-lineage-extractor-1.5b) (LoRA on Qwen2.5-Coder-Instruct, Python/Shell synthetic ETL scripts, zero real scripts); only the base size differs. ## This point's numbers (table-level, Convention A) | metric | synthetic held-out | real GitHub ETL | |---|---|---| | precision | 0.994 | **0.243** | | direction accuracy | — | **0.369** | | **verbatim memorization leak** | — | **37.4%** | ## Where it fits (scale curve) | scale | real precision | real direction | verbatim leak | |---|---|---|---| | **0.5B (this)** | **0.243** | **0.369** | **37.4%** | | 1.5B (main) | 0.270 | 0.496 | 22.4% | | 3B (synthetic) | 0.325 | 0.468 | 10.9% | | **3B (real corpus)** | **0.64** | — | **~0** | Leak shrinks with scale (capacity), but only **real training data** closes the real-world gap (bottom row). Direction confusion does not improve with size. ## Intended use - ✅ Reproducing / studying the synthetic-training memorization-leak failure at minimal scale. - ❌ Not for production lineage — use the [real-corpus 3B](https://huggingface.co/wallfacers/weft-lineage-extractor-3b). ## Usage, prompt format, training details, citation Identical to the main model (this variant uses `task_type: PYTHON | SHELL`; everything else the same): **[weft-lineage-extractor-1.5b](https://huggingface.co/wallfacers/weft-lineage-extractor-1.5b)**. - **Dataset & reports:** [wallfacers/weft-script-lineage-synth](https://huggingface.co/datasets/wallfacers/weft-script-lineage-synth) - **Base model:** [Qwen/Qwen2.5-Coder-0.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-0.5B-Instruct) - **Platform:** [Weft (data-weave)](https://github.com/wallfacers/data-weave)