--- license: other license_name: weft-research license_link: https://github.com/wallfacers/data-weave pipeline_tag: text-generation library_name: transformers base_model: Qwen/Qwen2.5-Coder-1.5B-Instruct datasets: - wallfacers/weft-script-lineage-synth language: - en tags: - research-artifact - negative-result - memorization - domain-shift - cross-language - data-lineage - etl - lora model-index: - name: weft-lineage-extractor-jvm-1.5b results: - task: type: table-level-lineage-extraction name: ETL table-level data-lineage extraction (Scala/Java) dataset: type: real-jvm-etl name: real Scala/Java ETL (human gold, n=141) metrics: - type: precision value: 0.165 name: Table precision (real JVM, out-of-distribution) - type: accuracy value: 0.418 name: Read/write direction accuracy (real JVM) --- # weft-lineage-extractor-jvm-1.5b — cross-language reproduction of a NEGATIVE RESULT > ## ⚠️ RESEARCH ARTIFACT — the Scala/Java (JVM) point of a *synthetic-only* training study. Not a production tool. > > ### ✅ Resolved by real data: use **[weft-lineage-extractor-3b](https://huggingface.co/wallfacers/weft-lineage-extractor-3b)** (real corpus, real precision **0.64**). Full study: **[weft-lineage-extractor-1.5b](https://huggingface.co/wallfacers/weft-lineage-extractor-1.5b)**. The **JVM (Scala/Java)** point of a study showing that **synthetic-only training** induces a **verbatim memorization leak** in small ETL table-lineage models — and that the failure **reproduces, worse, across languages**. Synthetic held-out looks near-perfect (~0.99); on real Scala/Java ETL it drops to **precision 0.165**, with **40.8%** of hallucinations being table names recited verbatim from the synthetic training pool — the **highest leak** of all scale/language points, because real JVM code is the most out-of-distribution. Same recipe as the [1.5B main model](https://huggingface.co/wallfacers/weft-lineage-extractor-1.5b) (LoRA on Qwen2.5-Coder-1.5B), extended with synthetic Scala/Java forms; `task_type: SCALA | JAVA`. ## This point's numbers (table-level, Convention A) | metric | synthetic held-out | real Scala/Java ETL (n=141, non-empty 28) | |---|---|---| | precision | ~0.99 | **0.165** | | direction accuracy | — | **0.418** | | **verbatim memorization leak** | — | **40.8%** | The leak metric is **gold-independent** and stable to full-file re-adjudication (verbatim 40.4% → 40.8%). Given a real `SqlDao.java` it cannot parse, the model emits training-pool names like `dws.dws_member_point_df` — memorized generator vocabulary, not a label artifact. ## Where it fits (scale / cross-language curve) | point | real precision | real direction | verbatim leak | |---|---|---|---| | 0.5B (Py/Sh) | 0.243 | 0.369 | 37.4% | | 1.5B (Py/Sh) | 0.270 | 0.496 | 22.4% | | 3B (Py/Sh, synthetic) | 0.325 | 0.468 | 10.9% | | **1.5B + JVM (this, real JVM)** | **0.165** | **0.418** | **40.8%** | | **3B (real corpus)** | **0.64** | — | **~0** | "Add another language" does not fix the real-world gap; **real training data does**. ## Intended use - ✅ Reproducing / studying the cross-language reproduction of the memorization-leak failure. - ❌ Not for production lineage — use the [real-corpus 3B](https://huggingface.co/wallfacers/weft-lineage-extractor-3b). ## Usage, prompt format, training details, citation Identical to the main model (this variant adds `task_type: SCALA | JAVA` and JVM system-prompt wording): **[weft-lineage-extractor-1.5b](https://huggingface.co/wallfacers/weft-lineage-extractor-1.5b)**. - **Dataset & reports:** [wallfacers/weft-script-lineage-synth](https://huggingface.co/datasets/wallfacers/weft-script-lineage-synth) - **Base model:** [Qwen/Qwen2.5-Coder-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct) - **Platform:** [Weft (data-weave)](https://github.com/wallfacers/data-weave)