初始化项目,由ModelHub XC社区提供模型
Model: wallfacers/weft-lineage-extractor-jvm-1.5b Source: Original Platform
This commit is contained in:
91
README.md
Normal file
91
README.md
Normal file
@@ -0,0 +1,91 @@
|
||||
---
|
||||
license: other
|
||||
license_name: weft-research
|
||||
license_link: https://github.com/wallfacers/data-weave
|
||||
pipeline_tag: text-generation
|
||||
library_name: transformers
|
||||
base_model: Qwen/Qwen2.5-Coder-1.5B-Instruct
|
||||
datasets:
|
||||
- wallfacers/weft-script-lineage-synth
|
||||
language:
|
||||
- en
|
||||
tags:
|
||||
- research-artifact
|
||||
- negative-result
|
||||
- memorization
|
||||
- domain-shift
|
||||
- cross-language
|
||||
- data-lineage
|
||||
- etl
|
||||
- lora
|
||||
model-index:
|
||||
- name: weft-lineage-extractor-jvm-1.5b
|
||||
results:
|
||||
- task:
|
||||
type: table-level-lineage-extraction
|
||||
name: ETL table-level data-lineage extraction (Scala/Java)
|
||||
dataset:
|
||||
type: real-jvm-etl
|
||||
name: real Scala/Java ETL (human gold, n=141)
|
||||
metrics:
|
||||
- type: precision
|
||||
value: 0.165
|
||||
name: Table precision (real JVM, out-of-distribution)
|
||||
- type: accuracy
|
||||
value: 0.418
|
||||
name: Read/write direction accuracy (real JVM)
|
||||
---
|
||||
|
||||
# weft-lineage-extractor-jvm-1.5b — cross-language reproduction of a NEGATIVE RESULT
|
||||
|
||||
> ## ⚠️ RESEARCH ARTIFACT — the Scala/Java (JVM) point of a *synthetic-only* training study. Not a production tool.
|
||||
>
|
||||
> ### ✅ Resolved by real data: use **[weft-lineage-extractor-3b](https://huggingface.co/wallfacers/weft-lineage-extractor-3b)** (real corpus, real precision **0.64**). Full study: **[weft-lineage-extractor-1.5b](https://huggingface.co/wallfacers/weft-lineage-extractor-1.5b)**.
|
||||
|
||||
The **JVM (Scala/Java)** point of a study showing that **synthetic-only training** induces a
|
||||
**verbatim memorization leak** in small ETL table-lineage models — and that the failure
|
||||
**reproduces, worse, across languages**. Synthetic held-out looks near-perfect (~0.99); on real
|
||||
Scala/Java ETL it drops to **precision 0.165**, with **40.8%** of hallucinations being table names
|
||||
recited verbatim from the synthetic training pool — the **highest leak** of all scale/language
|
||||
points, because real JVM code is the most out-of-distribution.
|
||||
|
||||
Same recipe as the [1.5B main model](https://huggingface.co/wallfacers/weft-lineage-extractor-1.5b)
|
||||
(LoRA on Qwen2.5-Coder-1.5B), extended with synthetic Scala/Java forms; `task_type: SCALA | JAVA`.
|
||||
|
||||
## This point's numbers (table-level, Convention A)
|
||||
|
||||
| metric | synthetic held-out | real Scala/Java ETL (n=141, non-empty 28) |
|
||||
|---|---|---|
|
||||
| precision | ~0.99 | **0.165** |
|
||||
| direction accuracy | — | **0.418** |
|
||||
| **verbatim memorization leak** | — | **40.8%** |
|
||||
|
||||
The leak metric is **gold-independent** and stable to full-file re-adjudication (verbatim
|
||||
40.4% → 40.8%). Given a real `SqlDao.java` it cannot parse, the model emits training-pool names
|
||||
like `dws.dws_member_point_df` — memorized generator vocabulary, not a label artifact.
|
||||
|
||||
## Where it fits (scale / cross-language curve)
|
||||
|
||||
| point | real precision | real direction | verbatim leak |
|
||||
|---|---|---|---|
|
||||
| 0.5B (Py/Sh) | 0.243 | 0.369 | 37.4% |
|
||||
| 1.5B (Py/Sh) | 0.270 | 0.496 | 22.4% |
|
||||
| 3B (Py/Sh, synthetic) | 0.325 | 0.468 | 10.9% |
|
||||
| **1.5B + JVM (this, real JVM)** | **0.165** | **0.418** | **40.8%** |
|
||||
| **3B (real corpus)** | **0.64** | — | **~0** |
|
||||
|
||||
"Add another language" does not fix the real-world gap; **real training data does**.
|
||||
|
||||
## Intended use
|
||||
|
||||
- ✅ Reproducing / studying the cross-language reproduction of the memorization-leak failure.
|
||||
- ❌ Not for production lineage — use the [real-corpus 3B](https://huggingface.co/wallfacers/weft-lineage-extractor-3b).
|
||||
|
||||
## Usage, prompt format, training details, citation
|
||||
|
||||
Identical to the main model (this variant adds `task_type: SCALA | JAVA` and JVM system-prompt wording):
|
||||
**[weft-lineage-extractor-1.5b](https://huggingface.co/wallfacers/weft-lineage-extractor-1.5b)**.
|
||||
|
||||
- **Dataset & reports:** [wallfacers/weft-script-lineage-synth](https://huggingface.co/datasets/wallfacers/weft-script-lineage-synth)
|
||||
- **Base model:** [Qwen/Qwen2.5-Coder-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct)
|
||||
- **Platform:** [Weft (data-weave)](https://github.com/wallfacers/data-weave)
|
||||
Reference in New Issue
Block a user