Files
DataForge-0.5B-SFT/README.md
ModelHub XC 36502d45fa 初始化项目,由ModelHub XC社区提供模型
Model: Praneshrajan15/DataForge-0.5B-SFT
Source: Original Platform
2026-06-30 04:43:16 +08:00

122 lines
4.4 KiB
Markdown

---
license: apache-2.0
base_model: Qwen/Qwen2.5-0.5B-Instruct
library_name: transformers
tags:
- dataforge
- data-quality
- supervised-fine-tuning
- qlora
- kaggle
datasets:
- Praneshrajan15/dataforge-sft-trajectories
metrics:
- f1
model-index:
- name: DataForge-0.5B-SFT
results:
- task:
type: text-generation
name: DataForge repair planning
dataset:
name: DataForge SFT trajectories
type: Praneshrajan15/dataforge-sft-trajectories
metrics:
- type: macro_f1
name: Held-out macro F1
value: null
---
# DataForge-0.5B-SFT
DataForge-0.5B-SFT is a supervised-fine-tuned warmup checkpoint for tabular
data-quality repair experiments. The current training path uses chunk-level
DataForge expert trajectories whose exact repairs are derived from audited
dirty/clean CSV diffs. The earlier `v0-smoke` release only proved the
Kaggle-to-Hugging-Face pipeline and should not be read as a performance claim.
## Intended Use
- Research on tabular data-quality agents and repair planning.
- Offline evaluation on DataForge-Bench-style Hospital, Flights, and Beers
tasks.
- Warm-starting later DataForge RL experiments.
This checkpoint is not intended for autonomous production data modification,
medical decision support, regulated data governance, or unsupervised repair of
private datasets.
## Training Data
- Dataset repo: `Praneshrajan15/dataforge-sft-trajectories`.
- Dataset repo SHA used for this run: `94e2dd556d4f1260c5123d93ca6bf4f9da9b160a`.
- Training examples: `1958` chunk-level `expert_v4.jsonl`
records.
- Data sources: Raha benchmark Hospital, Flights, and Beers datasets via the
BigDaMa/raha repository.
- Primary label source: `oracle_from_clean_diff` dirty/clean CSV diffs.
- Legacy teacher lineage: Groq-hosted `clean-diff-v1` ReAct smoke records may
remain for auditability, but exact repairs are not teacher-discovered labels.
- Flights schedule and actual-time repairs are supervised from dirty/clean
labels; they are not inferred from incomplete prompt context.
- Split safety: held-out rows are reserved before chunking and excluded from SFT
target rows, context rows, normalization candidates, fixes, and messages.
- Hard negatives: clean train chunks are retained as `finish` examples with
empty repairs so the model is penalized for unnecessary edits.
The trajectory JSONL includes state, tool calls, diagnosis text, proposed fixes,
teacher/oracle metadata, benchmark metrics, split metadata, and source
provenance for auditability.
## Training Procedure
- Base model: `Qwen/Qwen2.5-0.5B-Instruct`.
- Method: 4-bit QLoRA warmup, then LoRA merge into fp16 merged weights.
- Compute target: Kaggle or Hugging Face remote GPU only; no laptop model
training or full evaluation.
- Kaggle hours used: `1.147`.
- Epochs: 2.
- Batch size: 1 per device with gradient accumulation of 16.
- Learning rate: 2e-5.
## Evaluation
Evaluation is reported on held-out DataForge-Bench-style tasks sampled after the
training trajectory seeds. The release status generated by the notebook is
`quality_improved_verified`. Only `quality_improved_verified` should be treated as a
quality milestone. `diagnostic_complete_no_gain` means the run is authentic and
published, but not promoted.
| Model | Held-out macro F1 |
| --- | ---: |
| `Qwen/Qwen2.5-0.5B-Instruct` | `0.0` |
| `DataForge-0.5B-SFT` | `0.0077` |
Release gates:
- Parse success: `1.0`.
- Schema-case errors: `0`.
- Quality milestone: `True`.
These numbers are produced by the publishing notebook and should not be edited
manually. Re-run the notebook to regenerate them. Detailed per-dataset metrics
are stored in `training_metrics.json` under `base_eval` and `sft_eval`.
Bounded per-task failure evidence is stored in `eval_diagnostics.json`.
## Limitations
- The checkpoint is a Week 9 warmup model, not the final DataForge model family.
- It has only seen small chunk-level ReAct traces and may fail on larger schemas,
unseen domains, adversarial dirty values, or tasks requiring multi-step
database access.
- Legacy teacher traces can contain teacher errors; the primary current labels
come from exact dirty/clean diffs.
- The model should be used behind DataForge's safety, verifier, and transaction
layers before any real data changes.
## License
Weights are published as `apache-2.0` after verifying the base model
metadata for `Qwen/Qwen2.5-0.5B-Instruct`. Users must also comply with the source dataset
licenses/terms and the teacher model terms that governed trajectory generation.