33 lines
937 B
Markdown
33 lines
937 B
Markdown
|
|
---
|
||
|
|
license: apache-2.0
|
||
|
|
base_model: Qwen/Qwen3-4B
|
||
|
|
tags:
|
||
|
|
- hintflow
|
||
|
|
- dpo
|
||
|
|
- orchestrator
|
||
|
|
- qwen3
|
||
|
|
- lora-merged
|
||
|
|
pipeline_tag: text-generation
|
||
|
|
---
|
||
|
|
|
||
|
|
# pagent-hintflow-dpo-v2-merged
|
||
|
|
|
||
|
|
Merged full-weight **Qwen3-4B** HintFlow orchestrator after offline DPO on tree-exported preference pairs (v2 trees).
|
||
|
|
|
||
|
|
## Training snapshot
|
||
|
|
|
||
|
|
- Base: `Qwen/Qwen3-4B`
|
||
|
|
- Data: `hintflow_trees_v2` → 13565 DPO pairs (plan 466 / review 13099), τ=0.05
|
||
|
|
- Train: 3 epochs, LoRA r=16, β=0.1, lr=5e-6, world_size=3, effective_batch=9
|
||
|
|
- Val pair-acc (preference ranking, not EM): e1 50.2% / e2 50.2% / e3 50.9%
|
||
|
|
|
||
|
|
## Files
|
||
|
|
|
||
|
|
This repo is the **merged** checkpoint used for vLLM serve (`served-model-name` typically `qwen3-4b`).
|
||
|
|
|
||
|
|
Companion LoRA adapter (if uploaded): `ivaning0919/pagent-hintflow-dpo-v2-adapter`.
|
||
|
|
|
||
|
|
## Note
|
||
|
|
|
||
|
|
Backup / archival upload from the PAgent project. Preference-pair accuracy stayed near chance; use downstream HintFlow 128 EM for real evaluation.
|