license, base_model, tags, pipeline_tag
license base_model tags pipeline_tag
apache-2.0 Qwen/Qwen3-4B
hintflow
dpo
orchestrator
qwen3
lora-merged
text-generation

pagent-hintflow-dpo-v2-merged

Merged full-weight Qwen3-4B HintFlow orchestrator after offline DPO on tree-exported preference pairs (v2 trees).

Training snapshot

  • Base: Qwen/Qwen3-4B
  • Data: hintflow_trees_v2 → 13565 DPO pairs (plan 466 / review 13099), τ=0.05
  • Train: 3 epochs, LoRA r=16, β=0.1, lr=5e-6, world_size=3, effective_batch=9
  • Val pair-acc (preference ranking, not EM): e1 50.2% / e2 50.2% / e3 50.9%

Files

This repo is the merged checkpoint used for vLLM serve (served-model-name typically qwen3-4b).

Companion LoRA adapter (if uploaded): ivaning0919/pagent-hintflow-dpo-v2-adapter.

Note

Backup / archival upload from the PAgent project. Preference-pair accuracy stayed near chance; use downstream HintFlow 128 EM for real evaluation.

Description
Model synced from source: ivaning0919/pagent-hintflow-dpo-v2-merged
Readme 27 KiB
Languages
Jinja 100%