初始化项目,由ModelHub XC社区提供模型
Model: ivaning0919/pagent-hintflow-dpo-v2-merged Source: Original Platform
This commit is contained in:
32
README.md
Normal file
32
README.md
Normal file
@@ -0,0 +1,32 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
base_model: Qwen/Qwen3-4B
|
||||
tags:
|
||||
- hintflow
|
||||
- dpo
|
||||
- orchestrator
|
||||
- qwen3
|
||||
- lora-merged
|
||||
pipeline_tag: text-generation
|
||||
---
|
||||
|
||||
# pagent-hintflow-dpo-v2-merged
|
||||
|
||||
Merged full-weight **Qwen3-4B** HintFlow orchestrator after offline DPO on tree-exported preference pairs (v2 trees).
|
||||
|
||||
## Training snapshot
|
||||
|
||||
- Base: `Qwen/Qwen3-4B`
|
||||
- Data: `hintflow_trees_v2` → 13565 DPO pairs (plan 466 / review 13099), τ=0.05
|
||||
- Train: 3 epochs, LoRA r=16, β=0.1, lr=5e-6, world_size=3, effective_batch=9
|
||||
- Val pair-acc (preference ranking, not EM): e1 50.2% / e2 50.2% / e3 50.9%
|
||||
|
||||
## Files
|
||||
|
||||
This repo is the **merged** checkpoint used for vLLM serve (`served-model-name` typically `qwen3-4b`).
|
||||
|
||||
Companion LoRA adapter (if uploaded): `ivaning0919/pagent-hintflow-dpo-v2-adapter`.
|
||||
|
||||
## Note
|
||||
|
||||
Backup / archival upload from the PAgent project. Preference-pair accuracy stayed near chance; use downstream HintFlow 128 EM for real evaluation.
|
||||
Reference in New Issue
Block a user