945 B
945 B
library_name, base_model, license, tags, datasets
| library_name | base_model | license | tags | datasets | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| transformers | Qwen/Qwen3-8B | other |
|
|
qwen3-8b-dpo-avg-beta0p01-s42
This is a research checkpoint for the RONPO AAAI revision experiments.
- Method: DPO-avg on the averaged three-reward oracle
- Base model:
Qwen/Qwen3-8B, non-thinking mode - Training seed: 42
- DPO beta: 0.01
- Data split: the existing MNPO/RONPO UltraFeedback split
- Oracle construction: per-prompt min-max normalization over
Skywork/Skywork-Reward-V2-Llama-3.1-8B,Nexusflow/Athene-RM-8B, andRLHFlow/ArmoRM-Llama3-8B-v0.1, followed by an unweighted average - Local source checkpoint at upload time:
/ext_hdd/sjkim/mnpo/revision_qwen3_8b/full_iter1/train/dpo_avg_beta0p01_s42_odin2
Intended use: reproducibility and evaluation for the RONPO research paper. This model is not intended as a general-purpose production assistant.