8bc779b4d21ca5ed43b5e1541811d3b041bc35bf
Model: promotion/qwen3-8b-dpo-avg-beta0p05-s42 Source: Original Platform
library_name, base_model, license, tags, datasets
| library_name | base_model | license | tags | datasets | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| transformers | Qwen/Qwen3-8B | other |
|
|
qwen3-8b-dpo-avg-beta0p05-s42
This is a research checkpoint for the RONPO AAAI revision experiments.
- Method: DPO-avg on the averaged three-reward oracle
- Base model:
Qwen/Qwen3-8B, non-thinking mode - Training seed: 42
- DPO beta: 0.05
- Data split: the existing MNPO/RONPO UltraFeedback split
- Oracle construction: per-prompt min-max normalization over
Skywork/Skywork-Reward-V2-Llama-3.1-8B,Nexusflow/Athene-RM-8B, andRLHFlow/ArmoRM-Llama3-8B-v0.1, followed by an unweighted average - Local source checkpoint at upload time:
/ext_hdd/sjkim/mnpo/revision_qwen3_8b/full_iter1/train/dpo_avg_beta0p05_s42_odin2
Intended use: reproducibility and evaluation for the RONPO research paper. This model is not intended as a general-purpose production assistant.
Description
Languages
Jinja
100%