--- library_name: transformers base_model: Qwen/Qwen3-8B license: other tags: - ronpo - mnpo - dpo - qwen3 - preference-optimization datasets: - UltraFeedback --- # qwen3-8b-dpo-avg-beta0p05-s42 This is a research checkpoint for the RONPO AAAI revision experiments. - Method: DPO-avg on the averaged three-reward oracle - Base model: `Qwen/Qwen3-8B`, non-thinking mode - Training seed: 42 - DPO beta: 0.05 - Data split: the existing MNPO/RONPO UltraFeedback split - Oracle construction: per-prompt min-max normalization over `Skywork/Skywork-Reward-V2-Llama-3.1-8B`, `Nexusflow/Athene-RM-8B`, and `RLHFlow/ArmoRM-Llama3-8B-v0.1`, followed by an unweighted average - Local source checkpoint at upload time: `/ext_hdd/sjkim/mnpo/revision_qwen3_8b/full_iter1/train/dpo_avg_beta0p05_s42_odin2` Intended use: reproducibility and evaluation for the RONPO research paper. This model is not intended as a general-purpose production assistant.