ModelHub XC 49bc018110 初始化项目,由ModelHub XC社区提供模型
Model: promotion/qwen3-8b-dpo-avg-beta0p01-s42
Source: Original Platform
2026-07-29 05:32:20 +08:00

library_name, base_model, license, tags, datasets
library_name base_model license tags datasets
transformers Qwen/Qwen3-8B other
ronpo
mnpo
dpo
qwen3
preference-optimization
UltraFeedback

qwen3-8b-dpo-avg-beta0p01-s42

This is a research checkpoint for the RONPO AAAI revision experiments.

  • Method: DPO-avg on the averaged three-reward oracle
  • Base model: Qwen/Qwen3-8B, non-thinking mode
  • Training seed: 42
  • DPO beta: 0.01
  • Data split: the existing MNPO/RONPO UltraFeedback split
  • Oracle construction: per-prompt min-max normalization over Skywork/Skywork-Reward-V2-Llama-3.1-8B, Nexusflow/Athene-RM-8B, and RLHFlow/ArmoRM-Llama3-8B-v0.1, followed by an unweighted average
  • Local source checkpoint at upload time: /ext_hdd/sjkim/mnpo/revision_qwen3_8b/full_iter1/train/dpo_avg_beta0p01_s42_odin2

Intended use: reproducibility and evaluation for the RONPO research paper. This model is not intended as a general-purpose production assistant.

Description
Model synced from source: promotion/qwen3-8b-dpo-avg-beta0p01-s42
Readme 13 MiB
Languages
Jinja 100%