library_name, base_model, tags
library_name base_model tags
transformers Qwen/Qwen3-8B
ronpo
mnpo
qwen3
preference-optimization

qwen3-8b-kto-avg-beta0p05-s42

Research checkpoint for the RONPO AAAI revision experiments.

  • Method: KTO-avg beta=0.05
  • Base model: Qwen/Qwen3-8B with non-thinking generation protocol
  • Seed: 42
  • Local source at upload time: /NHNHOME/WORKSPACE/26msit001_A/BASE/aipr_lab_sjkim_eval/revision_qwen3_8b/full_iter1/train/kto_avg_beta0p05_s42_nhn-b200-q3-seq-07101430-kto/checkpoint-292
  • Uploaded at UTC: 2026-07-11T01:17:52Z
  • Run status metadata: not found

Qwen3-8B KTO baseline on the averaged three-reward oracle.

Intended use: reproducibility and evaluation for the RONPO paper. This checkpoint is not intended as a production assistant.

Description
Model synced from source: promotion/qwen3-8b-kto-avg-beta0p05-s42
Readme 31 KiB
Languages
Jinja 100%