--- library_name: transformers base_model: Qwen/Qwen3-8B tags: - ronpo - mnpo - qwen3 - preference-optimization --- # qwen3-8b-inpo-avg-eta0p005-s42 Research checkpoint for the RONPO AAAI revision experiments. - Method: INPO-avg eta=0.005 - Base model: `Qwen/Qwen3-8B` with non-thinking generation protocol - Seed: 42 - Local source at upload time: `/NHNHOME/WORKSPACE/26msit001_A/BASE/aipr_lab_sjkim_eval/revision_qwen3_8b/full_iter1/train/inpo_avg_eta0p005_s42_nhn-b200-q3-seq-07101430-inpo` - Uploaded at UTC: 2026-07-10T16:05:55Z - Run status metadata: `not found` Qwen3-8B INPO baseline on the averaged three-reward oracle. Intended use: reproducibility and evaluation for the RONPO paper. This checkpoint is not intended as a production assistant.