26 lines
769 B
Markdown
26 lines
769 B
Markdown
|
|
---
|
||
|
|
library_name: transformers
|
||
|
|
base_model: Qwen/Qwen3-8B
|
||
|
|
tags:
|
||
|
|
- ronpo
|
||
|
|
- mnpo
|
||
|
|
- qwen3
|
||
|
|
- preference-optimization
|
||
|
|
---
|
||
|
|
|
||
|
|
# qwen3-8b-kto-avg-beta0p05-s42
|
||
|
|
|
||
|
|
Research checkpoint for the RONPO AAAI revision experiments.
|
||
|
|
|
||
|
|
- Method: KTO-avg beta=0.05
|
||
|
|
- Base model: `Qwen/Qwen3-8B` with non-thinking generation protocol
|
||
|
|
- Seed: 42
|
||
|
|
- Local source at upload time: `/NHNHOME/WORKSPACE/26msit001_A/BASE/aipr_lab_sjkim_eval/revision_qwen3_8b/full_iter1/train/kto_avg_beta0p05_s42_nhn-b200-q3-seq-07101430-kto/checkpoint-292`
|
||
|
|
- Uploaded at UTC: 2026-07-11T01:17:52Z
|
||
|
|
- Run status metadata: `not found`
|
||
|
|
|
||
|
|
Qwen3-8B KTO baseline on the averaged three-reward oracle.
|
||
|
|
|
||
|
|
Intended use: reproducibility and evaluation for the RONPO paper. This
|
||
|
|
checkpoint is not intended as a production assistant.
|