--- library_name: transformers base_model: Qwen/Qwen3-8B tags: - ronpo - mnpo - preference-optimization --- # qwen3-8b-aaai27-flagship-inpo-avg-s43 Research checkpoint for the RONPO AAAI revision experiments. - Method: inpo_avg - Base model: `Qwen/Qwen3-8B` - Seed: 43 - Local source at upload time: `/NHNHOME/26msit001_A/BASE/AIPR/sjkim/revision_qwen3_8b/full_iter1/flagship_20260712/full/inpo_avg/seed43/attempt2` - Uploaded at UTC: 2026-07-13T13:52:08Z - Run status metadata: `{"attempt": 2, "effective_batch_size": 16, "method": "inpo_avg", "objective_protocol": "objective_protocol.json", "optimizer_steps": 900, "resource_profile": "resource_profile.json", "scope_amendment": "scope_amendment_20260713.json", "seed": 43, "split_manifest": "split_manifest.json", "stability_gate": "stability_gate.json", "status": "completed", "wandb_run_id": "051016e9c42d", "wandb_url": "https://wandb.ai/promotion-kim/mnpo/runs/051016e9c42d"}` AAAI-27 P1 matched budget: 900 optimizer steps, effective batch 16; passed non-thinking and collapse stability gates; sealed-test metrics are not used for checkpoint selection. Intended use: reproducibility and evaluation for the RONPO paper. This checkpoint is not intended as a production assistant.