初始化项目,由ModelHub XC社区提供模型
Model: promotion/qwen3-8b-aaai27-flagship-dpo-s42 Source: Original Platform
This commit is contained in:
24
README.md
Normal file
24
README.md
Normal file
@@ -0,0 +1,24 @@
|
||||
---
|
||||
library_name: transformers
|
||||
base_model: Qwen/Qwen3-8B
|
||||
tags:
|
||||
- ronpo
|
||||
- mnpo
|
||||
- preference-optimization
|
||||
---
|
||||
|
||||
# qwen3-8b-aaai27-flagship-dpo-s42
|
||||
|
||||
Research checkpoint for the RONPO AAAI revision experiments.
|
||||
|
||||
- Method: dpo
|
||||
- Base model: `Qwen/Qwen3-8B`
|
||||
- Seed: 42
|
||||
- Local source at upload time: `/NHNHOME/WORKSPACE/26msit001_A/BASE/aipr_lab_sjkim_eval/revision_qwen3_8b/full_iter1/flagship_20260712/full/dpo/seed42/attempt1`
|
||||
- Uploaded at UTC: 2026-07-12T17:21:45Z
|
||||
- Run status metadata: `{"attempt": 1, "effective_batch_size": 16, "method": "dpo", "objective_protocol": "objective_protocol.json", "optimizer_steps": 900, "seed": 42, "split_manifest": "split_manifest.json", "stability_gate": "stability_gate.json", "status": "completed", "wandb_run_id": "f11bcc5bf0a6", "wandb_url": "https://wandb.ai/promotion-kim/mnpo/runs/f11bcc5bf0a6"}`
|
||||
|
||||
AAAI-27 P1 matched budget: 900 optimizer steps, effective batch 16; passed non-thinking and collapse stability gates; sealed-test metrics are not used for checkpoint selection.
|
||||
|
||||
Intended use: reproducibility and evaluation for the RONPO paper. This
|
||||
checkpoint is not intended as a production assistant.
|
||||
Reference in New Issue
Block a user