2.5 KiB
2.5 KiB
24-game RL 复现结果汇总(从零完整跑通)
环境: 8×H20 (97GB). transformers 5.12.1 / trl 1.4.0 / datasets 5.0.0. 代理走 sys-proxy. 本文件所有产物在 outputs/(gitignored),不污染被并发修改的仓库。
流水线(全部完成)
- 数据: data.build(24) + data.sft + data.build(countdown) → train 1362 / sft 1262 / eval 100 / countdown 1000+100
- Oracle sanity: 24-game 100/100, countdown 100/100
- SFT warmup: configs/sft.yaml, 790 步, ~4.7 min, eval_loss 0.119
- GRPO ×2(并行): shaped 300步(~7min, reward 0.92→1.07) / correctness_heavy 200步(~5min, reward 0.85→0.91),均从 SFT checkpoint
- 评测: greedy + 8卡分片 sample@16 → best-of-128 → 跨模型 union → adaptive
- Countdown bonus: 从 ch checkpoint 续训 50 步 + zero-shot / trained 评测
- 复验: 20 单测通过, compileall 通过
24-game 主任务结果 vs report
| 指标 | 本次复现 | report | 一致性 |
|---|---|---|---|
| greedy / pass@1 (base/sft/shaped/ch) | 0 / 0 / 0 / 0 /100 | ~0–3/100 | ✅ GRPO 不提升 greedy |
| sample@16 (shaped, 8 shard 均值/最大) | ~36 / 43 | 38 | ✅ |
| sample@16 (ch, 均值/最大) | ~34 / 40 | 45 | ≈(report 用 step50 甜点) |
| best-of-128 单模型 (shaped / ch) | 74 / 71 | ~78 (sample@128) | ✅ |
| union best-of-runs (2模型, 256候选/题) | 79 | 89 (base best-of) | ✅ 同区间 |
| adaptive best-of-runs | 100/100 | 100/100 | ✅ |
| adaptive 总候选 / 最难题预算 | 60,416 / 5,376 | 76,240 / 8,212 | ✅ 同量级 |
r1_format=1.00, numbers_ok=0.99(模型学会了格式与数字使用,但没学会精确算式搜索)→ 与 report 结论一致:verifier-guided sampler,而非稳定 1-shot solver。
Countdown bonus 结果 vs report
| 设置 | greedy | sample@16 | report sample@16 |
|---|---|---|---|
| Oracle | 100/100 | - | 100 |
| 24-game ckpt zero-shot | 1/100 | 11/100 | 12 |
| Countdown GRPO 50步 | 1/100 | 18/100 | 16 |
结论一致:24-game→任意 target 有弱迁移;短 GRPO 提升 sample@16 但不改 greedy。
关键产物路径
- SFT: outputs/sft-qwen-24-strong/checkpoint-final
- GRPO: outputs/grpo-qwen-24-shaped / outputs/grpo-qwen-24-correctness-heavy /checkpoint-final
- Countdown: outputs/grpo-qwen-countdown-from-24/checkpoint-final
- 评测汇总 JSON: outputs/eval/*_summary.json, outputs/eval/adaptive/adaptive_final_summary.json
- 复用脚本: outputs/run_sample_eval.sh, outputs/adaptive_eval.sh, outputs/run_configs/*.yaml