45 lines
2.5 KiB
Markdown
45 lines
2.5 KiB
Markdown
# 24-game RL 复现结果汇总(从零完整跑通)
|
||
|
||
环境: 8×H20 (97GB). transformers 5.12.1 / trl 1.4.0 / datasets 5.0.0. 代理走 sys-proxy.
|
||
本文件所有产物在 outputs/(gitignored),不污染被并发修改的仓库。
|
||
|
||
## 流水线(全部完成)
|
||
1. 数据: data.build(24) + data.sft + data.build(countdown) → train 1362 / sft 1262 / eval 100 / countdown 1000+100
|
||
2. Oracle sanity: 24-game 100/100, countdown 100/100
|
||
3. SFT warmup: configs/sft.yaml, 790 步, ~4.7 min, eval_loss 0.119
|
||
4. GRPO ×2(并行): shaped 300步(~7min, reward 0.92→1.07) / correctness_heavy 200步(~5min, reward 0.85→0.91),均从 SFT checkpoint
|
||
5. 评测: greedy + 8卡分片 sample@16 → best-of-128 → 跨模型 union → adaptive
|
||
6. Countdown bonus: 从 ch checkpoint 续训 50 步 + zero-shot / trained 评测
|
||
7. 复验: 20 单测通过, compileall 通过
|
||
|
||
## 24-game 主任务结果 vs report
|
||
|
||
| 指标 | 本次复现 | report | 一致性 |
|
||
|---|---|---|---|
|
||
| greedy / pass@1 (base/sft/shaped/ch) | 0 / 0 / 0 / 0 /100 | ~0–3/100 | ✅ GRPO 不提升 greedy |
|
||
| sample@16 (shaped, 8 shard 均值/最大) | ~36 / 43 | 38 | ✅ |
|
||
| sample@16 (ch, 均值/最大) | ~34 / 40 | 45 | ≈(report 用 step50 甜点) |
|
||
| best-of-128 单模型 (shaped / ch) | 74 / 71 | ~78 (sample@128) | ✅ |
|
||
| union best-of-runs (2模型, 256候选/题) | 79 | 89 (base best-of) | ✅ 同区间 |
|
||
| adaptive best-of-runs | **100/100** | 100/100 | ✅ |
|
||
| adaptive 总候选 / 最难题预算 | 60,416 / 5,376 | 76,240 / 8,212 | ✅ 同量级 |
|
||
|
||
r1_format=1.00, numbers_ok=0.99(模型学会了格式与数字使用,但没学会精确算式搜索)→ 与 report 结论一致:**verifier-guided sampler,而非稳定 1-shot solver。**
|
||
|
||
## Countdown bonus 结果 vs report
|
||
|
||
| 设置 | greedy | sample@16 | report sample@16 |
|
||
|---|---|---|---|
|
||
| Oracle | 100/100 | - | 100 |
|
||
| 24-game ckpt zero-shot | 1/100 | 11/100 | 12 |
|
||
| Countdown GRPO 50步 | 1/100 | 18/100 | 16 |
|
||
|
||
结论一致:24-game→任意 target 有弱迁移;短 GRPO 提升 sample@16 但不改 greedy。
|
||
|
||
## 关键产物路径
|
||
- SFT: outputs/sft-qwen-24-strong/checkpoint-final
|
||
- GRPO: outputs/grpo-qwen-24-shaped / outputs/grpo-qwen-24-correctness-heavy /checkpoint-final
|
||
- Countdown: outputs/grpo-qwen-countdown-from-24/checkpoint-final
|
||
- 评测汇总 JSON: outputs/eval/*_summary.json, outputs/eval/adaptive/adaptive_final_summary.json
|
||
- 复用脚本: outputs/run_sample_eval.sh, outputs/adaptive_eval.sh, outputs/run_configs/*.yaml
|