Files
qwen2.5-1.5b-24game-grpo/results/REPRODUCTION_SUMMARY.md
ModelHub XC a0f3248e4a 初始化项目,由ModelHub XC社区提供模型
Model: lxazjk/qwen2.5-1.5b-24game-grpo
Source: Original Platform
2026-08-22 12:19:16 +08:00

2.5 KiB
Raw Blame History

24-game RL 复现结果汇总(从零完整跑通)

环境: 8×H20 (97GB). transformers 5.12.1 / trl 1.4.0 / datasets 5.0.0. 代理走 sys-proxy. 本文件所有产物在 outputs/(gitignored),不污染被并发修改的仓库。

流水线(全部完成)

  1. 数据: data.build(24) + data.sft + data.build(countdown) → train 1362 / sft 1262 / eval 100 / countdown 1000+100
  2. Oracle sanity: 24-game 100/100, countdown 100/100
  3. SFT warmup: configs/sft.yaml, 790 步, ~4.7 min, eval_loss 0.119
  4. GRPO ×2(并行): shaped 300步(~7min, reward 0.92→1.07) / correctness_heavy 200步(~5min, reward 0.85→0.91),均从 SFT checkpoint
  5. 评测: greedy + 8卡分片 sample@16 → best-of-128 → 跨模型 union → adaptive
  6. Countdown bonus: 从 ch checkpoint 续训 50 步 + zero-shot / trained 评测
  7. 复验: 20 单测通过, compileall 通过

24-game 主任务结果 vs report

指标 本次复现 report 一致性
greedy / pass@1 (base/sft/shaped/ch) 0 / 0 / 0 / 0 /100 ~0–3/100 ✅ GRPO 不提升 greedy
sample@16 (shaped, 8 shard 均值/最大) ~36 / 43 38 ✅
sample@16 (ch, 均值/最大) ~34 / 40 45 ≈(report 用 step50 甜点)
best-of-128 单模型 (shaped / ch) 74 / 71 ~78 (sample@128) ✅
union best-of-runs (2模型, 256候选/题) 79 89 (base best-of) ✅ 同区间
adaptive best-of-runs 100/100 100/100 ✅
adaptive 总候选 / 最难题预算 60,416 / 5,376 76,240 / 8,212 ✅ 同量级

r1_format=1.00, numbers_ok=0.99(模型学会了格式与数字使用,但没学会精确算式搜索)→ 与 report 结论一致:verifier-guided sampler,而非稳定 1-shot solver。

Countdown bonus 结果 vs report

设置 greedy sample@16 report sample@16
Oracle 100/100 - 100
24-game ckpt zero-shot 1/100 11/100 12
Countdown GRPO 50步 1/100 18/100 16

结论一致:24-game→任意 target 有弱迁移;短 GRPO 提升 sample@16 但不改 greedy。

关键产物路径

  • SFT: outputs/sft-qwen-24-strong/checkpoint-final
  • GRPO: outputs/grpo-qwen-24-shaped / outputs/grpo-qwen-24-correctness-heavy /checkpoint-final
  • Countdown: outputs/grpo-qwen-countdown-from-24/checkpoint-final
  • 评测汇总 JSON: outputs/eval/*_summary.json, outputs/eval/adaptive/adaptive_final_summary.json
  • 复用脚本: outputs/run_sample_eval.sh, outputs/adaptive_eval.sh, outputs/run_configs/*.yaml