Files
qwen2.5-1.5b-24game-grpo/results/REPRODUCTION_SUMMARY.md
ModelHub XC a0f3248e4a 初始化项目,由ModelHub XC社区提供模型
Model: lxazjk/qwen2.5-1.5b-24game-grpo
Source: Original Platform
2026-08-22 12:19:16 +08:00

45 lines
2.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# 24-game RL 复现结果汇总(从零完整跑通)
环境: 8×H20 (97GB). transformers 5.12.1 / trl 1.4.0 / datasets 5.0.0. 代理走 sys-proxy.
本文件所有产物在 outputs/(gitignored),不污染被并发修改的仓库。
## 流水线(全部完成)
1. 数据: data.build(24) + data.sft + data.build(countdown) → train 1362 / sft 1262 / eval 100 / countdown 1000+100
2. Oracle sanity: 24-game 100/100, countdown 100/100
3. SFT warmup: configs/sft.yaml, 790 步, ~4.7 min, eval_loss 0.119
4. GRPO ×2(并行): shaped 300步(~7min, reward 0.92→1.07) / correctness_heavy 200步(~5min, reward 0.85→0.91),均从 SFT checkpoint
5. 评测: greedy + 8卡分片 sample@16 → best-of-128 → 跨模型 union → adaptive
6. Countdown bonus: 从 ch checkpoint 续训 50 步 + zero-shot / trained 评测
7. 复验: 20 单测通过, compileall 通过
## 24-game 主任务结果 vs report
| 指标 | 本次复现 | report | 一致性 |
|---|---|---|---|
| greedy / pass@1 (base/sft/shaped/ch) | 0 / 0 / 0 / 0 /100 | ~0–3/100 | ✅ GRPO 不提升 greedy |
| sample@16 (shaped, 8 shard 均值/最大) | ~36 / 43 | 38 | ✅ |
| sample@16 (ch, 均值/最大) | ~34 / 40 | 45 | ≈(report 用 step50 甜点) |
| best-of-128 单模型 (shaped / ch) | 74 / 71 | ~78 (sample@128) | ✅ |
| union best-of-runs (2模型, 256候选/题) | 79 | 89 (base best-of) | ✅ 同区间 |
| adaptive best-of-runs | **100/100** | 100/100 | ✅ |
| adaptive 总候选 / 最难题预算 | 60,416 / 5,376 | 76,240 / 8,212 | ✅ 同量级 |
r1_format=1.00, numbers_ok=0.99(模型学会了格式与数字使用,但没学会精确算式搜索)→ 与 report 结论一致:**verifier-guided sampler,而非稳定 1-shot solver。**
## Countdown bonus 结果 vs report
| 设置 | greedy | sample@16 | report sample@16 |
|---|---|---|---|
| Oracle | 100/100 | - | 100 |
| 24-game ckpt zero-shot | 1/100 | 11/100 | 12 |
| Countdown GRPO 50步 | 1/100 | 18/100 | 16 |
结论一致:24-game→任意 target 有弱迁移;短 GRPO 提升 sample@16 但不改 greedy。
## 关键产物路径
- SFT: outputs/sft-qwen-24-strong/checkpoint-final
- GRPO: outputs/grpo-qwen-24-shaped / outputs/grpo-qwen-24-correctness-heavy /checkpoint-final
- Countdown: outputs/grpo-qwen-countdown-from-24/checkpoint-final
- 评测汇总 JSON: outputs/eval/*_summary.json, outputs/eval/adaptive/adaptive_final_summary.json
- 复用脚本: outputs/run_sample_eval.sh, outputs/adaptive_eval.sh, outputs/run_configs/*.yaml