Files
ModelHub XC 34b8f238f1 初始化项目,由ModelHub XC社区提供模型
Model: Parallel-R1/Parallel-R1-Unseen_Step_200
Source: Original Platform
2026-07-29 04:39:11 +08:00

517 B

license, datasets
license datasets
mit
Leo-Dai/dapo-math-17k_dedup

🧠 Parallel-R1-Unseen_Step_200

Mid-Training Checkpoint of Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
Stage: After 200 RL steps via alternating rewards — showing the adaptive parallel reasoning ability and serve as structure exploration stage.

This checkpoint aims to help you reproduce experimental results in Section 4.5: Extra Bonus: Parallel Thinking as a Mid-Training Exploration Strategy for RL Training.