初始化项目,由ModelHub XC社区提供模型
Model: Parallel-R1/Parallel-R1-Unseen_Step_200 Source: Original Platform
This commit is contained in:
12
README.md
Normal file
12
README.md
Normal file
@@ -0,0 +1,12 @@
|
||||
---
|
||||
license: mit
|
||||
datasets:
|
||||
- Leo-Dai/dapo-math-17k_dedup
|
||||
---
|
||||
# 🧠 Parallel-R1-Unseen_Step_200
|
||||
|
||||
> **Mid-Training Checkpoint of Parallel-R1: Towards Parallel Thinking via Reinforcement Learning**
|
||||
> Stage: **After 200 RL steps via alternating rewards** — showing the adaptive parallel reasoning ability and serve as structure exploration stage.
|
||||
|
||||
This checkpoint aims to help you reproduce experimental results in Section 4.5: Extra Bonus: Parallel Thinking as a Mid-Training Exploration Strategy for RL Training.
|
||||
|
||||
Reference in New Issue
Block a user