Files
qwen3-4b-grpo-dapo17k/README.md

14 lines
222 B
Markdown
Raw Permalink Normal View History

---
datasets:
- BytedTsinghua-SIA/DAPO-Math-17k
base_model:
- Qwen/Qwen3-4B
---
Trained on DAPO-Math-17k using GRPO implemented by VeRL.
- Batch Size: 32
- Group Size: 8
- Epoch: 1
- Step: 559
- Max Response Length: 8192