Files
qwen3-4b-grpo-dapo17k-invma…/README.md

13 lines
211 B
Markdown
Raw Normal View History

---
datasets:
- BytedTsinghua-SIA/DAPO-Math-17k
base_model:
- Qwen/Qwen3-4B
---
Trained on DAPO-Math-17k using GRPO implemented by VeRL.
- Batch Size: 32
- Group Size: 8
- Step: 500
- Max Response Length: 8192