初始化项目,由ModelHub XC社区提供模型

Model: hkr04/qwen3-4b-grpo-dapo17k-invmax-linear
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-18 19:29:14 +08:00
commit 25b8c898bc
14 changed files with 152322 additions and 0 deletions

13
README.md Normal file
View File

@@ -0,0 +1,13 @@
---
datasets:
- BytedTsinghua-SIA/DAPO-Math-17k
base_model:
- Qwen/Qwen3-4B
---
Trained on DAPO-Math-17k using GRPO implemented by VeRL.
- Batch Size: 32
- Group Size: 8
- Step: 500
- Max Response Length: 8192