Model: Linyuana/qwen3-0.6b-grpo-math-reasoning Source: Original Platform
license, base_model, tags, datasets, language, pipeline_tag
| license | base_model | tags | datasets | language | pipeline_tag | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| apache-2.0 | Qwen/Qwen3-0.6B-Base |
|
|
|
text-generation |
Qwen3-0.6B GRPO Math Reasoning
Qwen3-0.6B-Base fine-tuned with SFT cold-start followed by GRPO (Group Relative Policy Optimization) with a verifiable rule-based reward, on GSM8K + MATH. Part of a reproduction study on the DeepSeek-R1 recipe comparing GRPO / PPO / DPO under matched initialization — full writeup, ablations, and failure-mode analysis (including a PPO training collapse and fix) here: Linyuan30/llm-rl-reasoning.
Training Recipe
Qwen3-0.6B-Base --SFT (cold start)--> unified <think>/<answer> format --GRPO--> this checkpoint
- Reward: purely rule-based — regex-extract the
<answer>tag, normalize, compare to ground truth. No reward model. - Algorithm: GRPO, group-relative advantage (no critic/value model).
- Framework: veRL v0.4.0 + vLLM for rollout.
- Checkpoint corresponds to
global_step_116ofgrpo_qwen3_0.6b(best result in the sweep).
Results
Evaluated on 500 held-out samples/dataset, temperature=0.8, top_p=0.95, 8 samples/question, same rule-reward scorer used at both train and eval time.
| Method | GSM8K pass@1 | GSM8K pass@8 | MATH pass@1 | MATH pass@8 |
|---|---|---|---|---|
| Base (no post-training) | 4.8 | 30.6 | 3.7 | 24.8 |
| SFT only | 38.9 | 78.4 | 29.0 | 66.4 |
| SFT + GRPO (this model) | 67.7 | 85.0 | 48.8 | 79.2 |
GRPO clearly outperformed PPO (42.6 pass@1) and DPO (39.0 pass@1) under the same base model, data, and reward — see docs/grpo_analysis.md for the group-size ablation (n=4/8/16) and why GRPO's critic-free advantage estimation avoided the instability PPO ran into.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "Linyuana/qwen3-0.6b-grpo-math-reasoning"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="auto", device_map="auto")
prompt = "Natalia sold clips to 48 of her friends in April, and then she sold half as many clips in May. How many clips did Natalia sell altogether in April and May?"
messages = [{"role": "user", "content": prompt}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
output = model.generate(inputs, max_new_tokens=512, temperature=0.8, top_p=0.95, do_sample=True)
print(tokenizer.decode(output[0], skip_special_tokens=True))
The model outputs a <think>...</think><answer>...</answer> format; extract the final answer from within the <answer> tag.
Limitations
- 0.6B scale — solid gains from RL, but absolute accuracy is well below what larger models achieve on MATH.
strict_format_rate(exact<think>/<answer>tag closure) is lower than expected across all training stages despitehas_answer_rate>94%; this is a known open issue in the format-matching regex, not a correctness issue — see Open Question in the repo.- Trained/evaluated only on GSM8K and MATH; not tested on other reasoning benchmarks.
Citation / Acknowledgements
Built on Qwen3, trained with veRL and vLLM, following the DeepSeek-R1 RLVR recipe.