license, base_model, tags, datasets, language, pipeline_tag
license base_model tags datasets language pipeline_tag
apache-2.0 Qwen/Qwen3-0.6B-Base
reasoning
math
grpo
reinforcement-learning
rlvr
qwen3
gsm8k
math
en
text-generation

Qwen3-0.6B GRPO Math Reasoning

Qwen3-0.6B-Base fine-tuned with SFT cold-start followed by GRPO (Group Relative Policy Optimization) with a verifiable rule-based reward, on GSM8K + MATH. Part of a reproduction study on the DeepSeek-R1 recipe comparing GRPO / PPO / DPO under matched initialization — full writeup, ablations, and failure-mode analysis (including a PPO training collapse and fix) here: Linyuan30/llm-rl-reasoning.

Training Recipe

Qwen3-0.6B-Base --SFT (cold start)--> unified <think>/<answer> format --GRPO--> this checkpoint
  • Reward: purely rule-based — regex-extract the <answer> tag, normalize, compare to ground truth. No reward model.
  • Algorithm: GRPO, group-relative advantage (no critic/value model).
  • Framework: veRL v0.4.0 + vLLM for rollout.
  • Checkpoint corresponds to global_step_116 of grpo_qwen3_0.6b (best result in the sweep).

Results

Evaluated on 500 held-out samples/dataset, temperature=0.8, top_p=0.95, 8 samples/question, same rule-reward scorer used at both train and eval time.

Method GSM8K pass@1 GSM8K pass@8 MATH pass@1 MATH pass@8
Base (no post-training) 4.8 30.6 3.7 24.8
SFT only 38.9 78.4 29.0 66.4
SFT + GRPO (this model) 67.7 85.0 48.8 79.2

GRPO clearly outperformed PPO (42.6 pass@1) and DPO (39.0 pass@1) under the same base model, data, and reward — see docs/grpo_analysis.md for the group-size ablation (n=4/8/16) and why GRPO's critic-free advantage estimation avoided the instability PPO ran into.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "Linyuana/qwen3-0.6b-grpo-math-reasoning"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="auto", device_map="auto")

prompt = "Natalia sold clips to 48 of her friends in April, and then she sold half as many clips in May. How many clips did Natalia sell altogether in April and May?"
messages = [{"role": "user", "content": prompt}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
output = model.generate(inputs, max_new_tokens=512, temperature=0.8, top_p=0.95, do_sample=True)
print(tokenizer.decode(output[0], skip_special_tokens=True))

The model outputs a <think>...</think><answer>...</answer> format; extract the final answer from within the <answer> tag.

Limitations

  • 0.6B scale — solid gains from RL, but absolute accuracy is well below what larger models achieve on MATH.
  • strict_format_rate (exact <think>/<answer> tag closure) is lower than expected across all training stages despite has_answer_rate >94%; this is a known open issue in the format-matching regex, not a correctness issue — see Open Question in the repo.
  • Trained/evaluated only on GSM8K and MATH; not tested on other reasoning benchmarks.

Citation / Acknowledgements

Built on Qwen3, trained with veRL and vLLM, following the DeepSeek-R1 RLVR recipe.

Description
Model synced from source: Linyuana/qwen3-0.6b-grpo-math-reasoning
Readme 13 MiB