--- license: apache-2.0 base_model: Qwen/Qwen2.5-1.5B datasets: [openai/gsm8k] pipeline_tag: text-generation tags: [grpo, reinforcement-learning, nemo-rl, math] --- # Qwen2.5-1.5B GRPO GSM8K Qwen2.5-1.5B trained with GRPO using NVIDIA NeMo RL v0.7.0. ## Results | Model | GSM8K test (pass@1) | Correct | |---|---|---| | Qwen/Qwen2.5-1.5B (base) | 35.03% | 462 / 1319 | | This model | 73.84% | 974 / 1319 | +38.8 points, 2.1x relative. Both models evaluated identically on the full 1319-problem GSM8K test split, greedy decoding, same prompt template. ## Training - GRPO, 130 steps, 16 prompts x 8 generations (2,080 samples) - Reward: binary exact-match on the answer inside boxed tags - DTensor v2 trainer + vLLM generation, colocated - 1x H100 80GB, ~3 hours - lr 1e-6 AdamW, KL penalty 0.01, clip 0.2/0.2, seq len 1024 No human-written solutions. The model learned from a programmatic grader. ## Prompt format Trained with this template and expects it at inference. A bare question without the wrapper degrades results substantially. Think step-by-step to solve the following problem. Output your answer inside of \boxed{} tags.: {question} Let's think step-by-step