Model: s1lv3rj1nx/countdown-qwen2.5-0.5b-grpo-lr3e6 Source: Original Platform
license, base_model, library_name, pipeline_tag, tags
| license | base_model | library_name | pipeline_tag | tags | |||||
|---|---|---|---|---|---|---|---|---|---|
| apache-2.0 | Qwen/Qwen2.5-0.5B-Instruct | transformers | text-generation |
|
Countdown-Qwen2.5-0.5B-GRPO (lr 3e-6 ablation)
Qwen2.5-0.5B-Instruct fine-tuned with from-scratch GRPO (Group-Relative Policy
Optimization) on the Countdown number-puzzle task — combine the given numbers exactly
once with + - * / to hit a target. Rewards are verifiable (exact rational arithmetic
via a whitelisted-AST evaluator), so this is RLVR: no reward model, no critic.
This checkpoint is the best ablation (learning rate 3e-6, group 8, 1500 steps).
Results (dev_public, 300 puzzles, greedy, exact verifier)
| model | accuracy | hard (5-num) | format_rate | avg_tokens (correct) |
|---|---|---|---|---|
| base Qwen2.5-0.5B-Instruct (floor) | 0.33% | 0.00% | 0.00% | 20.0 |
| this model (GRPO, lr 3e-6) | 12.00% | 1.67% | 0.00% | 16.9 |
36x over the base floor (1/300 -> 36/300). Ablations: learning rate was the only lever
that moved accuracy (1e-6 -> 3e-6 gave 1% -> 12%); more steps (3000) and larger group (16)
both plateaued. Known failure mode: reasoning collapse — the shaped reward scores only the
<answer>, so the model emits bare answers with no <think> (format_rate = 0).
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("<your-username>/countdown-qwen2.5-0.5b-grpo-lr3e6")
tok = AutoTokenizer.from_pretrained("<your-username>/countdown-qwen2.5-0.5b-grpo-lr3e6")
Trained for the RLVR Arena capstone (RL in Production Bootcamp).