--- license: apache-2.0 base_model: Qwen/Qwen2.5-0.5B-Instruct library_name: transformers pipeline_tag: text-generation tags: - grpo - rlvr - reinforcement-learning - countdown - reasoning --- # Countdown-Qwen2.5-0.5B-GRPO (lr 3e-6 ablation) `Qwen2.5-0.5B-Instruct` fine-tuned with **from-scratch GRPO** (Group-Relative Policy Optimization) on the **Countdown** number-puzzle task — combine the given numbers exactly once with `+ - * /` to hit a target. Rewards are **verifiable** (exact rational arithmetic via a whitelisted-AST evaluator), so this is RLVR: no reward model, no critic. This checkpoint is the **best ablation** (learning rate `3e-6`, group `8`, `1500` steps). ## Results (dev_public, 300 puzzles, greedy, exact verifier) | model | accuracy | hard (5-num) | format_rate | avg_tokens (correct) | |---|---|---|---|---| | base Qwen2.5-0.5B-Instruct (floor) | 0.33% | 0.00% | 0.00% | 20.0 | | **this model (GRPO, lr 3e-6)** | **12.00%** | **1.67%** | 0.00% | 16.9 | **36x over the base floor** (1/300 -> 36/300). Ablations: learning rate was the only lever that moved accuracy (1e-6 -> 3e-6 gave 1% -> 12%); more steps (3000) and larger group (16) both plateaued. Known failure mode: reasoning collapse — the shaped reward scores only the ``, so the model emits bare answers with no `` (format_rate = 0). ## Usage ```python from transformers import AutoModelForCausalLM, AutoTokenizer model = AutoModelForCausalLM.from_pretrained("/countdown-qwen2.5-0.5b-grpo-lr3e6") tok = AutoTokenizer.from_pretrained("/countdown-qwen2.5-0.5b-grpo-lr3e6") ``` Trained for the RLVR Arena capstone (RL in Production Bootcamp).