Model: s1lv3rj1nx/countdown-qwen2.5-0.5b-grpo-dense Source: Original Platform
license, base_model, library_name, pipeline_tag, tags
| license | base_model | library_name | pipeline_tag | tags | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| apache-2.0 | Qwen/Qwen2.5-0.5B-Instruct | transformers | text-generation |
|
Countdown-Qwen2.5-0.5B-GRPO (dense reward)
Qwen2.5-0.5B-Instruct fine-tuned with from-scratch GRPO on the Countdown number-puzzle
task (combine the given numbers exactly once with + - * / to hit a target). Rewards are
verifiable (exact rational arithmetic, whitelisted-AST evaluator) — RLVR: no reward model,
no critic.
This is the best checkpoint: GRPO at lr 3e-6, group 8, 1500 steps, trained with a
dense closeness-shaped reward — for a right-numbers/wrong-value attempt the reward scales
with proximity to the target (0.10 + 0.85·exp(-|value-target|/10)) instead of a flat 0.10.
That gives GRPO a gradient on near-misses and broke the ~12% plateau of the step-function reward.
Results (dev_public, 300 puzzles, greedy, exact verifier)
| model | accuracy | easy | medium | hard | avg_tokens (correct) |
|---|---|---|---|---|---|
| base Qwen2.5-0.5B-Instruct (floor) | 0.33% | — | — | 0.00% | 20.0 |
| GRPO, original step reward | 12.00% | 25.83% | 3.33% | 1.67% | 16.9 |
| this model (GRPO + dense reward) | 14.67% | 29.17% | 7.50% | 0.00% | 17.1 |
44× over the base floor (1/300 → 44/300). Gain concentrated in easy/medium puzzles where the
model lands near the target. Known failure mode: reasoning collapse — the reward credits only the
<answer>, so format_rate stays 0 (no <think>).
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "s1lv3rj1nx/countdown-qwen2.5-0.5b-grpo-dense"
model = AutoModelForCausalLM.from_pretrained(repo)
tok = AutoTokenizer.from_pretrained(repo)
Trained for the RLVR Arena capstone (RL in Production Bootcamp).