license, base_model, library_name, pipeline_tag, tags
license base_model library_name pipeline_tag tags
apache-2.0 Qwen/Qwen2.5-0.5B-Instruct transformers text-generation
grpo
rlvr
reinforcement-learning
countdown
reasoning
reward-shaping

Countdown-Qwen2.5-0.5B-GRPO (dense reward)

Qwen2.5-0.5B-Instruct fine-tuned with from-scratch GRPO on the Countdown number-puzzle task (combine the given numbers exactly once with + - * / to hit a target). Rewards are verifiable (exact rational arithmetic, whitelisted-AST evaluator) — RLVR: no reward model, no critic.

This is the best checkpoint: GRPO at lr 3e-6, group 8, 1500 steps, trained with a dense closeness-shaped reward — for a right-numbers/wrong-value attempt the reward scales with proximity to the target (0.10 + 0.85·exp(-|value-target|/10)) instead of a flat 0.10. That gives GRPO a gradient on near-misses and broke the ~12% plateau of the step-function reward.

Results (dev_public, 300 puzzles, greedy, exact verifier)

model accuracy easy medium hard avg_tokens (correct)
base Qwen2.5-0.5B-Instruct (floor) 0.33% — — 0.00% 20.0
GRPO, original step reward 12.00% 25.83% 3.33% 1.67% 16.9
this model (GRPO + dense reward) 14.67% 29.17% 7.50% 0.00% 17.1

44× over the base floor (1/300 → 44/300). Gain concentrated in easy/medium puzzles where the model lands near the target. Known failure mode: reasoning collapse — the reward credits only the <answer>, so format_rate stays 0 (no <think>).

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "s1lv3rj1nx/countdown-qwen2.5-0.5b-grpo-dense"
model = AutoModelForCausalLM.from_pretrained(repo)
tok = AutoTokenizer.from_pretrained(repo)

Trained for the RLVR Arena capstone (RL in Production Bootcamp).

Description
Model synced from source: s1lv3rj1nx/countdown-qwen2.5-0.5b-grpo-dense
Readme 2 MiB