Files
countdown-qwen2.5-0.5b-grpo…/README.md
ModelHub XC fdd40a0b18 初始化项目,由ModelHub XC社区提供模型
Model: s1lv3rj1nx/countdown-qwen2.5-0.5b-grpo-lr3e6
Source: Original Platform
2026-07-19 07:47:11 +08:00

1.7 KiB

license, base_model, library_name, pipeline_tag, tags
license base_model library_name pipeline_tag tags
apache-2.0 Qwen/Qwen2.5-0.5B-Instruct transformers text-generation
grpo
rlvr
reinforcement-learning
countdown
reasoning

Countdown-Qwen2.5-0.5B-GRPO (lr 3e-6 ablation)

Qwen2.5-0.5B-Instruct fine-tuned with from-scratch GRPO (Group-Relative Policy Optimization) on the Countdown number-puzzle task — combine the given numbers exactly once with + - * / to hit a target. Rewards are verifiable (exact rational arithmetic via a whitelisted-AST evaluator), so this is RLVR: no reward model, no critic.

This checkpoint is the best ablation (learning rate 3e-6, group 8, 1500 steps).

Results (dev_public, 300 puzzles, greedy, exact verifier)

model accuracy hard (5-num) format_rate avg_tokens (correct)
base Qwen2.5-0.5B-Instruct (floor) 0.33% 0.00% 0.00% 20.0
this model (GRPO, lr 3e-6) 12.00% 1.67% 0.00% 16.9

36x over the base floor (1/300 -> 36/300). Ablations: learning rate was the only lever that moved accuracy (1e-6 -> 3e-6 gave 1% -> 12%); more steps (3000) and larger group (16) both plateaued. Known failure mode: reasoning collapse — the shaped reward scores only the <answer>, so the model emits bare answers with no <think> (format_rate = 0).

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("<your-username>/countdown-qwen2.5-0.5b-grpo-lr3e6")
tok = AutoTokenizer.from_pretrained("<your-username>/countdown-qwen2.5-0.5b-grpo-lr3e6")

Trained for the RLVR Arena capstone (RL in Production Bootcamp).