Files
ModelHub XC 5d09f285e4 初始化项目,由ModelHub XC社区提供模型
Model: s1lv3rj1nx/countdown-qwen2.5-0.5b-grpo-dense
Source: Original Platform
2026-07-15 04:50:12 +08:00

49 lines
1.8 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
license: apache-2.0
base_model: Qwen/Qwen2.5-0.5B-Instruct
library_name: transformers
pipeline_tag: text-generation
tags:
- grpo
- rlvr
- reinforcement-learning
- countdown
- reasoning
- reward-shaping
---
# Countdown-Qwen2.5-0.5B-GRPO (dense reward)
`Qwen2.5-0.5B-Instruct` fine-tuned with **from-scratch GRPO** on the **Countdown** number-puzzle
task (combine the given numbers exactly once with `+ - * /` to hit a target). Rewards are
**verifiable** (exact rational arithmetic, whitelisted-AST evaluator) — RLVR: no reward model,
no critic.
This is the **best checkpoint**: GRPO at lr `3e-6`, group `8`, `1500` steps, trained with a
**dense closeness-shaped reward** — for a right-numbers/wrong-value attempt the reward scales
with proximity to the target (`0.10 + 0.85·exp(-|value-target|/10)`) instead of a flat `0.10`.
That gives GRPO a gradient on near-misses and broke the ~12% plateau of the step-function reward.
## Results (dev_public, 300 puzzles, greedy, exact verifier)
| model | accuracy | easy | medium | hard | avg_tokens (correct) |
|---|---|---|---|---|---|
| base Qwen2.5-0.5B-Instruct (floor) | 0.33% | — | — | 0.00% | 20.0 |
| GRPO, original step reward | 12.00% | 25.83% | 3.33% | 1.67% | 16.9 |
| **this model (GRPO + dense reward)** | **14.67%** | 29.17% | 7.50% | 0.00% | 17.1 |
**44× over the base floor** (1/300 → 44/300). Gain concentrated in easy/medium puzzles where the
model lands near the target. Known failure mode: reasoning collapse — the reward credits only the
`<answer>`, so `format_rate` stays 0 (no `<think>`).
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "s1lv3rj1nx/countdown-qwen2.5-0.5b-grpo-dense"
model = AutoModelForCausalLM.from_pretrained(repo)
tok = AutoTokenizer.from_pretrained(repo)
```
Trained for the RLVR Arena capstone (RL in Production Bootcamp).