初始化项目,由ModelHub XC社区提供模型

Model: s1lv3rj1nx/countdown-qwen2.5-0.5b-grpo-lr3e6
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-19 07:47:11 +08:00
commit fdd40a0b18
11 changed files with 151780 additions and 0 deletions

44
README.md Normal file
View File

@@ -0,0 +1,44 @@
---
license: apache-2.0
base_model: Qwen/Qwen2.5-0.5B-Instruct
library_name: transformers
pipeline_tag: text-generation
tags:
- grpo
- rlvr
- reinforcement-learning
- countdown
- reasoning
---
# Countdown-Qwen2.5-0.5B-GRPO (lr 3e-6 ablation)
`Qwen2.5-0.5B-Instruct` fine-tuned with **from-scratch GRPO** (Group-Relative Policy
Optimization) on the **Countdown** number-puzzle task — combine the given numbers exactly
once with `+ - * /` to hit a target. Rewards are **verifiable** (exact rational arithmetic
via a whitelisted-AST evaluator), so this is RLVR: no reward model, no critic.
This checkpoint is the **best ablation** (learning rate `3e-6`, group `8`, `1500` steps).
## Results (dev_public, 300 puzzles, greedy, exact verifier)
| model | accuracy | hard (5-num) | format_rate | avg_tokens (correct) |
|---|---|---|---|---|
| base Qwen2.5-0.5B-Instruct (floor) | 0.33% | 0.00% | 0.00% | 20.0 |
| **this model (GRPO, lr 3e-6)** | **12.00%** | **1.67%** | 0.00% | 16.9 |
**36x over the base floor** (1/300 -> 36/300). Ablations: learning rate was the only lever
that moved accuracy (1e-6 -> 3e-6 gave 1% -> 12%); more steps (3000) and larger group (16)
both plateaued. Known failure mode: reasoning collapse — the shaped reward scores only the
`<answer>`, so the model emits bare answers with no `<think>` (format_rate = 0).
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("<your-username>/countdown-qwen2.5-0.5b-grpo-lr3e6")
tok = AutoTokenizer.from_pretrained("<your-username>/countdown-qwen2.5-0.5b-grpo-lr3e6")
```
Trained for the RLVR Arena capstone (RL in Production Bootcamp).