初始化项目,由ModelHub XC社区提供模型
Model: s1lv3rj1nx/countdown-qwen2.5-0.5b-grpo-lr3e6 Source: Original Platform
This commit is contained in:
44
README.md
Normal file
44
README.md
Normal file
@@ -0,0 +1,44 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
base_model: Qwen/Qwen2.5-0.5B-Instruct
|
||||
library_name: transformers
|
||||
pipeline_tag: text-generation
|
||||
tags:
|
||||
- grpo
|
||||
- rlvr
|
||||
- reinforcement-learning
|
||||
- countdown
|
||||
- reasoning
|
||||
---
|
||||
|
||||
# Countdown-Qwen2.5-0.5B-GRPO (lr 3e-6 ablation)
|
||||
|
||||
`Qwen2.5-0.5B-Instruct` fine-tuned with **from-scratch GRPO** (Group-Relative Policy
|
||||
Optimization) on the **Countdown** number-puzzle task — combine the given numbers exactly
|
||||
once with `+ - * /` to hit a target. Rewards are **verifiable** (exact rational arithmetic
|
||||
via a whitelisted-AST evaluator), so this is RLVR: no reward model, no critic.
|
||||
|
||||
This checkpoint is the **best ablation** (learning rate `3e-6`, group `8`, `1500` steps).
|
||||
|
||||
## Results (dev_public, 300 puzzles, greedy, exact verifier)
|
||||
|
||||
| model | accuracy | hard (5-num) | format_rate | avg_tokens (correct) |
|
||||
|---|---|---|---|---|
|
||||
| base Qwen2.5-0.5B-Instruct (floor) | 0.33% | 0.00% | 0.00% | 20.0 |
|
||||
| **this model (GRPO, lr 3e-6)** | **12.00%** | **1.67%** | 0.00% | 16.9 |
|
||||
|
||||
**36x over the base floor** (1/300 -> 36/300). Ablations: learning rate was the only lever
|
||||
that moved accuracy (1e-6 -> 3e-6 gave 1% -> 12%); more steps (3000) and larger group (16)
|
||||
both plateaued. Known failure mode: reasoning collapse — the shaped reward scores only the
|
||||
`<answer>`, so the model emits bare answers with no `<think>` (format_rate = 0).
|
||||
|
||||
## Usage
|
||||
|
||||
```python
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
|
||||
model = AutoModelForCausalLM.from_pretrained("<your-username>/countdown-qwen2.5-0.5b-grpo-lr3e6")
|
||||
tok = AutoTokenizer.from_pretrained("<your-username>/countdown-qwen2.5-0.5b-grpo-lr3e6")
|
||||
```
|
||||
|
||||
Trained for the RLVR Arena capstone (RL in Production Bootcamp).
|
||||
Reference in New Issue
Block a user