Files
grpo-finetune-walkthrough/README.md
ModelHub XC 956041038b 初始化项目,由ModelHub XC社区提供模型
Model: narinzar/grpo-finetune-walkthrough
Source: Original Platform
2026-08-25 08:33:19 +08:00

56 lines
2.0 KiB
Markdown

---
license: apache-2.0
base_model: Qwen/Qwen2.5-0.5B-Instruct
pipeline_tag: text-generation
tags:
- grpo
- reinforcement-learning
- rlvr
- text-generation
---
# grpo-finetune-walkthrough
A `Qwen/Qwen2.5-0.5B-Instruct` policy fine-tuned with GRPO (Group Relative
Policy Optimization) on a verifiable arithmetic task, using a from-scratch
implementation of the group-relative advantage and policy-gradient loss (no TRL
for the core math). Code: https://github.com/narinzar/grpo-finetune-walkthrough
## What this is
This is a minimal 20-step demonstration run on a single RTX 5090, not a
converged model. Its purpose is to show that the from-scratch GRPO training loop
works end to end and produces a rising reward signal. Run configuration: 20
steps, group size 4, 2 prompts per step, `max_new_tokens` 64, evaluated on 80
held-out arithmetic items.
## Results
| Metric | Value |
| -------------------------------- | ------------ |
| Mean reward, step 0 | 0.125 |
| Mean reward, step 19 | 0.250 |
| Eval exact-match accuracy before | 0.000 (0/80) |
| Eval exact-match accuracy after | 0.013 (1/80) |
The mean reward roughly doubles across the run (0.125 to 0.25), confirming the
GRPO update is optimizing the policy in the right direction. At this tiny scale
the eval accuracy barely moves (0/80 to 1/80) and format compliance has not
emerged yet. Real accuracy and format gains need far more steps, many more
prompts per step, and longer generations. Treat these numbers as a correctness
check on the training loop, not as a measure of how well GRPO can do on this
task.
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("narinzar/grpo-finetune-walkthrough")
tokenizer = AutoTokenizer.from_pretrained("narinzar/grpo-finetune-walkthrough")
```
- Base model: `Qwen/Qwen2.5-0.5B-Instruct`
- Task: text generation / GRPO RL fine-tuning on verifiable arithmetic
- License: apache-2.0