Files
qwen3-4b-instruct-2507-reto…/README.md
ModelHub XC 7d2ef7965b 初始化项目,由ModelHub XC社区提供模型
Model: BillyWang1/qwen3-4b-instruct-2507-retool-grpo
Source: Original Platform
2026-09-13 15:17:25 +08:00

61 lines
2.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
license: apache-2.0
base_model: Qwen/Qwen3-4B-Instruct-2507
pipeline_tag: text-generation
language:
- en
tags:
- reinforcement-learning
- grpo
- tool-use
- code-interpreter
- math
- retool
- slime
---
# qwen3-4b-instruct-2507-retool-grpo
**Qwen3-4B-Instruct-2507 trained with GRPO for tool-integrated math reasoning (ReTool-style code-interpreter RL).**
This is the GRPO control run of a 26summer series comparing RL objectives (GRPO, GFlowRL, and process-reward variants) on identical data, seed, and infrastructure. The model interleaves natural-language reasoning with native Qwen3 `code_interpreter` tool calls (JSON tool-call format from the tokenizer's own chat template — no custom tags) and executes Python to verify intermediate steps before committing to a final boxed answer.
## Training setup
| | |
|---|---|
| Base model | [Qwen/Qwen3-4B-Instruct-2507](https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507) |
| Algorithm | GRPO (group-normalized outcome advantages, no KL penalty) |
| Framework | [slime](https://github.com/THUDM/slime) 0.3.0 (Megatron-LM training + SGLang rollouts, async RL) |
| Data | dapo-math-17k, 1 epoch = 271 rollouts / 1084 optimizer steps |
| Rollout geometry | 64 prompts × 16 samples per rollout, global batch 256 |
| Reward | rule-based ±1 on the final `\boxed{...}` answer |
| Tool | sandboxed Python `code_interpreter`, multi-turn, native chat-template tool calls |
| Max response length | 8192 (train) / 16384 (eval) |
| LR / seed | 1e-6 constant / 42 |
| Hardware | 1 node × 8 B200 (2 training + 6 inference GPUs, disaggregated) |
## Results (final checkpoint, step 270)
Evaluated with the code interpreter at 16k response budget; AIME scores are mean accuracy over 16 samples/problem.
| benchmark | accuracy |
|---|---|
| AIME 2024 | 0.581 |
| AIME 2025 | 0.498 |
| MATH-500 | 0.956 |
## Usage
Standard Qwen3 instruct usage; to reproduce the tool-use behavior, serve with a `code_interpreter` tool in the chat template and the training system prompt:
> You are a helpful assistant that solves math problems step by step. You may call the code_interpreter tool to execute Python code whenever it helps your reasoning; use complete scripts including any imports. End your solution with the final answer in \boxed{...}.
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("BillyWang1/qwen3-4b-instruct-2507-retool-grpo", torch_dtype="bfloat16")
tokenizer = AutoTokenizer.from_pretrained("BillyWang1/qwen3-4b-instruct-2507-retool-grpo")
```
The model was trained purely with RL on top of the instruct model — no SFT stage — so it retains the base model's general chat ability.