初始化项目,由ModelHub XC社区提供模型
Model: BillyWang1/qwen3-4b-instruct-2507-retool-grpo Source: Original Platform
This commit is contained in:
60
README.md
Normal file
60
README.md
Normal file
@@ -0,0 +1,60 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
base_model: Qwen/Qwen3-4B-Instruct-2507
|
||||
pipeline_tag: text-generation
|
||||
language:
|
||||
- en
|
||||
tags:
|
||||
- reinforcement-learning
|
||||
- grpo
|
||||
- tool-use
|
||||
- code-interpreter
|
||||
- math
|
||||
- retool
|
||||
- slime
|
||||
---
|
||||
|
||||
# qwen3-4b-instruct-2507-retool-grpo
|
||||
|
||||
**Qwen3-4B-Instruct-2507 trained with GRPO for tool-integrated math reasoning (ReTool-style code-interpreter RL).**
|
||||
|
||||
This is the GRPO control run of a 26summer series comparing RL objectives (GRPO, GFlowRL, and process-reward variants) on identical data, seed, and infrastructure. The model interleaves natural-language reasoning with native Qwen3 `code_interpreter` tool calls (JSON tool-call format from the tokenizer's own chat template — no custom tags) and executes Python to verify intermediate steps before committing to a final boxed answer.
|
||||
|
||||
## Training setup
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| Base model | [Qwen/Qwen3-4B-Instruct-2507](https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507) |
|
||||
| Algorithm | GRPO (group-normalized outcome advantages, no KL penalty) |
|
||||
| Framework | [slime](https://github.com/THUDM/slime) 0.3.0 (Megatron-LM training + SGLang rollouts, async RL) |
|
||||
| Data | dapo-math-17k, 1 epoch = 271 rollouts / 1084 optimizer steps |
|
||||
| Rollout geometry | 64 prompts × 16 samples per rollout, global batch 256 |
|
||||
| Reward | rule-based ±1 on the final `\boxed{...}` answer |
|
||||
| Tool | sandboxed Python `code_interpreter`, multi-turn, native chat-template tool calls |
|
||||
| Max response length | 8192 (train) / 16384 (eval) |
|
||||
| LR / seed | 1e-6 constant / 42 |
|
||||
| Hardware | 1 node × 8 B200 (2 training + 6 inference GPUs, disaggregated) |
|
||||
|
||||
## Results (final checkpoint, step 270)
|
||||
|
||||
Evaluated with the code interpreter at 16k response budget; AIME scores are mean accuracy over 16 samples/problem.
|
||||
|
||||
| benchmark | accuracy |
|
||||
|---|---|
|
||||
| AIME 2024 | 0.581 |
|
||||
| AIME 2025 | 0.498 |
|
||||
| MATH-500 | 0.956 |
|
||||
|
||||
## Usage
|
||||
|
||||
Standard Qwen3 instruct usage; to reproduce the tool-use behavior, serve with a `code_interpreter` tool in the chat template and the training system prompt:
|
||||
|
||||
> You are a helpful assistant that solves math problems step by step. You may call the code_interpreter tool to execute Python code whenever it helps your reasoning; use complete scripts including any imports. End your solution with the final answer in \boxed{...}.
|
||||
|
||||
```python
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
model = AutoModelForCausalLM.from_pretrained("BillyWang1/qwen3-4b-instruct-2507-retool-grpo", torch_dtype="bfloat16")
|
||||
tokenizer = AutoTokenizer.from_pretrained("BillyWang1/qwen3-4b-instruct-2507-retool-grpo")
|
||||
```
|
||||
|
||||
The model was trained purely with RL on top of the instruct model — no SFT stage — so it retains the base model's general chat ability.
|
||||
Reference in New Issue
Block a user