Files
ModelHub XC 7ce12d72cf 初始化项目,由ModelHub XC社区提供模型
Model: BillyWang1/qwen2.5-7b-base-retool-slime-sft
Source: Original Platform
2026-07-17 16:33:13 +08:00

46 lines
1.5 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
license: apache-2.0
base_model: Qwen/Qwen2.5-7B
datasets:
- JoeYing/ReTool-SFT
language:
- en
tags:
- retool
- tool-use
- code-interpreter
- math
- sft
- slime
- megatron
library_name: transformers
pipeline_tag: text-generation
---
# qwen2.5-7b-base-retool-slime-sft
ReTool **SFT cold-start** of `Qwen/Qwen2.5-7B` (base, *not* -Instruct). This stage teaches
the multi-turn **`<code>` / `<interpreter>` / `\boxed{}`** tool-use format on
[`JoeYing/ReTool-SFT`](https://huggingface.co/datasets/JoeYing/ReTool-SFT) so the model can
write Python, read sandbox output, and finish with a boxed answer. It is the first stage of
the `retool-slime` pipeline (**base → tool-format SFT → RL (GRPO / GFlowRL)**).
## Training
- **Framework:** [slime](https://github.com/THUDM/slime) v0.3.0 (Megatron-LM backend).
- **Data:** `JoeYing/ReTool-SFT` (2,000 multi-turn `{messages}` trajectories).
- **Recipe:** `loss_type=sft_loss`, 3 epochs, rollout/global batch 128, Adam, lr 1e-5
(cosine → 1e-6, 10% warmup), weight decay 0.1, bf16, `max_tokens_per_gpu=9216`,
full activation recompute.
- **Parallelism:** 8×A100-40GB, tensor-parallel 4 (+ sequence parallel), PP 1.
- **Final train loss:** ≈ 0.37–0.40 (from ≈ 0.70 at start).
The prompt/response template is owned by the rollout (`generate_with_retool`); chat templates
are **not** applied on top during RL.
## Intended use
Consume directly as the `--hf-checkpoint` / `MODEL_PATH` for the GRPO / GFlowRL stages of
`retool-slime`, or for tool-augmented math inference using the ReTool `<code>`/`<interpreter>`
format.