初始化项目,由ModelHub XC社区提供模型

Model: BillyWang1/qwen2.5-7b-base-retool-slime-sft
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-17 16:33:13 +08:00
commit 7ce12d72cf
39 changed files with 455428 additions and 0 deletions

45
README.md Normal file
View File

@@ -0,0 +1,45 @@
---
license: apache-2.0
base_model: Qwen/Qwen2.5-7B
datasets:
- JoeYing/ReTool-SFT
language:
- en
tags:
- retool
- tool-use
- code-interpreter
- math
- sft
- slime
- megatron
library_name: transformers
pipeline_tag: text-generation
---
# qwen2.5-7b-base-retool-slime-sft
ReTool **SFT cold-start** of `Qwen/Qwen2.5-7B` (base, *not* -Instruct). This stage teaches
the multi-turn **`<code>` / `<interpreter>` / `\boxed{}`** tool-use format on
[`JoeYing/ReTool-SFT`](https://huggingface.co/datasets/JoeYing/ReTool-SFT) so the model can
write Python, read sandbox output, and finish with a boxed answer. It is the first stage of
the `retool-slime` pipeline (**base → tool-format SFT → RL (GRPO / GFlowRL)**).
## Training
- **Framework:** [slime](https://github.com/THUDM/slime) v0.3.0 (Megatron-LM backend).
- **Data:** `JoeYing/ReTool-SFT` (2,000 multi-turn `{messages}` trajectories).
- **Recipe:** `loss_type=sft_loss`, 3 epochs, rollout/global batch 128, Adam, lr 1e-5
(cosine → 1e-6, 10% warmup), weight decay 0.1, bf16, `max_tokens_per_gpu=9216`,
full activation recompute.
- **Parallelism:** 8×A100-40GB, tensor-parallel 4 (+ sequence parallel), PP 1.
- **Final train loss:** ≈ 0.37–0.40 (from ≈ 0.70 at start).
The prompt/response template is owned by the rollout (`generate_with_retool`); chat templates
are **not** applied on top during RL.
## Intended use
Consume directly as the `--hf-checkpoint` / `MODEL_PATH` for the GRPO / GFlowRL stages of
`retool-slime`, or for tool-augmented math inference using the ReTool `<code>`/`<interpreter>`
format.