ModelHub XC 7ce12d72cf 初始化项目,由ModelHub XC社区提供模型
Model: BillyWang1/qwen2.5-7b-base-retool-slime-sft
Source: Original Platform
2026-07-17 16:33:13 +08:00

license, base_model, datasets, language, tags, library_name, pipeline_tag
license base_model datasets language tags library_name pipeline_tag
apache-2.0 Qwen/Qwen2.5-7B
JoeYing/ReTool-SFT
en
retool
tool-use
code-interpreter
math
sft
slime
megatron
transformers text-generation

qwen2.5-7b-base-retool-slime-sft

ReTool SFT cold-start of Qwen/Qwen2.5-7B (base, not -Instruct). This stage teaches the multi-turn <code> / <interpreter> / \boxed{} tool-use format on JoeYing/ReTool-SFT so the model can write Python, read sandbox output, and finish with a boxed answer. It is the first stage of the retool-slime pipeline (base → tool-format SFT → RL (GRPO / GFlowRL)).

Training

  • Framework: slime v0.3.0 (Megatron-LM backend).
  • Data: JoeYing/ReTool-SFT (2,000 multi-turn {messages} trajectories).
  • Recipe: loss_type=sft_loss, 3 epochs, rollout/global batch 128, Adam, lr 1e-5 (cosine → 1e-6, 10% warmup), weight decay 0.1, bf16, max_tokens_per_gpu=9216, full activation recompute.
  • Parallelism: 8×A100-40GB, tensor-parallel 4 (+ sequence parallel), PP 1.
  • Final train loss: ≈ 0.37–0.40 (from ≈ 0.70 at start).

The prompt/response template is owned by the rollout (generate_with_retool); chat templates are not applied on top during RL.

Intended use

Consume directly as the --hf-checkpoint / MODEL_PATH for the GRPO / GFlowRL stages of retool-slime, or for tool-augmented math inference using the ReTool <code>/<interpreter> format.

Description
Model synced from source: BillyWang1/qwen2.5-7b-base-retool-slime-sft
Readme 4.2 MiB