Model: BillyWang1/qwen2.5-7b-base-retool-slime-sft Source: Original Platform
license, base_model, datasets, language, tags, library_name, pipeline_tag
| license | base_model | datasets | language | tags | library_name | pipeline_tag | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| apache-2.0 | Qwen/Qwen2.5-7B |
|
|
|
transformers | text-generation |
qwen2.5-7b-base-retool-slime-sft
ReTool SFT cold-start of Qwen/Qwen2.5-7B (base, not -Instruct). This stage teaches
the multi-turn <code> / <interpreter> / \boxed{} tool-use format on
JoeYing/ReTool-SFT so the model can
write Python, read sandbox output, and finish with a boxed answer. It is the first stage of
the retool-slime pipeline (base → tool-format SFT → RL (GRPO / GFlowRL)).
Training
- Framework: slime v0.3.0 (Megatron-LM backend).
- Data:
JoeYing/ReTool-SFT(2,000 multi-turn{messages}trajectories). - Recipe:
loss_type=sft_loss, 3 epochs, rollout/global batch 128, Adam, lr 1e-5 (cosine → 1e-6, 10% warmup), weight decay 0.1, bf16,max_tokens_per_gpu=9216, full activation recompute. - Parallelism: 8×A100-40GB, tensor-parallel 4 (+ sequence parallel), PP 1.
- Final train loss: ≈ 0.37–0.40 (from ≈ 0.70 at start).
The prompt/response template is owned by the rollout (generate_with_retool); chat templates
are not applied on top during RL.
Intended use
Consume directly as the --hf-checkpoint / MODEL_PATH for the GRPO / GFlowRL stages of
retool-slime, or for tool-augmented math inference using the ReTool <code>/<interpreter>
format.
Description