--- license: apache-2.0 base_model: Qwen/Qwen2.5-7B datasets: - JoeYing/ReTool-SFT language: - en tags: - retool - tool-use - code-interpreter - math - sft - slime - megatron library_name: transformers pipeline_tag: text-generation --- # qwen2.5-7b-base-retool-slime-sft ReTool **SFT cold-start** of `Qwen/Qwen2.5-7B` (base, *not* -Instruct). This stage teaches the multi-turn **`` / `` / `\boxed{}`** tool-use format on [`JoeYing/ReTool-SFT`](https://huggingface.co/datasets/JoeYing/ReTool-SFT) so the model can write Python, read sandbox output, and finish with a boxed answer. It is the first stage of the `retool-slime` pipeline (**base → tool-format SFT → RL (GRPO / GFlowRL)**). ## Training - **Framework:** [slime](https://github.com/THUDM/slime) v0.3.0 (Megatron-LM backend). - **Data:** `JoeYing/ReTool-SFT` (2,000 multi-turn `{messages}` trajectories). - **Recipe:** `loss_type=sft_loss`, 3 epochs, rollout/global batch 128, Adam, lr 1e-5 (cosine → 1e-6, 10% warmup), weight decay 0.1, bf16, `max_tokens_per_gpu=9216`, full activation recompute. - **Parallelism:** 8×A100-40GB, tensor-parallel 4 (+ sequence parallel), PP 1. - **Final train loss:** ≈ 0.37–0.40 (from ≈ 0.70 at start). The prompt/response template is owned by the rollout (`generate_with_retool`); chat templates are **not** applied on top during RL. ## Intended use Consume directly as the `--hf-checkpoint` / `MODEL_PATH` for the GRPO / GFlowRL stages of `retool-slime`, or for tool-augmented math inference using the ReTool ``/`` format.