--- license: apache-2.0 base_model: Qwen/Qwen2.5-7B tags: - retool - sft - tool-use - slime - megatron library_name: transformers pipeline_tag: text-generation --- # qwen2.5-7b-base-retool-slime-sft-v2 ReTool SFT cold-start of **Qwen/Qwen2.5-7B** (base, *not* Instruct), trained with [slime](https://github.com/THUDM/slime) on a Megatron backend. This is the tool-format SFT stage that teaches the base model the ReTool interleaved code/tool-call format before RL. ## Training configuration | Setting | Value | |---|---| | Base model | `Qwen/Qwen2.5-7B` (base) | | Dataset | [ReTool-SFT](https://huggingface.co/datasets/JoeYing/ReTool-SFT) (`messages` format) | | Loss | `sft_loss`, per-token, Qwen loss mask | | **Global batch size** | **32** | | **Epochs** | **6** (~371 optimizer steps over the ~2k-row corpus, ~62 steps/epoch) | | Optimizer | Adam (β1=0.9, β2=0.95), weight decay 0.01 | | LR schedule | 1e-5 → 1e-6, cosine, 10% warmup | | Precision | bf16, grads all-reduced in fp32 | | Parallelism | TP=4, PP=1, CP=1, sequence-parallel | | Hardware | 8× A100-40GB | | Final train loss | ~0.02 | The checkpoint corresponds to iteration **371** (end of epoch 6). ## Usage ```python from transformers import AutoModelForCausalLM, AutoTokenizer model = AutoModelForCausalLM.from_pretrained( "BillyWang1/qwen2.5-7b-base-retool-slime-sft-v2", torch_dtype="bfloat16", device_map="auto", ) tok = AutoTokenizer.from_pretrained("BillyWang1/qwen2.5-7b-base-retool-slime-sft-v2") ``` ## Notes This is an intermediate SFT cold-start checkpoint intended as the starting point for subsequent RL (GRPO / GFlow-RL) in the ReTool pipeline.