Model: BillyWang1/qwen2.5-7b-base-retool-slime-sft-v2 Source: Original Platform
license, base_model, tags, library_name, pipeline_tag
| license | base_model | tags | library_name | pipeline_tag | |||||
|---|---|---|---|---|---|---|---|---|---|
| apache-2.0 | Qwen/Qwen2.5-7B |
|
transformers | text-generation |
qwen2.5-7b-base-retool-slime-sft-v2
ReTool SFT cold-start of Qwen/Qwen2.5-7B (base, not Instruct), trained with slime on a Megatron backend. This is the tool-format SFT stage that teaches the base model the ReTool interleaved code/tool-call format before RL.
Training configuration
| Setting | Value |
|---|---|
| Base model | Qwen/Qwen2.5-7B (base) |
| Dataset | ReTool-SFT (messages format) |
| Loss | sft_loss, per-token, Qwen loss mask |
| Global batch size | 32 |
| Epochs | 6 (~371 optimizer steps over the ~2k-row corpus, ~62 steps/epoch) |
| Optimizer | Adam (β1=0.9, β2=0.95), weight decay 0.01 |
| LR schedule | 1e-5 → 1e-6, cosine, 10% warmup |
| Precision | bf16, grads all-reduced in fp32 |
| Parallelism | TP=4, PP=1, CP=1, sequence-parallel |
| Hardware | 8× A100-40GB |
| Final train loss | ~0.02 |
The checkpoint corresponds to iteration 371 (end of epoch 6).
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained(
"BillyWang1/qwen2.5-7b-base-retool-slime-sft-v2",
torch_dtype="bfloat16",
device_map="auto",
)
tok = AutoTokenizer.from_pretrained("BillyWang1/qwen2.5-7b-base-retool-slime-sft-v2")
Notes
This is an intermediate SFT cold-start checkpoint intended as the starting point for subsequent RL (GRPO / GFlow-RL) in the ReTool pipeline.
Description