Files
ModelHub XC 65288a8af6 初始化项目,由ModelHub XC社区提供模型
Model: BillyWang1/qwen2.5-7b-base-retool-slime-sft-v2
Source: Original Platform
2026-08-08 15:58:30 +08:00

1.6 KiB
Raw Permalink Blame History

license, base_model, tags, library_name, pipeline_tag
license base_model tags library_name pipeline_tag
apache-2.0 Qwen/Qwen2.5-7B
retool
sft
tool-use
slime
megatron
transformers text-generation

qwen2.5-7b-base-retool-slime-sft-v2

ReTool SFT cold-start of Qwen/Qwen2.5-7B (base, not Instruct), trained with slime on a Megatron backend. This is the tool-format SFT stage that teaches the base model the ReTool interleaved code/tool-call format before RL.

Training configuration

Setting Value
Base model Qwen/Qwen2.5-7B (base)
Dataset ReTool-SFT (messages format)
Loss sft_loss, per-token, Qwen loss mask
Global batch size 32
Epochs 6 (~371 optimizer steps over the ~2k-row corpus, ~62 steps/epoch)
Optimizer Adam (β1=0.9, β2=0.95), weight decay 0.01
LR schedule 1e-5 → 1e-6, cosine, 10% warmup
Precision bf16, grads all-reduced in fp32
Parallelism TP=4, PP=1, CP=1, sequence-parallel
Hardware 8× A100-40GB
Final train loss ~0.02

The checkpoint corresponds to iteration 371 (end of epoch 6).

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "BillyWang1/qwen2.5-7b-base-retool-slime-sft-v2",
    torch_dtype="bfloat16",
    device_map="auto",
)
tok = AutoTokenizer.from_pretrained("BillyWang1/qwen2.5-7b-base-retool-slime-sft-v2")

Notes

This is an intermediate SFT cold-start checkpoint intended as the starting point for subsequent RL (GRPO / GFlow-RL) in the ReTool pipeline.