Files

56 lines
1.6 KiB
Markdown
Raw Permalink Normal View History

---
license: apache-2.0
base_model: Qwen/Qwen2.5-7B
tags:
- retool
- sft
- tool-use
- slime
- megatron
library_name: transformers
pipeline_tag: text-generation
---
# qwen2.5-7b-base-retool-slime-sft-v2
ReTool SFT cold-start of **Qwen/Qwen2.5-7B** (base, *not* Instruct), trained with
[slime](https://github.com/THUDM/slime) on a Megatron backend. This is the tool-format
SFT stage that teaches the base model the ReTool interleaved code/tool-call format before
RL.
## Training configuration
| Setting | Value |
|---|---|
| Base model | `Qwen/Qwen2.5-7B` (base) |
| Dataset | [ReTool-SFT](https://huggingface.co/datasets/JoeYing/ReTool-SFT) (`messages` format) |
| Loss | `sft_loss`, per-token, Qwen loss mask |
| **Global batch size** | **32** |
| **Epochs** | **6** (~371 optimizer steps over the ~2k-row corpus, ~62 steps/epoch) |
| Optimizer | Adam (β1=0.9, β2=0.95), weight decay 0.01 |
| LR schedule | 1e-5 → 1e-6, cosine, 10% warmup |
| Precision | bf16, grads all-reduced in fp32 |
| Parallelism | TP=4, PP=1, CP=1, sequence-parallel |
| Hardware | 8× A100-40GB |
| Final train loss | ~0.02 |
The checkpoint corresponds to iteration **371** (end of epoch 6).
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained(
"BillyWang1/qwen2.5-7b-base-retool-slime-sft-v2",
torch_dtype="bfloat16",
device_map="auto",
)
tok = AutoTokenizer.from_pretrained("BillyWang1/qwen2.5-7b-base-retool-slime-sft-v2")
```
## Notes
This is an intermediate SFT cold-start checkpoint intended as the starting point for
subsequent RL (GRPO / GFlow-RL) in the ReTool pipeline.