56 lines
1.6 KiB
Markdown
56 lines
1.6 KiB
Markdown
|
|
---
|
|||
|
|
license: apache-2.0
|
|||
|
|
base_model: Qwen/Qwen2.5-7B
|
|||
|
|
tags:
|
|||
|
|
- retool
|
|||
|
|
- sft
|
|||
|
|
- tool-use
|
|||
|
|
- slime
|
|||
|
|
- megatron
|
|||
|
|
library_name: transformers
|
|||
|
|
pipeline_tag: text-generation
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# qwen2.5-7b-base-retool-slime-sft-v2
|
|||
|
|
|
|||
|
|
ReTool SFT cold-start of **Qwen/Qwen2.5-7B** (base, *not* Instruct), trained with
|
|||
|
|
[slime](https://github.com/THUDM/slime) on a Megatron backend. This is the tool-format
|
|||
|
|
SFT stage that teaches the base model the ReTool interleaved code/tool-call format before
|
|||
|
|
RL.
|
|||
|
|
|
|||
|
|
## Training configuration
|
|||
|
|
|
|||
|
|
| Setting | Value |
|
|||
|
|
|---|---|
|
|||
|
|
| Base model | `Qwen/Qwen2.5-7B` (base) |
|
|||
|
|
| Dataset | [ReTool-SFT](https://huggingface.co/datasets/JoeYing/ReTool-SFT) (`messages` format) |
|
|||
|
|
| Loss | `sft_loss`, per-token, Qwen loss mask |
|
|||
|
|
| **Global batch size** | **32** |
|
|||
|
|
| **Epochs** | **6** (~371 optimizer steps over the ~2k-row corpus, ~62 steps/epoch) |
|
|||
|
|
| Optimizer | Adam (β1=0.9, β2=0.95), weight decay 0.01 |
|
|||
|
|
| LR schedule | 1e-5 → 1e-6, cosine, 10% warmup |
|
|||
|
|
| Precision | bf16, grads all-reduced in fp32 |
|
|||
|
|
| Parallelism | TP=4, PP=1, CP=1, sequence-parallel |
|
|||
|
|
| Hardware | 8× A100-40GB |
|
|||
|
|
| Final train loss | ~0.02 |
|
|||
|
|
|
|||
|
|
The checkpoint corresponds to iteration **371** (end of epoch 6).
|
|||
|
|
|
|||
|
|
## Usage
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
from transformers import AutoModelForCausalLM, AutoTokenizer
|
|||
|
|
|
|||
|
|
model = AutoModelForCausalLM.from_pretrained(
|
|||
|
|
"BillyWang1/qwen2.5-7b-base-retool-slime-sft-v2",
|
|||
|
|
torch_dtype="bfloat16",
|
|||
|
|
device_map="auto",
|
|||
|
|
)
|
|||
|
|
tok = AutoTokenizer.from_pretrained("BillyWang1/qwen2.5-7b-base-retool-slime-sft-v2")
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
## Notes
|
|||
|
|
|
|||
|
|
This is an intermediate SFT cold-start checkpoint intended as the starting point for
|
|||
|
|
subsequent RL (GRPO / GFlow-RL) in the ReTool pipeline.
|