46 lines
1.5 KiB
Markdown
46 lines
1.5 KiB
Markdown
---
|
||
license: apache-2.0
|
||
base_model: Qwen/Qwen2.5-7B
|
||
datasets:
|
||
- JoeYing/ReTool-SFT
|
||
language:
|
||
- en
|
||
tags:
|
||
- retool
|
||
- tool-use
|
||
- code-interpreter
|
||
- math
|
||
- sft
|
||
- slime
|
||
- megatron
|
||
library_name: transformers
|
||
pipeline_tag: text-generation
|
||
---
|
||
|
||
# qwen2.5-7b-base-retool-slime-sft
|
||
|
||
ReTool **SFT cold-start** of `Qwen/Qwen2.5-7B` (base, *not* -Instruct). This stage teaches
|
||
the multi-turn **`<code>` / `<interpreter>` / `\boxed{}`** tool-use format on
|
||
[`JoeYing/ReTool-SFT`](https://huggingface.co/datasets/JoeYing/ReTool-SFT) so the model can
|
||
write Python, read sandbox output, and finish with a boxed answer. It is the first stage of
|
||
the `retool-slime` pipeline (**base → tool-format SFT → RL (GRPO / GFlowRL)**).
|
||
|
||
## Training
|
||
|
||
- **Framework:** [slime](https://github.com/THUDM/slime) v0.3.0 (Megatron-LM backend).
|
||
- **Data:** `JoeYing/ReTool-SFT` (2,000 multi-turn `{messages}` trajectories).
|
||
- **Recipe:** `loss_type=sft_loss`, 3 epochs, rollout/global batch 128, Adam, lr 1e-5
|
||
(cosine → 1e-6, 10% warmup), weight decay 0.1, bf16, `max_tokens_per_gpu=9216`,
|
||
full activation recompute.
|
||
- **Parallelism:** 8×A100-40GB, tensor-parallel 4 (+ sequence parallel), PP 1.
|
||
- **Final train loss:** ≈ 0.37–0.40 (from ≈ 0.70 at start).
|
||
|
||
The prompt/response template is owned by the rollout (`generate_with_retool`); chat templates
|
||
are **not** applied on top during RL.
|
||
|
||
## Intended use
|
||
|
||
Consume directly as the `--hf-checkpoint` / `MODEL_PATH` for the GRPO / GFlowRL stages of
|
||
`retool-slime`, or for tool-augmented math inference using the ReTool `<code>`/`<interpreter>`
|
||
format.
|