66 lines
1.8 KiB
Markdown
66 lines
1.8 KiB
Markdown
---
|
|
license: apache-2.0
|
|
base_model: Qwen/Qwen3-4B
|
|
library_name: transformers
|
|
tags:
|
|
- qwen3
|
|
- reinforcement-learning
|
|
- sdpo
|
|
- srpo
|
|
- ema
|
|
---
|
|
|
|
# qwen3-4b-tooluse-srpo-ema005
|
|
|
|
This repository contains the **last checkpoint** (`global_step_100`) for `qwen3gen-tooluse-SRPO-Qwen-Qwen3-4B-mbs32-ema0.05-dwtrue-train64-rollout8-lr5e-6-vllm0.8`,
|
|
converted to Hugging Face Transformers format.
|
|
|
|
## Evaluation
|
|
|
|
The reported headline score is the **best validation `mean@16` observed during training**.
|
|
It is not necessarily the score of the uploaded last checkpoint.
|
|
|
|
| Dataset | Method | Model | Uploaded checkpoint | Best val mean@16 | Best step | Final val mean@16 |
|
|
|---|---|---|---|---:|---:|---:|
|
|
| tooluse | SRPO | Qwen3-4B | global_step_100 | 62.59% | 40 | 57.26% |
|
|
|
|

|
|
|
|
Raw result files:
|
|
|
|
- `results/validation_mean16.csv`
|
|
- `results/training_scores.csv`
|
|
- `artifacts/config.yaml`
|
|
- `artifacts/wandb-summary.json`
|
|
|
|
## Training Setup
|
|
|
|
- Base model: `Qwen/Qwen3-4B`
|
|
- Dataset: `tooluse`
|
|
- Method: `SRPO`
|
|
- EMA teacher update rate: `0.05`
|
|
- Uploaded weights: last checkpoint, `global_step_100`
|
|
- Validation metric used for headline score: `val-aux/*/mean@16`
|
|
- Validation sampling: `n=16`
|
|
|
|
## Usage
|
|
|
|
```python
|
|
from transformers import AutoModelForCausalLM, AutoTokenizer
|
|
|
|
repo_id = "SeongryongJung/qwen3-4b-tooluse-srpo-ema005"
|
|
tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)
|
|
model = AutoModelForCausalLM.from_pretrained(
|
|
repo_id,
|
|
torch_dtype="auto",
|
|
device_map="auto",
|
|
trust_remote_code=True,
|
|
)
|
|
```
|
|
|
|
## Notes
|
|
|
|
Most best-validation intermediate checkpoints were not retained as full actor checkpoints because
|
|
training kept only the latest actor checkpoint. Therefore, this repository publishes the last
|
|
checkpoint and records the best validation score separately.
|