Files
ModelHub XC 9aabae373d 初始化项目,由ModelHub XC社区提供模型
Model: SeongryongJung/qwen3-4b-tooluse-srpo-ema005
Source: Original Platform
2026-09-20 16:14:29 +08:00

66 lines
1.8 KiB
Markdown

---
license: apache-2.0
base_model: Qwen/Qwen3-4B
library_name: transformers
tags:
- qwen3
- reinforcement-learning
- sdpo
- srpo
- ema
---
# qwen3-4b-tooluse-srpo-ema005
This repository contains the **last checkpoint** (`global_step_100`) for `qwen3gen-tooluse-SRPO-Qwen-Qwen3-4B-mbs32-ema0.05-dwtrue-train64-rollout8-lr5e-6-vllm0.8`,
converted to Hugging Face Transformers format.
## Evaluation
The reported headline score is the **best validation `mean@16` observed during training**.
It is not necessarily the score of the uploaded last checkpoint.
| Dataset | Method | Model | Uploaded checkpoint | Best val mean@16 | Best step | Final val mean@16 |
|---|---|---|---|---:|---:|---:|
| tooluse | SRPO | Qwen3-4B | global_step_100 | 62.59% | 40 | 57.26% |
![Training and validation scores](results/training_score.png)
Raw result files:
- `results/validation_mean16.csv`
- `results/training_scores.csv`
- `artifacts/config.yaml`
- `artifacts/wandb-summary.json`
## Training Setup
- Base model: `Qwen/Qwen3-4B`
- Dataset: `tooluse`
- Method: `SRPO`
- EMA teacher update rate: `0.05`
- Uploaded weights: last checkpoint, `global_step_100`
- Validation metric used for headline score: `val-aux/*/mean@16`
- Validation sampling: `n=16`
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
repo_id = "SeongryongJung/qwen3-4b-tooluse-srpo-ema005"
tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
repo_id,
torch_dtype="auto",
device_map="auto",
trust_remote_code=True,
)
```
## Notes
Most best-validation intermediate checkpoints were not retained as full actor checkpoints because
training kept only the latest actor checkpoint. Therefore, this repository publishes the last
checkpoint and records the best validation score separately.