初始化项目,由ModelHub XC社区提供模型

Model: SeongryongJung/qwen3-4b-tooluse-srpo-ema005
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-09-20 16:14:29 +08:00
commit 9aabae373d
19 changed files with 153324 additions and 0 deletions

65
README.md Normal file
View File

@@ -0,0 +1,65 @@
---
license: apache-2.0
base_model: Qwen/Qwen3-4B
library_name: transformers
tags:
- qwen3
- reinforcement-learning
- sdpo
- srpo
- ema
---
# qwen3-4b-tooluse-srpo-ema005
This repository contains the **last checkpoint** (`global_step_100`) for `qwen3gen-tooluse-SRPO-Qwen-Qwen3-4B-mbs32-ema0.05-dwtrue-train64-rollout8-lr5e-6-vllm0.8`,
converted to Hugging Face Transformers format.
## Evaluation
The reported headline score is the **best validation `mean@16` observed during training**.
It is not necessarily the score of the uploaded last checkpoint.
| Dataset | Method | Model | Uploaded checkpoint | Best val mean@16 | Best step | Final val mean@16 |
|---|---|---|---|---:|---:|---:|
| tooluse | SRPO | Qwen3-4B | global_step_100 | 62.59% | 40 | 57.26% |
![Training and validation scores](results/training_score.png)
Raw result files:
- `results/validation_mean16.csv`
- `results/training_scores.csv`
- `artifacts/config.yaml`
- `artifacts/wandb-summary.json`
## Training Setup
- Base model: `Qwen/Qwen3-4B`
- Dataset: `tooluse`
- Method: `SRPO`
- EMA teacher update rate: `0.05`
- Uploaded weights: last checkpoint, `global_step_100`
- Validation metric used for headline score: `val-aux/*/mean@16`
- Validation sampling: `n=16`
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
repo_id = "SeongryongJung/qwen3-4b-tooluse-srpo-ema005"
tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
repo_id,
torch_dtype="auto",
device_map="auto",
trust_remote_code=True,
)
```
## Notes
Most best-validation intermediate checkpoints were not retained as full actor checkpoints because
training kept only the latest actor checkpoint. Therefore, this repository publishes the last
checkpoint and records the best validation score separately.