初始化项目,由ModelHub XC社区提供模型
Model: SeongryongJung/qwen3-4b-tooluse-srpo-ema005 Source: Original Platform
This commit is contained in:
65
README.md
Normal file
65
README.md
Normal file
@@ -0,0 +1,65 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
base_model: Qwen/Qwen3-4B
|
||||
library_name: transformers
|
||||
tags:
|
||||
- qwen3
|
||||
- reinforcement-learning
|
||||
- sdpo
|
||||
- srpo
|
||||
- ema
|
||||
---
|
||||
|
||||
# qwen3-4b-tooluse-srpo-ema005
|
||||
|
||||
This repository contains the **last checkpoint** (`global_step_100`) for `qwen3gen-tooluse-SRPO-Qwen-Qwen3-4B-mbs32-ema0.05-dwtrue-train64-rollout8-lr5e-6-vllm0.8`,
|
||||
converted to Hugging Face Transformers format.
|
||||
|
||||
## Evaluation
|
||||
|
||||
The reported headline score is the **best validation `mean@16` observed during training**.
|
||||
It is not necessarily the score of the uploaded last checkpoint.
|
||||
|
||||
| Dataset | Method | Model | Uploaded checkpoint | Best val mean@16 | Best step | Final val mean@16 |
|
||||
|---|---|---|---|---:|---:|---:|
|
||||
| tooluse | SRPO | Qwen3-4B | global_step_100 | 62.59% | 40 | 57.26% |
|
||||
|
||||

|
||||
|
||||
Raw result files:
|
||||
|
||||
- `results/validation_mean16.csv`
|
||||
- `results/training_scores.csv`
|
||||
- `artifacts/config.yaml`
|
||||
- `artifacts/wandb-summary.json`
|
||||
|
||||
## Training Setup
|
||||
|
||||
- Base model: `Qwen/Qwen3-4B`
|
||||
- Dataset: `tooluse`
|
||||
- Method: `SRPO`
|
||||
- EMA teacher update rate: `0.05`
|
||||
- Uploaded weights: last checkpoint, `global_step_100`
|
||||
- Validation metric used for headline score: `val-aux/*/mean@16`
|
||||
- Validation sampling: `n=16`
|
||||
|
||||
## Usage
|
||||
|
||||
```python
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
|
||||
repo_id = "SeongryongJung/qwen3-4b-tooluse-srpo-ema005"
|
||||
tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)
|
||||
model = AutoModelForCausalLM.from_pretrained(
|
||||
repo_id,
|
||||
torch_dtype="auto",
|
||||
device_map="auto",
|
||||
trust_remote_code=True,
|
||||
)
|
||||
```
|
||||
|
||||
## Notes
|
||||
|
||||
Most best-validation intermediate checkpoints were not retained as full actor checkpoints because
|
||||
training kept only the latest actor checkpoint. Therefore, this repository publishes the last
|
||||
checkpoint and records the best validation score separately.
|
||||
Reference in New Issue
Block a user