license, library_name, pipeline_tag, tags, base_model
license
library_name
pipeline_tag
tags
base_model
apache-2.0
transformers
text-generation
qwen3
reinforcement-learning
grpo
text-generation
Qwen/Qwen3-4B
Qwen3-4B-Material-GRPO-TR
This repository contains the Qwen3-4B material GRPO batch-size-32 run. The repository name uses the project GRPO-TR naming convention, but the actual training method for this checkpoint is GRPO.
The repository root contains the best validation checkpoint, selected by validation mean@16. checkpoints/last/ contains the final checkpoint.
Performance
Dataset
Method
Base model
Train batch size
Best val mean@16
Best checkpoint
Final val mean@16
Final checkpoint
Material / SciKnowEval material
GRPO
Qwen3-4B
32
76.60%
60
76.26%
100
Validation Mean@16
step
val_mean16
percent
10
0.668882978723
66.89%
20
0.695478723404
69.55%
30
0.716755319149
71.68%
40
0.739361702128
73.94%
50
0.754654255319
75.47%
60
0.765957446809
76.60%
70
0.750000000000
75.00%
80
0.750664893617
75.07%
90
0.756648936170
75.66%
100
0.762632978723
76.26%
Detailed Training Hyperparameters
Section
Parameter
Value
Source
Run identity
Base model
Qwen/Qwen3-4B
queue/script override
Run identity
Dataset
Material / SciKnowEval material
run_qwen3_generalization.sh
Run identity
Method
GRPO
run_qwen3_generalization.sh
Run identity
Config
baseline_grpo
run_qwen3_generalization.sh
Run identity
Experiment
qwen3gen-material-GRPO-Qwen-Qwen3-4B-mbs8-train32-rollout8-lr1e-6-vllm0.8
run_qwen3_generalization.sh
Run identity
W&B run
run-20260702_125526-lnzvk3fv
wandb
Data
Train file
datasets/sciknoweval/material/train.parquet
script override
Data
Validation file
datasets/sciknoweval/material/test.parquet
script override
Data
Train batch size
32
queue/script override
Data
Train max samples
3200
queue/script override
Schedule
Total training steps
100
queue/script override
Schedule
Validation before train
False
queue/script override
Schedule
Save frequency
10
queue/script override
Schedule
Validation frequency
10
queue/script override
Sequence
Max prompt length
2048
queue/script override
Sequence
Max response length
8192
queue/script override
Sequence
Max model length
10240
queue/script override
Rollout
Train rollout n
8
queue/script override
Rollout
Validation rollout n
16
queue/script override
Rollout
vLLM GPU memory utilization
0.8
queue/script override
Optimization
Learning rate
1e-6
GRPO method override
Optimization
Weight decay
0.01
script override
PPO/GRPO
PPO mini batch size
8
queue/script override
PPO/GRPO
Normalize GRPO advantages by std
False
baseline_grpo.yaml / script override
Rollout correction
Importance sampling mode
token
script override
Rollout correction
IS threshold
2.0
script override
Checkpoint/Logging
Checkpoint root
checkpoints/datasets/sciknoweval/material/qwen3gen-material-GRPO-Qwen-Qwen3-4B-mbs8-train32-rollout8-lr1e-6-vllm0.8
script override
Checkpoint/Logging
Latest checkpointed iteration
100
latest_checkpointed_iteration.txt
Checkpoint/Logging
External actor archive
checkpoints/datasets/sciknoweval/material/qwen3gen-material-GRPO-Qwen-Qwen3-4B-mbs8-train32-rollout8-lr1e-6-vllm0.8/_actor_archive
preserve_actor_checkpoints.py
Checkpoint/Logging
Logger
console, wandb
ppo_trainer.yaml
PPO/GRPO
Policy loss mode
vanilla
method override
PPO/GRPO
Actor KL loss coef
0.0
method override
Raw result and artifact files:
results/validation_mean16.csv
results/training_scores.csv
results/hyperparameters.csv
results/training_score.png
results/training_score.svg
artifacts/config.yaml
artifacts/wandb-summary.json
artifacts/wandb-metadata.json
artifacts/output.log
artifacts/queue.log
Usage
Source
Checkpoint: checkpoints/datasets/sciknoweval/material/qwen3gen-material-GRPO-Qwen-Qwen3-4B-mbs8-train32-rollout8-lr1e-6-vllm0.8
Root actor checkpoint: checkpoints/datasets/sciknoweval/material/qwen3gen-material-GRPO-Qwen-Qwen3-4B-mbs8-train32-rollout8-lr1e-6-vllm0.8/_actor_archive/global_step_60/actor
Last actor checkpoint: checkpoints/datasets/sciknoweval/material/qwen3gen-material-GRPO-Qwen-Qwen3-4B-mbs8-train32-rollout8-lr1e-6-vllm0.8/global_step_100/actor
W&B run: run-20260702_125526-lnzvk3fv
Queue log: artifacts/queue.log