ModelHub XC 5a9933097b 初始化项目,由ModelHub XC社区提供模型
Model: SeongryongJung/Qwen3-4B-Material-GRPO-TR
Source: Original Platform
2026-08-04 08:21:18 +08:00

license, library_name, pipeline_tag, tags, base_model
license library_name pipeline_tag tags base_model
apache-2.0 transformers text-generation
qwen3
reinforcement-learning
grpo
text-generation
Qwen/Qwen3-4B

Qwen3-4B-Material-GRPO-TR

This repository contains the Qwen3-4B material GRPO batch-size-32 run. The repository name uses the project GRPO-TR naming convention, but the actual training method for this checkpoint is GRPO.

The repository root contains the best validation checkpoint, selected by validation mean@16. checkpoints/last/ contains the final checkpoint.

Performance

Dataset Method Base model Train batch size Best val mean@16 Best checkpoint Final val mean@16 Final checkpoint
Material / SciKnowEval material GRPO Qwen3-4B 32 76.60% 60 76.26% 100

Training and validation scores

Validation Mean@16

step val_mean16 percent
10 0.668882978723 66.89%
20 0.695478723404 69.55%
30 0.716755319149 71.68%
40 0.739361702128 73.94%
50 0.754654255319 75.47%
60 0.765957446809 76.60%
70 0.750000000000 75.00%
80 0.750664893617 75.07%
90 0.756648936170 75.66%
100 0.762632978723 76.26%

Detailed Training Hyperparameters

Section Parameter Value Source
Run identity Base model Qwen/Qwen3-4B queue/script override
Run identity Dataset Material / SciKnowEval material run_qwen3_generalization.sh
Run identity Method GRPO run_qwen3_generalization.sh
Run identity Config baseline_grpo run_qwen3_generalization.sh
Run identity Experiment qwen3gen-material-GRPO-Qwen-Qwen3-4B-mbs8-train32-rollout8-lr1e-6-vllm0.8 run_qwen3_generalization.sh
Run identity W&B run run-20260702_125526-lnzvk3fv wandb
Data Train file datasets/sciknoweval/material/train.parquet script override
Data Validation file datasets/sciknoweval/material/test.parquet script override
Data Train batch size 32 queue/script override
Data Train max samples 3200 queue/script override
Schedule Total training steps 100 queue/script override
Schedule Validation before train False queue/script override
Schedule Save frequency 10 queue/script override
Schedule Validation frequency 10 queue/script override
Sequence Max prompt length 2048 queue/script override
Sequence Max response length 8192 queue/script override
Sequence Max model length 10240 queue/script override
Rollout Train rollout n 8 queue/script override
Rollout Validation rollout n 16 queue/script override
Rollout vLLM GPU memory utilization 0.8 queue/script override
Optimization Learning rate 1e-6 GRPO method override
Optimization Weight decay 0.01 script override
PPO/GRPO PPO mini batch size 8 queue/script override
PPO/GRPO Normalize GRPO advantages by std False baseline_grpo.yaml / script override
Rollout correction Importance sampling mode token script override
Rollout correction IS threshold 2.0 script override
Checkpoint/Logging Checkpoint root checkpoints/datasets/sciknoweval/material/qwen3gen-material-GRPO-Qwen-Qwen3-4B-mbs8-train32-rollout8-lr1e-6-vllm0.8 script override
Checkpoint/Logging Latest checkpointed iteration 100 latest_checkpointed_iteration.txt
Checkpoint/Logging External actor archive checkpoints/datasets/sciknoweval/material/qwen3gen-material-GRPO-Qwen-Qwen3-4B-mbs8-train32-rollout8-lr1e-6-vllm0.8/_actor_archive preserve_actor_checkpoints.py
Checkpoint/Logging Logger console, wandb ppo_trainer.yaml
PPO/GRPO Policy loss mode vanilla method override
PPO/GRPO Actor KL loss coef 0.0 method override

Raw result and artifact files:

  • results/validation_mean16.csv
  • results/training_scores.csv
  • results/hyperparameters.csv
  • results/training_score.png
  • results/training_score.svg
  • artifacts/config.yaml
  • artifacts/wandb-summary.json
  • artifacts/wandb-metadata.json
  • artifacts/output.log
  • artifacts/queue.log

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

repo_id = "SeongryongJung/Qwen3-4B-Material-GRPO-TR"
tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    repo_id,
    torch_dtype="auto",
    device_map="auto",
    trust_remote_code=True,
)

Source

  • Checkpoint: checkpoints/datasets/sciknoweval/material/qwen3gen-material-GRPO-Qwen-Qwen3-4B-mbs8-train32-rollout8-lr1e-6-vllm0.8
  • Root actor checkpoint: checkpoints/datasets/sciknoweval/material/qwen3gen-material-GRPO-Qwen-Qwen3-4B-mbs8-train32-rollout8-lr1e-6-vllm0.8/_actor_archive/global_step_60/actor
  • Last actor checkpoint: checkpoints/datasets/sciknoweval/material/qwen3gen-material-GRPO-Qwen-Qwen3-4B-mbs8-train32-rollout8-lr1e-6-vllm0.8/global_step_100/actor
  • W&B run: run-20260702_125526-lnzvk3fv
  • Queue log: artifacts/queue.log
Description
Model synced from source: SeongryongJung/Qwen3-4B-Material-GRPO-TR
Readme 14 MiB
Languages
Jinja 100%