初始化项目,由ModelHub XC社区提供模型

Model: SeongryongJung/Qwen3-8B-Material-GRPO-TR
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-07 11:03:18 +08:00
commit aebb065570
42 changed files with 326704 additions and 0 deletions

View File

@@ -0,0 +1,33 @@
section,parameter,value,source
Run identity,Base model,Qwen/Qwen3-8B,queue/script override
Run identity,Dataset,Material / SciKnowEval material,run_qwen3_generalization.sh
Run identity,Method,GRPO,run_qwen3_generalization.sh
Run identity,Config,baseline_grpo,run_qwen3_generalization.sh
Run identity,Experiment,qwen3gen-material-GRPO-Qwen-Qwen3-8B-mbs8-train32-rollout8-lr1e-6-vllm0.8,run_qwen3_generalization.sh
Run identity,W&B run,run-20260703_105138-1ud3umaz,wandb
Data,Train file,datasets/sciknoweval/material/train.parquet,script override
Data,Validation file,datasets/sciknoweval/material/test.parquet,script override
Data,Train batch size,32,queue/script override
Data,Train max samples,3200,queue/script override
Schedule,Total training steps,100,queue/script override
Schedule,Validation before train,False,queue/script override
Schedule,Save frequency,10,queue/script override
Schedule,Validation frequency,10,queue/script override
Sequence,Max prompt length,2048,queue/script override
Sequence,Max response length,8192,queue/script override
Sequence,Max model length,10240,queue/script override
Rollout,Train rollout n,8,queue/script override
Rollout,Validation rollout n,16,queue/script override
Rollout,vLLM GPU memory utilization,0.8,queue/script override
Optimization,Learning rate,1e-6,GRPO method override
Optimization,Weight decay,0.01,script override
PPO/GRPO,PPO mini batch size,8,queue/script override
PPO/GRPO,Normalize GRPO advantages by std,False,baseline_grpo.yaml / script override
Rollout correction,Importance sampling mode,token,script override
Rollout correction,IS threshold,2.0,script override
Checkpoint/Logging,Checkpoint root,checkpoints/datasets/sciknoweval/material/qwen3gen-material-GRPO-Qwen-Qwen3-8B-mbs8-train32-rollout8-lr1e-6-vllm0.8,script override
Checkpoint/Logging,Latest checkpointed iteration,100,latest_checkpointed_iteration.txt
Checkpoint/Logging,External actor archive,checkpoints/datasets/sciknoweval/material/qwen3gen-material-GRPO-Qwen-Qwen3-8B-mbs8-train32-rollout8-lr1e-6-vllm0.8/_actor_archive,preserve_actor_checkpoints.py
Checkpoint/Logging,Logger,"console, wandb",ppo_trainer.yaml
PPO/GRPO,Policy loss mode,vanilla,method override
PPO/GRPO,Actor KL loss coef,0.0,method override
1 section parameter value source
2 Run identity Base model Qwen/Qwen3-8B queue/script override
3 Run identity Dataset Material / SciKnowEval material run_qwen3_generalization.sh
4 Run identity Method GRPO run_qwen3_generalization.sh
5 Run identity Config baseline_grpo run_qwen3_generalization.sh
6 Run identity Experiment qwen3gen-material-GRPO-Qwen-Qwen3-8B-mbs8-train32-rollout8-lr1e-6-vllm0.8 run_qwen3_generalization.sh
7 Run identity W&B run run-20260703_105138-1ud3umaz wandb
8 Data Train file datasets/sciknoweval/material/train.parquet script override
9 Data Validation file datasets/sciknoweval/material/test.parquet script override
10 Data Train batch size 32 queue/script override
11 Data Train max samples 3200 queue/script override
12 Schedule Total training steps 100 queue/script override
13 Schedule Validation before train False queue/script override
14 Schedule Save frequency 10 queue/script override
15 Schedule Validation frequency 10 queue/script override
16 Sequence Max prompt length 2048 queue/script override
17 Sequence Max response length 8192 queue/script override
18 Sequence Max model length 10240 queue/script override
19 Rollout Train rollout n 8 queue/script override
20 Rollout Validation rollout n 16 queue/script override
21 Rollout vLLM GPU memory utilization 0.8 queue/script override
22 Optimization Learning rate 1e-6 GRPO method override
23 Optimization Weight decay 0.01 script override
24 PPO/GRPO PPO mini batch size 8 queue/script override
25 PPO/GRPO Normalize GRPO advantages by std False baseline_grpo.yaml / script override
26 Rollout correction Importance sampling mode token script override
27 Rollout correction IS threshold 2.0 script override
28 Checkpoint/Logging Checkpoint root checkpoints/datasets/sciknoweval/material/qwen3gen-material-GRPO-Qwen-Qwen3-8B-mbs8-train32-rollout8-lr1e-6-vllm0.8 script override
29 Checkpoint/Logging Latest checkpointed iteration 100 latest_checkpointed_iteration.txt
30 Checkpoint/Logging External actor archive checkpoints/datasets/sciknoweval/material/qwen3gen-material-GRPO-Qwen-Qwen3-8B-mbs8-train32-rollout8-lr1e-6-vllm0.8/_actor_archive preserve_actor_checkpoints.py
31 Checkpoint/Logging Logger console, wandb ppo_trainer.yaml
32 PPO/GRPO Policy loss mode vanilla method override
33 PPO/GRPO Actor KL loss coef 0.0 method override

29
results/summary.json Normal file
View File

@@ -0,0 +1,29 @@
{
"repo_id": "SeongryongJung/Qwen3-8B-Material-GRPO-TR",
"output_dir": "/mnt/mole/SDPO/L2T/hf_upload_tr/Qwen3-8B-Material-GRPO-TR",
"experiment": "qwen3gen-material-GRPO-Qwen-Qwen3-8B-mbs8-train32-rollout8-lr1e-6-vllm0.8",
"best_step": 100,
"best_val_mean16": 0.7799202127659575,
"best_actor_dir": "/mnt/mole/SDPO/L2T/checkpoints/datasets/sciknoweval/material/qwen3gen-material-GRPO-Qwen-Qwen3-8B-mbs8-train32-rollout8-lr1e-6-vllm0.8/global_step_100/actor",
"final_step": 100,
"final_val_mean16": 0.7799202127659575,
"final_actor_dir": "/mnt/mole/SDPO/L2T/checkpoints/datasets/sciknoweval/material/qwen3gen-material-GRPO-Qwen-Qwen3-8B-mbs8-train32-rollout8-lr1e-6-vllm0.8/global_step_100/actor",
"train_rows": 90,
"val_rows": 10,
"hf_model_files": [
"added_tokens.json",
"chat_template.jinja",
"config.json",
"generation_config.json",
"merges.txt",
"model.safetensors.index.json",
"special_tokens_map.json",
"tokenizer.json",
"tokenizer_config.json",
"vocab.json",
"model-00001-of-00004.safetensors",
"model-00002-of-00004.safetensors",
"model-00003-of-00004.safetensors",
"model-00004-of-00004.safetensors"
]
}

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6429e63128c876d36ca1e4b3bb49dafdbf29c471f855a361579859981d4db61a
size 127678

1591
results/training_score.svg Normal file

File diff suppressed because it is too large Load Diff

After

Width:  |  Height:  |  Size: 43 KiB

View File

@@ -0,0 +1,91 @@
step,critic_score_mean
1,0.546875
2,0.421875
3,0.48828125
4,0.5
5,0.54296875
6,0.48828125
7,0.55078125
8,0.54296875
9,0.546875
11,0.50390625
12,0.60546875
13,0.546875
14,0.77734375
15,0.62109375
16,0.640625
17,0.56640625
18,0.6171875
19,0.62109375
21,0.56640625
22,0.671875
23,0.70703125
24,0.828125
25,0.63671875
26,0.734375
27,0.80078125
28,0.5859375
29,0.65625
31,0.625
32,0.61328125
33,0.76171875
34,0.79296875
35,0.76171875
36,0.59765625
37,0.796875
38,0.59765625
39,0.7265625
41,0.56640625
42,0.625
43,0.53515625
44,0.703125
45,0.7890625
46,0.71484375
47,0.65234375
48,0.85546875
49,0.6484375
51,0.67578125
52,0.765625
53,0.640625
54,0.54296875
55,0.75
56,0.78515625
57,0.7265625
58,0.796875
59,0.71484375
61,0.8046875
62,0.8203125
63,0.71484375
64,0.80859375
65,0.66015625
66,0.69140625
67,0.78125
68,0.7734375
69,0.82421875
71,0.7734375
72,0.796875
73,0.71484375
74,0.8359375
75,0.6796875
76,0.734375
77,0.53125
78,0.68359375
79,0.7734375
81,0.703125
82,0.72265625
83,0.73828125
84,0.83203125
85,0.70703125
86,0.83203125
87,0.59375
88,0.59765625
89,0.71875
91,0.84765625
92,0.72265625
93,0.7421875
94,0.68359375
95,0.6953125
96,0.55078125
97,0.83984375
98,0.9296875
99,0.76171875
1 step critic_score_mean
2 1 0.546875
3 2 0.421875
4 3 0.48828125
5 4 0.5
6 5 0.54296875
7 6 0.48828125
8 7 0.55078125
9 8 0.54296875
10 9 0.546875
11 11 0.50390625
12 12 0.60546875
13 13 0.546875
14 14 0.77734375
15 15 0.62109375
16 16 0.640625
17 17 0.56640625
18 18 0.6171875
19 19 0.62109375
20 21 0.56640625
21 22 0.671875
22 23 0.70703125
23 24 0.828125
24 25 0.63671875
25 26 0.734375
26 27 0.80078125
27 28 0.5859375
28 29 0.65625
29 31 0.625
30 32 0.61328125
31 33 0.76171875
32 34 0.79296875
33 35 0.76171875
34 36 0.59765625
35 37 0.796875
36 38 0.59765625
37 39 0.7265625
38 41 0.56640625
39 42 0.625
40 43 0.53515625
41 44 0.703125
42 45 0.7890625
43 46 0.71484375
44 47 0.65234375
45 48 0.85546875
46 49 0.6484375
47 51 0.67578125
48 52 0.765625
49 53 0.640625
50 54 0.54296875
51 55 0.75
52 56 0.78515625
53 57 0.7265625
54 58 0.796875
55 59 0.71484375
56 61 0.8046875
57 62 0.8203125
58 63 0.71484375
59 64 0.80859375
60 65 0.66015625
61 66 0.69140625
62 67 0.78125
63 68 0.7734375
64 69 0.82421875
65 71 0.7734375
66 72 0.796875
67 73 0.71484375
68 74 0.8359375
69 75 0.6796875
70 76 0.734375
71 77 0.53125
72 78 0.68359375
73 79 0.7734375
74 81 0.703125
75 82 0.72265625
76 83 0.73828125
77 84 0.83203125
78 85 0.70703125
79 86 0.83203125
80 87 0.59375
81 88 0.59765625
82 89 0.71875
83 91 0.84765625
84 92 0.72265625
85 93 0.7421875
86 94 0.68359375
87 95 0.6953125
88 96 0.55078125
89 97 0.83984375
90 98 0.9296875
91 99 0.76171875

View File

@@ -0,0 +1,11 @@
step,val_mean16
10,0.6795212765957447
20,0.706781914893617
30,0.7027925531914894
40,0.7240691489361702
50,0.7406914893617021
60,0.7533244680851063
70,0.7593085106382979
80,0.7506648936170213
90,0.7593085106382979
100,0.7799202127659575
1 step val_mean16
2 10 0.6795212765957447
3 20 0.706781914893617
4 30 0.7027925531914894
5 40 0.7240691489361702
6 50 0.7406914893617021
7 60 0.7533244680851063
8 70 0.7593085106382979
9 80 0.7506648936170213
10 90 0.7593085106382979
11 100 0.7799202127659575