初始化项目,由ModelHub XC社区提供模型

Model: SeongryongJung/Qwen3-4B-Chemical-RLSD-TR
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-14 23:14:45 +08:00
commit 110a0a7a0c
37 changed files with 311313 additions and 0 deletions

View File

@@ -0,0 +1,84 @@
section,parameter,value,source
Run identity,Base model,Qwen/Qwen3-4B,queue/script override
Run identity,Dataset,Chemical / SciKnowEval chemistry,run_qwen3_generalization.sh
Run identity,Method,RLSD_TR,run_qwen3_generalization.sh
Run identity,Config,rlsd,run_qwen3_generalization.sh
Run identity,Experiment,qwen3gen-chemistry-RLSD_TR-Qwen-Qwen3-4B-mbs8-decay0-tr0.1-train32-rollout8-lr1e-6-vllm0.8,run_qwen3_generalization.sh
Run identity,W&B run,run-20260702_062054-bz6p2yxy,wandb
Data,Train file,datasets/sciknoweval/chemistry/train.parquet,script override
Data,Validation file,datasets/sciknoweval/chemistry/test.parquet,script override
Data,Train batch size,32,queue/script override
Data,Train max samples,3200,queue/script override
Data,Prompt key,prompt,legacy_data.yaml default
Data,Reward key,data_source,legacy_data.yaml default
Data,Shuffle train data,True,user.yaml / legacy_data.yaml
Data,Validation shuffle,False,legacy_data.yaml default
Data,Filter overlong prompts,True,user.yaml
Data,Prompt truncation,error,legacy_data.yaml default
Data,enable_thinking,false,script override
Schedule,Total training steps,100,queue/script override
Schedule,Total epochs,30,ppo_trainer/user.yaml default
Schedule,Validation before train,False,queue/script override
Schedule,Save frequency,10,queue/script override
Schedule,Validation frequency,10,queue/script override
Sequence,Max prompt length,2048,queue/script override
Sequence,Max response length,8192,queue/script override
Sequence,Max model length,10240,queue/script override
Sequence,Actor max token length per GPU,10240,queue/script override
Rollout,Rollout engine,vllm,user.yaml
Rollout,Rollout dtype,bfloat16,rollout.yaml default
Rollout,Train rollout n,8,queue/script override
Rollout,Train rollout temperature,1.0,script override
Rollout,Train rollout top_p,1.0,script override
Rollout,Train rollout do_sample,True,rollout.yaml default
Rollout,Calculate rollout log probs,True,rlsd.yaml / script override
Rollout,Max num batched tokens,10240,queue/script override
Rollout,vLLM GPU memory utilization,0.8,queue/script override
Rollout,Tensor model parallel size,2,rollout.yaml default
Rollout,Free cache engine,True,rollout.yaml default
Validation,Validation rollout n,16,queue/script override
Validation,Validation temperature,0.6,queue/script override
Validation,Validation top_p,0.95,queue/script override
Validation,Validation do_sample,True,queue/script override
Optimization,Optimizer,AdamW,fsdp optimizer config
Optimization,Learning rate,1e-6,RLSD_TR method override
Optimization,LR scheduler,constant,W&B config
Optimization,LR warmup steps,10,script override
Optimization,Weight decay,0.01,script override
Optimization,Betas,"(0.9, 0.999)",W&B config
Optimization,Gradient clip,1.0,script override
PPO/GRPO,Advantage estimator,grpo,rlsd.yaml
PPO/GRPO,Normalize GRPO advantages by std,False,script override
PPO/GRPO,PPO epochs,1,W&B config
PPO/GRPO,PPO mini batch size,8,queue/script override
PPO/GRPO,PPO micro batch size per GPU,1,user.yaml
PPO/GRPO,Clip ratio low,0.2,script override
PPO/GRPO,Clip ratio high,0.28,script override
PPO/GRPO,Gamma,1.0,ppo_trainer.yaml default
PPO/GRPO,Lambda,1.0,ppo_trainer.yaml default
PPO/GRPO,Use KL in reward,False,ppo_trainer/user.yaml
PPO/GRPO,Actor KL loss observed,0.0,output.log
Rollout correction,Importance sampling mode,token,script override
Rollout correction,IS threshold,2.0,script override
RLSD_TR,Policy loss mode,rlsd,method override
RLSD_TR,Teacher regularization,trust-region,method override
RLSD_TR,Trust-region mix / teacher update rate,0.1,queue/script override
RLSD_TR,Token reweight lambda,0.5,queue/script override
RLSD_TR,Token reweight eps_w,0.2,queue/script override
RLSD_TR,Token reweight decay steps,0,queue/script override
RLSD_TR,Max reprompt length,10240,method override
RLSD_TR,Fused kernels,False,method override
FSDP/System,Actor strategy,fsdp,dp_actor.yaml
FSDP/System,FSDP dtype,bfloat16,W&B config
FSDP/System,FSDP model dtype,fp32,W&B config
FSDP/System,Use torch compile,True,W&B config
FSDP/System,GPUs per node,8,queue/script override
FSDP/System,Nodes,1,user.yaml
FSDP/System,GPU type,NVIDIA H200,wandb-metadata
Checkpoint/Logging,Checkpoint root,checkpoints/datasets/sciknoweval/chemistry,script override
Checkpoint/Logging,Latest checkpointed iteration,100,latest_checkpointed_iteration.txt
Checkpoint/Logging,Max actor checkpoints to keep,1,user.yaml
Checkpoint/Logging,Logger,"console, wandb",ppo_trainer.yaml
Checkpoint/Logging,W&B entity,seongryongjung-chung-ang-university,environment
Checkpoint/Logging,W&B project,SDPO-root,user.yaml project_name
Checkpoint/Logging,W&B group,QWEN3-RLSD-TR-GRPO-matched-generalization,method override
1 section parameter value source
2 Run identity Base model Qwen/Qwen3-4B queue/script override
3 Run identity Dataset Chemical / SciKnowEval chemistry run_qwen3_generalization.sh
4 Run identity Method RLSD_TR run_qwen3_generalization.sh
5 Run identity Config rlsd run_qwen3_generalization.sh
6 Run identity Experiment qwen3gen-chemistry-RLSD_TR-Qwen-Qwen3-4B-mbs8-decay0-tr0.1-train32-rollout8-lr1e-6-vllm0.8 run_qwen3_generalization.sh
7 Run identity W&B run run-20260702_062054-bz6p2yxy wandb
8 Data Train file datasets/sciknoweval/chemistry/train.parquet script override
9 Data Validation file datasets/sciknoweval/chemistry/test.parquet script override
10 Data Train batch size 32 queue/script override
11 Data Train max samples 3200 queue/script override
12 Data Prompt key prompt legacy_data.yaml default
13 Data Reward key data_source legacy_data.yaml default
14 Data Shuffle train data True user.yaml / legacy_data.yaml
15 Data Validation shuffle False legacy_data.yaml default
16 Data Filter overlong prompts True user.yaml
17 Data Prompt truncation error legacy_data.yaml default
18 Data enable_thinking false script override
19 Schedule Total training steps 100 queue/script override
20 Schedule Total epochs 30 ppo_trainer/user.yaml default
21 Schedule Validation before train False queue/script override
22 Schedule Save frequency 10 queue/script override
23 Schedule Validation frequency 10 queue/script override
24 Sequence Max prompt length 2048 queue/script override
25 Sequence Max response length 8192 queue/script override
26 Sequence Max model length 10240 queue/script override
27 Sequence Actor max token length per GPU 10240 queue/script override
28 Rollout Rollout engine vllm user.yaml
29 Rollout Rollout dtype bfloat16 rollout.yaml default
30 Rollout Train rollout n 8 queue/script override
31 Rollout Train rollout temperature 1.0 script override
32 Rollout Train rollout top_p 1.0 script override
33 Rollout Train rollout do_sample True rollout.yaml default
34 Rollout Calculate rollout log probs True rlsd.yaml / script override
35 Rollout Max num batched tokens 10240 queue/script override
36 Rollout vLLM GPU memory utilization 0.8 queue/script override
37 Rollout Tensor model parallel size 2 rollout.yaml default
38 Rollout Free cache engine True rollout.yaml default
39 Validation Validation rollout n 16 queue/script override
40 Validation Validation temperature 0.6 queue/script override
41 Validation Validation top_p 0.95 queue/script override
42 Validation Validation do_sample True queue/script override
43 Optimization Optimizer AdamW fsdp optimizer config
44 Optimization Learning rate 1e-6 RLSD_TR method override
45 Optimization LR scheduler constant W&B config
46 Optimization LR warmup steps 10 script override
47 Optimization Weight decay 0.01 script override
48 Optimization Betas (0.9, 0.999) W&B config
49 Optimization Gradient clip 1.0 script override
50 PPO/GRPO Advantage estimator grpo rlsd.yaml
51 PPO/GRPO Normalize GRPO advantages by std False script override
52 PPO/GRPO PPO epochs 1 W&B config
53 PPO/GRPO PPO mini batch size 8 queue/script override
54 PPO/GRPO PPO micro batch size per GPU 1 user.yaml
55 PPO/GRPO Clip ratio low 0.2 script override
56 PPO/GRPO Clip ratio high 0.28 script override
57 PPO/GRPO Gamma 1.0 ppo_trainer.yaml default
58 PPO/GRPO Lambda 1.0 ppo_trainer.yaml default
59 PPO/GRPO Use KL in reward False ppo_trainer/user.yaml
60 PPO/GRPO Actor KL loss observed 0.0 output.log
61 Rollout correction Importance sampling mode token script override
62 Rollout correction IS threshold 2.0 script override
63 RLSD_TR Policy loss mode rlsd method override
64 RLSD_TR Teacher regularization trust-region method override
65 RLSD_TR Trust-region mix / teacher update rate 0.1 queue/script override
66 RLSD_TR Token reweight lambda 0.5 queue/script override
67 RLSD_TR Token reweight eps_w 0.2 queue/script override
68 RLSD_TR Token reweight decay steps 0 queue/script override
69 RLSD_TR Max reprompt length 10240 method override
70 RLSD_TR Fused kernels False method override
71 FSDP/System Actor strategy fsdp dp_actor.yaml
72 FSDP/System FSDP dtype bfloat16 W&B config
73 FSDP/System FSDP model dtype fp32 W&B config
74 FSDP/System Use torch compile True W&B config
75 FSDP/System GPUs per node 8 queue/script override
76 FSDP/System Nodes 1 user.yaml
77 FSDP/System GPU type NVIDIA H200 wandb-metadata
78 Checkpoint/Logging Checkpoint root checkpoints/datasets/sciknoweval/chemistry script override
79 Checkpoint/Logging Latest checkpointed iteration 100 latest_checkpointed_iteration.txt
80 Checkpoint/Logging Max actor checkpoints to keep 1 user.yaml
81 Checkpoint/Logging Logger console, wandb ppo_trainer.yaml
82 Checkpoint/Logging W&B entity seongryongjung-chung-ang-university environment
83 Checkpoint/Logging W&B project SDPO-root user.yaml project_name
84 Checkpoint/Logging W&B group QWEN3-RLSD-TR-GRPO-matched-generalization method override

24
results/summary.json Normal file
View File

@@ -0,0 +1,24 @@
{
"repo_id": "SeongryongJung/Qwen3-4B-Chemical-RLSD-TR",
"output_dir": "/mnt/mole/SDPO/L2T/hf_upload_tr/Qwen3-4B-Chemical-RLSD-TR",
"best_step": 100,
"best_val_mean16": 0.6895833333333333,
"final_step": 100,
"final_val_mean16": 0.6895833333333333,
"train_rows": 90,
"val_rows": 10,
"hf_model_files": [
"added_tokens.json",
"chat_template.jinja",
"config.json",
"generation_config.json",
"merges.txt",
"model.safetensors.index.json",
"special_tokens_map.json",
"tokenizer.json",
"tokenizer_config.json",
"vocab.json",
"model-00001-of-00002.safetensors",
"model-00002-of-00002.safetensors"
]
}

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4fd721b8c29820fe41373b5c16ac993329102e09e022b7bd3a9f6f56b3ccee2a
size 123640

1630
results/training_score.svg Normal file

File diff suppressed because it is too large Load Diff

After

Width:  |  Height:  |  Size: 44 KiB

View File

@@ -0,0 +1,91 @@
step,critic_score_mean
1,0.5546875
2,0.3828125
3,0.453125
4,0.33984375
5,0.484375
6,0.44140625
7,0.40625
8,0.52734375
9,0.32421875
11,0.421875
12,0.4609375
13,0.47265625
14,0.453125
15,0.4921875
16,0.42578125
17,0.41015625
18,0.421875
19,0.515625
21,0.375
22,0.375
23,0.59375
24,0.5078125
25,0.48828125
26,0.62890625
27,0.5
28,0.5859375
29,0.34375
31,0.625
32,0.5625
33,0.65234375
34,0.515625
35,0.5078125
36,0.64453125
37,0.515625
38,0.59765625
39,0.59765625
41,0.59375
42,0.5859375
43,0.62890625
44,0.6171875
45,0.53515625
46,0.54296875
47,0.609375
48,0.5859375
49,0.65234375
51,0.5390625
52,0.59375
53,0.57421875
54,0.50390625
55,0.609375
56,0.64453125
57,0.578125
58,0.64453125
59,0.58203125
61,0.65234375
62,0.64453125
63,0.48828125
64,0.70703125
65,0.74609375
66,0.546875
67,0.66796875
68,0.4765625
69,0.56640625
71,0.5859375
72,0.71484375
73,0.546875
74,0.62109375
75,0.73046875
76,0.5625
77,0.7265625
78,0.65234375
79,0.59765625
81,0.640625
82,0.625
83,0.5625
84,0.68359375
85,0.69921875
86,0.5625
87,0.53125
88,0.57421875
89,0.625
91,0.60546875
92,0.625
93,0.87890625
94,0.73046875
95,0.6640625
96,0.6015625
97,0.52734375
98,0.58203125
99,0.640625
1 step critic_score_mean
2 1 0.5546875
3 2 0.3828125
4 3 0.453125
5 4 0.33984375
6 5 0.484375
7 6 0.44140625
8 7 0.40625
9 8 0.52734375
10 9 0.32421875
11 11 0.421875
12 12 0.4609375
13 13 0.47265625
14 14 0.453125
15 15 0.4921875
16 16 0.42578125
17 17 0.41015625
18 18 0.421875
19 19 0.515625
20 21 0.375
21 22 0.375
22 23 0.59375
23 24 0.5078125
24 25 0.48828125
25 26 0.62890625
26 27 0.5
27 28 0.5859375
28 29 0.34375
29 31 0.625
30 32 0.5625
31 33 0.65234375
32 34 0.515625
33 35 0.5078125
34 36 0.64453125
35 37 0.515625
36 38 0.59765625
37 39 0.59765625
38 41 0.59375
39 42 0.5859375
40 43 0.62890625
41 44 0.6171875
42 45 0.53515625
43 46 0.54296875
44 47 0.609375
45 48 0.5859375
46 49 0.65234375
47 51 0.5390625
48 52 0.59375
49 53 0.57421875
50 54 0.50390625
51 55 0.609375
52 56 0.64453125
53 57 0.578125
54 58 0.64453125
55 59 0.58203125
56 61 0.65234375
57 62 0.64453125
58 63 0.48828125
59 64 0.70703125
60 65 0.74609375
61 66 0.546875
62 67 0.66796875
63 68 0.4765625
64 69 0.56640625
65 71 0.5859375
66 72 0.71484375
67 73 0.546875
68 74 0.62109375
69 75 0.73046875
70 76 0.5625
71 77 0.7265625
72 78 0.65234375
73 79 0.59765625
74 81 0.640625
75 82 0.625
76 83 0.5625
77 84 0.68359375
78 85 0.69921875
79 86 0.5625
80 87 0.53125
81 88 0.57421875
82 89 0.625
83 91 0.60546875
84 92 0.625
85 93 0.87890625
86 94 0.73046875
87 95 0.6640625
88 96 0.6015625
89 97 0.52734375
90 98 0.58203125
91 99 0.640625

View File

@@ -0,0 +1,11 @@
step,val_mean16
10,0.4398809523809524
20,0.5101190476190476
30,0.5943452380952381
40,0.6357142857142857
50,0.6586309523809524
60,0.6735119047619048
70,0.6851190476190476
80,0.6788690476190476
90,0.6821428571428572
100,0.6895833333333333
1 step val_mean16
2 10 0.4398809523809524
3 20 0.5101190476190476
4 30 0.5943452380952381
5 40 0.6357142857142857
6 50 0.6586309523809524
7 60 0.6735119047619048
8 70 0.6851190476190476
9 80 0.6788690476190476
10 90 0.6821428571428572
11 100 0.6895833333333333