Files
qwen3-8b-chemistry-rlsd-ema005/README.md
ModelHub XC a4c3f6168e 初始化项目,由ModelHub XC社区提供模型
Model: SeongryongJung/qwen3-8b-chemistry-rlsd-ema005
Source: Original Platform
2026-09-20 22:02:23 +08:00

1.4 KiB

license, library_name, pipeline_tag, tags
license library_name pipeline_tag tags
apache-2.0 transformers text-generation
qwen3
reinforcement-learning
rlsd

qwen3-8b-chemistry-rlsd-ema005

Fine-tuned from Qwen/Qwen3-8B with RLSD (EMA 0.05) on the chemistry split.

Validation Performance

Metric: val-aux/sciknoweval/reward/mean@16 from 10-step validation logs.

best mean@16 best step final mean@16 final step
74.38% 100 74.38% 100

Validation mean@16

step mean@16
10 45.95%
20 60.57%
30 67.65%
40 68.75%
50 70.18%
60 72.32%
70 71.61%
80 72.20%
90 73.18%
100 74.38%

Files included with this repo:

  • metrics.json: parsed validation summary
  • eval_mean16.csv: step-level validation curve data
  • eval_mean16.png: validation curve plot

Important: the uploaded weights are the final global_step_100/actor checkpoint. If best step is earlier than 100, the best validation point is reported for tracking, but the corresponding actor weights may not be retained locally.

Checkpoint source: /mnt/mole/SDPO/L2T/checkpoints/datasets/sciknoweval/chemistry/qwen3gen-chemistry-RLSD-Qwen-Qwen3-8B-mbs8-decay0-ema0.05-train64-rollout8-lr1e-6-vllm0.8

W&B run: run-20260630_122838-yvmohps3

This upload uses global_step_100/actor converted from VERL FSDP shards to Hugging Face format.