--- license: apache-2.0 library_name: transformers pipeline_tag: text-generation tags: - qwen3 - reinforcement-learning - rlsd --- # qwen3-8b-physics-rlsd-ema005 Fine-tuned from `Qwen/Qwen3-8B` with RLSD (EMA 0.05) on the `physics` split. ## Validation Performance Metric: `val-aux/sciknoweval/reward/mean@16` from 10-step validation logs. | best mean@16 | best step | final mean@16 | final step | |---:|---:|---:|---:| | 68.83% | 70 | 67.73% | 100 | ![Validation mean@16](eval_mean16.png) | step | mean@16 | |---:|---:| | 10 | 61.25% | | 20 | 61.02% | | 30 | 64.38% | | 40 | 63.98% | | 50 | 66.72% | | 60 | 67.89% | | 70 | 68.83% | | 80 | 67.11% | | 90 | 66.33% | | 100 | 67.73% | Files included with this repo: - `metrics.json`: parsed validation summary - `eval_mean16.csv`: step-level validation curve data - `eval_mean16.png`: validation curve plot Important: the uploaded weights are the final `global_step_100/actor` checkpoint. If `best step` is earlier than 100, the best validation point is reported for tracking, but the corresponding actor weights may not be retained locally. Checkpoint source: `/mnt/mole/SDPO/L2T/checkpoints/datasets/sciknoweval/physics/qwen3gen-physics-RLSD-Qwen-Qwen3-8B-mbs8-decay0-ema0.05-train64-rollout8-lr1e-6-vllm0.8` W&B run: `run-20260630_183038-9m3e0rqb` This upload uses `global_step_100/actor` converted from VERL FSDP shards to Hugging Face format.