--- license: apache-2.0 base_model: laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink tags: - reinforcement-learning - rl - swe - agent - skyrl datasets: - SankalpKJ/swesmith-oracle-filtered --- # a3-rl-SankalpKJ_swesmith-oracle-filtered-40-8B RL-tuned (SkyRL) agentic SWE model. This is the **global_step 40** checkpoint, selected as the best checkpoint by EMA reward (EMA reward 0.0877). - **Base model:** [laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink](https://huggingface.co/laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink) (a Qwen3-8B SFT) - **Training dataset:** [SankalpKJ/swesmith-oracle-filtered](https://huggingface.co/datasets/SankalpKJ/swesmith-oracle-filtered) - **Architecture:** Qwen3-8B - **Training framework:** SkyRL (agentic RL, a3 series) The RL training configuration is included in this repo as `rl_config.yaml`. Parsed training metrics and the raw training logs are in `training_logs/`. ## Training Traces The full agentic rollout traces from this RL run are published as a dataset: - **Traces:** [penfever/a3-rl-SankalpKJ_swesmith-oracle-filtered](https://huggingface.co/datasets/penfever/a3-rl-SankalpKJ_swesmith-oracle-filtered) ## Note Published to the `penfever/` namespace pending a `laion/` write-role bump (may be re-homed to `laion/` later).