Files

37 lines
1.4 KiB
Markdown
Raw Permalink Normal View History

---
license: apache-2.0
base_model: laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink
tags:
- reinforcement-learning
- rl
- swe
- agent
- skyrl
datasets:
- SankalpKJ/swesmith-oracle-filtered
---
# a3-rl-SankalpKJ_swesmith-oracle-filtered-40-8B
RL-tuned (SkyRL) agentic SWE model. This is the **global_step 40** checkpoint,
selected as the best checkpoint by EMA reward (EMA reward 0.0877).
- **Base model:** [laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink](https://huggingface.co/laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink) (a Qwen3-8B SFT)
- **Training dataset:** [SankalpKJ/swesmith-oracle-filtered](https://huggingface.co/datasets/SankalpKJ/swesmith-oracle-filtered)
- **Architecture:** Qwen3-8B
- **Training framework:** SkyRL (agentic RL, a3 series)
The RL training configuration is included in this repo as `rl_config.yaml`.
Parsed training metrics and the raw training logs are in `training_logs/`.
## Training Traces
The full agentic rollout traces from this RL run are published as a dataset:
- **Traces:** [penfever/a3-rl-SankalpKJ_swesmith-oracle-filtered](https://huggingface.co/datasets/penfever/a3-rl-SankalpKJ_swesmith-oracle-filtered)
## Note
Published to the `penfever/` namespace pending a `laion/` write-role bump
(may be re-homed to `laion/` later).