ModelHub XC 01e586b9fd 初始化项目,由ModelHub XC社区提供模型
Model: penfever/a3-rl-SankalpKJ_swesmith-oracle-filtered-40-8B
Source: Original Platform
2026-08-08 03:17:19 +08:00

license, base_model, tags, datasets
license base_model tags datasets
apache-2.0 laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink
reinforcement-learning
rl
swe
agent
skyrl
SankalpKJ/swesmith-oracle-filtered

a3-rl-SankalpKJ_swesmith-oracle-filtered-40-8B

RL-tuned (SkyRL) agentic SWE model. This is the global_step 40 checkpoint, selected as the best checkpoint by EMA reward (EMA reward 0.0877).

The RL training configuration is included in this repo as rl_config.yaml. Parsed training metrics and the raw training logs are in training_logs/.

Training Traces

The full agentic rollout traces from this RL run are published as a dataset:

Note

Published to the penfever/ namespace pending a laion/ write-role bump (may be re-homed to laion/ later).

Description
Model synced from source: penfever/a3-rl-SankalpKJ_swesmith-oracle-filtered-40-8B
Readme 15 MiB
Languages
Jinja 100%