初始化项目,由ModelHub XC社区提供模型

Model: penfever/a3-rl-SankalpKJ_swesmith-oracle-filtered-40-8B
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-08 03:17:19 +08:00
commit 01e586b9fd
40 changed files with 238471 additions and 0 deletions

36
README.md Normal file
View File

@@ -0,0 +1,36 @@
---
license: apache-2.0
base_model: laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink
tags:
- reinforcement-learning
- rl
- swe
- agent
- skyrl
datasets:
- SankalpKJ/swesmith-oracle-filtered
---
# a3-rl-SankalpKJ_swesmith-oracle-filtered-40-8B
RL-tuned (SkyRL) agentic SWE model. This is the **global_step 40** checkpoint,
selected as the best checkpoint by EMA reward (EMA reward 0.0877).
- **Base model:** [laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink](https://huggingface.co/laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink) (a Qwen3-8B SFT)
- **Training dataset:** [SankalpKJ/swesmith-oracle-filtered](https://huggingface.co/datasets/SankalpKJ/swesmith-oracle-filtered)
- **Architecture:** Qwen3-8B
- **Training framework:** SkyRL (agentic RL, a3 series)
The RL training configuration is included in this repo as `rl_config.yaml`.
Parsed training metrics and the raw training logs are in `training_logs/`.
## Training Traces
The full agentic rollout traces from this RL run are published as a dataset:
- **Traces:** [penfever/a3-rl-SankalpKJ_swesmith-oracle-filtered](https://huggingface.co/datasets/penfever/a3-rl-SankalpKJ_swesmith-oracle-filtered)
## Note
Published to the `penfever/` namespace pending a `laion/` write-role bump
(may be re-homed to `laion/` later).