Files
ModelHub XC a2beab1ad3 初始化项目,由ModelHub XC社区提供模型
Model: laion/a3-rl-DCAgent_exp_rpt_curriculum-medium-10-8B
Source: Original Platform
2026-07-30 10:45:18 +08:00

1.1 KiB
Raw Permalink Blame History

a3-rl-DCAgent_exp_rpt_curriculum-medium-10-8B

RL (SkyRL/terminus-2) fine-tune of laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink on DCAgent/exp_rpt_curriculum-medium (512-task set, 2 epochs).

Checkpoint = global_step_10, selected by EMA-best of reward/avg_raw_reward (5-step EMA, α=1/3) across the full chain (job 591150). EMA at step 10 = 0.3106 (reward 0.2812). Data-limited run; training metrics emitted through step 16, checkpoints exported through step 21.

Training Traces

Training-time Daytona/Harbor rollouts for this run are uploaded as a companion dataset: penfever/a3-rl-DCAgent_exp_rpt_curriculum-medium

The dataset contains the last episode of each trial (per make_and_upload_trace_dataset --episodes last) — the same rollouts the policy was trained on after rollback / truncation.

Training logs

See training_logs/ for per-step metrics CSVs, the metrics report, and the reward-vs-steps plot (Jupiter has no W&B).