Files
ablation-pymethods2test-seq…/training_logs/20260612_220124_metrics_report.md
ModelHub XC 0c293fc953 初始化项目,由ModelHub XC社区提供模型
Model: laion/ablation-pymethods2test-seqmean-arm0-30-8B
Source: Original Platform
2026-07-18 18:02:02 +08:00

21 KiB

SkyRL Training Metrics Analysis

Generated from 4 log files

Overview

Log File Total Steps Metric Blocks Final Reward (mean) Final Reward (max) Total Time (s)
job_630124 28 28 0.4467 0.7695 41354.5
job_630127 56 28 0.4329 0.5703 40185.7
job_630130 70 14 0.3858 0.5449 36451.3
job_661735 78 8 0.3747 0.5566 13910.7

Async Metrics

Mean Std Min Max Count
async/discard_rate 0 0 0 0 78
async/discarded_count 0 0 0 0 78
async/effective_batch_groups 64 0 64 64 78
async/effective_batch_samples 512 0 512 512 78
async/staleness_max 3.52564 1.00291 0 7 78
async/staleness_mean 1.76964 0.525328 0 4 78
async/staleness_min 0.153846 0.625786 0 4 78
async/staleness_ratio 0.872991 0.178479 0 1 78

Generate Metrics

Mean Std Min Max Count
generate/avg_num_tokens 4953.5 781.064 1687.9 6119.26 78
generate/avg_tokens_non_zero_rewards 4928.65 524.786 3823.39 6195.01 78
generate/avg_tokens_zero_rewards 5030.87 1001.93 1110.32 6614.23 78
generate/max_num_tokens 31246.8 2087.2 19335 31885 78
generate/min_num_tokens 17.1795 142.893 1 1263 78
generate/std_num_tokens 4459.29 622.221 2744.51 5829.72 78

Loss Metrics

Mean Std Min Max Count
loss/avg_final_rewards 0.423427 0.0901758 0.2051 0.7695 78
loss/avg_raw_advantages 0.0121513 0.0324204 -0.0162 0.1959 78
loss/avg_raw_advantages_abs 0.106845 0.0320619 0.0431 0.2355 78

Policy Metrics

Mean Std Min Max Count
policy/final_loss -9.23077e-05 0.000277638 -0.0015 0 78
policy/log_ratio_abs_max 0 0 0 0 78
policy/log_ratio_abs_mean 0 0 0 0 78
policy/log_ratio_abs_p99 0 0 0 0 78
policy/log_ratio_abs_pos00 0 0 0 0 78
policy/log_ratio_abs_pos10 0 0 0 0 78
policy/log_ratio_abs_pos20 0 0 0 0 78
policy/log_ratio_abs_pos30 0 0 0 0 78
policy/log_ratio_abs_pos40 0 0 0 0 78
policy/log_ratio_abs_pos50 0 0 0 0 78
policy/log_ratio_abs_pos60 0 0 0 0 78
policy/log_ratio_abs_pos70 0 0 0 0 78
policy/log_ratio_abs_pos80 0 0 0 0 78
policy/log_ratio_abs_pos90 0 0 0 0 78
policy/n_tokens_dp_gt_10pct 0 0 0 0 78
policy/n_tokens_dp_gt_1pct 0 0 0 0 78
policy/n_tokens_dp_gt_50pct 0 0 0 0 78
policy/policy_entropy 0.119653 0.0179549 0.0511 0.1411 78
policy/policy_loss -0.00595641 0.017477 -0.0959 0 78
policy/policy_lr 0 0 0 0 78
policy/policy_update_steps 1 0 1 1 78
policy/ppo_clip_ratio 0 0 0 0 78
policy/raw_grad_norm 0.0149538 0.00982649 0.0068 0.0967 78

Reward Metrics

Mean Std Min Max Count
reward/avg_pass_at_8 0.543153 0.079973 0.375 0.8125 78
reward/avg_raw_reward 0.423427 0.0901758 0.2051 0.7695 78

System Metrics

Mean Std Min Max Count
system/process_rss_gb 14.2314 1.13198 9.9914 16.2064 78
system/process_vms_gb 51.7172 0.827814 50.291 53.4671 78
system/ram_available_gb 570.794 11.4823 551.846 594.612 78
system/ram_percent 33.4705 1.34049 30.7 35.7 78
system/ram_total_gb 857.968 3.05352e-05 857.967 857.968 78
system/ram_used_gb 287.174 11.4823 263.356 306.122 78

Timing Metrics

Mean Std Min Max Count
timing/compute_advantages_and_returns 0.102533 0.107328 0.0486 0.8661 78
timing/convert_to_training_input 4.70007 0.352461 2.7178 5.493 78
timing/fwd_logprobs_values_reward 35.7722 9.32356 15.1498 104.946 78
timing/policy_train 207.902 31.034 108.532 269.709 78
timing/run_training 244.034 34.7885 123.918 312.363 78
timing/step 1691.05 1167.37 558.623 10105.4 78
timing/sync_weights 22.7971 1.88627 17.1266 26.2516 78
timing/train_critic_and_policy 208.159 31.0529 108.708 270.02 78
timing/wait_for_generation_buffer 1419.52 1176.94 247.59 9825.37 78
timing/cleanup_old_checkpoints 4.0377 18.4601 0.0067 113.745 39
timing/save_checkpoints 14.6624 18.3571 7.9437 123.126 39
timing/save_hf_model 24.8864 69.3195 5.8995 275.429 15

Trainer Metrics

Mean Std Min Max Count
trainer/epoch 0 0 0 0 78
trainer/global_step 39.5 22.6605 1 78 78

Batch_Errors Metrics

Mean Std Min Max Count
batch_errors/total_batches 61.5256 7.45676 44 119 78
batch_errors/total_instances 492.205 59.6541 352 952 78
batch_errors/total_successful 432.179 76.3362 191 918 78
batch_errors/total_failed 20.7308 9.45897 8 58 78
batch_errors/total_masked 49.3333 27.27 6 160 78
batch_errors/avg_RuntimeError 0.10335 0.0403443 0.0420168 0.15873 6
batch_errors/total_RuntimeError 6.83333 1.94079 5 10 6
batch_errors/avg_VerifierTimeoutError 0.200918 0.399779 0.015625 1.94915 61
batch_errors/total_VerifierTimeoutError 12.1639 24.1283 1 115 61
batch_errors/avg_AgentTimeoutError 0.184673 0.486823 0.00840336 2.55932 62
batch_errors/total_AgentTimeoutError 11.1613 29.1948 1 151 62
batch_errors/avg_InvalidChatHistory 0.549849 0.239037 0.10084 1.41379 78
batch_errors/total_InvalidChatHistory 33.5256 14.2343 7 82 78
batch_errors/avg_ContextLengthExceededError 0.161599 0.087392 0.015873 0.460317 77
batch_errors/total_ContextLengthExceededError 9.84416 5.14287 1 29 77
batch_errors/avg_DaytonaError 0.0555526 0.0370314 0.016129 0.114754 10
batch_errors/total_DaytonaError 3.4 2.27058 1 7 10
batch_errors/avg_AgentSetupTimeoutError 0.0930756 0.121065 0.0153846 0.372881 18
batch_errors/total_AgentSetupTimeoutError 5.5 7.15583 1 22 18
batch_errors/avg_DaytonaAuthenticationError 0.0222833 0.0083652 0.015873 0.031746 3
batch_errors/total_DaytonaAuthenticationError 1.33333 0.57735 1 2 3

Training Progression by Log

job_630124

Step Reward Pass@8 KL Loss Step Time (s) Gen Wait (s)
1 0.7695 0.8125 0.000000 0.0000 3372.0 3167.0
2 0.4199 0.5781 0.000000 -0.0001 1718.0 1388.3
3 0.3457 0.4688 0.000000 -0.0000 1220.7 885.5
4 0.4141 0.4844 0.000000 -0.0001 1357.9 1059.7
5 0.5156 0.6250 0.000000 -0.0001 1224.4 955.5
6 0.4023 0.4844 0.000000 -0.0000 1281.0 982.2
7 0.4492 0.5469 0.000000 -0.0000 1559.3 1299.2
8 0.4766 0.6250 0.000000 0.0000 1610.3 1270.4
9 0.4453 0.5469 0.000000 0.0000 1321.3 1032.5
10 0.3730 0.5312 0.000000 0.0000 1294.9 1011.8
11 0.4922 0.6562 0.000000 -0.0001 1475.0 1214.5
12 0.3477 0.4219 0.000000 0.0000 1559.4 1264.5
13 0.5020 0.6250 0.000000 -0.0000 1431.9 1188.5
14 0.4355 0.5469 0.000000 -0.0000 1142.8 885.9
15 0.4902 0.6094 0.000000 -0.0000 1314.0 1035.8
16 0.4512 0.5781 0.000000 -0.0000 1308.8 1024.4
17 0.3887 0.5156 0.000000 -0.0000 1220.9 899.1
18 0.4551 0.5625 0.000000 -0.0000 1582.7 1299.8
19 0.3691 0.4531 0.000000 -0.0000 1533.3 1254.4
20 0.4023 0.5156 0.000000 0.0000 1501.3 1228.6
21 0.3633 0.4688 0.000000 -0.0001 1384.7 1125.9
22 0.4473 0.5781 0.000000 0.0000 1341.1 1063.0
23 0.4160 0.5781 0.000000 -0.0000 1636.8 1328.4
24 0.3535 0.4531 0.000000 0.0000 1813.1 1525.8
25 0.4121 0.4844 0.000000 -0.0000 988.4 747.1
26 0.5898 0.6875 0.000000 -0.0001 1031.5 783.3
27 0.4199 0.5469 0.000000 -0.0002 2131.0 1853.8
28 0.5605 0.6250 0.000000 -0.0003 998.4 778.3

job_630127

Step Reward Pass@8 KL Loss Step Time (s) Gen Wait (s)
29 0.5430 0.6610 0.000000 -0.0001 3907.1 3661.5
30 0.3984 0.5156 0.000000 -0.0000 1561.1 1238.2
31 0.2578 0.3906 0.000000 -0.0001 558.6 247.6
32 0.4316 0.5469 0.000000 -0.0001 1881.7 1636.3
33 0.3945 0.5469 0.000000 0.0000 1516.1 1212.1
34 0.4160 0.5156 0.000000 -0.0000 1068.5 739.5
35 0.4102 0.5000 0.000000 -0.0000 1384.1 1112.5
36 0.4609 0.6406 0.000000 -0.0000 1464.7 1177.2
37 0.5703 0.6875 0.000000 -0.0000 1077.8 785.6
38 0.4570 0.6094 0.000000 -0.0000 1281.5 984.5
39 0.4219 0.5625 0.000000 -0.0000 1220.2 969.3
40 0.4766 0.5938 0.000000 0.0000 1608.5 1354.9
41 0.4980 0.5781 0.000000 -0.0000 1037.8 802.5
42 0.3770 0.4531 0.000000 0.0000 1385.1 1098.6
43 0.4043 0.4844 0.000000 -0.0000 1474.0 1198.9
44 0.4980 0.6250 0.000000 -0.0000 995.3 726.5
45 0.4707 0.5156 0.000000 0.0000 1030.1 747.8
46 0.4766 0.5625 0.000000 0.0000 1116.6 848.4
47 0.5098 0.6250 0.000000 0.0000 1436.6 1153.8
48 0.3965 0.5000 0.000000 -0.0000 1388.9 1118.5
49 0.3613 0.4844 0.000000 0.0000 1674.2 1335.1
50 0.4473 0.5469 0.000000 0.0000 1060.2 789.3
51 0.3848 0.4531 0.000000 0.0000 1686.7 1378.9
52 0.4102 0.5312 0.000000 -0.0000 1600.4 1319.3
53 0.4473 0.5312 0.000000 0.0000 1082.9 829.7
54 0.3887 0.5156 0.000000 0.0000 1827.2 1523.5
55 0.4004 0.5156 0.000000 -0.0000 1484.0 1185.3
56 0.4121 0.5156 0.000000 -0.0000 1375.8 1104.5

job_630130

Step Reward Pass@8 KL Loss Step Time (s) Gen Wait (s)
57 0.2051 0.3898 0.000000 -0.0007 10105.4 9825.4
58 0.2266 0.5000 0.000000 -0.0008 3407.9 3181.4
59 0.2520 0.5000 0.000000 -0.0015 3636.1 3468.6
60 0.2129 0.4531 0.000000 -0.0008 3855.0 3710.5
61 0.4590 0.7188 0.000000 -0.0015 1882.4 1715.1
62 0.5117 0.6094 0.000000 -0.0003 1318.3 1070.7
63 0.4746 0.5469 0.000000 -0.0001 1359.1 1102.8
64 0.3809 0.4844 0.000000 0.0000 1636.6 1345.1
65 0.4668 0.6094 0.000000 0.0000 1536.8 1260.2
66 0.3965 0.5781 0.000000 -0.0000 1507.0 1239.0
67 0.4414 0.5625 0.000000 -0.0000 1571.3 1328.7
68 0.4180 0.5312 0.000000 -0.0000 1442.2 1172.7
69 0.5449 0.6562 0.000000 -0.0000 1413.2 1162.3
70 0.4102 0.5312 0.000000 -0.0000 1779.8 1503.9

job_661735

Step Reward Pass@8 KL Loss Step Time (s) Gen Wait (s)
71 0.5566 0.5965 0.000000 -0.0000 3675.9 3449.1
72 0.2773 0.4219 0.000000 0.0000 1642.3 1345.3
73 0.3320 0.4062 0.000000 0.0000 875.4 618.9
74 0.4941 0.5938 0.000000 -0.0000 1734.0 1436.9
75 0.3906 0.5625 0.000000 -0.0000 1567.8 1296.9
76 0.3164 0.4531 0.000000 0.0000 1415.5 1155.5
77 0.4043 0.4688 0.000000 -0.0001 1161.9 906.7
78 0.2266 0.3750 0.000000 0.0000 1837.8 1593.9

Timing Analysis

Average Time Breakdown (% of step time)

Component Avg % of Step Time
wait_for_generation_buffer 80.8%
run_training 17.3%
train_critic_and_policy 14.8%
policy_train 14.8%
fwd_logprobs_values_reward 2.5%
sync_weights 1.6%
save_hf_model 1.0%
save_checkpoints 0.9%
convert_to_training_input 0.3%
cleanup_old_checkpoints 0.2%
compute_advantages_and_returns 0.0%

Cross-Log Comparison

Log Avg Reward Pass@8 Step Time (s) Gen Wait Time (s) Avg Tokens Staleness
job_630124 0.4467 0.5575 1476.9481 1198.3201 5107.1387 1.7701
job_630127 0.4329 0.5431 1435.2019 1152.8499 5301.5774 1.8175
job_630130 0.3858 0.5479 2603.6675 2363.3124 4025.6459 1.5424
job_661735 0.3747 0.4847 1738.8407 1475.4061 4821.2654 1.9980

vLLM Inference Engine Analysis

Metrics from vLLM stat loggers (V1LoggingStatLoggerFixed).

Note

: Ray deduplicates similar log messages with [repeated Nx across cluster], so we typically capture stats from one engine per timestamp. The stats shown are per-engine values. Multiply by num_inference_engines for cluster-wide estimates.

Summary by Log (Per-Engine Stats)

Log Avg Running/Engine Avg Waiting/Engine Avg Gen Throughput/Engine Avg KV Cache % Avg Prefix Hit %
job_630124 11.5 0.4 494.3 tok/s 51.4% 67.8%
job_630127 12.0 0.6 510.3 tok/s 53.0% 69.6%
job_630130 6.4 0.0 287.3 tok/s 28.1% 85.0%
job_661734 8.1 0.0 344.0 tok/s 34.5% 83.8%
job_661735 10.2 0.2 442.6 tok/s 44.3% 69.2%

Utilization Analysis (Per-Engine)

Key indicators of inference engine utilization:

  • Running requests/engine: Concurrent requests being processed by each engine
  • Waiting requests: Requests queued (0 = engine not saturated, has spare capacity)
  • Generation throughput: Decode tokens/sec per engine
    • 8B model on H100 can do 1000+ tok/s when saturated
    • If seeing <300 tok/s with 0 waiting, engine is starved for requests

job_630124

  • Running requests/engine: avg=11.5, max=24
  • Waiting requests: avg=0.4, max=43
  • Generation throughput/engine: avg=494.3 tok/s, max=984.1 tok/s
  • KV cache usage: avg=51.4%
  • Prefix cache hit rate: avg=67.8%
  • Well-utilized: Engines saturated (waiting > 0)

job_630127

  • Running requests/engine: avg=12.0, max=24
  • Waiting requests: avg=0.6, max=41
  • Generation throughput/engine: avg=510.3 tok/s, max=1005.9 tok/s
  • KV cache usage: avg=53.0%
  • Prefix cache hit rate: avg=69.6%
  • Well-utilized: Engines saturated (waiting > 0)

job_630130

  • Running requests/engine: avg=6.4, max=24
  • Waiting requests: avg=0.0, max=19
  • Generation throughput/engine: avg=287.3 tok/s, max=952.9 tok/s
  • KV cache usage: avg=28.1%
  • Prefix cache hit rate: avg=85.0%
  • Well-utilized: Engines saturated (waiting > 0)

job_661734

  • Running requests/engine: avg=8.1, max=24
  • Waiting requests: avg=0.0, max=4
  • Generation throughput/engine: avg=344.0 tok/s, max=862.0 tok/s
  • KV cache usage: avg=34.5%
  • Prefix cache hit rate: avg=83.8%
  • Well-utilized: Engines saturated (waiting > 0)

job_661735

  • Running requests/engine: avg=10.2, max=24
  • Waiting requests: avg=0.2, max=18
  • Generation throughput/engine: avg=442.6 tok/s, max=962.9 tok/s
  • KV cache usage: avg=44.3%
  • Prefix cache hit rate: avg=69.2%
  • Well-utilized: Engines saturated (waiting > 0)

Trial-Level Analysis (from result.json)

Total trials parsed: 38123

Turn Count Statistics

Metric Value
Mean 5.0
Median 5.0
Std 2.4
Min 1
Max 33
Count 37976

Exception Distribution

Exception Type Count %
No exception 26113 68.5%
AgentTimeoutError 10372 27.2%
VerifierTimeoutError 802 2.1%
ContextLengthExceededError 679 1.8%
AgentSetupTimeoutError 105 0.3%
RuntimeError 41 0.1%
DaytonaError 7 0.0%
DaytonaAuthenticationError 4 0.0%

Turn Count by Exception Type

Exception Type Mean Turns Median Turns Count
No exception 5.6 5.0 26113
VerifierTimeoutError 5.6 5.0 802
ContextLengthExceededError 3.8 4.0 679
AgentTimeoutError 3.4 3.0 10372
DaytonaError 3.2 4.0 6
DaytonaAuthenticationError 1.8 1.5 4

Turn Count by Outcome

Outcome Mean Turns Median Turns Count
Success 5.3 5.0 16218
Failure 4.8 5.0 19449

Reward Summary

  • Mean reward: 0.4547
  • Success rate: 45.5%
  • Trials with reward data: 35667