Files
ablation-pymethods2test-sha…/training_logs/20260612_220132_metrics_report.md
ModelHub XC 68d3be9a84 初始化项目,由ModelHub XC社区提供模型
Model: laion/ablation-pymethods2test-shaped-45-8B
Source: Original Platform
2026-07-18 17:31:01 +08:00

21 KiB

SkyRL Training Metrics Analysis

Generated from 4 log files

Overview

Log File Total Steps Metric Blocks Final Reward (mean) Final Reward (max) Total Time (s)
job_628823 27 27 0.4447 0.7695 40599.0
job_628824 53 27 0.4415 0.5918 41370.3
job_628825 62 10 0.2861 0.4746 40760.4
job_628826 80 18 0.4392 0.7715 31912.3

Async Metrics

Mean Std Min Max Count
async/discard_rate 0 0 0 0 82
async/discarded_count 0 0 0 0 82
async/effective_batch_groups 64 0 64 64 82
async/effective_batch_samples 512 0 512 512 82
async/staleness_max 3.28049 1.06891 0 6 82
async/staleness_mean 1.68046 0.550336 0 4 82
async/staleness_min 0.134146 0.603737 0 4 82
async/staleness_ratio 0.861278 0.19627 0 1 82

Generate Metrics

Mean Std Min Max Count
generate/avg_num_tokens 4140.68 964.063 1476.57 5649.51 82
generate/avg_tokens_non_zero_rewards 4359.18 680.881 2851.3 5970.64 82
generate/avg_tokens_zero_rewards 4079.29 1145.8 1023.09 5838.63 82
generate/max_num_tokens 30066.1 4389.61 13369 31844 82
generate/min_num_tokens 32.2195 198.741 1 1316 82
generate/std_num_tokens 3878.95 724.206 2062.28 5375.27 82

Loss Metrics

Mean Std Min Max Count
loss/avg_final_rewards 0.423117 0.108373 0.1191 0.7715 82
loss/avg_raw_advantages 0.0113988 0.021723 -0.0129 0.0952 82
loss/avg_raw_advantages_abs 0.104518 0.0323332 0.0106 0.2013 82

Policy Metrics

Mean Std Min Max Count
policy/final_loss -8.65854e-05 0.000207133 -0.0009 0 82
policy/log_ratio_abs_max 0 0 0 0 82
policy/log_ratio_abs_mean 0 0 0 0 82
policy/log_ratio_abs_p99 0 0 0 0 82
policy/log_ratio_abs_pos00 0 0 0 0 82
policy/log_ratio_abs_pos10 0 0 0 0 82
policy/log_ratio_abs_pos20 0 0 0 0 82
policy/log_ratio_abs_pos30 0 0 0 0 82
policy/log_ratio_abs_pos40 0 0 0 0 82
policy/log_ratio_abs_pos50 0 0 0 0 82
policy/log_ratio_abs_pos60 0 0 0 0 82
policy/log_ratio_abs_pos70 0 0 0 0 82
policy/log_ratio_abs_pos80 0 0 0 0 82
policy/log_ratio_abs_pos90 0 0 0 0 82
policy/n_tokens_dp_gt_10pct 0 0 0 0 82
policy/n_tokens_dp_gt_1pct 0 0 0 0 82
policy/n_tokens_dp_gt_50pct 0 0 0 0 82
policy/policy_entropy 0.113544 0.0196945 0.054 0.1366 82
policy/policy_loss -0.00551951 0.0127705 -0.0567 0 82
policy/policy_lr 0 0 0 0 82
policy/policy_update_steps 1 0 1 1 82
policy/ppo_clip_ratio 0 0 0 0 82
policy/raw_grad_norm 0.0170537 0.00539994 0.008 0.0536 82

Reward Metrics

Mean Std Min Max Count
reward/avg_pass_at_8 0.541405 0.0905809 0.3281 0.8281 82
reward/avg_raw_reward 0.423117 0.108373 0.1191 0.7715 82

System Metrics

Mean Std Min Max Count
system/process_rss_gb 14.4693 1.40257 9.7812 17.1838 82
system/process_vms_gb 51.5833 0.759579 50.2029 53.3813 82
system/ram_available_gb 568.401 23.2428 536.837 626.327 82
system/ram_percent 33.7439 2.70948 27 37.4 82
system/ram_total_gb 857.968 1.14386e-13 857.968 857.968 82
system/ram_used_gb 289.566 23.2428 231.64 321.13 82

Timing Metrics

Mean Std Min Max Count
timing/compute_advantages_and_returns 0.0824878 0.0302697 0.0389 0.2346 82
timing/convert_to_training_input 4.44468 0.639779 1.9872 5.7156 82
timing/fwd_logprobs_values_reward 31.1934 14.3675 13.9888 148.12 82
timing/policy_train 178.7 35.655 89.1352 249.186 82
timing/run_training 210.243 42.58 103.356 306.846 82
timing/step 1885.88 1071.01 710.547 7430.96 82
timing/sync_weights 22.7287 2.34342 17.454 28.164 82
timing/train_critic_and_policy 178.967 35.6753 89.3191 249.423 82
timing/wait_for_generation_buffer 1648.46 1089.83 448.748 7102.07 82
timing/cleanup_old_checkpoints 20.7478 124.52 0.0077 788.255 40
timing/save_checkpoints 35.0582 136.626 7.892 873.951 40
timing/save_hf_model 8.11335 6.74447 5.7822 33.3153 16

Trainer Metrics

Mean Std Min Max Count
trainer/epoch 0.0243902 0.155207 0 1 82
trainer/global_step 40.4878 23.0401 1 80 82

Batch_Errors Metrics

Mean Std Min Max Count
batch_errors/total_batches 61.378 10.9518 16 122 82
batch_errors/total_instances 491.024 87.6143 128 976 82
batch_errors/total_successful 414.378 106.966 85 943 82
batch_errors/total_failed 22.8902 12.5502 10 59 82
batch_errors/total_masked 65.6098 48.9638 13 239 82
batch_errors/avg_VerifierTimeoutError 0.212868 0.287732 0.0153846 1.1129 62
batch_errors/total_VerifierTimeoutError 11.9194 16.382 1 69 62
batch_errors/avg_RuntimeError 0.0780424 0.0380127 0.015625 0.140351 9
batch_errors/total_RuntimeError 5 2 1 8 9
batch_errors/avg_ContextLengthExceededError 0.102943 0.0715027 0.015625 0.403509 80
batch_errors/total_ContextLengthExceededError 6.325 4.32472 1 23 80
batch_errors/avg_InvalidChatHistory 0.709029 0.397476 0.0793651 1.80702 82
batch_errors/total_InvalidChatHistory 43 24.6957 5 103 82
batch_errors/avg_AgentTimeoutError 0.404978 0.812804 0.015625 3.06452 70
batch_errors/total_AgentTimeoutError 23.1429 49.3504 1 190 70
batch_errors/avg_AgentSetupTimeoutError 0.0897322 0.129154 0.015625 0.507246 29
batch_errors/total_AgentSetupTimeoutError 4.65517 7.22731 1 35 29
batch_errors/avg_DaytonaAuthenticationError 0.0464813 0.0260413 0.016129 0.101695 12
batch_errors/total_DaytonaAuthenticationError 2.33333 1.66969 1 6 12
batch_errors/avg_DaytonaNotFoundError 0.016129 nan 0.016129 0.016129 1
batch_errors/total_DaytonaNotFoundError 1 nan 1 1 1
batch_errors/avg_DaytonaError 0.0164326 0.000994804 0.015625 0.0175439 3
batch_errors/total_DaytonaError 1 0 1 1 3

Training Progression by Log

job_628823

Step Reward Pass@8 KL Loss Step Time (s) Gen Wait (s)
1 0.7695 0.8281 0.000000 0.0000 3057.6 2847.4
2 0.4316 0.5781 0.000000 -0.0001 1974.6 1674.5
3 0.3223 0.4375 0.000000 -0.0000 764.6 448.7
4 0.4414 0.5625 0.000000 -0.0000 1688.5 1436.7
5 0.3984 0.5000 0.000000 -0.0000 1372.0 1100.7
6 0.4727 0.5781 0.000000 0.0000 1345.3 1049.0
7 0.5352 0.6250 0.000000 -0.0000 1358.4 1111.0
8 0.4551 0.5625 0.000000 -0.0000 1520.6 1224.2
9 0.3867 0.4844 0.000000 -0.0000 1325.5 1083.5
10 0.4590 0.5625 0.000000 -0.0000 1527.1 1243.8
11 0.4961 0.6875 0.000000 -0.0000 1372.1 1128.0
12 0.4102 0.5000 0.000000 0.0000 1553.8 1251.6
13 0.4023 0.5781 0.000000 -0.0000 1401.1 1096.5
14 0.4160 0.5625 0.000000 -0.0000 1386.3 1113.5
15 0.4844 0.5625 0.000000 0.0000 1415.7 1160.7
16 0.4512 0.5938 0.000000 0.0000 1435.3 1156.5
17 0.3633 0.5156 0.000000 0.0000 1238.7 930.7
18 0.4941 0.6094 0.000000 0.0000 1604.5 1338.2
19 0.3633 0.4844 0.000000 -0.0000 1704.6 1415.3
20 0.3359 0.4219 0.000000 -0.0000 1244.0 998.0
21 0.4414 0.5781 0.000000 -0.0000 1850.9 1594.7
22 0.4355 0.5156 0.000000 -0.0000 1397.3 1080.5
23 0.3359 0.5000 0.000000 -0.0000 1504.8 1261.4
24 0.4746 0.5625 0.000000 -0.0000 1700.5 1460.9
25 0.3633 0.5000 0.000000 -0.0001 1624.9 1359.2
26 0.5430 0.6719 0.000000 -0.0001 995.4 773.4
27 0.5254 0.6406 0.000000 -0.0001 1234.8 993.9

job_628824

Step Reward Pass@8 KL Loss Step Time (s) Gen Wait (s)
27 0.5879 0.6250 0.000000 -0.0001 4078.7 3842.9
28 0.3770 0.5000 0.000000 -0.0001 1263.0 1008.9
29 0.4277 0.5469 0.000000 0.0000 1068.8 792.5
30 0.4375 0.5156 0.000000 -0.0001 1539.5 1315.8
31 0.4531 0.5312 0.000000 0.0000 1535.1 1274.5
32 0.3320 0.4844 0.000000 -0.0001 1701.1 1466.6
33 0.4023 0.4844 0.000000 0.0000 1470.0 1220.7
34 0.4180 0.4844 0.000000 -0.0000 1363.7 1119.2
35 0.4316 0.5781 0.000000 0.0000 1421.4 1186.7
36 0.4141 0.5469 0.000000 -0.0000 1687.0 1439.7
37 0.4512 0.5469 0.000000 -0.0000 1634.9 1394.5
38 0.5918 0.7344 0.000000 0.0000 1204.9 971.6
39 0.4766 0.5469 0.000000 0.0000 1350.2 1096.7
40 0.4180 0.5625 0.000000 -0.0000 1518.6 1271.4
41 0.3906 0.5156 0.000000 -0.0000 1695.8 1427.1
42 0.3965 0.5469 0.000000 0.0000 1186.0 926.3
43 0.4590 0.5312 0.000000 -0.0000 1288.4 1028.0
44 0.5645 0.6406 0.000000 -0.0000 1113.2 888.1
45 0.4570 0.5625 0.000000 0.0000 1357.3 1111.4
46 0.4883 0.5625 0.000000 0.0000 1268.2 1044.4
47 0.4434 0.5781 0.000000 0.0000 1545.0 1273.0
48 0.4355 0.5938 0.000000 -0.0000 1623.9 1389.0
49 0.4336 0.5156 0.000000 0.0000 1295.4 1047.7
50 0.4141 0.5000 0.000000 -0.0000 1512.6 1282.3
51 0.4121 0.5312 0.000000 -0.0000 1530.0 1323.1
52 0.3945 0.5312 0.000000 -0.0000 1451.4 1226.7
53 0.4121 0.4844 0.000000 0.0000 1666.3 1423.3

job_628825

Step Reward Pass@8 KL Loss Step Time (s) Gen Wait (s)
53 0.4746 0.5818 0.000000 -0.0005 7431.0 7102.1
54 0.1191 0.3281 0.000000 -0.0002 2972.7 2812.8
55 0.3262 0.5156 0.000000 -0.0009 4313.7 4135.8
56 0.1934 0.3906 0.000000 -0.0005 2256.5 2086.4
57 0.2754 0.5000 0.000000 -0.0007 4721.0 4568.8
58 0.2930 0.5000 0.000000 -0.0008 4293.2 4168.3
59 0.2480 0.4375 0.000000 -0.0005 3191.9 3068.6
60 0.3379 0.4688 0.000000 -0.0002 3332.0 3174.5
61 0.2676 0.4375 0.000000 -0.0006 4540.5 4396.9
62 0.3262 0.5781 0.000000 -0.0009 3708.0 3576.0

job_628826

Step Reward Pass@8 KL Loss Step Time (s) Gen Wait (s)
63 0.7168 0.8136 0.000000 -0.0002 2578.9 2397.1
64 0.6172 0.6562 0.000000 -0.0000 2364.1 2140.7
65 0.2656 0.4062 0.000000 0.0000 710.5 470.5
66 0.5352 0.6875 0.000000 -0.0000 1419.1 1185.4
67 0.3438 0.4219 0.000000 -0.0001 2122.2 1843.5
68 0.3828 0.5000 0.000000 -0.0000 1248.7 1014.2
69 0.4609 0.5469 0.000000 -0.0000 1332.3 1118.8
70 0.4473 0.5469 0.000000 0.0000 1573.6 1352.4
71 0.4707 0.6250 0.000000 -0.0000 1325.1 1115.1
72 0.3750 0.4688 0.000000 0.0000 1530.9 1311.4
73 0.4512 0.5312 0.000000 0.0000 1586.4 1350.9
74 0.4180 0.5469 0.000000 0.0000 1485.3 1225.2
75 0.4512 0.5156 0.000000 -0.0001 1332.1 1131.4
76 0.3008 0.4062 0.000000 0.0000 1875.4 1654.7
77 0.3457 0.4219 0.000000 -0.0000 1368.0 1162.8
78 0.2754 0.3594 0.000000 -0.0000 1663.1 1456.8
79 0.7715 0.7812 0.000000 0.0000 3065.3 2902.5
80 0.2773 0.4375 0.000000 -0.0001 3331.0 3125.1

Timing Analysis

Average Time Breakdown (% of step time)

Component Avg % of Step Time
wait_for_generation_buffer 84.5%
run_training 13.7%
train_critic_and_policy 11.8%
policy_train 11.7%
fwd_logprobs_values_reward 2.0%
sync_weights 1.5%
save_checkpoints 1.5%
cleanup_old_checkpoints 0.7%
save_hf_model 0.5%
convert_to_training_input 0.3%
compute_advantages_and_returns 0.0%

Cross-Log Comparison

Log Avg Reward Pass@8 Step Time (s) Gen Wait Time (s) Avg Tokens Staleness
job_628823 0.4447 0.5631 1503.6669 1234.5434 4883.9745 1.7604
job_628824 0.4415 0.5475 1532.2337 1288.5914 4351.9871 1.7622
job_628825 0.2861 0.4738 4076.0445 3909.0126 2019.2750 1.3391
job_628826 0.4392 0.5374 1772.9040 1553.2566 3887.3310 1.6276

vLLM Inference Engine Analysis

Metrics from vLLM stat loggers (V1LoggingStatLoggerFixed).

Note

: Ray deduplicates similar log messages with [repeated Nx across cluster], so we typically capture stats from one engine per timestamp. The stats shown are per-engine values. Multiply by num_inference_engines for cluster-wide estimates.

Summary by Log (Per-Engine Stats)

Log Avg Running/Engine Avg Waiting/Engine Avg Gen Throughput/Engine Avg KV Cache % Avg Prefix Hit %
job_628823 12.0 0.6 514.1 tok/s 53.3% 67.4%
job_628824 12.4 0.5 534.0 tok/s 54.1% 70.2%
job_628825 4.4 0.0 199.9 tok/s 19.2% 86.4%
job_628826 11.8 1.2 499.9 tok/s 52.8% 66.7%

Utilization Analysis (Per-Engine)

Key indicators of inference engine utilization:

  • Running requests/engine: Concurrent requests being processed by each engine
  • Waiting requests: Requests queued (0 = engine not saturated, has spare capacity)
  • Generation throughput: Decode tokens/sec per engine
    • 8B model on H100 can do 1000+ tok/s when saturated
    • If seeing <300 tok/s with 0 waiting, engine is starved for requests

job_628823

  • Running requests/engine: avg=12.0, max=24
  • Waiting requests: avg=0.6, max=28
  • Generation throughput/engine: avg=514.1 tok/s, max=952.0 tok/s
  • KV cache usage: avg=53.3%
  • Prefix cache hit rate: avg=67.4%
  • Well-utilized: Engines saturated (waiting > 0)

job_628824

  • Running requests/engine: avg=12.4, max=24
  • Waiting requests: avg=0.5, max=44
  • Generation throughput/engine: avg=534.0 tok/s, max=977.8 tok/s
  • KV cache usage: avg=54.1%
  • Prefix cache hit rate: avg=70.2%
  • Well-utilized: Engines saturated (waiting > 0)

job_628825

  • Running requests/engine: avg=4.4, max=17
  • Waiting requests: avg=0.0, max=0
  • Generation throughput/engine: avg=199.9 tok/s, max=712.4 tok/s
  • KV cache usage: avg=19.2%
  • Prefix cache hit rate: avg=86.4%
  • ⚠️ Underutilized: Engines starved for requests (0 waiting, avg 4.4 running)
    • Bottleneck is likely upstream (environment execution, not inference)

job_628826

  • Running requests/engine: avg=11.8, max=24
  • Waiting requests: avg=1.2, max=52
  • Generation throughput/engine: avg=499.9 tok/s, max=986.9 tok/s
  • KV cache usage: avg=52.8%
  • Prefix cache hit rate: avg=66.7%
  • Well-utilized: Engines saturated (waiting > 0)

Trial-Level Analysis (from result.json)

Total trials parsed: 44409

Turn Count Statistics

Metric Value
Mean 4.7
Median 4.0
Std 2.3
Min 1
Max 30
Count 44161

Exception Distribution

Exception Type Count %
No exception 29002 65.3%
AgentTimeoutError 13737 30.9%
VerifierTimeoutError 880 2.0%
ContextLengthExceededError 511 1.2%
AgentSetupTimeoutError 202 0.5%
RuntimeError 45 0.1%
DaytonaAuthenticationError 28 0.1%
DaytonaError 3 0.0%
DaytonaNotFoundError 1 0.0%

Turn Count by Exception Type

Exception Type Mean Turns Median Turns Count
DaytonaError 6.5 6.5 2
No exception 5.4 5.0 29002
VerifierTimeoutError 5.2 5.0 880
ContextLengthExceededError 4.0 4.0 511
AgentTimeoutError 3.1 3.0 13737
DaytonaAuthenticationError 2.6 2.5 28
DaytonaNotFoundError 2.0 2.0 1

Turn Count by Outcome

Outcome Mean Turns Median Turns Count
Success 5.1 5.0 18561
Failure 4.4 4.0 22465

Reward Summary

  • Mean reward: 0.4524
  • Success rate: 45.2%
  • Trials with reward data: 41026