Files
ablation-pymethods2test-seq…/training_logs/20260608_210110_metrics_report.md
ModelHub XC 67edc8d5a6 初始化项目,由ModelHub XC社区提供模型
Model: laion/ablation-pymethods2test-seqnorm-15-8B
Source: Original Platform
2026-08-08 06:25:19 +08:00

23 KiB
Raw Permalink Blame History

SkyRL Training Metrics Analysis

Generated from 5 log files

Overview

Log File Total Steps Metric Blocks Final Reward (mean) Final Reward (max) Total Time (s)
job_639470 28 28 0.4418 0.7793 41831.8
job_639471 49 21 0.3985 0.5566 40830.2
job_639473 65 17 0.4187 0.5000 32146.5
job_639474 80 16 0.3705 0.7148 38247.4
job_639475 82 1 0.7676 0.7676 2771.7

Async Metrics

Mean Std Min Max Count
async/discard_rate 0 0 0 0 83
async/discarded_count 0 0 0 0 83
async/effective_batch_groups 64 0 64 64 83
async/effective_batch_samples 512 0 512 512 83
async/staleness_max 3.3494 1.22408 0 7 83
async/staleness_mean 1.71348 0.591401 0 4 83
async/staleness_min 0.156627 0.614498 0 4 83
async/staleness_ratio 0.851277 0.224556 0 1 83

Generate Metrics

Mean Std Min Max Count
generate/avg_num_tokens 4458.85 862.392 1713.21 6032.85 83
generate/avg_tokens_non_zero_rewards 4543.78 578.607 3120.78 5796.55 83
generate/avg_tokens_zero_rewards 4503.36 1081.35 1084.97 6253.61 83
generate/max_num_tokens 29333.3 4980.39 12999 31847 83
generate/min_num_tokens 1 0 1 1 83
generate/std_num_tokens 4009.1 673.611 2205.1 5542.49 83

Loss Metrics

Mean Std Min Max Count
loss/avg_final_rewards 0.416272 0.112108 0.1445 0.7793 83
loss/avg_raw_advantages 0.0180807 0.0320128 -0.0169 0.1768 83
loss/avg_raw_advantages_abs 0.115995 0.0339654 0.0566 0.2262 83

Policy Metrics

Mean Std Min Max Count
policy/final_loss 0 0 -0 0 83
policy/log_ratio_abs_max 0 0 0 0 83
policy/log_ratio_abs_mean 0 0 0 0 83
policy/log_ratio_abs_p99 0 0 0 0 83
policy/log_ratio_abs_pos00 0 0 0 0 83
policy/log_ratio_abs_pos10 0 0 0 0 83
policy/log_ratio_abs_pos20 0 0 0 0 83
policy/log_ratio_abs_pos30 0 0 0 0 83
policy/log_ratio_abs_pos40 0 0 0 0 83
policy/log_ratio_abs_pos50 0 0 0 0 83
policy/log_ratio_abs_pos60 0 0 0 0 83
policy/log_ratio_abs_pos70 0 0 0 0 83
policy/log_ratio_abs_pos80 0 0 0 0 83
policy/log_ratio_abs_pos90 0 0 0 0 83
policy/n_tokens_dp_gt_10pct 0 0 0 0 83
policy/n_tokens_dp_gt_1pct 0 0 0 0 83
policy/n_tokens_dp_gt_50pct 0 0 0 0 83
policy/policy_entropy 0.123722 0.0189336 0.057 0.1469 83
policy/policy_loss 0 0 -0 0 83
policy/policy_lr 0 0 0 0 83
policy/policy_update_steps 1 0 1 1 83
policy/ppo_clip_ratio 0 0 0 0 83
policy/raw_grad_norm 0 0 0 0 83

Reward Metrics

Mean Std Min Max Count
reward/avg_pass_at_8 0.54622 0.0939465 0.3438 0.8438 83
reward/avg_raw_reward 0.416272 0.112108 0.1445 0.7793 83

System Metrics

Mean Std Min Max Count
system/process_rss_gb 13.911 1.27121 7.261 15.3689 83
system/process_vms_gb 51.4434 0.673945 50.0898 52.7578 83
system/ram_available_gb 575.799 19.8564 554.735 627.205 83
system/ram_percent 32.8867 2.31763 26.9 35.3 83
system/ram_total_gb 857.968 4.75679e-05 857.967 857.968 83
system/ram_used_gb 282.169 19.8564 230.763 303.233 83

Timing Metrics

Mean Std Min Max Count
timing/compute_advantages_and_returns 0.0802976 0.026931 0.0314 0.1765 83
timing/convert_to_training_input 4.31686 0.687255 1.9562 5.1247 83
timing/fwd_logprobs_values_reward 32.4571 6.87409 14.1649 69.6756 83
timing/policy_train 187.948 32.7461 98.0173 266.819 83
timing/run_training 220.717 37.8457 113.461 308.663 83
timing/step 1877.44 1202.31 605.126 9181.46 83
timing/sync_weights 22.2263 2.01847 17.9428 27.5955 83
timing/train_critic_and_policy 188.179 32.7593 98.173 267.078 83
timing/wait_for_generation_buffer 1630.18 1218.35 302.928 9003.63 83
timing/cleanup_old_checkpoints 3.75404 13.3134 0.007 73.2928 41
timing/save_checkpoints 16.8853 18.0611 7.6926 82.7069 41
timing/save_hf_model 9.786 7.68081 5.8565 36.4673 17

Trainer Metrics

Mean Std Min Max Count
trainer/epoch 0.0361446 0.187784 0 1 83
trainer/global_step 41.3976 23.4229 1 82 83

Batch_Errors Metrics

Mean Std Min Max Count
batch_errors/total_batches 60.8554 10.6749 9 125 83
batch_errors/total_instances 486.843 85.3996 72 1000 83
batch_errors/total_successful 429.048 104.407 72 979 83
batch_errors/total_failed 21.1807 13.1716 0 57 83
batch_errors/total_masked 38.6506 26.3954 0 126 83
batch_errors/avg_VerifierTimeoutError 0.309534 0.530966 0.015625 2.57895 72
batch_errors/total_VerifierTimeoutError 18.1389 30.3984 1 147 72
batch_errors/avg_ContextLengthExceededError 0.129623 0.092254 0.0163934 0.640625 82
batch_errors/total_ContextLengthExceededError 7.93902 5.80623 1 41 82
batch_errors/avg_AgentTimeoutError 0.298439 0.547499 0.008 2.20339 69
batch_errors/total_AgentTimeoutError 17.5362 32.125 1 130 69
batch_errors/avg_InvalidChatHistory 0.312427 0.162863 0.04 0.891892 82
batch_errors/total_InvalidChatHistory 18.6585 8.79284 4 41 82
batch_errors/avg_AgentSetupTimeoutError 0.168594 0.226175 0.008 0.719298 35
batch_errors/total_AgentSetupTimeoutError 10.0857 13.6411 1 41 35
batch_errors/avg_DaytonaAuthenticationError 0.0210748 0.00742164 0.015625 0.0350877 13
batch_errors/total_DaytonaAuthenticationError 1.23077 0.438529 1 2 13
batch_errors/avg_RuntimeError 0.0166667 nan 0.0166667 0.0166667 1
batch_errors/total_RuntimeError 1 nan 1 1 1
batch_errors/avg_Timeout 0.015873 nan 0.015873 0.015873 1
batch_errors/total_Timeout 1 nan 1 1 1
batch_errors/avg_DaytonaError 0.0377139 0.0400871 0.015625 0.171875 18
batch_errors/total_DaytonaError 2.38889 2.54694 1 11 18

Training Progression by Log

job_639470

Step Reward Pass@8 KL Loss Step Time (s) Gen Wait (s)
1 0.7793 0.8438 0.000000 0.0000 2948.9 2739.0
2 0.4785 0.6250 0.000000 -0.0000 1838.1 1543.4
3 0.2559 0.3594 0.000000 -0.0000 828.6 530.3
4 0.5254 0.6250 0.000000 0.0000 1643.3 1379.0
5 0.3672 0.4844 0.000000 -0.0000 1592.0 1316.7
6 0.4902 0.5781 0.000000 -0.0000 968.7 657.1
7 0.4492 0.5625 0.000000 0.0000 1394.4 1111.7
8 0.4629 0.6250 0.000000 0.0000 1826.3 1488.3
9 0.4219 0.5625 0.000000 0.0000 1383.6 1101.3
10 0.4609 0.5938 0.000000 -0.0000 1131.1 858.3
11 0.5020 0.6719 0.000000 -0.0000 1502.1 1254.5
12 0.3535 0.4375 0.000000 -0.0000 1268.0 1001.2
13 0.4434 0.5938 0.000000 0.0000 1478.8 1216.7
14 0.4199 0.6094 0.000000 -0.0000 1529.1 1277.9
15 0.5527 0.6250 0.000000 -0.0000 1397.2 1161.4
16 0.3379 0.5000 0.000000 -0.0000 2136.4 1877.7
17 0.3711 0.4531 0.000000 -0.0000 1116.8 883.8
18 0.4590 0.6094 0.000000 -0.0000 2036.8 1792.5
19 0.3379 0.4688 0.000000 0.0000 1344.7 1086.8
20 0.4277 0.5000 0.000000 -0.0000 1843.2 1606.6
21 0.3633 0.5156 0.000000 0.0000 1421.9 1169.5
22 0.3945 0.5312 0.000000 0.0000 1494.0 1237.3
23 0.4062 0.5156 0.000000 -0.0000 1294.0 1041.8
24 0.3438 0.4531 0.000000 -0.0000 1641.5 1393.6
25 0.5234 0.6406 0.000000 -0.0000 1090.5 876.4
26 0.5078 0.5781 0.000000 -0.0000 1169.9 878.8
27 0.5137 0.6406 0.000000 -0.0000 1096.8 860.3
28 0.4199 0.5469 0.000000 0.0000 1415.0 1171.5

job_639471

Step Reward Pass@8 KL Loss Step Time (s) Gen Wait (s)
29 0.5566 0.6481 0.000000 0.0000 4371.3 4105.3
30 0.3203 0.4219 0.000000 -0.0000 695.9 443.3
31 0.3984 0.5312 0.000000 -0.0000 1107.9 856.4
32 0.4180 0.5000 0.000000 0.0000 2021.9 1750.4
33 0.2852 0.4062 0.000000 -0.0000 1289.5 999.2
34 0.5000 0.6406 0.000000 -0.0000 1301.8 1072.4
35 0.3945 0.5781 0.000000 -0.0000 1785.1 1510.3
36 0.5352 0.6250 0.000000 0.0000 1218.1 941.4
37 0.5000 0.6250 0.000000 -0.0000 1048.3 802.1
38 0.3906 0.5312 0.000000 -0.0000 1732.0 1470.6
39 0.4941 0.6250 0.000000 -0.0000 1244.4 951.3
40 0.4492 0.5625 0.000000 0.0000 1248.5 996.7
41 0.4199 0.5469 0.000000 -0.0000 1448.2 1228.0
42 0.3848 0.4688 0.000000 0.0000 1518.1 1247.8
43 0.5371 0.6250 0.000000 -0.0000 1281.3 1041.2
44 0.4004 0.5000 0.000000 -0.0000 1070.9 816.1
45 0.3652 0.4688 0.000000 -0.0000 2427.1 2152.5
46 0.3203 0.5156 0.000000 -0.0000 2780.7 2597.7
47 0.1445 0.4375 0.000000 -0.0000 4012.7 3830.9
48 0.3086 0.5312 0.000000 -0.0000 3641.6 3508.2
49 0.2461 0.4531 0.000000 -0.0000 3585.2 3403.3

job_639473

Step Reward Pass@8 KL Loss Step Time (s) Gen Wait (s)
49 0.3770 0.5625 0.000000 -0.0000 9181.5 9003.6
50 0.3066 0.4531 0.000000 -0.0000 611.4 379.4
51 0.4395 0.5312 0.000000 -0.0000 1837.5 1632.0
52 0.4648 0.5781 0.000000 -0.0000 1732.7 1494.7
53 0.3184 0.5000 0.000000 -0.0000 1262.3 1029.2
54 0.4043 0.4844 0.000000 0.0000 1903.4 1623.4
55 0.3945 0.4844 0.000000 -0.0000 1296.1 1039.0
56 0.3691 0.4844 0.000000 0.0000 1565.5 1274.6
57 0.4434 0.5781 0.000000 0.0000 1726.5 1440.4
58 0.4336 0.5781 0.000000 -0.0000 1407.8 1164.7
59 0.4609 0.5625 0.000000 -0.0000 1490.2 1213.3
60 0.4902 0.6250 0.000000 -0.0000 1325.7 1073.0
61 0.5000 0.6562 0.000000 -0.0000 1404.2 1158.4
62 0.4824 0.5625 0.000000 0.0000 1122.7 856.3
63 0.3809 0.4844 0.000000 -0.0000 1242.4 982.7
64 0.3887 0.4844 0.000000 -0.0000 1523.1 1218.0
65 0.4629 0.5469 0.000000 -0.0000 1513.4 1192.9

job_639474

Step Reward Pass@8 KL Loss Step Time (s) Gen Wait (s)
65 0.5820 0.6727 0.000000 -0.0000 4741.5 4442.7
66 0.3535 0.5000 0.000000 -0.0000 605.1 302.9
67 0.4102 0.5156 0.000000 -0.0000 1183.0 936.1
68 0.2285 0.3594 0.000000 -0.0000 5067.2 4796.8
69 0.2480 0.5000 0.000000 -0.0000 2845.6 2663.1
70 0.2461 0.5469 0.000000 -0.0000 2600.3 2461.9
71 0.4375 0.6094 0.000000 -0.0000 1848.2 1645.6
72 0.3438 0.4531 0.000000 -0.0000 3286.0 3069.1
73 0.2227 0.3594 0.000000 -0.0000 2698.0 2502.3
74 0.2812 0.4844 0.000000 -0.0000 1731.4 1534.3
75 0.4922 0.6406 0.000000 -0.0000 1591.8 1391.3
76 0.3574 0.5312 0.000000 -0.0000 1915.9 1706.4
77 0.3340 0.4375 0.000000 0.0000 1171.7 945.1
78 0.1719 0.3438 0.000000 -0.0000 1996.0 1752.6
79 0.7148 0.7812 0.000000 -0.0000 3207.1 3026.0
80 0.5039 0.6250 0.000000 -0.0000 1758.6 1521.0

job_639475

Step Reward Pass@8 KL Loss Step Time (s) Gen Wait (s)
82 0.7676 0.8281 0.000000 0.0000 2771.7 2593.2

Timing Analysis

Average Time Breakdown (% of step time)

Component Avg % of Step Time
wait_for_generation_buffer 83.2%
run_training 15.0%
train_critic_and_policy 12.8%
policy_train 12.8%
fwd_logprobs_values_reward 2.2%
sync_weights 1.5%
save_checkpoints 1.3%
save_hf_model 0.7%
cleanup_old_checkpoints 0.5%
convert_to_training_input 0.3%
compute_advantages_and_returns 0.0%

Cross-Log Comparison

Log Avg Reward Pass@8 Step Time (s) Gen Wait Time (s) Avg Tokens Staleness
job_639470 0.4418 0.5625 1493.9938 1232.6167 4864.4295 1.7768
job_639471 0.3985 0.5353 1944.2947 1701.1955 4306.7650 1.7485
job_639473 0.4187 0.5386 1890.9678 1633.8600 4557.1723 1.6259
job_639474 0.3705 0.5225 2390.4636 2168.5831 3890.0232 1.7568
job_639475 0.7676 0.8281 2771.6856 2593.2438 3726.2617 0.0000

vLLM Inference Engine Analysis

Metrics from vLLM stat loggers (V1LoggingStatLoggerFixed).

Note

: Ray deduplicates similar log messages with [repeated Nx across cluster], so we typically capture stats from one engine per timestamp. The stats shown are per-engine values. Multiply by num_inference_engines for cluster-wide estimates.

Summary by Log (Per-Engine Stats)

Log Avg Running/Engine Avg Waiting/Engine Avg Gen Throughput/Engine Avg KV Cache % Avg Prefix Hit %
job_632522 8.5 0.0 380.3 tok/s 35.3% 86.8%
job_632523 9.7 0.0 433.1 tok/s 40.9% 87.6%
job_632524 10.4 0.0 457.7 tok/s 43.2% 86.4%
job_632525 9.1 0.0 404.4 tok/s 37.4% 88.5%
job_639470 11.5 0.4 492.9 tok/s 51.0% 73.8%
job_639471 8.8 0.4 374.1 tok/s 39.0% 72.1%
job_639473 10.4 0.3 443.9 tok/s 46.0% 75.4%
job_639474 7.2 0.1 315.1 tok/s 31.4% 81.6%
job_639475 7.6 0.0 338.0 tok/s 31.2% 84.6%
job_639476 3.4 0.0 150.5 tok/s 8.7% 67.5%

Utilization Analysis (Per-Engine)

Key indicators of inference engine utilization:

  • Running requests/engine: Concurrent requests being processed by each engine
  • Waiting requests: Requests queued (0 = engine not saturated, has spare capacity)
  • Generation throughput: Decode tokens/sec per engine
    • 8B model on H100 can do 1000+ tok/s when saturated
    • If seeing <300 tok/s with 0 waiting, engine is starved for requests

job_632522

  • Running requests/engine: avg=8.5, max=21
  • Waiting requests: avg=0.0, max=0
  • Generation throughput/engine: avg=380.3 tok/s, max=731.5 tok/s
  • KV cache usage: avg=35.3%
  • Prefix cache hit rate: avg=86.8%
  • Moderate utilization

job_632523

  • Running requests/engine: avg=9.7, max=21
  • Waiting requests: avg=0.0, max=1
  • Generation throughput/engine: avg=433.1 tok/s, max=779.6 tok/s
  • KV cache usage: avg=40.9%
  • Prefix cache hit rate: avg=87.6%
  • Well-utilized: Engines saturated (waiting > 0)

job_632524

  • Running requests/engine: avg=10.4, max=24
  • Waiting requests: avg=0.0, max=8
  • Generation throughput/engine: avg=457.7 tok/s, max=1016.5 tok/s
  • KV cache usage: avg=43.2%
  • Prefix cache hit rate: avg=86.4%
  • Well-utilized: Engines saturated (waiting > 0)

job_632525

  • Running requests/engine: avg=9.1, max=24
  • Waiting requests: avg=0.0, max=0
  • Generation throughput/engine: avg=404.4 tok/s, max=845.3 tok/s
  • KV cache usage: avg=37.4%
  • Prefix cache hit rate: avg=88.5%
  • Moderate utilization

job_639470

  • Running requests/engine: avg=11.5, max=32
  • Waiting requests: avg=0.4, max=31
  • Generation throughput/engine: avg=492.9 tok/s, max=1265.8 tok/s
  • KV cache usage: avg=51.0%
  • Prefix cache hit rate: avg=73.8%
  • Well-utilized: Engines saturated (waiting > 0)

job_639471

  • Running requests/engine: avg=8.8, max=24
  • Waiting requests: avg=0.4, max=39
  • Generation throughput/engine: avg=374.1 tok/s, max=939.4 tok/s
  • KV cache usage: avg=39.0%
  • Prefix cache hit rate: avg=72.1%
  • Well-utilized: Engines saturated (waiting > 0)

job_639473

  • Running requests/engine: avg=10.4, max=24
  • Waiting requests: avg=0.3, max=22
  • Generation throughput/engine: avg=443.9 tok/s, max=958.4 tok/s
  • KV cache usage: avg=46.0%
  • Prefix cache hit rate: avg=75.4%
  • Well-utilized: Engines saturated (waiting > 0)

job_639474

  • Running requests/engine: avg=7.2, max=24
  • Waiting requests: avg=0.1, max=16
  • Generation throughput/engine: avg=315.1 tok/s, max=974.2 tok/s
  • KV cache usage: avg=31.4%
  • Prefix cache hit rate: avg=81.6%
  • Well-utilized: Engines saturated (waiting > 0)

job_639475

  • Running requests/engine: avg=7.6, max=20
  • Waiting requests: avg=0.0, max=0
  • Generation throughput/engine: avg=338.0 tok/s, max=782.0 tok/s
  • KV cache usage: avg=31.2%
  • Prefix cache hit rate: avg=84.6%
  • Moderate utilization

job_639476

  • Running requests/engine: avg=3.4, max=9
  • Waiting requests: avg=0.0, max=0
  • Generation throughput/engine: avg=150.5 tok/s, max=425.9 tok/s
  • KV cache usage: avg=8.7%
  • Prefix cache hit rate: avg=67.5%
  • ⚠️ Underutilized: Engines starved for requests (0 waiting, avg 3.4 running)
    • Bottleneck is likely upstream (environment execution, not inference)

Trial-Level Analysis (from result.json)

Total trials parsed: 46308

Turn Count Statistics

Metric Value
Mean 5.4
Median 5.0
Std 2.6
Min 1
Max 137
Count 45918

Exception Distribution

Exception Type Count %
No exception 31188 67.3%
AgentTimeoutError 12483 27.0%
VerifierTimeoutError 1536 3.3%
ContextLengthExceededError 671 1.4%
AgentSetupTimeoutError 366 0.8%
DaytonaError 46 0.1%
DaytonaAuthenticationError 16 0.0%
Timeout 1 0.0%
RuntimeError 1 0.0%

Turn Count by Exception Type

Exception Type Mean Turns Median Turns Count
RuntimeError 13.0 13.0 1
No exception 6.0 5.0 31188
VerifierTimeoutError 5.7 5.0 1536
ContextLengthExceededError 4.8 4.0 671
AgentTimeoutError 3.9 4.0 12483
DaytonaAuthenticationError 3.3 3.0 16
DaytonaError 3.1 2.5 22
Timeout 2.0 2.0 1

Turn Count by Outcome

Outcome Mean Turns Median Turns Count
Success 5.6 5.0 19074
Failure 5.3 5.0 23341

Reward Summary

  • Mean reward: 0.4497
  • Success rate: 45.0%
  • Trials with reward data: 42415