Files
ablation-pymethods2test-seq…/training_logs/20260609_214045_metrics_report.md
ModelHub XC 9f22667c70 初始化项目,由ModelHub XC社区提供模型
Model: laion/ablation-pymethods2test-seqmean-arm0-tis-15-8B
Source: Original Platform
2026-08-09 05:27:18 +08:00

23 KiB

SkyRL Training Metrics Analysis

Generated from 5 log files

Overview

Log File Total Steps Metric Blocks Final Reward (mean) Final Reward (max) Total Time (s)
job_653585 14 14 0.4019 0.5645 35492.1
job_653586 34 20 0.3792 0.6133 42208.0
job_653587 60 26 0.4408 0.5781 41323.7
job_653588 80 20 0.4295 0.7598 35422.1
job_653589 82 1 0.7832 0.7832 3316.0

Async Metrics

Mean Std Min Max Count
async/discard_rate 0 0 0 0 81
async/discarded_count 0 0 0 0 81
async/effective_batch_groups 64 0 64 64 81
async/effective_batch_samples 512 0 512 512 81
async/staleness_max 3.4321 1.28392 0 8 81
async/staleness_mean 1.65336 0.609775 0 4.4219 81
async/staleness_min 0.148148 0.614636 0 4 81
async/staleness_ratio 0.841627 0.223298 0 1 81

Generate Metrics

Mean Std Min Max Count
generate/avg_num_tokens 3986.2 804.39 1484.36 5270.01 81
generate/avg_tokens_non_zero_rewards 3999.99 543.159 2738.59 5697.48 81
generate/avg_tokens_zero_rewards 4077.34 1100.09 775.223 5697.65 81
generate/max_num_tokens 28006.5 5837.94 13362 31858 81
generate/min_num_tokens 1 0 1 1 81
generate/std_num_tokens 3737.82 710.017 1838.67 5068.11 81
generate/tis/aligned_tokens 1.40771e+06 292189 540011 1.95353e+06 81
generate/tis/alignment_fail_count 27.6173 24.1674 2 119 81
generate/tis/exact_match_fraction 0.948404 0.0347641 0.8463 0.9993 81
generate/tis/lcs_fallback_fraction 0 0 0 0 81
generate/tis/lcs_fallback_messages 0 0 0 0 81
generate/tis/unaligned_fraction 0.0515963 0.0347641 0.0007 0.1537 81

Loss Metrics

Mean Std Min Max Count
loss/avg_final_rewards 0.420311 0.108774 0.1621 0.7832 81
loss/avg_raw_advantages 0.0198173 0.046916 -0.0268 0.2099 81
loss/avg_raw_advantages_abs 0.10669 0.045766 0.0294 0.2501 81

Policy Metrics

Mean Std Min Max Count
policy/final_loss -0.000169136 0.000402692 -0.0019 0 81
policy/log_ratio_abs_max 0 0 0 0 81
policy/log_ratio_abs_mean 0 0 0 0 81
policy/log_ratio_abs_p99 0 0 0 0 81
policy/log_ratio_abs_pos00 0 0 0 0 81
policy/log_ratio_abs_pos10 0 0 0 0 81
policy/log_ratio_abs_pos20 0 0 0 0 81
policy/log_ratio_abs_pos30 0 0 0 0 81
policy/log_ratio_abs_pos40 0 0 0 0 81
policy/log_ratio_abs_pos50 0 0 0 0 81
policy/log_ratio_abs_pos60 0 0 0 0 81
policy/log_ratio_abs_pos70 0 0 0 0 81
policy/log_ratio_abs_pos80 0 0 0 0 81
policy/log_ratio_abs_pos90 0 0 0 0 81
policy/n_tokens_dp_gt_10pct 0 0 0 0 81
policy/n_tokens_dp_gt_1pct 0 0 0 0 81
policy/n_tokens_dp_gt_50pct 0 0 0 0 81
policy/policy_entropy 0.0500136 0.0150327 0.0251 0.0938 81
policy/policy_loss -0.0109025 0.0256364 -0.1214 0.0003 81
policy/policy_lr 0 0 0 0 81
policy/policy_update_steps 1 0 1 1 81
policy/ppo_clip_ratio 0 0 0 0 81
policy/raw_grad_norm 0.017084 0.00422949 0.0092 0.0312 81
policy/rollout_train_prob_diff_mean 3.20105e+22 1.82246e+23 1.0002 1.39973e+24 81
policy/rollout_train_prob_diff_std 3.58965e+25 2.09602e+26 1.312 1.70849e+27 81
policy/tis/imp_ratio_capped_fraction 1.23457e-06 1.11111e-05 0 0.0001 81
policy/tis/imp_ratio_mean 0.857601 0.128126 0.4014 1.0023 81
policy/tis/log_ratio_abs_mean 0.0188667 0.00900744 0.0083 0.0461 81

Reward Metrics

Mean Std Min Max Count
reward/avg_pass_at_8 0.540064 0.0939115 0.3438 0.8125 81
reward/avg_raw_reward 0.420311 0.108774 0.1621 0.7832 81

System Metrics

Mean Std Min Max Count
system/process_rss_gb 15.1484 1.86533 7.5378 17.3769 81
system/process_vms_gb 52.1818 1.04916 49.9637 54.5306 81
system/ram_available_gb 595.168 19.2446 551.906 634.681 81
system/ram_percent 30.6296 2.24552 26 35.7 81
system/ram_total_gb 857.968 4.98454e-05 857.967 857.968 81
system/ram_used_gb 262.8 19.2446 223.287 306.061 81

Timing Metrics

Mean Std Min Max Count
timing/compute_advantages_and_returns 0.0897247 0.105922 0.0323 0.8639 81
timing/convert_to_training_input 5.42322 1.13896 2.4443 7.6651 81
timing/fwd_logprobs_values_reward 31.8642 19.5819 14.0718 199.362 81
timing/policy_train 171.947 31.4651 89.3179 235.705 81
timing/run_training 204.138 36.2341 111.54 288.842 81
timing/step 1947.68 1358.35 608.871 11028.1 81
timing/sync_weights 31.7567 6.30151 18.1382 48.8291 81
timing/train_critic_and_policy 172.183 31.4877 89.4422 235.953 81
timing/wait_for_generation_buffer 1706.35 1363.3 352.922 10718.7 81
timing/cleanup_old_checkpoints 4.49109 21.9401 0.0072 139.958 41
timing/save_checkpoints 14.6602 22.1707 8.0671 149.994 41
timing/save_hf_model 7.06378 2.67737 5.7884 16.893 16

Tis Metrics

Mean Std Min Max Count
tis/batch_skipped_no_logprobs 0 0 0 0 81
tis/skipped_fraction 0 0 0 0 81

Trainer Metrics

Mean Std Min Max Count
trainer/epoch 0.037037 0.190029 0 1 81
trainer/global_step 41.0123 23.5481 1 82 81

Batch_Errors Metrics

Mean Std Min Max Count
batch_errors/total_batches 61.284 11.5198 3 121 81
batch_errors/total_instances 490.272 92.1584 24 968 81
batch_errors/total_successful 415.506 86.5525 24 690 81
batch_errors/total_failed 22.8889 15.9421 0 114 81
batch_errors/total_masked 53.6173 32.8792 0 216 81
batch_errors/avg_VerifierTimeoutError 0.367127 0.644218 0.015625 2.7 63
batch_errors/total_VerifierTimeoutError 24.1587 45.4107 1 231 63
batch_errors/avg_AgentTimeoutError 0.221598 0.442882 0.015625 2.01587 73
batch_errors/total_AgentTimeoutError 15.0685 33.7895 1 200 73
batch_errors/avg_AgentSetupTimeoutError 0.231264 0.20643 0.016129 0.730159 16
batch_errors/total_AgentSetupTimeoutError 15.125 13.4604 1 46 16
batch_errors/avg_InvalidChatHistory 0.573631 0.281376 0.0169492 1.25926 80
batch_errors/total_InvalidChatHistory 34.8 16.5073 1 68 80
batch_errors/avg_ContextLengthExceededError 0.091597 0.0813106 0.015625 0.580645 74
batch_errors/total_ContextLengthExceededError 5.72973 5.20615 1 36 74
batch_errors/avg_DaytonaAuthenticationError 0.0275673 0.0259453 0.015625 0.0952381 9
batch_errors/total_DaytonaAuthenticationError 1.88889 1.69148 1 6 9
batch_errors/avg_Timeout 0.015873 nan 0.015873 0.015873 1
batch_errors/total_Timeout 1 nan 1 1 1
batch_errors/avg_DaytonaError 0.0940174 0.08989 0.015625 0.380952 38
batch_errors/total_DaytonaError 5.76316 5.66847 1 24 38

Training Progression by Log

job_653585

Step Reward Pass@8 KL Loss Step Time (s) Gen Wait (s)
1 0.3398 0.6406 0.000000 -0.0019 11028.1 10718.7
2 0.2090 0.4844 0.000000 -0.0009 3675.2 3494.5
3 0.2266 0.5156 0.000000 -0.0013 3304.1 3148.8
4 0.2754 0.5312 0.000000 -0.0013 3238.5 3089.0
5 0.4766 0.6250 0.000000 -0.0006 1062.3 897.3
6 0.5332 0.6562 0.000000 -0.0004 1290.7 1072.8
7 0.4629 0.5781 0.000000 -0.0001 2018.6 1788.6
8 0.3965 0.4688 0.000000 0.0000 1224.5 965.7
9 0.4121 0.5469 0.000000 0.0000 1481.3 1223.5
10 0.5645 0.7031 0.000000 -0.0000 1344.8 1126.3
11 0.4258 0.5312 0.000000 -0.0000 1608.6 1359.8
12 0.3945 0.5156 0.000000 0.0000 1419.0 1162.0
13 0.5293 0.6719 0.000000 0.0000 1174.5 939.0
14 0.3809 0.5000 0.000000 0.0000 1621.8 1379.8

job_653586

Step Reward Pass@8 KL Loss Step Time (s) Gen Wait (s)
15 0.6133 0.6500 0.000000 -0.0000 4387.1 4155.6
16 0.2305 0.3438 0.000000 -0.0000 940.4 626.3
17 0.4629 0.6094 0.000000 -0.0000 866.8 587.6
18 0.4062 0.5312 0.000000 -0.0000 2168.7 1866.0
19 0.4102 0.5156 0.000000 0.0000 1196.7 924.2
20 0.3809 0.4531 0.000000 -0.0000 1352.0 1092.1
21 0.3828 0.5000 0.000000 -0.0001 1653.9 1381.6
22 0.3711 0.5156 0.000000 -0.0000 1708.2 1478.8
23 0.4746 0.5781 0.000000 -0.0002 2873.6 2651.6
24 0.1816 0.3750 0.000000 -0.0005 4553.2 4339.2
25 0.1621 0.3906 0.000000 -0.0006 2911.5 2775.5
26 0.3945 0.6875 0.000000 -0.0017 2183.5 2048.2
27 0.4180 0.5938 0.000000 -0.0010 1456.9 1257.3
28 0.3145 0.5312 0.000000 -0.0011 3692.4 3488.9
29 0.3105 0.4375 0.000000 -0.0007 2121.6 1954.8
30 0.4941 0.7031 0.000000 -0.0008 1800.9 1599.0
31 0.4023 0.5312 0.000000 -0.0002 1635.7 1413.2
32 0.3926 0.5156 0.000000 -0.0000 1896.1 1646.5
33 0.4258 0.4844 0.000000 -0.0000 1165.0 952.0
34 0.3555 0.4844 0.000000 0.0000 1643.8 1436.7

job_653587

Step Reward Pass@8 KL Loss Step Time (s) Gen Wait (s)
35 0.4922 0.6066 0.000000 -0.0000 4940.9 4656.1
36 0.3652 0.4844 0.000000 0.0000 608.9 352.9
37 0.5430 0.6250 0.000000 0.0000 1222.0 998.8
38 0.4531 0.5781 0.000000 -0.0000 1845.9 1590.6
39 0.4766 0.5781 0.000000 -0.0000 1216.3 944.2
40 0.4668 0.5312 0.000000 0.0000 1338.9 1117.5
41 0.4609 0.5781 0.000000 -0.0000 1768.8 1538.6
42 0.3809 0.4375 0.000000 0.0000 1528.5 1260.4
43 0.4141 0.5000 0.000000 0.0000 1429.6 1137.8
44 0.5078 0.6094 0.000000 0.0000 1259.6 1037.9
45 0.4785 0.5312 0.000000 0.0000 1402.9 1151.6
46 0.5059 0.5781 0.000000 -0.0000 1235.1 992.3
47 0.5352 0.6250 0.000000 0.0000 1375.9 1112.2
48 0.5000 0.5938 0.000000 0.0000 1598.9 1377.6
49 0.4395 0.5781 0.000000 0.0000 1442.6 1224.6
50 0.3223 0.3906 0.000000 -0.0000 1863.3 1570.9
51 0.3555 0.4844 0.000000 -0.0000 1460.0 1228.4
52 0.3750 0.4688 0.000000 -0.0000 1339.8 1090.1
53 0.4746 0.5625 0.000000 0.0000 1705.9 1492.1
54 0.4082 0.4844 0.000000 -0.0000 1372.4 1169.5
55 0.3613 0.4688 0.000000 0.0000 1613.6 1341.7
56 0.2969 0.3750 0.000000 -0.0001 1633.4 1368.8
57 0.5000 0.6094 0.000000 -0.0000 1501.6 1212.1
58 0.3770 0.5000 0.000000 -0.0000 1448.7 1204.0
59 0.5781 0.6875 0.000000 -0.0000 1663.2 1421.5
60 0.3926 0.4688 0.000000 -0.0000 1507.2 1234.0

job_653588

Step Reward Pass@8 KL Loss Step Time (s) Gen Wait (s)
61 0.6797 0.7544 0.000000 0.0000 4895.5 4705.7
62 0.3574 0.4688 0.000000 0.0000 884.6 631.4
63 0.4160 0.4531 0.000000 0.0000 1021.0 742.7
64 0.4395 0.5938 0.000000 -0.0000 1993.6 1729.3
65 0.4531 0.5781 0.000000 0.0000 1228.5 950.0
66 0.4023 0.5312 0.000000 -0.0000 1499.4 1214.1
67 0.3652 0.4844 0.000000 -0.0001 1811.2 1536.1
68 0.3848 0.4688 0.000000 -0.0000 1248.6 975.7
69 0.5664 0.6406 0.000000 0.0000 1399.8 1145.2
70 0.4062 0.5000 0.000000 -0.0000 1712.3 1441.7
71 0.3496 0.5000 0.000000 -0.0000 1436.6 1191.9
72 0.4590 0.5000 0.000000 0.0000 1574.2 1291.8
73 0.4453 0.5469 0.000000 -0.0000 1411.8 1149.8
74 0.3574 0.4844 0.000000 0.0000 1584.8 1289.8
75 0.4121 0.5312 0.000000 -0.0000 1374.3 1132.6
76 0.3516 0.4531 0.000000 -0.0001 1688.1 1401.4
77 0.3320 0.4219 0.000000 -0.0000 1562.0 1308.2
78 0.2500 0.3594 0.000000 0.0000 1882.9 1623.0
79 0.7598 0.8125 0.000000 -0.0000 2970.0 2804.2
80 0.4023 0.5312 0.000000 -0.0000 2243.0 1957.2

job_653589

Step Reward Pass@8 KL Loss Step Time (s) Gen Wait (s)
82 0.7832 0.7969 0.000000 -0.0000 3316.0 3125.4

Timing Analysis

Average Time Breakdown (% of step time)

Component Avg % of Step Time
wait_for_generation_buffer 84.5%
run_training 13.1%
train_critic_and_policy 11.2%
policy_train 11.2%
sync_weights 2.1%
fwd_logprobs_values_reward 1.9%
save_checkpoints 0.8%
save_hf_model 0.4%
convert_to_training_input 0.4%
cleanup_old_checkpoints 0.2%
compute_advantages_and_returns 0.0%

Cross-Log Comparison

Log Avg Reward Pass@8 Step Time (s) Gen Wait Time (s) Avg Tokens Staleness
job_653585 0.4019 0.5692 2535.1497 2311.8403 3748.4641 1.6038
job_653586 0.3792 0.5216 2110.4003 1883.7625 3771.8087 1.6047
job_653587 0.4408 0.5360 1589.3730 1339.4724 4128.8501 1.7007
job_653588 0.4295 0.5307 1771.1060 1511.0877 4207.9546 1.7578
job_653589 0.7832 0.7969 3316.0311 3125.4410 3458.0977 0.0000

vLLM Inference Engine Analysis

Metrics from vLLM stat loggers (V1LoggingStatLoggerFixed).

Note

: Ray deduplicates similar log messages with [repeated Nx across cluster], so we typically capture stats from one engine per timestamp. The stats shown are per-engine values. Multiply by num_inference_engines for cluster-wide estimates.

Summary by Log (Per-Engine Stats)

Log Avg Running/Engine Avg Waiting/Engine Avg Gen Throughput/Engine Avg KV Cache % Avg Prefix Hit %
job_653585 5.9 0.0 256.7 tok/s 25.8% 86.8%
job_653586 8.3 0.2 353.1 tok/s 36.7% 72.5%
job_653587 13.1 1.0 535.5 tok/s 58.8% 62.0%
job_653588 11.9 1.3 488.1 tok/s 53.3% 66.6%
job_653589 10.5 0.3 431.1 tok/s 43.4% 80.4%

Utilization Analysis (Per-Engine)

Key indicators of inference engine utilization:

  • Running requests/engine: Concurrent requests being processed by each engine
  • Waiting requests: Requests queued (0 = engine not saturated, has spare capacity)
  • Generation throughput: Decode tokens/sec per engine
    • 8B model on H100 can do 1000+ tok/s when saturated
    • If seeing <300 tok/s with 0 waiting, engine is starved for requests

job_653585

  • Running requests/engine: avg=5.9, max=24
  • Waiting requests: avg=0.0, max=10
  • Generation throughput/engine: avg=256.7 tok/s, max=991.1 tok/s
  • KV cache usage: avg=25.8%
  • Prefix cache hit rate: avg=86.8%
  • Well-utilized: Engines saturated (waiting > 0)

job_653586

  • Running requests/engine: avg=8.3, max=24
  • Waiting requests: avg=0.2, max=31
  • Generation throughput/engine: avg=353.1 tok/s, max=961.0 tok/s
  • KV cache usage: avg=36.7%
  • Prefix cache hit rate: avg=72.5%
  • Well-utilized: Engines saturated (waiting > 0)

job_653587

  • Running requests/engine: avg=13.1, max=24
  • Waiting requests: avg=1.0, max=35
  • Generation throughput/engine: avg=535.5 tok/s, max=1039.7 tok/s
  • KV cache usage: avg=58.8%
  • Prefix cache hit rate: avg=62.0%
  • Well-utilized: Engines saturated (waiting > 0)

job_653588

  • Running requests/engine: avg=11.9, max=24
  • Waiting requests: avg=1.3, max=37
  • Generation throughput/engine: avg=488.1 tok/s, max=995.5 tok/s
  • KV cache usage: avg=53.3%
  • Prefix cache hit rate: avg=66.6%
  • Well-utilized: Engines saturated (waiting > 0)

job_653589

  • Running requests/engine: avg=10.5, max=24
  • Waiting requests: avg=0.3, max=17
  • Generation throughput/engine: avg=431.1 tok/s, max=943.7 tok/s
  • KV cache usage: avg=43.4%
  • Prefix cache hit rate: avg=80.4%
  • Well-utilized: Engines saturated (waiting > 0)

Trial-Level Analysis (from result.json)

Total trials parsed: 44755

Turn Count Statistics

Metric Value
Mean 4.9
Median 5.0
Std 2.3
Min 1
Max 70
Count 44394

Exception Distribution

Exception Type Count %
No exception 29160 65.2%
AgentTimeoutError 13102 29.3%
VerifierTimeoutError 1579 3.5%
ContextLengthExceededError 428 1.0%
AgentSetupTimeoutError 242 0.5%
DaytonaError 226 0.5%
DaytonaAuthenticationError 17 0.0%
Timeout 1 0.0%

Turn Count by Exception Type

Exception Type Mean Turns Median Turns Count
No exception 5.5 5.0 29160
VerifierTimeoutError 5.3 5.0 1579
Timeout 5.0 5.0 1
ContextLengthExceededError 4.8 4.0 428
AgentTimeoutError 3.5 3.0 13102
DaytonaError 3.2 2.0 108
DaytonaAuthenticationError 3.0 3.0 16

Turn Count by Outcome

Outcome Mean Turns Median Turns Count
Success 5.1 5.0 18816
Failure 4.8 4.0 22337

Reward Summary

  • Mean reward: 0.4572
  • Success rate: 45.7%
  • Trials with reward data: 41153