Files
ablation-pymethods2test-sha…/training_logs/20260612_220132_metrics_report.md
ModelHub XC 68d3be9a84 初始化项目,由ModelHub XC社区提供模型
Model: laion/ablation-pymethods2test-shaped-45-8B
Source: Original Platform
2026-07-18 17:31:01 +08:00

393 lines
21 KiB
Markdown

# SkyRL Training Metrics Analysis
Generated from 4 log files
## Overview
| Log File | Total Steps | Metric Blocks | Final Reward (mean) | Final Reward (max) | Total Time (s) |
|----------|-------------|---------------|---------------------|-------------------|----------------|
| job_628823 | 27 | 27 | 0.4447 | 0.7695 | 40599.0 |
| job_628824 | 53 | 27 | 0.4415 | 0.5918 | 41370.3 |
| job_628825 | 62 | 10 | 0.2861 | 0.4746 | 40760.4 |
| job_628826 | 80 | 18 | 0.4392 | 0.7715 | 31912.3 |
## Async Metrics
| | Mean | Std | Min | Max | Count |
|:------------------------------|-----------:|---------:|------:|------:|--------:|
| async/discard_rate | 0 | 0 | 0 | 0 | 82 |
| async/discarded_count | 0 | 0 | 0 | 0 | 82 |
| async/effective_batch_groups | 64 | 0 | 64 | 64 | 82 |
| async/effective_batch_samples | 512 | 0 | 512 | 512 | 82 |
| async/staleness_max | 3.28049 | 1.06891 | 0 | 6 | 82 |
| async/staleness_mean | 1.68046 | 0.550336 | 0 | 4 | 82 |
| async/staleness_min | 0.134146 | 0.603737 | 0 | 4 | 82 |
| async/staleness_ratio | 0.861278 | 0.19627 | 0 | 1 | 82 |
## Generate Metrics
| | Mean | Std | Min | Max | Count |
|:-------------------------------------|-----------:|---------:|---------:|---------:|--------:|
| generate/avg_num_tokens | 4140.68 | 964.063 | 1476.57 | 5649.51 | 82 |
| generate/avg_tokens_non_zero_rewards | 4359.18 | 680.881 | 2851.3 | 5970.64 | 82 |
| generate/avg_tokens_zero_rewards | 4079.29 | 1145.8 | 1023.09 | 5838.63 | 82 |
| generate/max_num_tokens | 30066.1 | 4389.61 | 13369 | 31844 | 82 |
| generate/min_num_tokens | 32.2195 | 198.741 | 1 | 1316 | 82 |
| generate/std_num_tokens | 3878.95 | 724.206 | 2062.28 | 5375.27 | 82 |
## Loss Metrics
| | Mean | Std | Min | Max | Count |
|:----------------------------|----------:|----------:|--------:|-------:|--------:|
| loss/avg_final_rewards | 0.423117 | 0.108373 | 0.1191 | 0.7715 | 82 |
| loss/avg_raw_advantages | 0.0113988 | 0.021723 | -0.0129 | 0.0952 | 82 |
| loss/avg_raw_advantages_abs | 0.104518 | 0.0323332 | 0.0106 | 0.2013 | 82 |
## Policy Metrics
| | Mean | Std | Min | Max | Count |
|:----------------------------|-------------:|------------:|--------:|-------:|--------:|
| policy/final_loss | -8.65854e-05 | 0.000207133 | -0.0009 | 0 | 82 |
| policy/log_ratio_abs_max | 0 | 0 | 0 | 0 | 82 |
| policy/log_ratio_abs_mean | 0 | 0 | 0 | 0 | 82 |
| policy/log_ratio_abs_p99 | 0 | 0 | 0 | 0 | 82 |
| policy/log_ratio_abs_pos00 | 0 | 0 | 0 | 0 | 82 |
| policy/log_ratio_abs_pos10 | 0 | 0 | 0 | 0 | 82 |
| policy/log_ratio_abs_pos20 | 0 | 0 | 0 | 0 | 82 |
| policy/log_ratio_abs_pos30 | 0 | 0 | 0 | 0 | 82 |
| policy/log_ratio_abs_pos40 | 0 | 0 | 0 | 0 | 82 |
| policy/log_ratio_abs_pos50 | 0 | 0 | 0 | 0 | 82 |
| policy/log_ratio_abs_pos60 | 0 | 0 | 0 | 0 | 82 |
| policy/log_ratio_abs_pos70 | 0 | 0 | 0 | 0 | 82 |
| policy/log_ratio_abs_pos80 | 0 | 0 | 0 | 0 | 82 |
| policy/log_ratio_abs_pos90 | 0 | 0 | 0 | 0 | 82 |
| policy/n_tokens_dp_gt_10pct | 0 | 0 | 0 | 0 | 82 |
| policy/n_tokens_dp_gt_1pct | 0 | 0 | 0 | 0 | 82 |
| policy/n_tokens_dp_gt_50pct | 0 | 0 | 0 | 0 | 82 |
| policy/policy_entropy | 0.113544 | 0.0196945 | 0.054 | 0.1366 | 82 |
| policy/policy_loss | -0.00551951 | 0.0127705 | -0.0567 | 0 | 82 |
| policy/policy_lr | 0 | 0 | 0 | 0 | 82 |
| policy/policy_update_steps | 1 | 0 | 1 | 1 | 82 |
| policy/ppo_clip_ratio | 0 | 0 | 0 | 0 | 82 |
| policy/raw_grad_norm | 0.0170537 | 0.00539994 | 0.008 | 0.0536 | 82 |
## Reward Metrics
| | Mean | Std | Min | Max | Count |
|:----------------------|---------:|----------:|-------:|-------:|--------:|
| reward/avg_pass_at_8 | 0.541405 | 0.0905809 | 0.3281 | 0.8281 | 82 |
| reward/avg_raw_reward | 0.423117 | 0.108373 | 0.1191 | 0.7715 | 82 |
## System Metrics
| | Mean | Std | Min | Max | Count |
|:------------------------|---------:|-------------:|---------:|---------:|--------:|
| system/process_rss_gb | 14.4693 | 1.40257 | 9.7812 | 17.1838 | 82 |
| system/process_vms_gb | 51.5833 | 0.759579 | 50.2029 | 53.3813 | 82 |
| system/ram_available_gb | 568.401 | 23.2428 | 536.837 | 626.327 | 82 |
| system/ram_percent | 33.7439 | 2.70948 | 27 | 37.4 | 82 |
| system/ram_total_gb | 857.968 | 1.14386e-13 | 857.968 | 857.968 | 82 |
| system/ram_used_gb | 289.566 | 23.2428 | 231.64 | 321.13 | 82 |
## Timing Metrics
| | Mean | Std | Min | Max | Count |
|:--------------------------------------|-------------:|-------------:|---------:|----------:|--------:|
| timing/compute_advantages_and_returns | 0.0824878 | 0.0302697 | 0.0389 | 0.2346 | 82 |
| timing/convert_to_training_input | 4.44468 | 0.639779 | 1.9872 | 5.7156 | 82 |
| timing/fwd_logprobs_values_reward | 31.1934 | 14.3675 | 13.9888 | 148.12 | 82 |
| timing/policy_train | 178.7 | 35.655 | 89.1352 | 249.186 | 82 |
| timing/run_training | 210.243 | 42.58 | 103.356 | 306.846 | 82 |
| timing/step | 1885.88 | 1071.01 | 710.547 | 7430.96 | 82 |
| timing/sync_weights | 22.7287 | 2.34342 | 17.454 | 28.164 | 82 |
| timing/train_critic_and_policy | 178.967 | 35.6753 | 89.3191 | 249.423 | 82 |
| timing/wait_for_generation_buffer | 1648.46 | 1089.83 | 448.748 | 7102.07 | 82 |
| timing/cleanup_old_checkpoints | 20.7478 | 124.52 | 0.0077 | 788.255 | 40 |
| timing/save_checkpoints | 35.0582 | 136.626 | 7.892 | 873.951 | 40 |
| timing/save_hf_model | 8.11335 | 6.74447 | 5.7822 | 33.3153 | 16 |
## Trainer Metrics
| | Mean | Std | Min | Max | Count |
|:--------------------|-----------:|----------:|------:|------:|--------:|
| trainer/epoch | 0.0243902 | 0.155207 | 0 | 1 | 82 |
| trainer/global_step | 40.4878 | 23.0401 | 1 | 80 | 82 |
## Batch_Errors Metrics
| | Mean | Std | Min | Max | Count |
|:----------------------------------------------|------------:|--------------:|------------:|------------:|--------:|
| batch_errors/total_batches | 61.378 | 10.9518 | 16 | 122 | 82 |
| batch_errors/total_instances | 491.024 | 87.6143 | 128 | 976 | 82 |
| batch_errors/total_successful | 414.378 | 106.966 | 85 | 943 | 82 |
| batch_errors/total_failed | 22.8902 | 12.5502 | 10 | 59 | 82 |
| batch_errors/total_masked | 65.6098 | 48.9638 | 13 | 239 | 82 |
| batch_errors/avg_VerifierTimeoutError | 0.212868 | 0.287732 | 0.0153846 | 1.1129 | 62 |
| batch_errors/total_VerifierTimeoutError | 11.9194 | 16.382 | 1 | 69 | 62 |
| batch_errors/avg_RuntimeError | 0.0780424 | 0.0380127 | 0.015625 | 0.140351 | 9 |
| batch_errors/total_RuntimeError | 5 | 2 | 1 | 8 | 9 |
| batch_errors/avg_ContextLengthExceededError | 0.102943 | 0.0715027 | 0.015625 | 0.403509 | 80 |
| batch_errors/total_ContextLengthExceededError | 6.325 | 4.32472 | 1 | 23 | 80 |
| batch_errors/avg_InvalidChatHistory | 0.709029 | 0.397476 | 0.0793651 | 1.80702 | 82 |
| batch_errors/total_InvalidChatHistory | 43 | 24.6957 | 5 | 103 | 82 |
| batch_errors/avg_AgentTimeoutError | 0.404978 | 0.812804 | 0.015625 | 3.06452 | 70 |
| batch_errors/total_AgentTimeoutError | 23.1429 | 49.3504 | 1 | 190 | 70 |
| batch_errors/avg_AgentSetupTimeoutError | 0.0897322 | 0.129154 | 0.015625 | 0.507246 | 29 |
| batch_errors/total_AgentSetupTimeoutError | 4.65517 | 7.22731 | 1 | 35 | 29 |
| batch_errors/avg_DaytonaAuthenticationError | 0.0464813 | 0.0260413 | 0.016129 | 0.101695 | 12 |
| batch_errors/total_DaytonaAuthenticationError | 2.33333 | 1.66969 | 1 | 6 | 12 |
| batch_errors/avg_DaytonaNotFoundError | 0.016129 | nan | 0.016129 | 0.016129 | 1 |
| batch_errors/total_DaytonaNotFoundError | 1 | nan | 1 | 1 | 1 |
| batch_errors/avg_DaytonaError | 0.0164326 | 0.000994804 | 0.015625 | 0.0175439 | 3 |
| batch_errors/total_DaytonaError | 1 | 0 | 1 | 1 | 3 |
## Training Progression by Log
### job_628823
| Step | Reward | Pass@8 | KL | Loss | Step Time (s) | Gen Wait (s) |
|------|--------|--------|-----|------|---------------|-------------|
| 1 | 0.7695 | 0.8281 | 0.000000 | 0.0000 | 3057.6 | 2847.4 |
| 2 | 0.4316 | 0.5781 | 0.000000 | -0.0001 | 1974.6 | 1674.5 |
| 3 | 0.3223 | 0.4375 | 0.000000 | -0.0000 | 764.6 | 448.7 |
| 4 | 0.4414 | 0.5625 | 0.000000 | -0.0000 | 1688.5 | 1436.7 |
| 5 | 0.3984 | 0.5000 | 0.000000 | -0.0000 | 1372.0 | 1100.7 |
| 6 | 0.4727 | 0.5781 | 0.000000 | 0.0000 | 1345.3 | 1049.0 |
| 7 | 0.5352 | 0.6250 | 0.000000 | -0.0000 | 1358.4 | 1111.0 |
| 8 | 0.4551 | 0.5625 | 0.000000 | -0.0000 | 1520.6 | 1224.2 |
| 9 | 0.3867 | 0.4844 | 0.000000 | -0.0000 | 1325.5 | 1083.5 |
| 10 | 0.4590 | 0.5625 | 0.000000 | -0.0000 | 1527.1 | 1243.8 |
| 11 | 0.4961 | 0.6875 | 0.000000 | -0.0000 | 1372.1 | 1128.0 |
| 12 | 0.4102 | 0.5000 | 0.000000 | 0.0000 | 1553.8 | 1251.6 |
| 13 | 0.4023 | 0.5781 | 0.000000 | -0.0000 | 1401.1 | 1096.5 |
| 14 | 0.4160 | 0.5625 | 0.000000 | -0.0000 | 1386.3 | 1113.5 |
| 15 | 0.4844 | 0.5625 | 0.000000 | 0.0000 | 1415.7 | 1160.7 |
| 16 | 0.4512 | 0.5938 | 0.000000 | 0.0000 | 1435.3 | 1156.5 |
| 17 | 0.3633 | 0.5156 | 0.000000 | 0.0000 | 1238.7 | 930.7 |
| 18 | 0.4941 | 0.6094 | 0.000000 | 0.0000 | 1604.5 | 1338.2 |
| 19 | 0.3633 | 0.4844 | 0.000000 | -0.0000 | 1704.6 | 1415.3 |
| 20 | 0.3359 | 0.4219 | 0.000000 | -0.0000 | 1244.0 | 998.0 |
| 21 | 0.4414 | 0.5781 | 0.000000 | -0.0000 | 1850.9 | 1594.7 |
| 22 | 0.4355 | 0.5156 | 0.000000 | -0.0000 | 1397.3 | 1080.5 |
| 23 | 0.3359 | 0.5000 | 0.000000 | -0.0000 | 1504.8 | 1261.4 |
| 24 | 0.4746 | 0.5625 | 0.000000 | -0.0000 | 1700.5 | 1460.9 |
| 25 | 0.3633 | 0.5000 | 0.000000 | -0.0001 | 1624.9 | 1359.2 |
| 26 | 0.5430 | 0.6719 | 0.000000 | -0.0001 | 995.4 | 773.4 |
| 27 | 0.5254 | 0.6406 | 0.000000 | -0.0001 | 1234.8 | 993.9 |
### job_628824
| Step | Reward | Pass@8 | KL | Loss | Step Time (s) | Gen Wait (s) |
|------|--------|--------|-----|------|---------------|-------------|
| 27 | 0.5879 | 0.6250 | 0.000000 | -0.0001 | 4078.7 | 3842.9 |
| 28 | 0.3770 | 0.5000 | 0.000000 | -0.0001 | 1263.0 | 1008.9 |
| 29 | 0.4277 | 0.5469 | 0.000000 | 0.0000 | 1068.8 | 792.5 |
| 30 | 0.4375 | 0.5156 | 0.000000 | -0.0001 | 1539.5 | 1315.8 |
| 31 | 0.4531 | 0.5312 | 0.000000 | 0.0000 | 1535.1 | 1274.5 |
| 32 | 0.3320 | 0.4844 | 0.000000 | -0.0001 | 1701.1 | 1466.6 |
| 33 | 0.4023 | 0.4844 | 0.000000 | 0.0000 | 1470.0 | 1220.7 |
| 34 | 0.4180 | 0.4844 | 0.000000 | -0.0000 | 1363.7 | 1119.2 |
| 35 | 0.4316 | 0.5781 | 0.000000 | 0.0000 | 1421.4 | 1186.7 |
| 36 | 0.4141 | 0.5469 | 0.000000 | -0.0000 | 1687.0 | 1439.7 |
| 37 | 0.4512 | 0.5469 | 0.000000 | -0.0000 | 1634.9 | 1394.5 |
| 38 | 0.5918 | 0.7344 | 0.000000 | 0.0000 | 1204.9 | 971.6 |
| 39 | 0.4766 | 0.5469 | 0.000000 | 0.0000 | 1350.2 | 1096.7 |
| 40 | 0.4180 | 0.5625 | 0.000000 | -0.0000 | 1518.6 | 1271.4 |
| 41 | 0.3906 | 0.5156 | 0.000000 | -0.0000 | 1695.8 | 1427.1 |
| 42 | 0.3965 | 0.5469 | 0.000000 | 0.0000 | 1186.0 | 926.3 |
| 43 | 0.4590 | 0.5312 | 0.000000 | -0.0000 | 1288.4 | 1028.0 |
| 44 | 0.5645 | 0.6406 | 0.000000 | -0.0000 | 1113.2 | 888.1 |
| 45 | 0.4570 | 0.5625 | 0.000000 | 0.0000 | 1357.3 | 1111.4 |
| 46 | 0.4883 | 0.5625 | 0.000000 | 0.0000 | 1268.2 | 1044.4 |
| 47 | 0.4434 | 0.5781 | 0.000000 | 0.0000 | 1545.0 | 1273.0 |
| 48 | 0.4355 | 0.5938 | 0.000000 | -0.0000 | 1623.9 | 1389.0 |
| 49 | 0.4336 | 0.5156 | 0.000000 | 0.0000 | 1295.4 | 1047.7 |
| 50 | 0.4141 | 0.5000 | 0.000000 | -0.0000 | 1512.6 | 1282.3 |
| 51 | 0.4121 | 0.5312 | 0.000000 | -0.0000 | 1530.0 | 1323.1 |
| 52 | 0.3945 | 0.5312 | 0.000000 | -0.0000 | 1451.4 | 1226.7 |
| 53 | 0.4121 | 0.4844 | 0.000000 | 0.0000 | 1666.3 | 1423.3 |
### job_628825
| Step | Reward | Pass@8 | KL | Loss | Step Time (s) | Gen Wait (s) |
|------|--------|--------|-----|------|---------------|-------------|
| 53 | 0.4746 | 0.5818 | 0.000000 | -0.0005 | 7431.0 | 7102.1 |
| 54 | 0.1191 | 0.3281 | 0.000000 | -0.0002 | 2972.7 | 2812.8 |
| 55 | 0.3262 | 0.5156 | 0.000000 | -0.0009 | 4313.7 | 4135.8 |
| 56 | 0.1934 | 0.3906 | 0.000000 | -0.0005 | 2256.5 | 2086.4 |
| 57 | 0.2754 | 0.5000 | 0.000000 | -0.0007 | 4721.0 | 4568.8 |
| 58 | 0.2930 | 0.5000 | 0.000000 | -0.0008 | 4293.2 | 4168.3 |
| 59 | 0.2480 | 0.4375 | 0.000000 | -0.0005 | 3191.9 | 3068.6 |
| 60 | 0.3379 | 0.4688 | 0.000000 | -0.0002 | 3332.0 | 3174.5 |
| 61 | 0.2676 | 0.4375 | 0.000000 | -0.0006 | 4540.5 | 4396.9 |
| 62 | 0.3262 | 0.5781 | 0.000000 | -0.0009 | 3708.0 | 3576.0 |
### job_628826
| Step | Reward | Pass@8 | KL | Loss | Step Time (s) | Gen Wait (s) |
|------|--------|--------|-----|------|---------------|-------------|
| 63 | 0.7168 | 0.8136 | 0.000000 | -0.0002 | 2578.9 | 2397.1 |
| 64 | 0.6172 | 0.6562 | 0.000000 | -0.0000 | 2364.1 | 2140.7 |
| 65 | 0.2656 | 0.4062 | 0.000000 | 0.0000 | 710.5 | 470.5 |
| 66 | 0.5352 | 0.6875 | 0.000000 | -0.0000 | 1419.1 | 1185.4 |
| 67 | 0.3438 | 0.4219 | 0.000000 | -0.0001 | 2122.2 | 1843.5 |
| 68 | 0.3828 | 0.5000 | 0.000000 | -0.0000 | 1248.7 | 1014.2 |
| 69 | 0.4609 | 0.5469 | 0.000000 | -0.0000 | 1332.3 | 1118.8 |
| 70 | 0.4473 | 0.5469 | 0.000000 | 0.0000 | 1573.6 | 1352.4 |
| 71 | 0.4707 | 0.6250 | 0.000000 | -0.0000 | 1325.1 | 1115.1 |
| 72 | 0.3750 | 0.4688 | 0.000000 | 0.0000 | 1530.9 | 1311.4 |
| 73 | 0.4512 | 0.5312 | 0.000000 | 0.0000 | 1586.4 | 1350.9 |
| 74 | 0.4180 | 0.5469 | 0.000000 | 0.0000 | 1485.3 | 1225.2 |
| 75 | 0.4512 | 0.5156 | 0.000000 | -0.0001 | 1332.1 | 1131.4 |
| 76 | 0.3008 | 0.4062 | 0.000000 | 0.0000 | 1875.4 | 1654.7 |
| 77 | 0.3457 | 0.4219 | 0.000000 | -0.0000 | 1368.0 | 1162.8 |
| 78 | 0.2754 | 0.3594 | 0.000000 | -0.0000 | 1663.1 | 1456.8 |
| 79 | 0.7715 | 0.7812 | 0.000000 | 0.0000 | 3065.3 | 2902.5 |
| 80 | 0.2773 | 0.4375 | 0.000000 | -0.0001 | 3331.0 | 3125.1 |
## Timing Analysis
### Average Time Breakdown (% of step time)
| Component | Avg % of Step Time |
|-----------|-------------------|
| wait_for_generation_buffer | 84.5% |
| run_training | 13.7% |
| train_critic_and_policy | 11.8% |
| policy_train | 11.7% |
| fwd_logprobs_values_reward | 2.0% |
| sync_weights | 1.5% |
| save_checkpoints | 1.5% |
| cleanup_old_checkpoints | 0.7% |
| save_hf_model | 0.5% |
| convert_to_training_input | 0.3% |
| compute_advantages_and_returns | 0.0% |
## Cross-Log Comparison
| Log | Avg Reward | Pass@8 | Step Time (s) | Gen Wait Time (s) | Avg Tokens | Staleness |
|-----|------|------|------|------|------|------|
| job_628823 | 0.4447 | 0.5631 | 1503.6669 | 1234.5434 | 4883.9745 | 1.7604 |
| job_628824 | 0.4415 | 0.5475 | 1532.2337 | 1288.5914 | 4351.9871 | 1.7622 |
| job_628825 | 0.2861 | 0.4738 | 4076.0445 | 3909.0126 | 2019.2750 | 1.3391 |
| job_628826 | 0.4392 | 0.5374 | 1772.9040 | 1553.2566 | 3887.3310 | 1.6276 |
## vLLM Inference Engine Analysis
Metrics from vLLM stat loggers (V1LoggingStatLoggerFixed).
> **Note**: Ray deduplicates similar log messages with `[repeated Nx across cluster]`,
> so we typically capture stats from one engine per timestamp. The stats shown are
> **per-engine** values. Multiply by num_inference_engines for cluster-wide estimates.
### Summary by Log (Per-Engine Stats)
| Log | Avg Running/Engine | Avg Waiting/Engine | Avg Gen Throughput/Engine | Avg KV Cache % | Avg Prefix Hit % |
|-----|-------------------|-------------------|--------------------------|----------------|------------------|
| job_628823 | 12.0 | 0.6 | 514.1 tok/s | 53.3% | 67.4% |
| job_628824 | 12.4 | 0.5 | 534.0 tok/s | 54.1% | 70.2% |
| job_628825 | 4.4 | 0.0 | 199.9 tok/s | 19.2% | 86.4% |
| job_628826 | 11.8 | 1.2 | 499.9 tok/s | 52.8% | 66.7% |
### Utilization Analysis (Per-Engine)
Key indicators of inference engine utilization:
- **Running requests/engine**: Concurrent requests being processed by each engine
- **Waiting requests**: Requests queued (0 = engine not saturated, has spare capacity)
- **Generation throughput**: Decode tokens/sec per engine
- 8B model on H100 can do **1000+ tok/s** when saturated
- If seeing <300 tok/s with 0 waiting, engine is **starved for requests**
#### job_628823
- **Running requests/engine**: avg=12.0, max=24
- **Waiting requests**: avg=0.6, max=28
- **Generation throughput/engine**: avg=514.1 tok/s, max=952.0 tok/s
- **KV cache usage**: avg=53.3%
- **Prefix cache hit rate**: avg=67.4%
- **Well-utilized**: Engines saturated (waiting > 0)
#### job_628824
- **Running requests/engine**: avg=12.4, max=24
- **Waiting requests**: avg=0.5, max=44
- **Generation throughput/engine**: avg=534.0 tok/s, max=977.8 tok/s
- **KV cache usage**: avg=54.1%
- **Prefix cache hit rate**: avg=70.2%
-**Well-utilized**: Engines saturated (waiting > 0)
#### job_628825
- **Running requests/engine**: avg=4.4, max=17
- **Waiting requests**: avg=0.0, max=0
- **Generation throughput/engine**: avg=199.9 tok/s, max=712.4 tok/s
- **KV cache usage**: avg=19.2%
- **Prefix cache hit rate**: avg=86.4%
- ⚠️ **Underutilized**: Engines starved for requests (0 waiting, avg 4.4 running)
- Bottleneck is likely upstream (environment execution, not inference)
#### job_628826
- **Running requests/engine**: avg=11.8, max=24
- **Waiting requests**: avg=1.2, max=52
- **Generation throughput/engine**: avg=499.9 tok/s, max=986.9 tok/s
- **KV cache usage**: avg=52.8%
- **Prefix cache hit rate**: avg=66.7%
-**Well-utilized**: Engines saturated (waiting > 0)
## Trial-Level Analysis (from result.json)
Total trials parsed: 44409
### Turn Count Statistics
| Metric | Value |
|--------|-------|
| Mean | 4.7 |
| Median | 4.0 |
| Std | 2.3 |
| Min | 1 |
| Max | 30 |
| Count | 44161 |
### Exception Distribution
| Exception Type | Count | % |
|---------------|-------|---|
| No exception | 29002 | 65.3% |
| AgentTimeoutError | 13737 | 30.9% |
| VerifierTimeoutError | 880 | 2.0% |
| ContextLengthExceededError | 511 | 1.2% |
| AgentSetupTimeoutError | 202 | 0.5% |
| RuntimeError | 45 | 0.1% |
| DaytonaAuthenticationError | 28 | 0.1% |
| DaytonaError | 3 | 0.0% |
| DaytonaNotFoundError | 1 | 0.0% |
### Turn Count by Exception Type
| Exception Type | Mean Turns | Median Turns | Count |
|---------------|-----------|-------------|-------|
| DaytonaError | 6.5 | 6.5 | 2 |
| No exception | 5.4 | 5.0 | 29002 |
| VerifierTimeoutError | 5.2 | 5.0 | 880 |
| ContextLengthExceededError | 4.0 | 4.0 | 511 |
| AgentTimeoutError | 3.1 | 3.0 | 13737 |
| DaytonaAuthenticationError | 2.6 | 2.5 | 28 |
| DaytonaNotFoundError | 2.0 | 2.0 | 1 |
### Turn Count by Outcome
| Outcome | Mean Turns | Median Turns | Count |
|---------|-----------|-------------|-------|
| Success | 5.1 | 5.0 | 18561 |
| Failure | 4.4 | 4.0 | 22465 |
### Reward Summary
- Mean reward: 0.4524
- Success rate: 45.2%
- Trials with reward data: 41026