# SkyRL Training Metrics Analysis Generated from 6 log files ## Overview | Log File | Total Steps | Metric Blocks | Final Reward (mean) | Final Reward (max) | Total Time (s) | |----------|-------------|---------------|---------------------|-------------------|----------------| | job_485099 | 5 | 5 | 0.0621 | 0.1152 | 23691.2 | | job_485100 | 13 | 9 | 0.0610 | 0.1465 | 39612.5 | | job_485101 | 19 | 7 | 0.0948 | 0.1230 | 31256.0 | | job_496420 | 26 | 8 | 0.0935 | 0.2188 | 40891.0 | | job_496421 | 36 | 10 | 0.1310 | 0.2422 | 42300.2 | | job_496422 | 45 | 9 | 0.1430 | 0.2285 | 42232.9 | ## Async Metrics | | Mean | Std | Min | Max | Count | |:------------------------------|-----------:|---------:|------:|---------:|--------:| | async/discard_rate | 0 | 0 | 0 | 0 | 48 | | async/discarded_count | 0 | 0 | 0 | 0 | 48 | | async/effective_batch_groups | 64 | 0 | 64 | 64 | 48 | | async/effective_batch_samples | 512 | 0 | 512 | 512 | 48 | | async/staleness_max | 1.91667 | 0.794485 | 0 | 4 | 48 | | async/staleness_mean | 0.861646 | 0.370024 | 0 | 1.5625 | 48 | | async/staleness_min | 0.208333 | 0.410414 | 0 | 1 | 48 | | async/staleness_ratio | 0.69629 | 0.320514 | 0 | 1 | 48 | ## Generate Metrics | | Mean | Std | Min | Max | Count | |:-------------------------------------|---------:|---------:|---------:|----------:|--------:| | generate/avg_num_tokens | 300.937 | 132.434 | 53.543 | 653.519 | 48 | | generate/avg_tokens_non_zero_rewards | 3164.85 | 608.453 | 2069.75 | 4650.14 | 48 | | generate/avg_tokens_zero_rewards | 1 | 0 | 1 | 1 | 48 | | generate/max_num_tokens | 9563.94 | 4742.53 | 4538 | 31001 | 48 | | generate/min_num_tokens | 1 | 0 | 1 | 1 | 48 | | generate/std_num_tokens | 1025.02 | 261.051 | 409.607 | 1779.2 | 48 | ## Loss Metrics | | Mean | Std | Min | Max | Count | |:----------------------------|---------:|----------:|-------:|-------:|--------:| | loss/avg_final_rewards | 0.101433 | 0.0552715 | 0.0137 | 0.2422 | 48 | | loss/avg_raw_advantages | 0.223456 | 0.116288 | 0.0251 | 0.4541 | 48 | | loss/avg_raw_advantages_abs | 0.223596 | 0.116357 | 0.0252 | 0.4545 | 48 | ## Policy Metrics | | Mean | Std | Min | Max | Count | |:----------------------------|-------------:|------------:|--------:|--------:|--------:| | policy/final_loss | -0.000310417 | 0.000183627 | -0.0008 | -0 | 48 | | policy/log_ratio_abs_max | 0 | 0 | 0 | 0 | 48 | | policy/log_ratio_abs_mean | 0 | 0 | 0 | 0 | 48 | | policy/log_ratio_abs_p99 | 0 | 0 | 0 | 0 | 48 | | policy/log_ratio_abs_pos00 | 0 | 0 | 0 | 0 | 48 | | policy/log_ratio_abs_pos10 | 0 | 0 | 0 | 0 | 48 | | policy/log_ratio_abs_pos20 | 0 | 0 | 0 | 0 | 48 | | policy/log_ratio_abs_pos30 | 0 | 0 | 0 | 0 | 48 | | policy/log_ratio_abs_pos40 | 0 | 0 | 0 | 0 | 48 | | policy/log_ratio_abs_pos50 | 0 | 0 | 0 | 0 | 48 | | policy/log_ratio_abs_pos60 | 0 | 0 | 0 | 0 | 48 | | policy/log_ratio_abs_pos70 | 0 | 0 | 0 | 0 | 48 | | policy/log_ratio_abs_pos80 | 0 | 0 | 0 | 0 | 48 | | policy/log_ratio_abs_pos90 | 0 | 0 | 0 | 0 | 48 | | policy/n_tokens_dp_gt_10pct | 0 | 0 | 0 | 0 | 48 | | policy/n_tokens_dp_gt_1pct | 0 | 0 | 0 | 0 | 48 | | policy/n_tokens_dp_gt_50pct | 0 | 0 | 0 | 0 | 48 | | policy/policy_entropy | 0.00732917 | 0.0036161 | 0.0018 | 0.0182 | 48 | | policy/policy_loss | -0.0198583 | 0.0114544 | -0.0498 | -0.002 | 48 | | policy/policy_lr | 0 | 0 | 0 | 0 | 48 | | policy/policy_update_steps | 1 | 0 | 1 | 1 | 48 | | policy/ppo_clip_ratio | 0 | 0 | 0 | 0 | 48 | | policy/raw_grad_norm | 0.0109021 | 0.00505137 | 0.0028 | 0.0244 | 48 | ## Reward Metrics | | Mean | Std | Min | Max | Count | |:----------------------|---------:|----------:|-------:|-------:|--------:| | reward/avg_pass_at_8 | 0.261562 | 0.0939986 | 0.0938 | 0.4898 | 48 | | reward/avg_raw_reward | 0.101433 | 0.0552715 | 0.0137 | 0.2422 | 48 | ## System Metrics | | Mean | Std | Min | Max | Count | |:------------------------|---------:|------------:|---------:|---------:|--------:| | system/process_rss_gb | 9.5798 | 3.43573 | 5.3781 | 17.2552 | 48 | | system/process_vms_gb | 81.1127 | 5.44727 | 78.0531 | 94.4252 | 48 | | system/ram_available_gb | 606.245 | 15.2876 | 574.636 | 627.878 | 48 | | system/ram_percent | 29.3396 | 1.78847 | 26.8 | 33 | 48 | | system/ram_total_gb | 857.969 | 0.00033958 | 857.968 | 857.969 | 48 | | system/ram_used_gb | 251.724 | 15.2877 | 230.09 | 283.333 | 48 | ## Timing Metrics | | Mean | Std | Min | Max | Count | |:--------------------------------------|-------------:|-------------:|---------:|----------:|--------:| | timing/compute_advantages_and_returns | 0.0384646 | 0.0348721 | 0.0175 | 0.1923 | 48 | | timing/convert_to_training_input | 5.4315 | 26.8172 | 0.8755 | 187.298 | 48 | | timing/fwd_logprobs_values_reward | 7.50122 | 2.76571 | 3.0976 | 15.4184 | 48 | | timing/policy_train | 59.4245 | 7.15104 | 44.4851 | 76.4064 | 48 | | timing/run_training | 67.108 | 9.54666 | 48.0435 | 90.3725 | 48 | | timing/step | 4583 | 2394.44 | 351.854 | 8192.84 | 48 | | timing/sync_weights | 18.6968 | 1.59914 | 17.1829 | 28.5894 | 48 | | timing/train_critic_and_policy | 59.5679 | 7.15635 | 44.6233 | 76.5655 | 48 | | timing/wait_for_generation_buffer | 4491.76 | 2392.46 | 271.614 | 8105.89 | 48 | | timing/cleanup_old_checkpoints | 4.13983 | 7.40386 | 0.0075 | 24.0683 | 22 | | timing/save_checkpoints | 15.0911 | 11.7658 | 7.3633 | 44.8581 | 22 | | timing/save_hf_model | 6.85727 | 1.37645 | 5.8497 | 10.5996 | 10 | ## Trainer Metrics | | Mean | Std | Min | Max | Count | |:--------------------|--------:|--------:|------:|------:|--------:| | trainer/epoch | 0 | 0 | 0 | 0 | 48 | | trainer/global_step | 22.3333 | 13.0536 | 1 | 45 | 48 | ## Batch_Errors Metrics | | Mean | Std | Min | Max | Count | |:------------------------------------------------|------------:|------------:|-----------:|-------------:|--------:| | batch_errors/total_batches | 59.6875 | 18.9346 | 2 | 128 | 48 | | batch_errors/total_instances | 477.5 | 151.476 | 16 | 1024 | 48 | | batch_errors/total_successful | 43.8333 | 25.6634 | 0 | 117 | 48 | | batch_errors/total_failed | 59.0208 | 18.8617 | 2 | 128 | 48 | | batch_errors/total_masked | 357.688 | 113.847 | 15 | 732 | 48 | | batch_errors/avg_DaytonaAuthenticationError | 1.00014 | 0.399696 | 0.28125 | 1.76562 | 48 | | batch_errors/total_DaytonaAuthenticationError | 62.2708 | 36.1967 | 1 | 208 | 48 | | batch_errors/avg_AgentEnvironmentTimeoutError | 6.00745 | 0.573894 | 4.82812 | 7.5 | 48 | | batch_errors/total_AgentEnvironmentTimeoutError | 355.104 | 112.817 | 15 | 725 | 48 | | batch_errors/avg_DaytonaNotFoundError | 0.0269537 | 0.0189922 | 0.0144928 | 0.09375 | 32 | | batch_errors/total_DaytonaNotFoundError | 1.75 | 1.34404 | 1 | 6 | 32 | | batch_errors/avg_VerifierRuntimeError | 0.20424 | 0.127208 | 0.03125 | 0.5625 | 46 | | batch_errors/total_VerifierRuntimeError | 12.5217 | 8.05188 | 2 | 36 | 46 | | batch_errors/avg_DaytonaError | 0.0224232 | 0.0110744 | 0.0147059 | 0.046875 | 17 | | batch_errors/total_DaytonaError | 1.52941 | 0.799816 | 1 | 3 | 17 | | batch_errors/avg_RuntimeError | 0.015625 | nan | 0.015625 | 0.015625 | 1 | | batch_errors/total_RuntimeError | 2 | nan | 2 | 2 | 1 | | batch_errors/avg_DaytonaValidationError | 0.0078125 | nan | 0.0078125 | 0.0078125 | 1 | | batch_errors/total_DaytonaValidationError | 1 | nan | 1 | 1 | 1 | | batch_errors/avg_ContextLengthExceededError | 0.0284136 | 0.0238451 | 0.0078125 | 0.1 | 13 | | batch_errors/total_ContextLengthExceededError | 1.61538 | 0.767948 | 1 | 3 | 13 | | batch_errors/avg_AgentTimeoutError | 0.0445227 | 0.0358926 | 0.015625 | 0.1875 | 27 | | batch_errors/total_AgentTimeoutError | 2.81481 | 2.30446 | 1 | 12 | 27 | | batch_errors/avg_AgentSetupTimeoutError | 0.0386788 | 0.0257621 | 0.0147059 | 0.078125 | 6 | | batch_errors/total_AgentSetupTimeoutError | 2.33333 | 1.75119 | 1 | 5 | 6 | | batch_errors/avg_EnvironmentStartTimeoutError | 0.0153846 | nan | 0.0153846 | 0.0153846 | 1 | | batch_errors/total_EnvironmentStartTimeoutError | 1 | nan | 1 | 1 | 1 | | batch_errors/avg_VerifierTimeoutError | 0.0400638 | 0.0229369 | 0.0144928 | 0.0588235 | 3 | | batch_errors/total_VerifierTimeoutError | 2.66667 | 1.52753 | 1 | 4 | 3 | | batch_errors/avg_OSError | 0.0144928 | nan | 0.0144928 | 0.0144928 | 1 | | batch_errors/total_OSError | 1 | nan | 1 | 1 | 1 | ## Training Progression by Log ### job_485099 | Step | Reward | Pass@8 | KL | Loss | Step Time (s) | Gen Wait (s) | |------|--------|--------|-----|------|---------------|-------------| | 1 | 0.0410 | 0.1719 | 0.000000 | -0.0003 | 8192.8 | 8105.9 | | 2 | 0.0664 | 0.2344 | 0.000000 | -0.0002 | 7137.7 | 7051.2 | | 3 | 0.0742 | 0.2656 | 0.000000 | -0.0001 | 614.4 | 526.3 | | 4 | 0.0137 | 0.0938 | 0.000000 | -0.0000 | 7011.6 | 6938.8 | | 5 | 0.1152 | 0.2344 | 0.000000 | -0.0001 | 734.7 | 652.9 | ### job_485100 | Step | Reward | Pass@8 | KL | Loss | Step Time (s) | Gen Wait (s) | |------|--------|--------|-----|------|---------------|-------------| | 5 | 0.1465 | 0.3276 | 0.000000 | -0.0005 | 8132.7 | 8036.1 | | 6 | 0.0332 | 0.1562 | 0.000000 | -0.0003 | 3465.1 | 3390.7 | | 7 | 0.0234 | 0.1406 | 0.000000 | -0.0002 | 3987.0 | 3915.0 | | 8 | 0.0176 | 0.0938 | 0.000000 | -0.0001 | 6656.3 | 6578.8 | | 9 | 0.0234 | 0.1250 | 0.000000 | -0.0001 | 1287.2 | 1213.6 | | 10 | 0.0820 | 0.2344 | 0.000000 | -0.0005 | 6403.6 | 6309.4 | | 11 | 0.0781 | 0.1719 | 0.000000 | -0.0002 | 1527.3 | 1417.6 | | 12 | 0.0801 | 0.2500 | 0.000000 | -0.0004 | 6140.9 | 6053.0 | | 13 | 0.0645 | 0.2344 | 0.000000 | -0.0002 | 2012.4 | 1931.4 | ### job_485101 | Step | Reward | Pass@8 | KL | Loss | Step Time (s) | Gen Wait (s) | |------|--------|--------|-----|------|---------------|-------------| | 13 | 0.1230 | 0.2742 | 0.000000 | -0.0004 | 8113.5 | 8023.7 | | 14 | 0.0371 | 0.1250 | 0.000000 | -0.0002 | 6886.3 | 6818.0 | | 15 | 0.1191 | 0.3750 | 0.000000 | -0.0005 | 791.9 | 701.5 | | 16 | 0.1191 | 0.3125 | 0.000000 | -0.0003 | 6575.8 | 6487.1 | | 17 | 0.0918 | 0.2500 | 0.000000 | -0.0005 | 920.2 | 834.8 | | 18 | 0.0703 | 0.2344 | 0.000000 | -0.0002 | 6610.6 | 6533.4 | | 19 | 0.1035 | 0.2812 | 0.000000 | -0.0001 | 1357.7 | 1273.3 | ### job_496420 | Step | Reward | Pass@8 | KL | Loss | Step Time (s) | Gen Wait (s) | |------|--------|--------|-----|------|---------------|-------------| | 19 | 0.2188 | 0.4407 | 0.000000 | -0.0007 | 8067.2 | 7961.5 | | 20 | 0.0801 | 0.2031 | 0.000000 | -0.0003 | 3380.6 | 3297.4 | | 21 | 0.0996 | 0.3281 | 0.000000 | -0.0005 | 4118.8 | 4027.4 | | 22 | 0.0703 | 0.2500 | 0.000000 | -0.0004 | 5101.1 | 5011.8 | | 23 | 0.0938 | 0.2656 | 0.000000 | -0.0003 | 2467.0 | 2383.2 | | 24 | 0.0781 | 0.2344 | 0.000000 | -0.0001 | 6306.5 | 6217.1 | | 25 | 0.0547 | 0.2188 | 0.000000 | -0.0002 | 4522.1 | 4243.7 | | 26 | 0.0527 | 0.1875 | 0.000000 | -0.0000 | 6927.7 | 6847.7 | ### job_496421 | Step | Reward | Pass@8 | KL | Loss | Step Time (s) | Gen Wait (s) | |------|--------|--------|-----|------|---------------|-------------| | 27 | 0.2422 | 0.4898 | 0.000000 | -0.0007 | 8113.5 | 8003.3 | | 28 | 0.0586 | 0.2500 | 0.000000 | -0.0005 | 351.9 | 271.6 | | 29 | 0.1211 | 0.2656 | 0.000000 | -0.0004 | 7096.6 | 7009.6 | | 30 | 0.0527 | 0.1875 | 0.000000 | -0.0001 | 2113.7 | 2031.7 | | 31 | 0.1680 | 0.3750 | 0.000000 | -0.0004 | 5328.9 | 5229.5 | | 32 | 0.1035 | 0.2031 | 0.000000 | -0.0002 | 3341.1 | 3249.0 | | 33 | 0.1777 | 0.3438 | 0.000000 | -0.0004 | 4248.8 | 4152.3 | | 34 | 0.0918 | 0.1875 | 0.000000 | -0.0002 | 3744.8 | 3661.6 | | 35 | 0.1133 | 0.2188 | 0.000000 | -0.0002 | 3951.6 | 3868.0 | | 36 | 0.1816 | 0.2969 | 0.000000 | -0.0003 | 4009.3 | 3927.9 | ### job_496422 | Step | Reward | Pass@8 | KL | Loss | Step Time (s) | Gen Wait (s) | |------|--------|--------|-----|------|---------------|-------------| | 37 | 0.1953 | 0.4444 | 0.000000 | -0.0005 | 8145.7 | 8035.3 | | 38 | 0.0781 | 0.2344 | 0.000000 | -0.0003 | 3328.9 | 3242.5 | | 39 | 0.1680 | 0.4375 | 0.000000 | -0.0008 | 6397.1 | 6310.3 | | 40 | 0.1191 | 0.3281 | 0.000000 | -0.0004 | 5420.0 | 5330.6 | | 41 | 0.1211 | 0.2812 | 0.000000 | -0.0002 | 2226.2 | 2142.9 | | 42 | 0.1523 | 0.3281 | 0.000000 | -0.0003 | 4730.9 | 4645.8 | | 43 | 0.2285 | 0.4688 | 0.000000 | -0.0006 | 3828.9 | 3733.7 | | 44 | 0.1113 | 0.2656 | 0.000000 | -0.0002 | 4075.9 | 3983.0 | | 45 | 0.1133 | 0.2344 | 0.000000 | -0.0003 | 4079.3 | 3993.7 | ## Timing Analysis ### Average Time Breakdown (% of step time) | Component | Avg % of Step Time | |-----------|-------------------| | wait_for_generation_buffer | 96.5% | | run_training | 2.6% | | train_critic_and_policy | 2.3% | | policy_train | 2.3% | | sync_weights | 0.7% | | save_checkpoints | 0.5% | | save_hf_model | 0.4% | | fwd_logprobs_values_reward | 0.3% | | cleanup_old_checkpoints | 0.2% | | convert_to_training_input | 0.1% | | compute_advantages_and_returns | 0.0% | ## Cross-Log Comparison | Log | Avg Reward | Pass@8 | Step Time (s) | Gen Wait Time (s) | Avg Tokens | Staleness | |-----|------|------|------|------|------|------| | job_485099 | 0.0621 | 0.2000 | 4738.2360 | 4655.0062 | 202.7762 | 0.8125 | | job_485100 | 0.0610 | 0.1927 | 4401.3874 | 4316.1800 | 213.5232 | 0.7847 | | job_485101 | 0.0948 | 0.2646 | 4465.1476 | 4381.6743 | 292.8345 | 0.9732 | | job_496420 | 0.0935 | 0.2660 | 5111.3807 | 4998.7086 | 309.4238 | 0.8438 | | job_496421 | 0.1310 | 0.2818 | 4230.0172 | 4140.4545 | 357.4424 | 0.7937 | | job_496422 | 0.1430 | 0.3358 | 4692.5499 | 4601.9845 | 378.8585 | 0.9705 | ## vLLM Inference Engine Analysis Metrics from vLLM stat loggers (V1LoggingStatLoggerFixed). > **Note**: Ray deduplicates similar log messages with `[repeated Nx across cluster]`, > so we typically capture stats from one engine per timestamp. The stats shown are > **per-engine** values. Multiply by num_inference_engines for cluster-wide estimates. ### Summary by Log (Per-Engine Stats) | Log | Avg Running/Engine | Avg Waiting/Engine | Avg Gen Throughput/Engine | Avg KV Cache % | Avg Prefix Hit % | |-----|-------------------|-------------------|--------------------------|----------------|------------------| | job_485099 | 3.8 | 0.0 | 175.2 tok/s | 16.5% | 78.6% | | job_485100 | 3.6 | 0.0 | 165.7 tok/s | 15.7% | 81.1% | | job_485101 | 3.7 | 0.0 | 168.8 tok/s | 15.6% | 80.5% | | job_496420 | 3.2 | 0.0 | 145.1 tok/s | 13.9% | 79.2% | | job_496421 | 3.8 | 0.0 | 173.4 tok/s | 17.0% | 77.6% | | job_496422 | 3.2 | 0.0 | 145.2 tok/s | 13.7% | 74.7% | ### Utilization Analysis (Per-Engine) Key indicators of inference engine utilization: - **Running requests/engine**: Concurrent requests being processed by each engine - **Waiting requests**: Requests queued (0 = engine not saturated, has spare capacity) - **Generation throughput**: Decode tokens/sec per engine - 8B model on H100 can do **1000+ tok/s** when saturated - If seeing <300 tok/s with 0 waiting, engine is **starved for requests** #### job_485099 - **Running requests/engine**: avg=3.8, max=15 - **Waiting requests**: avg=0.0, max=0 - **Generation throughput/engine**: avg=175.2 tok/s, max=700.8 tok/s - **KV cache usage**: avg=16.5% - **Prefix cache hit rate**: avg=78.6% - ⚠️ **Underutilized**: Engines starved for requests (0 waiting, avg 3.8 running) - Bottleneck is likely upstream (environment execution, not inference) #### job_485100 - **Running requests/engine**: avg=3.6, max=17 - **Waiting requests**: avg=0.0, max=0 - **Generation throughput/engine**: avg=165.7 tok/s, max=682.9 tok/s - **KV cache usage**: avg=15.7% - **Prefix cache hit rate**: avg=81.1% - ⚠️ **Underutilized**: Engines starved for requests (0 waiting, avg 3.6 running) - Bottleneck is likely upstream (environment execution, not inference) #### job_485101 - **Running requests/engine**: avg=3.7, max=13 - **Waiting requests**: avg=0.0, max=0 - **Generation throughput/engine**: avg=168.8 tok/s, max=601.8 tok/s - **KV cache usage**: avg=15.6% - **Prefix cache hit rate**: avg=80.5% - ⚠️ **Underutilized**: Engines starved for requests (0 waiting, avg 3.7 running) - Bottleneck is likely upstream (environment execution, not inference) #### job_496420 - **Running requests/engine**: avg=3.2, max=11 - **Waiting requests**: avg=0.0, max=0 - **Generation throughput/engine**: avg=145.1 tok/s, max=485.1 tok/s - **KV cache usage**: avg=13.9% - **Prefix cache hit rate**: avg=79.2% - ⚠️ **Underutilized**: Engines starved for requests (0 waiting, avg 3.2 running) - Bottleneck is likely upstream (environment execution, not inference) #### job_496421 - **Running requests/engine**: avg=3.8, max=15 - **Waiting requests**: avg=0.0, max=0 - **Generation throughput/engine**: avg=173.4 tok/s, max=665.8 tok/s - **KV cache usage**: avg=17.0% - **Prefix cache hit rate**: avg=77.6% - ⚠️ **Underutilized**: Engines starved for requests (0 waiting, avg 3.8 running) - Bottleneck is likely upstream (environment execution, not inference) #### job_496422 - **Running requests/engine**: avg=3.2, max=16 - **Waiting requests**: avg=0.0, max=0 - **Generation throughput/engine**: avg=145.2 tok/s, max=682.5 tok/s - **KV cache usage**: avg=13.7% - **Prefix cache hit rate**: avg=74.7% - ⚠️ **Underutilized**: Engines starved for requests (0 waiting, avg 3.2 running) - Bottleneck is likely upstream (environment execution, not inference) ## Trial-Level Analysis (from result.json) Total trials parsed: 25855 ### Turn Count Statistics | Metric | Value | |--------|-------| | Mean | 1.7 | | Median | 1.0 | | Std | 1.3 | | Min | 1 | | Max | 33 | | Count | 25830 | ### Exception Distribution | Exception Type | Count | % | |---------------|-------|---| | AgentEnvironmentTimeoutError | 18399 | 71.2% | | DaytonaAuthenticationError | 3597 | 13.9% | | No exception | 2770 | 10.7% | | VerifierRuntimeError | 778 | 3.0% | | AgentTimeoutError | 158 | 0.6% | | DaytonaNotFoundError | 72 | 0.3% | | DaytonaError | 29 | 0.1% | | ContextLengthExceededError | 23 | 0.1% | | AgentSetupTimeoutError | 17 | 0.1% | | VerifierTimeoutError | 8 | 0.0% | | RuntimeError | 2 | 0.0% | | DaytonaValidationError | 1 | 0.0% | | EnvironmentStartTimeoutError | 1 | 0.0% | ### Turn Count by Exception Type | Exception Type | Mean Turns | Median Turns | Count | |---------------|-----------|-------------|-------| | AgentTimeoutError | 9.7 | 8.0 | 158 | | No exception | 3.2 | 3.0 | 2770 | | VerifierRuntimeError | 3.1 | 3.0 | 778 | | VerifierTimeoutError | 2.6 | 2.0 | 8 | | DaytonaValidationError | 2.0 | 2.0 | 1 | | DaytonaNotFoundError | 1.7 | 2.0 | 72 | | DaytonaError | 1.6 | 2.0 | 24 | | AgentEnvironmentTimeoutError | 1.5 | 1.0 | 18399 | | ContextLengthExceededError | 1.4 | 1.0 | 23 | | DaytonaAuthenticationError | 1.2 | 1.0 | 3597 | ### Turn Count by Outcome | Outcome | Mean Turns | Median Turns | Count | |---------|-----------|-------------|-------| | Success | 3.2 | 3.0 | 4834 | ### Reward Summary - Mean reward: 1.0000 - Success rate: 100.0% - Trials with reward data: 4834