# SkyRL Training Metrics Analysis Generated from 3 log files ## Overview | Log File | Total Steps | Metric Blocks | Final Reward (mean) | Final Reward (max) | Total Time (s) | |----------|-------------|---------------|---------------------|-------------------|----------------| | job_567548 | 9 | 9 | 0.4008 | 0.5059 | 6840.2 | | job_567550 | 50 | 42 | 0.3437 | 0.4668 | 38753.5 | | job_567551 | 80 | 30 | 0.3892 | 0.5586 | 25084.9 | ## Async Metrics | | Mean | Std | Min | Max | Count | |:------------------------------|-----------:|---------:|------:|---------:|--------:| | async/discard_rate | 0 | 0 | 0 | 0 | 81 | | async/discarded_count | 0 | 0 | 0 | 0 | 81 | | async/effective_batch_groups | 64 | 0 | 64 | 64 | 81 | | async/effective_batch_samples | 512 | 0 | 512 | 512 | 81 | | async/staleness_max | 5.48148 | 1.62874 | 0 | 8 | 81 | | async/staleness_mean | 2.52044 | 0.596524 | 0 | 3.2344 | 81 | | async/staleness_min | 0.135802 | 0.344713 | 0 | 1 | 81 | | async/staleness_ratio | 0.921498 | 0.127923 | 0 | 1 | 81 | ## Generate Metrics | | Mean | Std | Min | Max | Count | |:-------------------------------------|---------:|---------:|---------:|---------:|--------:| | generate/avg_num_tokens | 6219.84 | 593.767 | 3631.05 | 7508.99 | 81 | | generate/avg_tokens_non_zero_rewards | 6723.53 | 885.765 | 3316.3 | 8407.4 | 81 | | generate/avg_tokens_zero_rewards | 5945.13 | 669.261 | 3953.26 | 7989.15 | 81 | | generate/max_num_tokens | 30304.3 | 1297.57 | 25303 | 31890 | 81 | | generate/min_num_tokens | 1 | 0 | 1 | 1 | 81 | | generate/std_num_tokens | 6150.02 | 623.638 | 3243.07 | 7148.25 | 81 | ## Loss Metrics | | Mean | Std | Min | Max | Count | |:----------------------------|-----------:|-----------:|--------:|-------:|--------:| | loss/avg_final_rewards | 0.36687 | 0.0610739 | 0.2559 | 0.5586 | 81 | | loss/avg_raw_advantages | 0.00432716 | 0.00984494 | -0.0145 | 0.0461 | 81 | | loss/avg_raw_advantages_abs | 0.160136 | 0.0385764 | 0.0852 | 0.2488 | 81 | ## Policy Metrics | | Mean | Std | Min | Max | Count | |:----------------------------|-------------:|------------:|--------:|-------:|--------:| | policy/final_loss | -1.48148e-05 | 4.50309e-05 | -0.0003 | 0 | 81 | | policy/log_ratio_abs_max | 0 | 0 | 0 | 0 | 81 | | policy/log_ratio_abs_mean | 0 | 0 | 0 | 0 | 81 | | policy/log_ratio_abs_p99 | 0 | 0 | 0 | 0 | 81 | | policy/log_ratio_abs_pos00 | 0 | 0 | 0 | 0 | 81 | | policy/log_ratio_abs_pos10 | 0 | 0 | 0 | 0 | 81 | | policy/log_ratio_abs_pos20 | 0 | 0 | 0 | 0 | 81 | | policy/log_ratio_abs_pos30 | 0 | 0 | 0 | 0 | 81 | | policy/log_ratio_abs_pos40 | 0 | 0 | 0 | 0 | 81 | | policy/log_ratio_abs_pos50 | 0 | 0 | 0 | 0 | 81 | | policy/log_ratio_abs_pos60 | 0 | 0 | 0 | 0 | 81 | | policy/log_ratio_abs_pos70 | 0 | 0 | 0 | 0 | 81 | | policy/log_ratio_abs_pos80 | 0 | 0 | 0 | 0 | 81 | | policy/log_ratio_abs_pos90 | 0 | 0 | 0 | 0 | 81 | | policy/n_tokens_dp_gt_10pct | 0 | 0 | 0 | 0 | 81 | | policy/n_tokens_dp_gt_1pct | 0 | 0 | 0 | 0 | 81 | | policy/n_tokens_dp_gt_50pct | 0 | 0 | 0 | 0 | 81 | | policy/policy_entropy | 0.135762 | 0.0098146 | 0.1121 | 0.1546 | 81 | | policy/policy_loss | -0.00113086 | 0.00263399 | -0.0202 | 0 | 81 | | policy/policy_lr | 0 | 0 | 0 | 0 | 81 | | policy/policy_update_steps | 1 | 0 | 1 | 1 | 81 | | policy/ppo_clip_ratio | 0 | 0 | 0 | 0 | 81 | | policy/raw_grad_norm | 0.0180111 | 0.00588504 | 0.0124 | 0.064 | 81 | ## Reward Metrics | | Mean | Std | Min | Max | Count | |:----------------------|---------:|----------:|-------:|-------:|--------:| | reward/avg_pass_at_8 | 0.522458 | 0.0740283 | 0.3438 | 0.6875 | 81 | | reward/avg_raw_reward | 0.36687 | 0.0610739 | 0.2559 | 0.5586 | 81 | ## System Metrics | | Mean | Std | Min | Max | Count | |:------------------------|---------:|---------:|---------:|---------:|--------:| | system/process_rss_gb | 15.5018 | 1.58876 | 10.6846 | 17.8517 | 81 | | system/process_vms_gb | 52.6829 | 1.43311 | 50.2918 | 54.9248 | 81 | | system/ram_available_gb | 620.096 | 12.2909 | 588.791 | 634.891 | 81 | | system/ram_percent | 27.7247 | 1.43436 | 26 | 31.4 | 81 | | system/ram_total_gb | 857.968 | 0 | 857.968 | 857.968 | 81 | | system/ram_used_gb | 237.871 | 12.2909 | 223.076 | 269.176 | 81 | ## Timing Metrics | | Mean | Std | Min | Max | Count | |:--------------------------------------|------------:|------------:|---------:|----------:|--------:| | timing/compute_advantages_and_returns | 0.0666617 | 0.0122347 | 0.0519 | 0.1531 | 81 | | timing/convert_to_training_input | 4.72614 | 0.280405 | 3.6373 | 5.4807 | 81 | | timing/fwd_logprobs_values_reward | 49.3127 | 5.50948 | 28.9658 | 61.4679 | 81 | | timing/policy_train | 290.478 | 30.592 | 151.591 | 354.228 | 81 | | timing/run_training | 340.145 | 35.2621 | 180.815 | 408.019 | 81 | | timing/step | 872.574 | 400.434 | 507.833 | 3374.16 | 81 | | timing/sync_weights | 20.0706 | 1.03089 | 17.0362 | 23.2754 | 81 | | timing/train_critic_and_policy | 290.765 | 30.6031 | 151.737 | 354.553 | 81 | | timing/wait_for_generation_buffer | 507.629 | 406.806 | 131.502 | 3072.09 | 81 | | timing/cleanup_old_checkpoints | 2.02637 | 7.98054 | 0.0063 | 38.5904 | 40 | | timing/save_checkpoints | 23.1826 | 62.9803 | 7.7944 | 396.822 | 40 | | timing/save_hf_model | 6.37928 | 0.425326 | 5.9537 | 7.5975 | 16 | ## Trainer Metrics | | Mean | Std | Min | Max | Count | |:--------------------|--------:|--------:|------:|------:|--------:| | trainer/epoch | 0 | 0 | 0 | 0 | 81 | | trainer/global_step | 40.1111 | 23.3559 | 1 | 80 | 81 | ## Batch_Errors Metrics | | Mean | Std | Min | Max | Count | |:----------------------------------------------|------------:|--------------:|------------:|------------:|--------:| | batch_errors/total_batches | 60.9136 | 8.52379 | 42 | 121 | 81 | | batch_errors/total_instances | 487.309 | 68.1903 | 336 | 968 | 81 | | batch_errors/total_successful | 403.728 | 74.9498 | 283 | 939 | 81 | | batch_errors/total_failed | 19.8642 | 5.64525 | 5 | 32 | 81 | | batch_errors/total_masked | 79.2716 | 25.8404 | 25 | 134 | 81 | | batch_errors/avg_RewardFileNotFoundError | 0.531055 | 0.271125 | 0.03125 | 1.42593 | 81 | | batch_errors/total_RewardFileNotFoundError | 31.8272 | 15.4934 | 2 | 77 | 81 | | batch_errors/avg_DaytonaValidationError | 0.0672515 | 0.0494668 | 0.0106383 | 0.186441 | 24 | | batch_errors/total_DaytonaValidationError | 4 | 2.87417 | 1 | 11 | 24 | | batch_errors/avg_ContextLengthExceededError | 0.779734 | 0.294334 | 0.16129 | 1.55556 | 79 | | batch_errors/total_ContextLengthExceededError | 46.9873 | 18.3376 | 10 | 98 | 79 | | batch_errors/avg_RuntimeError | 0.0324855 | 0.023494 | 0.015625 | 0.101695 | 46 | | batch_errors/total_RuntimeError | 1.97826 | 1.40616 | 1 | 6 | 46 | | batch_errors/avg_DaytonaError | 0.0612569 | 0.0447787 | 0.0106383 | 0.206897 | 52 | | batch_errors/total_DaytonaError | 3.65385 | 2.65599 | 1 | 12 | 52 | | batch_errors/avg_AgentTimeoutError | 0.0810606 | 0.0876512 | 0.015625 | 0.625 | 66 | | batch_errors/total_AgentTimeoutError | 4.74242 | 4.54147 | 1 | 30 | 66 | | batch_errors/avg_VerifierTimeoutError | 0.0785312 | 0.05595 | 0.015873 | 0.225806 | 23 | | batch_errors/total_VerifierTimeoutError | 4.65217 | 3.2975 | 1 | 14 | 23 | | batch_errors/avg_VerifierOutputParseError | 0.137931 | 0 | 0.137931 | 0.137931 | 2 | | batch_errors/total_VerifierOutputParseError | 8 | 0 | 8 | 8 | 2 | | batch_errors/avg_DaytonaNotFoundError | 0.0559985 | 0.0495709 | 0.015873 | 0.145455 | 9 | | batch_errors/total_DaytonaNotFoundError | 3.22222 | 2.68225 | 1 | 8 | 9 | | batch_errors/avg_InvalidChatHistory | 0.0198765 | 0.00788726 | 0.015625 | 0.0338983 | 5 | | batch_errors/total_InvalidChatHistory | 1.2 | 0.447214 | 1 | 2 | 5 | | batch_errors/avg_DaytonaAuthenticationError | 0.0163291 | 0.000790025 | 0.015873 | 0.0172414 | 3 | | batch_errors/total_DaytonaAuthenticationError | 1 | 0 | 1 | 1 | 3 | | batch_errors/avg_RewardFileEmptyError | 0.015873 | nan | 0.015873 | 0.015873 | 1 | | batch_errors/total_RewardFileEmptyError | 1 | nan | 1 | 1 | 1 | | batch_errors/avg_AddTestsDirError | 0.015873 | nan | 0.015873 | 0.015873 | 1 | | batch_errors/total_AddTestsDirError | 1 | nan | 1 | 1 | 1 | | batch_errors/avg_AgentSetupTimeoutError | 0.206955 | 0.203236 | 0.016129 | 0.479167 | 4 | | batch_errors/total_AgentSetupTimeoutError | 11 | 9.62635 | 1 | 23 | 4 | | batch_errors/avg_OSError | 0.0164111 | 0.000760944 | 0.015873 | 0.0169492 | 2 | | batch_errors/total_OSError | 1 | 0 | 1 | 1 | 2 | ## Training Progression by Log ### job_567548 | Step | Reward | Pass@8 | KL | Loss | Step Time (s) | Gen Wait (s) | |------|--------|--------|-----|------|---------------|-------------| | 1 | 0.5059 | 0.6250 | 0.000000 | -0.0000 | 1358.1 | 1155.0 | | 2 | 0.5000 | 0.6719 | 0.000000 | 0.0000 | 833.5 | 586.8 | | 3 | 0.4199 | 0.5938 | 0.000000 | 0.0000 | 703.1 | 332.8 | | 4 | 0.4609 | 0.6406 | 0.000000 | 0.0000 | 924.0 | 548.2 | | 5 | 0.3359 | 0.4531 | 0.000000 | -0.0001 | 609.8 | 230.0 | | 6 | 0.3613 | 0.5000 | 0.000000 | -0.0000 | 528.5 | 145.4 | | 7 | 0.3340 | 0.4844 | 0.000000 | -0.0000 | 589.5 | 248.8 | | 8 | 0.4082 | 0.5312 | 0.000000 | -0.0000 | 605.9 | 228.5 | | 9 | 0.2812 | 0.4844 | 0.000000 | 0.0000 | 687.8 | 372.3 | ### job_567550 | Step | Reward | Pass@8 | KL | Loss | Step Time (s) | Gen Wait (s) | |------|--------|--------|-----|------|---------------|-------------| | 9 | 0.3418 | 0.5556 | 0.000000 | 0.0000 | 3374.2 | 3072.1 | | 10 | 0.3867 | 0.5938 | 0.000000 | 0.0000 | 1168.0 | 795.3 | | 11 | 0.2988 | 0.4531 | 0.000000 | -0.0000 | 507.8 | 131.5 | | 12 | 0.3145 | 0.4844 | 0.000000 | 0.0000 | 516.4 | 177.0 | | 13 | 0.3301 | 0.4688 | 0.000000 | -0.0000 | 610.9 | 302.7 | | 14 | 0.2559 | 0.4375 | 0.000000 | -0.0000 | 839.0 | 513.9 | | 15 | 0.3613 | 0.5000 | 0.000000 | -0.0000 | 643.1 | 282.0 | | 16 | 0.4551 | 0.5625 | 0.000000 | -0.0000 | 563.6 | 251.9 | | 17 | 0.3164 | 0.4688 | 0.000000 | -0.0000 | 640.3 | 279.4 | | 18 | 0.3438 | 0.4531 | 0.000000 | 0.0000 | 762.5 | 436.3 | | 19 | 0.3184 | 0.4688 | 0.000000 | 0.0000 | 836.6 | 483.8 | | 20 | 0.2637 | 0.3750 | 0.000000 | -0.0000 | 714.5 | 341.4 | | 21 | 0.3750 | 0.5781 | 0.000000 | 0.0000 | 829.4 | 430.9 | | 22 | 0.3418 | 0.5000 | 0.000000 | -0.0000 | 763.6 | 362.2 | | 23 | 0.3262 | 0.5156 | 0.000000 | 0.0000 | 1049.2 | 675.6 | | 24 | 0.2793 | 0.4219 | 0.000000 | 0.0000 | 704.3 | 389.0 | | 25 | 0.3574 | 0.4844 | 0.000000 | 0.0000 | 868.9 | 495.4 | | 26 | 0.3320 | 0.4688 | 0.000000 | -0.0000 | 854.4 | 467.6 | | 27 | 0.3223 | 0.5000 | 0.000000 | 0.0000 | 590.1 | 255.9 | | 28 | 0.3398 | 0.4844 | 0.000000 | -0.0000 | 880.2 | 456.9 | | 29 | 0.3613 | 0.5625 | 0.000000 | 0.0000 | 884.7 | 516.6 | | 30 | 0.3281 | 0.5156 | 0.000000 | 0.0000 | 1129.6 | 742.6 | | 31 | 0.2891 | 0.3750 | 0.000000 | -0.0000 | 880.0 | 508.9 | | 32 | 0.3418 | 0.5312 | 0.000000 | 0.0000 | 799.6 | 456.1 | | 33 | 0.3105 | 0.4688 | 0.000000 | -0.0000 | 801.0 | 438.2 | | 34 | 0.3457 | 0.5312 | 0.000000 | -0.0001 | 946.0 | 580.7 | | 35 | 0.4668 | 0.6875 | 0.000000 | 0.0000 | 848.1 | 465.5 | | 36 | 0.3984 | 0.5625 | 0.000000 | -0.0001 | 924.9 | 576.2 | | 37 | 0.3438 | 0.5312 | 0.000000 | -0.0000 | 883.1 | 530.8 | | 38 | 0.3926 | 0.5156 | 0.000000 | -0.0001 | 1456.7 | 1081.1 | | 39 | 0.3887 | 0.5312 | 0.000000 | -0.0000 | 854.4 | 500.4 | | 40 | 0.3066 | 0.4531 | 0.000000 | -0.0000 | 893.6 | 560.5 | | 41 | 0.3281 | 0.4688 | 0.000000 | -0.0003 | 2041.8 | 1642.7 | | 42 | 0.4082 | 0.5938 | 0.000000 | -0.0001 | 767.9 | 466.6 | | 43 | 0.3691 | 0.4531 | 0.000000 | -0.0000 | 869.6 | 506.1 | | 44 | 0.3926 | 0.5781 | 0.000000 | -0.0001 | 783.8 | 429.7 | | 45 | 0.3242 | 0.4844 | 0.000000 | -0.0000 | 673.9 | 300.3 | | 46 | 0.3359 | 0.5312 | 0.000000 | -0.0000 | 1024.1 | 663.9 | | 47 | 0.3203 | 0.4844 | 0.000000 | 0.0000 | 1112.0 | 706.6 | | 48 | 0.3535 | 0.5156 | 0.000000 | -0.0000 | 928.2 | 543.8 | | 49 | 0.4023 | 0.5938 | 0.000000 | -0.0000 | 759.6 | 398.2 | | 50 | 0.2656 | 0.3438 | 0.000000 | 0.0000 | 774.0 | 416.6 | ### job_567551 | Step | Reward | Pass@8 | KL | Loss | Step Time (s) | Gen Wait (s) | |------|--------|--------|-----|------|---------------|-------------| | 51 | 0.4785 | 0.6226 | 0.000000 | -0.0001 | 2184.4 | 1845.0 | | 52 | 0.5586 | 0.6875 | 0.000000 | -0.0000 | 1651.9 | 1245.4 | | 53 | 0.3867 | 0.5000 | 0.000000 | 0.0000 | 652.1 | 292.6 | | 54 | 0.3047 | 0.5156 | 0.000000 | 0.0000 | 560.9 | 186.2 | | 55 | 0.3281 | 0.4531 | 0.000000 | 0.0000 | 683.3 | 332.7 | | 56 | 0.4629 | 0.5938 | 0.000000 | -0.0000 | 854.5 | 488.2 | | 57 | 0.4434 | 0.5781 | 0.000000 | -0.0001 | 854.4 | 440.7 | | 58 | 0.3535 | 0.5156 | 0.000000 | -0.0000 | 785.5 | 392.6 | | 59 | 0.3613 | 0.5469 | 0.000000 | -0.0000 | 643.7 | 253.1 | | 60 | 0.3848 | 0.5469 | 0.000000 | 0.0000 | 674.6 | 305.4 | | 61 | 0.3457 | 0.5156 | 0.000000 | 0.0000 | 684.6 | 252.5 | | 62 | 0.3867 | 0.5312 | 0.000000 | 0.0000 | 624.4 | 254.2 | | 63 | 0.3574 | 0.5938 | 0.000000 | -0.0001 | 868.7 | 474.5 | | 64 | 0.3066 | 0.4375 | 0.000000 | -0.0000 | 692.7 | 263.9 | | 65 | 0.3711 | 0.5000 | 0.000000 | -0.0000 | 634.6 | 271.4 | | 66 | 0.3906 | 0.5469 | 0.000000 | -0.0000 | 709.9 | 303.6 | | 67 | 0.4883 | 0.6875 | 0.000000 | 0.0000 | 517.5 | 135.4 | | 68 | 0.4043 | 0.6094 | 0.000000 | 0.0000 | 671.0 | 285.3 | | 69 | 0.4531 | 0.6562 | 0.000000 | -0.0000 | 730.7 | 348.2 | | 70 | 0.3359 | 0.5000 | 0.000000 | -0.0000 | 928.1 | 566.1 | | 71 | 0.3750 | 0.4688 | 0.000000 | 0.0000 | 809.2 | 469.7 | | 72 | 0.4609 | 0.6562 | 0.000000 | 0.0000 | 775.9 | 388.3 | | 73 | 0.3965 | 0.5625 | 0.000000 | 0.0000 | 736.8 | 355.8 | | 74 | 0.2988 | 0.3594 | 0.000000 | 0.0000 | 639.6 | 270.0 | | 75 | 0.3301 | 0.4531 | 0.000000 | -0.0000 | 987.1 | 598.9 | | 76 | 0.4746 | 0.5938 | 0.000000 | 0.0000 | 1003.0 | 623.9 | | 77 | 0.3789 | 0.5625 | 0.000000 | -0.0000 | 827.6 | 440.6 | | 78 | 0.3594 | 0.4688 | 0.000000 | -0.0000 | 1027.6 | 645.9 | | 79 | 0.3223 | 0.4688 | 0.000000 | -0.0000 | 864.1 | 469.0 | | 80 | 0.3770 | 0.5156 | 0.000000 | 0.0000 | 806.3 | 438.0 | ## Timing Analysis ### Average Time Breakdown (% of step time) | Component | Avg % of Step Time | |-----------|-------------------| | wait_for_generation_buffer | 53.6% | | run_training | 43.2% | | train_critic_and_policy | 37.0% | | policy_train | 36.9% | | fwd_logprobs_values_reward | 6.3% | | save_checkpoints | 2.7% | | sync_weights | 2.6% | | save_hf_model | 0.8% | | convert_to_training_input | 0.6% | | cleanup_old_checkpoints | 0.2% | | compute_advantages_and_returns | 0.0% | ## Cross-Log Comparison | Log | Avg Reward | Pass@8 | Step Time (s) | Gen Wait Time (s) | Avg Tokens | Staleness | |-----|------|------|------|------|------|------| | job_567548 | 0.4008 | 0.5538 | 760.0180 | 427.5351 | 5697.2672 | 2.0174 | | job_567550 | 0.3437 | 0.5021 | 922.7027 | 562.6884 | 6171.1224 | 2.5595 | | job_567551 | 0.3892 | 0.5416 | 836.1619 | 454.5744 | 6444.8237 | 2.6167 | ## vLLM Inference Engine Analysis Metrics from vLLM stat loggers (V1LoggingStatLoggerFixed). > **Note**: Ray deduplicates similar log messages with `[repeated Nx across cluster]`, > so we typically capture stats from one engine per timestamp. The stats shown are > **per-engine** values. Multiply by num_inference_engines for cluster-wide estimates. ### Summary by Log (Per-Engine Stats) | Log | Avg Running/Engine | Avg Waiting/Engine | Avg Gen Throughput/Engine | Avg KV Cache % | Avg Prefix Hit % | |-----|-------------------|-------------------|--------------------------|----------------|------------------| | job_567548 | 8.7 | 0.0 | 383.7 tok/s | 36.5% | 91.9% | | job_567550 | 7.6 | 0.0 | 347.5 tok/s | 33.3% | 92.4% | | job_567551 | 9.0 | 0.0 | 399.0 tok/s | 39.7% | 91.1% | ### Utilization Analysis (Per-Engine) Key indicators of inference engine utilization: - **Running requests/engine**: Concurrent requests being processed by each engine - **Waiting requests**: Requests queued (0 = engine not saturated, has spare capacity) - **Generation throughput**: Decode tokens/sec per engine - 8B model on H100 can do **1000+ tok/s** when saturated - If seeing <300 tok/s with 0 waiting, engine is **starved for requests** #### job_567548 - **Running requests/engine**: avg=8.7, max=22 - **Waiting requests**: avg=0.0, max=4 - **Generation throughput/engine**: avg=383.7 tok/s, max=790.8 tok/s - **KV cache usage**: avg=36.5% - **Prefix cache hit rate**: avg=91.9% - ✅ **Well-utilized**: Engines saturated (waiting > 0) #### job_567550 - **Running requests/engine**: avg=7.6, max=24 - **Waiting requests**: avg=0.0, max=10 - **Generation throughput/engine**: avg=347.5 tok/s, max=878.7 tok/s - **KV cache usage**: avg=33.3% - **Prefix cache hit rate**: avg=92.4% - ✅ **Well-utilized**: Engines saturated (waiting > 0) #### job_567551 - **Running requests/engine**: avg=9.0, max=24 - **Waiting requests**: avg=0.0, max=19 - **Generation throughput/engine**: avg=399.0 tok/s, max=851.4 tok/s - **KV cache usage**: avg=39.7% - **Prefix cache hit rate**: avg=91.1% - ✅ **Well-utilized**: Engines saturated (waiting > 0) ## Trial-Level Analysis (from result.json) Total trials parsed: 44526 ### Turn Count Statistics | Metric | Value | |--------|-------| | Mean | 11.5 | | Median | 9.0 | | Std | 8.2 | | Min | 1 | | Max | 116 | | Count | 44464 | ### Exception Distribution | Exception Type | Count | % | |---------------|-------|---| | No exception | 35187 | 79.0% | | ContextLengthExceededError | 3811 | 8.6% | | RewardFileNotFoundError | 2794 | 6.3% | | AgentTimeoutError | 2128 | 4.8% | | DaytonaError | 204 | 0.5% | | VerifierTimeoutError | 112 | 0.3% | | DaytonaValidationError | 103 | 0.2% | | RuntimeError | 93 | 0.2% | | AgentSetupTimeoutError | 44 | 0.1% | | DaytonaNotFoundError | 29 | 0.1% | | VerifierOutputParseError | 16 | 0.0% | | DaytonaAuthenticationError | 3 | 0.0% | | RewardFileEmptyError | 1 | 0.0% | | AddTestsDirError | 1 | 0.0% | ### Turn Count by Exception Type | Exception Type | Mean Turns | Median Turns | Count | |---------------|-----------|-------------|-------| | RewardFileEmptyError | 21.0 | 21.0 | 1 | | AddTestsDirError | 19.0 | 19.0 | 1 | | AgentTimeoutError | 17.8 | 11.0 | 2128 | | ContextLengthExceededError | 16.4 | 13.0 | 3811 | | VerifierTimeoutError | 13.9 | 11.5 | 112 | | RewardFileNotFoundError | 13.2 | 12.0 | 2794 | | DaytonaAuthenticationError | 12.7 | 7.0 | 3 | | VerifierOutputParseError | 11.5 | 10.0 | 16 | | DaytonaError | 11.2 | 8.0 | 186 | | RuntimeError | 10.7 | 10.0 | 93 | | No exception | 10.4 | 8.0 | 35187 | | DaytonaValidationError | 8.5 | 7.0 | 103 | | DaytonaNotFoundError | 8.0 | 6.0 | 29 | ### Turn Count by Outcome | Outcome | Mean Turns | Median Turns | Count | |---------|-----------|-------------|-------| | Success | 9.8 | 8.0 | 16178 | | Failure | 11.6 | 9.0 | 20820 | ### Reward Summary - Mean reward: 0.4373 - Success rate: 43.7% - Trials with reward data: 36998