Files
a3-rl-laion_nemotron-gym-kn…/training_logs/20260604_042832_metrics_report.md
ModelHub XC e38b6cff77 初始化项目,由ModelHub XC社区提供模型
Model: laion/a3-rl-laion_nemotron-gym-knowledge-web-search-mcqa-25-8B
Source: Original Platform
2026-08-10 00:48:55 +08:00

19 KiB

SkyRL Training Metrics Analysis

Generated from 2 log files

Overview

Log File Total Steps Metric Blocks Final Reward (mean) Final Reward (max) Total Time (s)
job_561773 47 47 0.4832 0.6699 40348.4
job_561774 80 34 0.4752 0.6641 40570.1

Async Metrics

Mean Std Min Max Count
async/discard_rate 0 0 0 0 81
async/discarded_count 0 0 0 0 81
async/effective_batch_groups 64 0 64 64 81
async/effective_batch_samples 512 0 512 512 81
async/staleness_max 5.03704 1.78497 0 9 81
async/staleness_mean 2.78221 0.838605 0 5 81
async/staleness_min 0.259259 0.833333 0 5 81
async/staleness_ratio 0.909143 0.17908 0 1 81

Generate Metrics

Mean Std Min Max Count
generate/avg_num_tokens 1679.98 115.733 1419.4 1972.39 81
generate/avg_tokens_non_zero_rewards 1823.91 152.503 1391.42 2257.26 81
generate/avg_tokens_zero_rewards 1561.41 173.17 1143.64 1901.63 81
generate/max_num_tokens 14893.9 8449.88 4339 31942 81
generate/min_num_tokens 82.8272 233.383 1 792 81
generate/std_num_tokens 1248.78 367.594 474.896 2683.1 81

Loss Metrics

Mean Std Min Max Count
loss/avg_final_rewards 0.479843 0.0654987 0.377 0.6699 81
loss/avg_raw_advantages -0.0011 0.0157427 -0.0713 0.0435 81
loss/avg_raw_advantages_abs 0.210109 0.0365231 0.109 0.3151 81

Policy Metrics

Mean Std Min Max Count
policy/final_loss -1.85185e-05 9.09823e-05 -0.0006 0 81
policy/log_ratio_abs_max 0 0 0 0 81
policy/log_ratio_abs_mean 0 0 0 0 81
policy/log_ratio_abs_p99 0 0 0 0 81
policy/log_ratio_abs_pos00 0 0 0 0 81
policy/log_ratio_abs_pos10 0 0 0 0 81
policy/log_ratio_abs_pos20 0 0 0 0 81
policy/log_ratio_abs_pos30 0 0 0 0 81
policy/log_ratio_abs_pos40 0 0 0 0 81
policy/log_ratio_abs_pos50 0 0 0 0 81
policy/log_ratio_abs_pos60 0 0 0 0 81
policy/log_ratio_abs_pos70 0 0 0 0 81
policy/log_ratio_abs_pos80 0 0 0 0 81
policy/log_ratio_abs_pos90 0 0 0 0 81
policy/n_tokens_dp_gt_10pct 0 0 0 0 81
policy/n_tokens_dp_gt_1pct 0 0 0 0 81
policy/n_tokens_dp_gt_50pct 0 0 0 0 81
policy/policy_entropy 0.213363 0.0123506 0.1847 0.2375 81
policy/policy_loss -0.00124074 0.00556657 -0.0354 0 81
policy/policy_lr 0 0 0 0 81
policy/policy_update_steps 1 0 1 1 81
policy/ppo_clip_ratio 0 0 0 0 81
policy/raw_grad_norm 0.0342728 0.0147022 0.0239 0.1338 81

Reward Metrics

Mean Std Min Max Count
reward/avg_pass_at_8 0.742557 0.0605992 0.6094 0.8906 81
reward/avg_raw_reward 0.479843 0.0654987 0.377 0.6699 81

System Metrics

Mean Std Min Max Count
system/process_rss_gb 12.7194 2.95903 3.9384 15.7029 81
system/process_vms_gb 51.2557 0.655464 49.7056 52.769 81
system/ram_available_gb 596.999 9.85882 581.762 622.734 81
system/ram_percent 30.4185 1.15316 27.4 32.2 81
system/ram_total_gb 857.968 4.96593e-05 857.967 857.968 81
system/ram_used_gb 260.969 9.85882 235.234 276.205 81

Timing Metrics

Mean Std Min Max Count
timing/compute_advantages_and_returns 0.044784 0.041453 0.0171 0.3713 81
timing/convert_to_training_input 2.11839 1.02368 0.8234 4.4063 81
timing/fwd_logprobs_values_reward 11.5596 1.63923 8.4123 16.8231 81
timing/policy_train 71.9472 8.18653 53.9066 101.06 81
timing/run_training 83.7054 9.52006 62.4678 116.346 81
timing/step 998.994 405.534 465.143 3038.23 81
timing/sync_weights 23.0452 1.86714 17.5989 27.0223 81
timing/train_critic_and_policy 72.1007 8.19559 54.0381 101.242 81
timing/wait_for_generation_buffer 890.121 405.78 354.775 2914.47 81
timing/cleanup_old_checkpoints 0.513695 2.26344 0.0058 11.6815 40
timing/save_checkpoints 10.0415 2.31378 7.9137 19.2498 40
timing/save_hf_model 6.39602 0.444487 5.8828 7.7434 16

Trainer Metrics

Mean Std Min Max Count
trainer/epoch 0.444444 0.5 0 1 81
trainer/global_step 40.5802 23.1035 1 80 81

Batch_Errors Metrics

Mean Std Min Max Count
batch_errors/total_batches 57.1111 5.66569 46 96 81
batch_errors/total_instances 456.889 45.3255 368 768 81
batch_errors/total_successful 414.506 51.6568 335 768 81
batch_errors/total_failed 19.4568 10.1797 0 40 81
batch_errors/total_masked 40.6914 24.1866 0 95 81
batch_errors/avg_InvalidChatHistory 0.838266 0.38179 0.0175439 1.74074 74
batch_errors/total_InvalidChatHistory 47.6622 21.9491 1 96 74
batch_errors/avg_DaytonaError 0.0271896 0.0157125 0.0163934 0.0714286 28
batch_errors/total_DaytonaError 1.57143 0.920087 1 4 28
batch_errors/avg_ContextLengthExceededError 0.0319596 0.0406469 0.016129 0.160714 13
batch_errors/total_ContextLengthExceededError 1.76923 2.24179 1 9 13
batch_errors/avg_RuntimeError 0.0178571 nan 0.0178571 0.0178571 1
batch_errors/total_RuntimeError 1 nan 1 1 1
batch_errors/avg_AgentTimeoutError 0.0885766 0.141711 0.0166667 0.44 9
batch_errors/total_AgentTimeoutError 4.66667 7.1239 1 22 9
batch_errors/avg_AgentSetupTimeoutError 0.249893 0.30173 0.015873 0.74 8
batch_errors/total_AgentSetupTimeoutError 13.125 16.0128 1 38 8
batch_errors/avg_VerifierTimeoutError 0.105889 0.046451 0.0181818 0.16 7
batch_errors/total_VerifierTimeoutError 5.71429 2.42997 1 8 7
batch_errors/avg_DaytonaAuthenticationError 0.0324767 0.0194287 0.015873 0.0555556 5
batch_errors/total_DaytonaAuthenticationError 1.8 1.09545 1 3 5
batch_errors/avg_Timeout 0.0181818 nan 0.0181818 0.0181818 1
batch_errors/total_Timeout 1 nan 1 1 1

Training Progression by Log

job_561773

Step Reward Pass@8 KL Loss Step Time (s) Gen Wait (s)
1 0.5410 0.7031 0.000000 -0.0000 1394.8 1292.8
2 0.5488 0.7031 0.000000 -0.0000 486.3 393.9
3 0.4961 0.7812 0.000000 -0.0000 658.8 557.2
4 0.6094 0.8125 0.000000 0.0000 692.6 588.1
5 0.4883 0.7500 0.000000 0.0000 954.4 847.3
6 0.4922 0.8438 0.000000 0.0000 465.1 354.8
7 0.4453 0.6719 0.000000 0.0000 603.9 493.2
8 0.4570 0.7812 0.000000 0.0000 684.1 570.8
9 0.4121 0.7031 0.000000 0.0000 868.4 755.1
10 0.4492 0.7656 0.000000 0.0000 756.6 643.7
11 0.4824 0.6875 0.000000 0.0000 707.0 597.8
12 0.4199 0.7031 0.000000 0.0000 739.9 632.2
13 0.4805 0.8281 0.000000 0.0000 710.4 582.3
14 0.4180 0.7344 0.000000 0.0000 765.8 648.4
15 0.4805 0.7500 0.000000 0.0000 674.8 564.9
16 0.3770 0.6250 0.000000 0.0000 887.2 775.3
17 0.4922 0.7500 0.000000 -0.0000 831.8 712.7
18 0.3965 0.7031 0.000000 0.0000 880.0 779.4
19 0.4316 0.6875 0.000000 0.0000 683.0 582.3
20 0.4199 0.7188 0.000000 0.0000 637.4 534.5
21 0.5020 0.7812 0.000000 0.0000 857.8 751.1
22 0.4727 0.6719 0.000000 -0.0000 667.7 560.6
23 0.6172 0.8594 0.000000 0.0000 597.0 492.3
24 0.4648 0.6875 0.000000 0.0000 896.7 771.3
25 0.5918 0.8594 0.000000 -0.0000 597.1 471.1
26 0.4551 0.7031 0.000000 -0.0000 807.6 710.1
27 0.4316 0.6562 0.000000 0.0000 696.8 586.7
28 0.5566 0.7812 0.000000 -0.0000 784.6 682.0
29 0.5000 0.7188 0.000000 0.0000 707.0 608.3
30 0.4316 0.6719 0.000000 0.0000 610.6 494.2
31 0.4922 0.8125 0.000000 0.0000 589.5 486.7
32 0.4141 0.6562 0.000000 0.0000 750.3 643.1
33 0.4141 0.7656 0.000000 0.0000 777.1 651.1
34 0.5137 0.7500 0.000000 0.0000 979.9 866.6
35 0.5059 0.7812 0.000000 0.0000 713.4 609.5
36 0.4980 0.7969 0.000000 0.0000 775.4 660.8
37 0.4492 0.7500 0.000000 0.0000 757.6 655.7
38 0.5078 0.7812 0.000000 0.0000 692.2 581.6
39 0.4961 0.7656 0.000000 -0.0000 848.6 742.0
40 0.5488 0.8438 0.000000 -0.0000 893.1 787.3
41 0.5039 0.7656 0.000000 0.0000 885.2 785.7
42 0.4121 0.7188 0.000000 0.0000 1271.1 1141.7
43 0.4453 0.7500 0.000000 0.0000 1012.1 917.3
44 0.3984 0.7188 0.000000 0.0000 1307.2 1200.5
45 0.4492 0.7500 0.000000 0.0000 3038.2 2914.5
46 0.6309 0.7344 0.000000 0.0000 1619.0 1533.7
47 0.6699 0.8750 0.000000 0.0000 1133.4 1037.4

job_561774

Step Reward Pass@8 KL Loss Step Time (s) Gen Wait (s)
47 0.6641 0.8033 0.000000 -0.0000 2265.5 2157.0
48 0.6445 0.8438 0.000000 0.0000 1245.0 1145.6
49 0.5156 0.7344 0.000000 0.0000 1351.3 1239.3
50 0.4531 0.7969 0.000000 0.0000 749.7 624.8
51 0.4004 0.6719 0.000000 0.0000 1181.1 1081.8
52 0.5059 0.7969 0.000000 -0.0000 1068.0 959.9
53 0.4297 0.6875 0.000000 0.0000 890.2 780.5
54 0.4668 0.7500 0.000000 0.0000 1276.9 1155.4
55 0.4434 0.7500 0.000000 0.0000 1121.7 1011.1
56 0.4844 0.7188 0.000000 -0.0000 1158.0 1039.8
57 0.4473 0.7344 0.000000 0.0000 992.3 865.2
58 0.4473 0.6719 0.000000 0.0000 992.3 857.4
59 0.4648 0.7031 0.000000 -0.0000 986.5 884.0
60 0.4473 0.6719 0.000000 0.0000 1122.7 1016.5
61 0.4141 0.7188 0.000000 0.0000 1132.9 1038.0
62 0.4453 0.7812 0.000000 0.0000 1084.0 986.1
63 0.4414 0.7344 0.000000 0.0000 1180.4 1080.5
64 0.6113 0.8906 0.000000 -0.0000 1079.8 964.2
65 0.4746 0.7344 0.000000 0.0000 1055.0 931.4
66 0.4102 0.6250 0.000000 0.0000 1124.4 1024.0
67 0.4941 0.7969 0.000000 0.0000 1156.0 1049.4
68 0.4336 0.7031 0.000000 0.0000 1049.8 936.4
69 0.4688 0.7656 0.000000 0.0000 1134.5 1024.6
70 0.4570 0.7500 0.000000 0.0000 1080.6 967.3
71 0.5117 0.7344 0.000000 0.0000 1198.9 1097.3
72 0.5098 0.7656 0.000000 -0.0000 1520.3 1421.6
73 0.4277 0.6562 0.000000 -0.0002 955.5 837.1
74 0.3770 0.6094 0.000000 -0.0000 785.7 674.7
75 0.4434 0.7500 0.000000 -0.0006 2415.0 2317.4
76 0.5137 0.7031 0.000000 -0.0005 1109.8 1012.8
77 0.4316 0.7188 0.000000 -0.0002 1127.2 1025.8
78 0.4277 0.7500 0.000000 -0.0000 1029.7 934.0
79 0.4590 0.6562 0.000000 -0.0000 1479.8 1383.9
80 0.5898 0.8594 0.000000 0.0000 1469.6 1325.4

Timing Analysis

Average Time Breakdown (% of step time)

Component Avg % of Step Time
wait_for_generation_buffer 87.8%
run_training 9.4%
train_critic_and_policy 8.1%
policy_train 8.1%
sync_weights 2.6%
fwd_logprobs_values_reward 1.3%
save_checkpoints 1.2%
save_hf_model 0.7%
convert_to_training_input 0.2%
cleanup_old_checkpoints 0.1%
compute_advantages_and_returns 0.0%

Cross-Log Comparison

Log Avg Reward Pass@8 Step Time (s) Gen Wait Time (s) Avg Tokens Staleness
job_561773 0.4832 0.7470 858.4760 749.9947 1716.6914 2.8946
job_561774 0.4752 0.7364 1193.2385 1083.8262 1629.2433 2.6268

vLLM Inference Engine Analysis

Metrics from vLLM stat loggers (V1LoggingStatLoggerFixed).

Note

: Ray deduplicates similar log messages with [repeated Nx across cluster], so we typically capture stats from one engine per timestamp. The stats shown are per-engine values. Multiply by num_inference_engines for cluster-wide estimates.

Summary by Log (Per-Engine Stats)

Log Avg Running/Engine Avg Waiting/Engine Avg Gen Throughput/Engine Avg KV Cache % Avg Prefix Hit %
job_561773 11.1 7.4 470.8 tok/s 48.0% 84.3%
job_561774 11.8 8.5 493.2 tok/s 51.0% 82.4%

Utilization Analysis (Per-Engine)

Key indicators of inference engine utilization:

  • Running requests/engine: Concurrent requests being processed by each engine
  • Waiting requests: Requests queued (0 = engine not saturated, has spare capacity)
  • Generation throughput: Decode tokens/sec per engine
    • 8B model on H100 can do 1000+ tok/s when saturated
    • If seeing <300 tok/s with 0 waiting, engine is starved for requests

job_561773

  • Running requests/engine: avg=11.1, max=24
  • Waiting requests: avg=7.4, max=382
  • Generation throughput/engine: avg=470.8 tok/s, max=988.5 tok/s
  • KV cache usage: avg=48.0%
  • Prefix cache hit rate: avg=84.3%
  • Well-utilized: Engines saturated (waiting > 0)

job_561774

  • Running requests/engine: avg=11.8, max=24
  • Waiting requests: avg=8.5, max=274
  • Generation throughput/engine: avg=493.2 tok/s, max=972.1 tok/s
  • KV cache usage: avg=51.0%
  • Prefix cache hit rate: avg=82.4%
  • Well-utilized: Engines saturated (waiting > 0)

Trial-Level Analysis (from result.json)

Total trials parsed: 43795

Turn Count Statistics

Metric Value
Mean 2.1
Median 2.0
Std 0.7
Min 1
Max 18
Count 43632

Exception Distribution

Exception Type Count %
No exception 38610 88.2%
AgentTimeoutError 4917 11.2%
AgentSetupTimeoutError 133 0.3%
VerifierTimeoutError 56 0.1%
DaytonaError 44 0.1%
ContextLengthExceededError 24 0.1%
DaytonaAuthenticationError 9 0.0%
RuntimeError 1 0.0%
Timeout 1 0.0%

Turn Count by Exception Type

Exception Type Mean Turns Median Turns Count
DaytonaError 2.6 2.0 15
ContextLengthExceededError 2.2 1.0 24
No exception 2.2 2.0 38610
VerifierTimeoutError 2.0 2.0 56
AgentTimeoutError 1.3 1.0 4917
DaytonaAuthenticationError 1.2 1.0 9
Timeout 1.0 1.0 1

Turn Count by Outcome

Outcome Mean Turns Median Turns Count
Success 2.2 2.0 21100
Failure 2.0 2.0 22382

Reward Summary

  • Mean reward: 0.4853
  • Success rate: 48.5%
  • Trials with reward data: 43482