2055 lines
411 KiB
Plaintext
2055 lines
411 KiB
Plaintext
|
||
The following have been reloaded with a version change:
|
||
1) GCCcore/.14.3.0 => GCCcore/14.3.0
|
||
|
||
|
||
Lmod is automatically replacing "GCC/14.3.0" with
|
||
"nvidia-compilers/25.9-CUDA-13".
|
||
|
||
Deactivating conda environment: /e/scratch/jureap59/feuer1/miniforge3/envs/otagent
|
||
Activating RL environment: /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl
|
||
Python executable: /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/bin/python
|
||
Python path check: /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/bin/python
|
||
[ray] RAY_TMPDIR=/tmp/ray/ray_630129
|
||
[triton_cache] Triton cache: /tmp/triton_cache_feuer1_630129
|
||
[triton_cache] TorchInductor cache: /tmp/torchinductor_cache_feuer1_630129
|
||
[proxy] ✓ Found proxychains binary at /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4
|
||
[proxy] Setting up SSH tunnel to jpbl-s01-01
|
||
[proxy] SSH key: /e/home/jusers/feuer1/jupiter/.ssh/authorized_keys/id_ed25519_jsc
|
||
[proxy] Tunnel port: 7003
|
||
[proxy] Node IP: 10.128.32.33 (workers will connect here)
|
||
[proxy] ✓ SSH tunnel started successfully
|
||
[proxy] ✓ Generated proxychains config at /e/home/jusers/feuer1/jupiter/.proxychains/proxychains_630129.conf
|
||
[proxy] - Internal traffic (10.x.x.x, 172.x.x.x, 169.254.x.x) → DIRECT
|
||
[proxy] - External traffic (internet) → PROXY via tunnel
|
||
[proxy] ✓ Daytona timeout settings configured
|
||
[proxy] Testing proxy connectivity...
|
||
[proxychains] config file found: /e/home/jusers/feuer1/jupiter/.proxychains/proxychains_630129.conf
|
||
[proxychains] preloading /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/lib/libproxychains4.so
|
||
[proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e
|
||
[proxy] ✓ Proxy connectivity test passed (huggingface.co reachable via wrapped binary)
|
||
[proxy] ⚠ Tunnel not accessible at 10.128.32.33:7003 (workers may fail)
|
||
[proxy] ✓ Proxy setup complete (using wrapped binary for Ray workers)
|
||
[container_runtime] Using cloud backend: daytona (no local container setup)
|
||
=== Universal RL Training Runner ===
|
||
Config: /e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqmean-arm0/configs/ablation-pymethods2test-seqmean-arm0_rl_config.json
|
||
Working directory: /e/scratch/jureap59/feuer1/OpenThoughts-Agent
|
||
Python: /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/bin/python
|
||
Python version: Python 3.12.12
|
||
UV_USE_IO_URING: 0
|
||
Proxy: DISABLED (direct internet or not configured)
|
||
========================================
|
||
=== RLJobRunner: ablation-pymethods2test-seqmean-arm0 ===
|
||
[wandb_utils] Fixing permissions on: /e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqmean-arm0/wandb
|
||
[wandb_utils] WandB directory ready: /e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqmean-arm0/wandb
|
||
HF_TOKEN=****pDbg
|
||
HF_HUB_CACHE=/e/data1/datasets/playground/ot-baf/hf_hub
|
||
SUPABASE_URL=https://rpzmyuapoqilpghynmza.s... (direct Supabase config)
|
||
Environment configured:
|
||
TENSOR_PARALLEL_SIZE=1
|
||
NUM_INFERENCE_ENGINES=56
|
||
POLICY_NUM_NODES=14
|
||
WANDB_DIR=/e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqmean-arm0/wandb
|
||
Starting Ray cluster with 14 nodes, 4 GPUs/node
|
||
Cleaning up existing Ray instances...
|
||
=== Starting Ray Cluster ===
|
||
Nodes: 14
|
||
GPUs per node: 4
|
||
CPUs per node: 288
|
||
Head node: jpbo-041-33 (10.128.32.33)
|
||
Ray port: 6379
|
||
============================
|
||
Starting Ray head on jpbo-041-33 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_head_jpbo-041-33.log)...
|
||
Command: srun --export=ALL,VLLM_HOST_IP=10.128.32.33 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-041-33 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --head --node-ip-address=10.128.32.33 --port=6379 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960
|
||
Started Ray head on jpbo-041-33
|
||
Starting Ray worker on jpbo-041-34 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-041-34.log)...
|
||
Command: srun --export=ALL,VLLM_HOST_IP=10.128.32.34 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-041-34 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.32.33:6379 --node-ip-address=10.128.32.34 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960
|
||
Started Ray worker 1 on jpbo-041-34
|
||
Starting Ray worker on jpbo-041-35 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-041-35.log)...
|
||
Command: srun --export=ALL,VLLM_HOST_IP=10.128.32.35 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-041-35 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.32.33:6379 --node-ip-address=10.128.32.35 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960
|
||
Started Ray worker 2 on jpbo-041-35
|
||
Starting Ray worker on jpbo-041-36 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-041-36.log)...
|
||
Command: srun --export=ALL,VLLM_HOST_IP=10.128.32.36 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-041-36 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.32.33:6379 --node-ip-address=10.128.32.36 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960
|
||
Started Ray worker 3 on jpbo-041-36
|
||
Starting Ray worker on jpbo-041-37 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-041-37.log)...
|
||
Command: srun --export=ALL,VLLM_HOST_IP=10.128.32.37 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-041-37 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.32.33:6379 --node-ip-address=10.128.32.37 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960
|
||
Started Ray worker 4 on jpbo-041-37
|
||
Starting Ray worker on jpbo-041-38 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-041-38.log)...
|
||
Command: srun --export=ALL,VLLM_HOST_IP=10.128.32.38 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-041-38 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.32.33:6379 --node-ip-address=10.128.32.38 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960
|
||
Started Ray worker 5 on jpbo-041-38
|
||
Starting Ray worker on jpbo-041-39 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-041-39.log)...
|
||
Command: srun --export=ALL,VLLM_HOST_IP=10.128.32.39 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-041-39 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.32.33:6379 --node-ip-address=10.128.32.39 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960
|
||
Started Ray worker 6 on jpbo-041-39
|
||
Starting Ray worker on jpbo-041-40 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-041-40.log)...
|
||
Command: srun --export=ALL,VLLM_HOST_IP=10.128.32.40 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-041-40 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.32.33:6379 --node-ip-address=10.128.32.40 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960
|
||
Started Ray worker 7 on jpbo-041-40
|
||
Starting Ray worker on jpbo-041-41 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-041-41.log)...
|
||
Command: srun --export=ALL,VLLM_HOST_IP=10.128.32.41 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-041-41 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.32.33:6379 --node-ip-address=10.128.32.41 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960
|
||
Started Ray worker 8 on jpbo-041-41
|
||
Starting Ray worker on jpbo-041-42 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-041-42.log)...
|
||
Command: srun --export=ALL,VLLM_HOST_IP=10.128.32.42 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-041-42 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.32.33:6379 --node-ip-address=10.128.32.42 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960
|
||
Started Ray worker 9 on jpbo-041-42
|
||
Starting Ray worker on jpbo-041-43 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-041-43.log)...
|
||
Command: srun --export=ALL,VLLM_HOST_IP=10.128.32.43 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-041-43 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.32.33:6379 --node-ip-address=10.128.32.43 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960
|
||
Started Ray worker 10 on jpbo-041-43
|
||
Starting Ray worker on jpbo-041-44 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-041-44.log)...
|
||
Command: srun --export=ALL,VLLM_HOST_IP=10.128.32.44 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-041-44 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.32.33:6379 --node-ip-address=10.128.32.44 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960
|
||
Started Ray worker 11 on jpbo-041-44
|
||
Starting Ray worker on jpbo-041-45 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-041-45.log)...
|
||
Command: srun --export=ALL,VLLM_HOST_IP=10.128.32.45 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-041-45 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.32.33:6379 --node-ip-address=10.128.32.45 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960
|
||
Started Ray worker 12 on jpbo-041-45
|
||
Starting Ray worker on jpbo-041-46 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-041-46.log)...
|
||
Command: srun --export=ALL,VLLM_HOST_IP=10.128.32.46 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-041-46 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.32.33:6379 --node-ip-address=10.128.32.46 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960
|
||
Started Ray worker 13 on jpbo-041-46
|
||
Waiting for cluster (56 GPUs, 14 nodes)...
|
||
Connecting to Ray at 10.128.32.33:6379 (expecting 14 nodes, 56.0 GPUs)
|
||
Ray connection established, polling for resources...
|
||
[Ray wait] nodes=12/14 GPUs=48.0/56.0 resources={'accelerator_type:GH200': 12.0, 'GPU': 48.0, 'CPU': 3456.0, 'memory': 9095679574016.0, 'object_store_memory': 515396075520.0, 'node:10.128.32.35': 1.0, 'node:10.128.32.42': 1.0, 'node:10.128.32.33': 1.0, 'node:__internal_head__': 1.0, 'node:10.128.32.34': 1.0, 'node:10.128.32.40': 1.0, 'node:10.128.32.43': 1.0, 'node:10.128.32.38': 1.0, 'node:10.128.32.44': 1.0, 'node:10.128.32.36': 1.0, 'node:10.128.32.41': 1.0, 'node:10.128.32.45': 1.0, 'node:10.128.32.39': 1.0}
|
||
[Ray wait] nodes=13/14 GPUs=52.0/56.0 resources={'accelerator_type:GH200': 13.0, 'GPU': 52.0, 'object_store_memory': 558345748480.0, 'memory': 9872149315584.0, 'CPU': 3744.0, 'node:10.128.32.35': 1.0, 'node:10.128.32.42': 1.0, 'node:10.128.32.33': 1.0, 'node:__internal_head__': 1.0, 'node:10.128.32.34': 1.0, 'node:10.128.32.40': 1.0, 'node:10.128.32.43': 1.0, 'node:10.128.32.38': 1.0, 'node:10.128.32.44': 1.0, 'node:10.128.32.46': 1.0, 'node:10.128.32.36': 1.0, 'node:10.128.32.41': 1.0, 'node:10.128.32.45': 1.0, 'node:10.128.32.39': 1.0}
|
||
[Ray wait] nodes=13/14 GPUs=52.0/56.0 resources={'accelerator_type:GH200': 13.0, 'GPU': 52.0, 'object_store_memory': 558345748480.0, 'memory': 9872149315584.0, 'CPU': 3744.0, 'node:10.128.32.35': 1.0, 'node:10.128.32.42': 1.0, 'node:10.128.32.33': 1.0, 'node:__internal_head__': 1.0, 'node:10.128.32.34': 1.0, 'node:10.128.32.40': 1.0, 'node:10.128.32.43': 1.0, 'node:10.128.32.38': 1.0, 'node:10.128.32.44': 1.0, 'node:10.128.32.46': 1.0, 'node:10.128.32.36': 1.0, 'node:10.128.32.41': 1.0, 'node:10.128.32.45': 1.0, 'node:10.128.32.39': 1.0}
|
||
[Ray wait] nodes=13/14 GPUs=52.0/56.0 resources={'accelerator_type:GH200': 13.0, 'GPU': 52.0, 'CPU': 3744.0, 'memory': 9872149315584.0, 'object_store_memory': 558345748480.0, 'node:10.128.32.35': 1.0, 'node:10.128.32.42': 1.0, 'node:10.128.32.33': 1.0, 'node:__internal_head__': 1.0, 'node:10.128.32.34': 1.0, 'node:10.128.32.40': 1.0, 'node:10.128.32.43': 1.0, 'node:10.128.32.38': 1.0, 'node:10.128.32.44': 1.0, 'node:10.128.32.46': 1.0, 'node:10.128.32.36': 1.0, 'node:10.128.32.41': 1.0, 'node:10.128.32.45': 1.0, 'node:10.128.32.39': 1.0}
|
||
[Ray wait] nodes=14/14 GPUs=56.0/56.0 resources={'accelerator_type:GH200': 14.0, 'GPU': 56.0, 'object_store_memory': 601295421440.0, 'memory': 10634851450880.0, 'node:10.128.32.41': 1.0, 'CPU': 4032.0, 'node:10.128.32.45': 1.0, 'node:10.128.32.34': 1.0, 'node:10.128.32.35': 1.0, 'node:10.128.32.43': 1.0, 'node:10.128.32.44': 1.0, 'node:10.128.32.40': 1.0, 'node:10.128.32.36': 1.0, 'node:10.128.32.37': 1.0, 'node:10.128.32.39': 1.0, 'node:10.128.32.33': 1.0, 'node:__internal_head__': 1.0, 'node:10.128.32.42': 1.0, 'node:10.128.32.46': 1.0, 'node:10.128.32.38': 1.0}
|
||
✓ Ray cluster ready
|
||
=== Ray Cluster Ready ===
|
||
Address: 10.128.32.33:6379
|
||
Total GPUs: 56
|
||
=========================
|
||
Ray cluster ready at 10.128.32.33:6379
|
||
Total GPUs available: 56
|
||
[RLJobRunner] Pinggy check: url=False, token=False, needs_tunnel=False (agent=terminus-2, env=daytona)
|
||
[RLJobRunner] No Pinggy tunnel needed, using local vLLM
|
||
|
||
Running SkyRL:
|
||
Python: /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/bin/python
|
||
Entrypoint: examples.terminal_bench.entrypoints.main_tbench
|
||
Args: 120 Hydra arguments
|
||
Working dir: /e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train
|
||
Using proxychains binary: /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4
|
||
|
||
Executing command with srun: /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f $PROXYCHAINS_CONF_FILE /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/bin/python -m examples.terminal_bench.entrypoints.main_tbench +terminal_bench_config=terminal_bench trainer.strategy=fsdp2 trainer.algorithm.advantage_estimator=rloo_n trainer.algorithm.use_kl_loss=false trainer.algorithm.kl_loss_coef=0.0 trainer.algorithm.eps_clip_low=0.2 trainer.algorithm.eps_clip_high=0.05 trainer.algorithm.loss_reduction=sequence_mean trainer.epochs=2 trainer.max_steps=80 trainer.update_epochs_per_batch=1 trainer.train_batch_size=64 trainer.policy_mini_batch_size=64 trainer.eval_batch_size=64 trainer.micro_forward_batch_size_per_gpu=4 trainer.micro_train_batch_size_per_gpu=1 trainer.max_prompt_length=999999 trainer.eval_interval=999999 trainer.eval_before_train=false trainer.ckpt_interval=2 trainer.resume_mode=latest trainer.hf_save_interval=5 ++trainer.hf_hub_repo_id=laion/ablation-pymethods2test-seqmean-arm0 ++trainer.hf_hub_private=false ++trainer.hf_hub_revision=main ++trainer.enable_db_registration=false trainer.project_name=OpenThoughts-Agent trainer.log_level=INFO trainer.tracker_commit_each_step=true trainer.logger=console trainer.run_name=ablation-pymethods2test-seqmean-arm0 trainer.ckpt_path=/e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqmean-arm0/ablation-pymethods2test-seqmean-arm0/checkpoints trainer.export_path=/e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqmean-arm0/ablation-pymethods2test-seqmean-arm0/exports trainer.policy.optimizer_config.lr=8e-6 trainer.policy.optimizer_config.weight_decay=0.0 trainer.policy.optimizer_config.adam_betas=[0.9,0.999] trainer.policy.optimizer_config.max_grad_norm=0.9 trainer.policy.fsdp_config.cpu_offload=false trainer.policy.fsdp_config.reshard_after_forward=true trainer.policy.fsdp_config.fsdp_size=4 trainer.policy.model.path=/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6 trainer.ref.fsdp_config.cpu_offload=false trainer.ref.fsdp_config.reshard_after_forward=true trainer.ref.fsdp_config.fsdp_size=4 trainer.placement.colocate_all=false trainer.placement.policy_num_nodes=2 trainer.placement.ref_num_nodes=2 trainer.placement.policy_num_gpus_per_node=4 trainer.placement.ref_num_gpus_per_node=4 trainer.fully_async.max_staleness_steps=16 trainer.fully_async.num_parallel_generation_workers=338 generator.backend=vllm generator.timeout_multiplier=1.0 generator.model_dtype=bfloat16 generator.inference_engine_tensor_parallel_size=1 generator.num_inference_engines=48 generator.n_samples_per_prompt=8 generator.eval_n_samples_per_prompt=8 generator.gpu_memory_utilization=0.75 generator.max_num_seqs=24 generator.max_num_batched_tokens=65536 generator.enable_prefix_caching=true generator.enable_chunked_prefill=true generator.run_engines_locally=true generator.weight_sync_backend=nccl generator.async_engine=true generator.batched=false generator.enable_http_endpoint=true generator.enable_ray_prometheus_stats=false generator.vllm_stats_interval=1 generator.append_eos_token_after_stop_str_in_multi_turn=true generator.max_turns=999999 generator.sampling_params.max_generate_length=4096 generator.sampling_params.temperature=0.7 generator.sampling_params.top_p=0.95 generator.sampling_params.top_k=20 ++generator.engine_init_kwargs.max_model_len=32768 ++generator.engine_init_kwargs.custom_chat_template_chat_completion_path=chat_templates/qwen3_thinking_acc.jinja2 ++generator.engine_init_kwargs.served_model_name=0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6 data.train_data=["/e/scratch/jureap59/feuer1/tasks/exp_rpt_pymethods2test-large"] data.val_data=[] +terminal_bench_config.trials_dir=/e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqmean-arm0/ablation-pymethods2test-seqmean-arm0/trace_jobs +terminal_bench_config.harbor.name=terminus-2 +terminal_bench_config.harbor.max_episodes=999999 +terminal_bench_config.harbor.enable_summarize=false +terminal_bench_config.harbor.store_all_messages=true +terminal_bench_config.harbor.trajectory_config.raw_content=true +terminal_bench_config.harbor.enable_episode_logging=false +terminal_bench_config.harbor.record_terminal_session=false +terminal_bench_config.harbor.enable_pane_logging=false +terminal_bench_config.harbor.strict_json_parser=true +terminal_bench_config.harbor.interleaved_thinking=true +terminal_bench_config.harbor.extra_body.chat_template_kwargs.enable_thinking=true +terminal_bench_config.harbor.override_timeout_sec=900 +terminal_bench_config.harbor.override_cpus=1 +terminal_bench_config.harbor.override_memory_mb=2048 +terminal_bench_config.harbor.override_storage_mb=2048 +terminal_bench_config.harbor.auto_snapshot=true +terminal_bench_config.harbor.verifier_override_timeout_sec=120 +terminal_bench_config.harbor.max_retries=3 +terminal_bench_config.harbor.min_wait_sec=60.0 +terminal_bench_config.harbor.max_wait_sec=600.0 +terminal_bench_config.harbor.wait_multiplier=2.0 +terminal_bench_config.harbor.exclude_exceptions=["VerifierTimeoutError","VerifierRuntimeError","RewardFileNotFoundError","RewardFileEmptyError","VerifierOutputParseError"] +terminal_bench_config.harbor.n_concurrent_trials=675 +terminal_bench_config.harbor.log_level=INFO +terminal_bench_config.harbor.enable_reward_shaping=false +terminal_bench_config.harbor.enable_error_classification=true +terminal_bench_config.harbor.mask_exceptions=["DaytonaError","EnvironmentStartTimeoutError","NetworkError","ConnectionError","RewardFileNotFoundError","RewardFileEmptyError","AgentEnvironmentTimeoutError","ContextLengthExceededError"] +terminal_bench_config.harbor.default_error_treatment=zero +terminal_bench_config.harbor.passthrough_exceptions=["AgentTimeoutError"] +terminal_bench_config.harbor.zero_exceptions=[] +terminal_bench_config.model_info.max_input_tokens=32000 +terminal_bench_config.model_info.max_output_tokens=4096 +terminal_bench_config.archiving.enabled=false +terminal_bench_config.trace_upload.enabled=true +terminal_bench_config.trace_upload.repo_org=DCAgent +terminal_bench_config.trace_upload.episodes=last +terminal_bench_config.trace_upload.dataset_type=SFT +terminal_bench_config.trace_upload.cleanup=true
|
||
[proxychains] config file found: /e/home/jusers/feuer1/jupiter/.proxychains/proxychains_630129.conf
|
||
[proxychains] preloading /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/lib/libproxychains4.so
|
||
[proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e
|
||
2026-06-07 16:02:29.538 | INFO | skyrl_train.utils.utils:prepare_runtime_environment:753 - Exporting wandb api key to ray runtime env
|
||
2026-06-07 16:02:29.538 | INFO | skyrl_train.utils.utils:prepare_runtime_environment:772 - Exporting RAY_ADDRESS to ray runtime env
|
||
2026-06-07 16:02:29.538 | INFO | skyrl_train.utils.utils:prepare_runtime_environment:797 - Exporting `NCCL_SOCKET_IFNAME` to ray runtime env: ib0
|
||
2026-06-07 16:02:29.538 | INFO | skyrl_train.utils.utils:prepare_runtime_environment:797 - Exporting `NCCL_SOCKET_FAMILY` to ray runtime env: AF_INET
|
||
2026-06-07 16:02:29.538 | INFO | skyrl_train.utils.utils:prepare_runtime_environment:797 - Exporting `NCCL_DEBUG` to ray runtime env: WARN
|
||
2026-06-07 16:02:29,539 INFO worker.py:1680 -- Using address 10.128.32.33:6379 set in the environment variable RAY_ADDRESS
|
||
2026-06-07 16:02:29,574 INFO worker.py:1821 -- Connecting to existing Ray cluster at address: 10.128.32.33:6379...
|
||
2026-06-07 16:02:29,585 INFO worker.py:2007 -- Connected to Ray cluster.
|
||
/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/ray/_private/worker.py:2046: FutureWarning: Tip: In future versions of Ray, Ray will no longer override accelerator visible devices env var if num_gpus=0 or num_gpus=None (default). To enable this behavior and turn off this error message, set RAY_ACCEL_ENV_VAR_OVERRIDE_ON_ZERO=0
|
||
warnings.warn(
|
||
[33m(raylet, ip=10.128.32.41)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e
|
||
2026-06-07 16:02:32.338 | INFO | skyrl_train.utils.ppo_utils:sync_registries:546 - Synced registries to ray actor
|
||
[36m(skyrl_entrypoint pid=391889)[0m [32m2026-06-07 16:02:41.646[0m | [1mINFO [0m | [36mskyrl_train.entrypoints.main_base[0m:[36m_configure_log_level[0m:[36m212[0m - [1mSkyRL log level set to: INFO[0m
|
||
[33m(raylet)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 17x across cluster] (Ray deduplicates logs by default. Set RAY_DEDUP_LOGS=0 to disable log deduplication, or see https://docs.ray.io/en/master/ray-observability/user-guides/configure-logging.html#log-deduplication for more options.)[0m
|
||
[36m(skyrl_entrypoint pid=391889)[0m The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.
|
||
[36m(skyrl_entrypoint pid=391889)[0m [32m2026-06-07 16:02:42.007[0m | [1mINFO [0m | [36mexamples.terminal_bench.dataset[0m:[36m_load_data_files[0m:[36m40[0m - [1mLoading data from: /e/scratch/jureap59/feuer1/tasks/exp_rpt_pymethods2test-large[0m
|
||
[2026-06-07 16:02:59,789 E 391407 391863] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14
|
||
[36m(RegistryActor pid=2613231, ip=10.128.32.41)[0m [2026-06-07 16:03:00,820 E 2613231 2613271] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14
|
||
[36m(skyrl_entrypoint pid=391889)[0m [32m2026-06-07 16:03:39.460[0m | [1mINFO [0m | [36mexamples.terminal_bench.dataset[0m:[36m_load_data_files[0m:[36m50[0m - [1mFound 5000 valid task directories out of 5000 total directories[0m
|
||
[36m(skyrl_entrypoint pid=391889)[0m [32m2026-06-07 16:03:39.461[0m | [1mINFO [0m | [36mexamples.terminal_bench.dataset[0m:[36m__init__[0m:[36m27[0m - [1mTerminalBenchTaskDataset initialized with 5000 task paths[0m
|
||
[36m(skyrl_entrypoint pid=391889)[0m [32m2026-06-07 16:03:39.473[0m | [1mINFO [0m | [36mskyrl_train.entrypoints.main_base[0m:[36m_setup_trainer[0m:[36m405[0m - [1mdata:
|
||
[36m(skyrl_entrypoint pid=391889)[0m train_data:
|
||
[36m(skyrl_entrypoint pid=391889)[0m - /e/scratch/jureap59/feuer1/tasks/exp_rpt_pymethods2test-large
|
||
[36m(skyrl_entrypoint pid=391889)[0m val_data: []
|
||
[36m(skyrl_entrypoint pid=391889)[0m trainer:
|
||
[36m(skyrl_entrypoint pid=391889)[0m placement:
|
||
[36m(skyrl_entrypoint pid=391889)[0m colocate_all: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m colocate_policy_ref: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m policy_num_nodes: 2
|
||
[36m(skyrl_entrypoint pid=391889)[0m policy_num_gpus_per_node: 4
|
||
[36m(skyrl_entrypoint pid=391889)[0m critic_num_nodes: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m critic_num_gpus_per_node: 4
|
||
[36m(skyrl_entrypoint pid=391889)[0m ref_num_nodes: 2
|
||
[36m(skyrl_entrypoint pid=391889)[0m ref_num_gpus_per_node: 4
|
||
[36m(skyrl_entrypoint pid=391889)[0m policy_strict_spread_pg: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m policy_per_gpu_bundles: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m policy_force_cvd_mask: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m sequence_parallel_backend: ulysses
|
||
[36m(skyrl_entrypoint pid=391889)[0m strategy: fsdp2
|
||
[36m(skyrl_entrypoint pid=391889)[0m policy:
|
||
[36m(skyrl_entrypoint pid=391889)[0m model:
|
||
[36m(skyrl_entrypoint pid=391889)[0m path: /e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6
|
||
[36m(skyrl_entrypoint pid=391889)[0m lora:
|
||
[36m(skyrl_entrypoint pid=391889)[0m rank: 0
|
||
[36m(skyrl_entrypoint pid=391889)[0m alpha: 16
|
||
[36m(skyrl_entrypoint pid=391889)[0m dropout: 0
|
||
[36m(skyrl_entrypoint pid=391889)[0m lora_sync_path: /tmp/skyrl_lora_sync
|
||
[36m(skyrl_entrypoint pid=391889)[0m target_modules: all-linear
|
||
[36m(skyrl_entrypoint pid=391889)[0m exclude_modules: null
|
||
[36m(skyrl_entrypoint pid=391889)[0m deepspeed_config: ${deepspeed_config.train}
|
||
[36m(skyrl_entrypoint pid=391889)[0m optimizer_config:
|
||
[36m(skyrl_entrypoint pid=391889)[0m optimizer: AdamW
|
||
[36m(skyrl_entrypoint pid=391889)[0m lr: 8.0e-06
|
||
[36m(skyrl_entrypoint pid=391889)[0m adam_betas:
|
||
[36m(skyrl_entrypoint pid=391889)[0m - 0.9
|
||
[36m(skyrl_entrypoint pid=391889)[0m - 0.999
|
||
[36m(skyrl_entrypoint pid=391889)[0m weight_decay: 0.0
|
||
[36m(skyrl_entrypoint pid=391889)[0m max_grad_norm: 0.9
|
||
[36m(skyrl_entrypoint pid=391889)[0m offload_after_step: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m num_warmup_steps: 0
|
||
[36m(skyrl_entrypoint pid=391889)[0m scheduler: constant_with_warmup
|
||
[36m(skyrl_entrypoint pid=391889)[0m optimizer_kwargs: {}
|
||
[36m(skyrl_entrypoint pid=391889)[0m fsdp_config:
|
||
[36m(skyrl_entrypoint pid=391889)[0m cpu_offload: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m reshard_after_forward: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m fsdp_size: 4
|
||
[36m(skyrl_entrypoint pid=391889)[0m expert_model_parallel_size: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m expert_tensor_parallel_size: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m moe_token_dispatcher_type: alltoall
|
||
[36m(skyrl_entrypoint pid=391889)[0m moe_router_replay: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m moe_grouped_gemm: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m ep_comm_backend: torch
|
||
[36m(skyrl_entrypoint pid=391889)[0m deepep_num_sms: 20
|
||
[36m(skyrl_entrypoint pid=391889)[0m deepep_token_chunk_size: null
|
||
[36m(skyrl_entrypoint pid=391889)[0m sequence_parallel_size: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m use_torch_compile: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m record_memory: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m megatron_config:
|
||
[36m(skyrl_entrypoint pid=391889)[0m tensor_model_parallel_size: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m pipeline_model_parallel_size: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m context_parallel_size: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m expert_model_parallel_size: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m expert_tensor_parallel_size: null
|
||
[36m(skyrl_entrypoint pid=391889)[0m ddp_config:
|
||
[36m(skyrl_entrypoint pid=391889)[0m grad_reduce_in_fp32: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m overlap_grad_reduce: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m overlap_param_gather: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m average_in_collective: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m model_config_kwargs: {}
|
||
[36m(skyrl_entrypoint pid=391889)[0m torch_profiler_config:
|
||
[36m(skyrl_entrypoint pid=391889)[0m enable: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m ranks: []
|
||
[36m(skyrl_entrypoint pid=391889)[0m save_path: null
|
||
[36m(skyrl_entrypoint pid=391889)[0m optimizer_config_kwargs:
|
||
[36m(skyrl_entrypoint pid=391889)[0m overlap_cpu_optimizer_d2h_h2d: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m use_precision_aware_optimizer: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m optimizer_cpu_offload: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m optimizer_offload_fraction: 0.0
|
||
[36m(skyrl_entrypoint pid=391889)[0m transformer_config_kwargs:
|
||
[36m(skyrl_entrypoint pid=391889)[0m recompute_granularity: full
|
||
[36m(skyrl_entrypoint pid=391889)[0m recompute_modules:
|
||
[36m(skyrl_entrypoint pid=391889)[0m - core_attn
|
||
[36m(skyrl_entrypoint pid=391889)[0m recompute_method: uniform
|
||
[36m(skyrl_entrypoint pid=391889)[0m recompute_num_layers: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m empty_cuda_cache: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m ref:
|
||
[36m(skyrl_entrypoint pid=391889)[0m model:
|
||
[36m(skyrl_entrypoint pid=391889)[0m path: ${trainer.policy.model.path}
|
||
[36m(skyrl_entrypoint pid=391889)[0m sequence_parallel_size: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m deepspeed_config: ${deepspeed_config.eval}
|
||
[36m(skyrl_entrypoint pid=391889)[0m fsdp_config:
|
||
[36m(skyrl_entrypoint pid=391889)[0m cpu_offload: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m reshard_after_forward: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m fsdp_size: 4
|
||
[36m(skyrl_entrypoint pid=391889)[0m expert_model_parallel_size: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m expert_tensor_parallel_size: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m moe_token_dispatcher_type: alltoall
|
||
[36m(skyrl_entrypoint pid=391889)[0m moe_router_replay: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m moe_grouped_gemm: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m ep_comm_backend: torch
|
||
[36m(skyrl_entrypoint pid=391889)[0m deepep_num_sms: 20
|
||
[36m(skyrl_entrypoint pid=391889)[0m deepep_token_chunk_size: null
|
||
[36m(skyrl_entrypoint pid=391889)[0m megatron_config:
|
||
[36m(skyrl_entrypoint pid=391889)[0m tensor_model_parallel_size: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m pipeline_model_parallel_size: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m context_parallel_size: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m expert_model_parallel_size: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m expert_tensor_parallel_size: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m model_config_kwargs: {}
|
||
[36m(skyrl_entrypoint pid=391889)[0m transformer_config_kwargs: {}
|
||
[36m(skyrl_entrypoint pid=391889)[0m critic:
|
||
[36m(skyrl_entrypoint pid=391889)[0m model:
|
||
[36m(skyrl_entrypoint pid=391889)[0m path: null
|
||
[36m(skyrl_entrypoint pid=391889)[0m lora:
|
||
[36m(skyrl_entrypoint pid=391889)[0m rank: 0
|
||
[36m(skyrl_entrypoint pid=391889)[0m alpha: 16
|
||
[36m(skyrl_entrypoint pid=391889)[0m dropout: 0
|
||
[36m(skyrl_entrypoint pid=391889)[0m target_modules: all-linear
|
||
[36m(skyrl_entrypoint pid=391889)[0m exclude_modules: null
|
||
[36m(skyrl_entrypoint pid=391889)[0m deepspeed_config: ${deepspeed_config.train}
|
||
[36m(skyrl_entrypoint pid=391889)[0m optimizer_config:
|
||
[36m(skyrl_entrypoint pid=391889)[0m optimizer: AdamW
|
||
[36m(skyrl_entrypoint pid=391889)[0m lr: 5.0e-06
|
||
[36m(skyrl_entrypoint pid=391889)[0m adam_betas:
|
||
[36m(skyrl_entrypoint pid=391889)[0m - 0.9
|
||
[36m(skyrl_entrypoint pid=391889)[0m - 0.999
|
||
[36m(skyrl_entrypoint pid=391889)[0m weight_decay: 0.01
|
||
[36m(skyrl_entrypoint pid=391889)[0m max_grad_norm: 1.0
|
||
[36m(skyrl_entrypoint pid=391889)[0m offload_after_step: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m num_warmup_steps: 0
|
||
[36m(skyrl_entrypoint pid=391889)[0m scheduler: constant_with_warmup
|
||
[36m(skyrl_entrypoint pid=391889)[0m optimizer_kwargs: {}
|
||
[36m(skyrl_entrypoint pid=391889)[0m fsdp_config:
|
||
[36m(skyrl_entrypoint pid=391889)[0m cpu_offload: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m reshard_after_forward: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m fsdp_size: -1
|
||
[36m(skyrl_entrypoint pid=391889)[0m expert_model_parallel_size: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m expert_tensor_parallel_size: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m moe_token_dispatcher_type: alltoall
|
||
[36m(skyrl_entrypoint pid=391889)[0m moe_router_replay: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m moe_grouped_gemm: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m ep_comm_backend: torch
|
||
[36m(skyrl_entrypoint pid=391889)[0m deepep_num_sms: 20
|
||
[36m(skyrl_entrypoint pid=391889)[0m deepep_token_chunk_size: null
|
||
[36m(skyrl_entrypoint pid=391889)[0m sequence_parallel_size: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m algorithm:
|
||
[36m(skyrl_entrypoint pid=391889)[0m advantage_estimator: rloo_n
|
||
[36m(skyrl_entrypoint pid=391889)[0m kl_ctrl:
|
||
[36m(skyrl_entrypoint pid=391889)[0m type: fixed
|
||
[36m(skyrl_entrypoint pid=391889)[0m kl_target: 0.1
|
||
[36m(skyrl_entrypoint pid=391889)[0m horizon: 10000
|
||
[36m(skyrl_entrypoint pid=391889)[0m kl_estimator_type: k3
|
||
[36m(skyrl_entrypoint pid=391889)[0m use_kl_estimator_k3: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m use_abs_kl: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m use_kl_in_reward: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m use_kl_loss: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m kl_loss_coef: 0.0
|
||
[36m(skyrl_entrypoint pid=391889)[0m use_entropy_loss: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m entropy_loss_coef: 0.01
|
||
[36m(skyrl_entrypoint pid=391889)[0m advantage_batch_normalize: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m value_head_prefix: value_head
|
||
[36m(skyrl_entrypoint pid=391889)[0m policy_loss_type: regular
|
||
[36m(skyrl_entrypoint pid=391889)[0m loss_reduction: sequence_mean
|
||
[36m(skyrl_entrypoint pid=391889)[0m global_loss_denom: null
|
||
[36m(skyrl_entrypoint pid=391889)[0m grpo_norm_by_std: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m rloo_n_min_group_size: 4
|
||
[36m(skyrl_entrypoint pid=391889)[0m rloo_n_filter_zero_reward_groups: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m lambd: 1.0
|
||
[36m(skyrl_entrypoint pid=391889)[0m gamma: 1.0
|
||
[36m(skyrl_entrypoint pid=391889)[0m eps_clip_low: 0.2
|
||
[36m(skyrl_entrypoint pid=391889)[0m eps_clip_high: 0.05
|
||
[36m(skyrl_entrypoint pid=391889)[0m clip_ratio_c: 3.0
|
||
[36m(skyrl_entrypoint pid=391889)[0m tis_imp_ratio_cap: -1.0
|
||
[36m(skyrl_entrypoint pid=391889)[0m use_tis: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m sapo:
|
||
[36m(skyrl_entrypoint pid=391889)[0m tau_pos: 1.0
|
||
[36m(skyrl_entrypoint pid=391889)[0m tau_neg: 1.05
|
||
[36m(skyrl_entrypoint pid=391889)[0m value_clip: 0.2
|
||
[36m(skyrl_entrypoint pid=391889)[0m dynamic_sampling:
|
||
[36m(skyrl_entrypoint pid=391889)[0m type: null
|
||
[36m(skyrl_entrypoint pid=391889)[0m max_sample_batches: 30
|
||
[36m(skyrl_entrypoint pid=391889)[0m min_replace_ratio: 0.3
|
||
[36m(skyrl_entrypoint pid=391889)[0m clip_cov:
|
||
[36m(skyrl_entrypoint pid=391889)[0m clip_ratio: 0.0002
|
||
[36m(skyrl_entrypoint pid=391889)[0m clip_cov_lb: 1.0
|
||
[36m(skyrl_entrypoint pid=391889)[0m clip_cov_ub: 5.0
|
||
[36m(skyrl_entrypoint pid=391889)[0m kl_cov:
|
||
[36m(skyrl_entrypoint pid=391889)[0m kl_cov_frac: 0.2
|
||
[36m(skyrl_entrypoint pid=391889)[0m ppo_kl_coef: 1.0
|
||
[36m(skyrl_entrypoint pid=391889)[0m cispo:
|
||
[36m(skyrl_entrypoint pid=391889)[0m cispo_eps_clip_low: 0
|
||
[36m(skyrl_entrypoint pid=391889)[0m cispo_eps_clip_high: 5
|
||
[36m(skyrl_entrypoint pid=391889)[0m z_clip:
|
||
[36m(skyrl_entrypoint pid=391889)[0m enabled: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m alpha: 0.97
|
||
[36m(skyrl_entrypoint pid=391889)[0m z_thresh: 2.5
|
||
[36m(skyrl_entrypoint pid=391889)[0m warmup_steps: 25
|
||
[36m(skyrl_entrypoint pid=391889)[0m mode: zscore
|
||
[36m(skyrl_entrypoint pid=391889)[0m clip_option: adaptive_scaling
|
||
[36m(skyrl_entrypoint pid=391889)[0m clip_factor: 1.0
|
||
[36m(skyrl_entrypoint pid=391889)[0m skip_update_on_spike: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m stale_clip:
|
||
[36m(skyrl_entrypoint pid=391889)[0m enabled: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m alpha: 0.3
|
||
[36m(skyrl_entrypoint pid=391889)[0m entropy_threshold: 0.15
|
||
[36m(skyrl_entrypoint pid=391889)[0m entropy_window: 10
|
||
[36m(skyrl_entrypoint pid=391889)[0m min_lr_scale: 0.1
|
||
[36m(skyrl_entrypoint pid=391889)[0m max_seq_len: 1004095
|
||
[36m(skyrl_entrypoint pid=391889)[0m fully_async:
|
||
[36m(skyrl_entrypoint pid=391889)[0m max_staleness_steps: 16
|
||
[36m(skyrl_entrypoint pid=391889)[0m num_parallel_generation_workers: 338
|
||
[36m(skyrl_entrypoint pid=391889)[0m gradient_checkpointing: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m gradient_checkpointing_use_reentrant: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m seed: 42
|
||
[36m(skyrl_entrypoint pid=391889)[0m resume_mode: latest
|
||
[36m(skyrl_entrypoint pid=391889)[0m resume_path: null
|
||
[36m(skyrl_entrypoint pid=391889)[0m ckpt_path: /e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqmean-arm0/ablation-pymethods2test-seqmean-arm0/checkpoints
|
||
[36m(skyrl_entrypoint pid=391889)[0m max_ckpts_to_keep: -1
|
||
[36m(skyrl_entrypoint pid=391889)[0m ckpt_interval: 2
|
||
[36m(skyrl_entrypoint pid=391889)[0m hf_save_interval: 5
|
||
[36m(skyrl_entrypoint pid=391889)[0m hf_upload_mode: latest
|
||
[36m(skyrl_entrypoint pid=391889)[0m export_path: /e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqmean-arm0/ablation-pymethods2test-seqmean-arm0/exports
|
||
[36m(skyrl_entrypoint pid=391889)[0m bf16: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m epochs: 2
|
||
[36m(skyrl_entrypoint pid=391889)[0m max_steps: 80
|
||
[36m(skyrl_entrypoint pid=391889)[0m update_epochs_per_batch: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m train_batch_size: 64
|
||
[36m(skyrl_entrypoint pid=391889)[0m policy_mini_batch_size: 64
|
||
[36m(skyrl_entrypoint pid=391889)[0m critic_mini_batch_size: 256
|
||
[36m(skyrl_entrypoint pid=391889)[0m micro_train_batch_size_per_gpu: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m micro_forward_batch_size_per_gpu: 4
|
||
[36m(skyrl_entrypoint pid=391889)[0m update_ref_every_epoch: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m use_sample_packing: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m eval_batch_size: 64
|
||
[36m(skyrl_entrypoint pid=391889)[0m eval_before_train: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m eval_interval: 999999
|
||
[36m(skyrl_entrypoint pid=391889)[0m max_prompt_length: 999999
|
||
[36m(skyrl_entrypoint pid=391889)[0m flash_attn: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m disable_fast_tokenizer: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m target_modules: null
|
||
[36m(skyrl_entrypoint pid=391889)[0m exclude_modules: null
|
||
[36m(skyrl_entrypoint pid=391889)[0m project_name: OpenThoughts-Agent
|
||
[36m(skyrl_entrypoint pid=391889)[0m run_name: ablation-pymethods2test-seqmean-arm0
|
||
[36m(skyrl_entrypoint pid=391889)[0m logger: console
|
||
[36m(skyrl_entrypoint pid=391889)[0m tracker_commit_each_step: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m dump_data_batch: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m dump_eval_results: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m log_level: INFO
|
||
[36m(skyrl_entrypoint pid=391889)[0m rope_scaling: null
|
||
[36m(skyrl_entrypoint pid=391889)[0m rope_theta: null
|
||
[36m(skyrl_entrypoint pid=391889)[0m step_wise_training: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m hf_hub_repo_id: laion/ablation-pymethods2test-seqmean-arm0
|
||
[36m(skyrl_entrypoint pid=391889)[0m hf_hub_private: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m hf_hub_revision: main
|
||
[36m(skyrl_entrypoint pid=391889)[0m enable_db_registration: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m generator:
|
||
[36m(skyrl_entrypoint pid=391889)[0m model_name: ${trainer.policy.model.path}
|
||
[36m(skyrl_entrypoint pid=391889)[0m model_dtype: bfloat16
|
||
[36m(skyrl_entrypoint pid=391889)[0m timeout_multiplier: 1.0
|
||
[36m(skyrl_entrypoint pid=391889)[0m run_engines_locally: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m num_inference_engines: 48
|
||
[36m(skyrl_entrypoint pid=391889)[0m backend: vllm
|
||
[36m(skyrl_entrypoint pid=391889)[0m weight_sync_backend: nccl
|
||
[36m(skyrl_entrypoint pid=391889)[0m fuse_weights: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m weight_transfer_threshold_cuda_ipc_GB: 1.0
|
||
[36m(skyrl_entrypoint pid=391889)[0m inference_engine_tensor_parallel_size: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m inference_engine_pipeline_parallel_size: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m inference_engine_expert_parallel_size: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m inference_engine_data_parallel_size: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m n_samples_per_prompt: 8
|
||
[36m(skyrl_entrypoint pid=391889)[0m async_engine: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m batched: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m max_input_length: ${trainer.max_prompt_length}
|
||
[36m(skyrl_entrypoint pid=391889)[0m vllm_v1_disable_multiproc: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m enable_prefix_caching: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m enable_chunked_prefill: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m max_num_batched_tokens: 65536
|
||
[36m(skyrl_entrypoint pid=391889)[0m enforce_eager: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m fully_sharded_loras: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m enable_ray_prometheus_stats: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m vllm_stats_interval: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m gpu_memory_utilization: 0.75
|
||
[36m(skyrl_entrypoint pid=391889)[0m max_num_seqs: 24
|
||
[36m(skyrl_entrypoint pid=391889)[0m remote_inference_engine_urls:
|
||
[36m(skyrl_entrypoint pid=391889)[0m - 127.0.0.1:8001
|
||
[36m(skyrl_entrypoint pid=391889)[0m enable_http_endpoint: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m http_endpoint_host: 127.0.0.1
|
||
[36m(skyrl_entrypoint pid=391889)[0m http_endpoint_port: 8000
|
||
[36m(skyrl_entrypoint pid=391889)[0m max_turns: 999999
|
||
[36m(skyrl_entrypoint pid=391889)[0m chat_template:
|
||
[36m(skyrl_entrypoint pid=391889)[0m source: name
|
||
[36m(skyrl_entrypoint pid=391889)[0m name_or_path: null
|
||
[36m(skyrl_entrypoint pid=391889)[0m chat_template_kwargs: {}
|
||
[36m(skyrl_entrypoint pid=391889)[0m engine_init_kwargs:
|
||
[36m(skyrl_entrypoint pid=391889)[0m max_model_len: 32768
|
||
[36m(skyrl_entrypoint pid=391889)[0m custom_chat_template_chat_completion_path: chat_templates/qwen3_thinking_acc.jinja2
|
||
[36m(skyrl_entrypoint pid=391889)[0m served_model_name: 0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6
|
||
[36m(skyrl_entrypoint pid=391889)[0m override_existing_update_group: disable
|
||
[36m(skyrl_entrypoint pid=391889)[0m sampling_params:
|
||
[36m(skyrl_entrypoint pid=391889)[0m max_generate_length: 4096
|
||
[36m(skyrl_entrypoint pid=391889)[0m repetition_penalty: 1.0
|
||
[36m(skyrl_entrypoint pid=391889)[0m temperature: 0.7
|
||
[36m(skyrl_entrypoint pid=391889)[0m top_p: 0.95
|
||
[36m(skyrl_entrypoint pid=391889)[0m min_p: 0.0
|
||
[36m(skyrl_entrypoint pid=391889)[0m top_k: 20
|
||
[36m(skyrl_entrypoint pid=391889)[0m logprobs: null
|
||
[36m(skyrl_entrypoint pid=391889)[0m stop: null
|
||
[36m(skyrl_entrypoint pid=391889)[0m use_conversation_multi_turn: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m append_eos_token_after_stop_str_in_multi_turn: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m eval_sampling_params:
|
||
[36m(skyrl_entrypoint pid=391889)[0m max_generate_length: ${generator.sampling_params.max_generate_length}
|
||
[36m(skyrl_entrypoint pid=391889)[0m repetition_penalty: 1.0
|
||
[36m(skyrl_entrypoint pid=391889)[0m temperature: 0.0
|
||
[36m(skyrl_entrypoint pid=391889)[0m top_p: 1.0
|
||
[36m(skyrl_entrypoint pid=391889)[0m min_p: 0.0
|
||
[36m(skyrl_entrypoint pid=391889)[0m top_k: -1
|
||
[36m(skyrl_entrypoint pid=391889)[0m logprobs: null
|
||
[36m(skyrl_entrypoint pid=391889)[0m stop: null
|
||
[36m(skyrl_entrypoint pid=391889)[0m eval_n_samples_per_prompt: 8
|
||
[36m(skyrl_entrypoint pid=391889)[0m zero_reward_on_non_stop: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m apply_overlong_filtering: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m rope_scaling: ${trainer.rope_scaling}
|
||
[36m(skyrl_entrypoint pid=391889)[0m rope_theta: ${trainer.rope_theta}
|
||
[36m(skyrl_entrypoint pid=391889)[0m teacher:
|
||
[36m(skyrl_entrypoint pid=391889)[0m model_path: null
|
||
[36m(skyrl_entrypoint pid=391889)[0m top_k_logprobs: 256
|
||
[36m(skyrl_entrypoint pid=391889)[0m num_inference_engines: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m inference_engine_tensor_parallel_size: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m inference_engine_pipeline_parallel_size: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m gpu_memory_utilization: 0.9
|
||
[36m(skyrl_entrypoint pid=391889)[0m enforce_eager: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m backend: vllm
|
||
[36m(skyrl_entrypoint pid=391889)[0m engine_init_kwargs: {}
|
||
[36m(skyrl_entrypoint pid=391889)[0m environment:
|
||
[36m(skyrl_entrypoint pid=391889)[0m env_class: gsm8k
|
||
[36m(skyrl_entrypoint pid=391889)[0m skyrl_gym:
|
||
[36m(skyrl_entrypoint pid=391889)[0m max_env_workers: 32
|
||
[36m(skyrl_entrypoint pid=391889)[0m text2sql:
|
||
[36m(skyrl_entrypoint pid=391889)[0m db_path: /home/ray/default/sql_data
|
||
[36m(skyrl_entrypoint pid=391889)[0m llm_as_a_judge:
|
||
[36m(skyrl_entrypoint pid=391889)[0m model: gpt-4o-mini
|
||
[36m(skyrl_entrypoint pid=391889)[0m base_url: null
|
||
[36m(skyrl_entrypoint pid=391889)[0m search:
|
||
[36m(skyrl_entrypoint pid=391889)[0m log_requests: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m search_url: http://127.0.0.1:8000/retrieve
|
||
[36m(skyrl_entrypoint pid=391889)[0m topk: 3
|
||
[36m(skyrl_entrypoint pid=391889)[0m timeout: 30
|
||
[36m(skyrl_entrypoint pid=391889)[0m rollout:
|
||
[36m(skyrl_entrypoint pid=391889)[0m fanout:
|
||
[36m(skyrl_entrypoint pid=391889)[0m enabled: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m num_coordinators: 4
|
||
[36m(skyrl_entrypoint pid=391889)[0m cpus_per_coordinator: 8
|
||
[36m(skyrl_entrypoint pid=391889)[0m deepspeed_config:
|
||
[36m(skyrl_entrypoint pid=391889)[0m train:
|
||
[36m(skyrl_entrypoint pid=391889)[0m zero_optimization:
|
||
[36m(skyrl_entrypoint pid=391889)[0m stage: 3
|
||
[36m(skyrl_entrypoint pid=391889)[0m offload_param:
|
||
[36m(skyrl_entrypoint pid=391889)[0m device: none
|
||
[36m(skyrl_entrypoint pid=391889)[0m offload_optimizer:
|
||
[36m(skyrl_entrypoint pid=391889)[0m device: none
|
||
[36m(skyrl_entrypoint pid=391889)[0m pin_memory: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m sub_group_size: auto
|
||
[36m(skyrl_entrypoint pid=391889)[0m reduce_bucket_size: auto
|
||
[36m(skyrl_entrypoint pid=391889)[0m stage3_param_persistence_threshold: auto
|
||
[36m(skyrl_entrypoint pid=391889)[0m stage3_prefetch_bucket_size: auto
|
||
[36m(skyrl_entrypoint pid=391889)[0m stage3_max_live_parameters: auto
|
||
[36m(skyrl_entrypoint pid=391889)[0m stage3_max_reuse_distance: auto
|
||
[36m(skyrl_entrypoint pid=391889)[0m round_robin_gradients: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m zero_hpz_partition_size: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m zero_quantized_weights: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m zero_quantized_gradients: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m torch_autocast:
|
||
[36m(skyrl_entrypoint pid=391889)[0m enabled: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m dtype: bfloat16
|
||
[36m(skyrl_entrypoint pid=391889)[0m disable_trace_cache: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m data_types:
|
||
[36m(skyrl_entrypoint pid=391889)[0m grad_accum_dtype: fp32
|
||
[36m(skyrl_entrypoint pid=391889)[0m gradient_clipping: 1.0
|
||
[36m(skyrl_entrypoint pid=391889)[0m wall_clock_breakdown: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m prescale_gradient: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m eval:
|
||
[36m(skyrl_entrypoint pid=391889)[0m zero_optimization:
|
||
[36m(skyrl_entrypoint pid=391889)[0m stage: 3
|
||
[36m(skyrl_entrypoint pid=391889)[0m stage3_param_persistence_threshold: auto
|
||
[36m(skyrl_entrypoint pid=391889)[0m offload_param:
|
||
[36m(skyrl_entrypoint pid=391889)[0m device: cpu
|
||
[36m(skyrl_entrypoint pid=391889)[0m pin_memory: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m torch_autocast:
|
||
[36m(skyrl_entrypoint pid=391889)[0m enabled: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m dtype: bfloat16
|
||
[36m(skyrl_entrypoint pid=391889)[0m gradient_clipping: 1.0
|
||
[36m(skyrl_entrypoint pid=391889)[0m prescale_gradient: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m wall_clock_breakdown: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m terminal_bench_config:
|
||
[36m(skyrl_entrypoint pid=391889)[0m trials_dir: /e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqmean-arm0/ablation-pymethods2test-seqmean-arm0/trace_jobs
|
||
[36m(skyrl_entrypoint pid=391889)[0m harbor:
|
||
[36m(skyrl_entrypoint pid=391889)[0m name: terminus-2
|
||
[36m(skyrl_entrypoint pid=391889)[0m max_episodes: 999999
|
||
[36m(skyrl_entrypoint pid=391889)[0m enable_summarize: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m store_all_messages: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m trajectory_config:
|
||
[36m(skyrl_entrypoint pid=391889)[0m raw_content: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m enable_episode_logging: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m record_terminal_session: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m enable_pane_logging: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m strict_json_parser: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m interleaved_thinking: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m extra_body:
|
||
[36m(skyrl_entrypoint pid=391889)[0m chat_template_kwargs:
|
||
[36m(skyrl_entrypoint pid=391889)[0m enable_thinking: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m override_timeout_sec: 900
|
||
[36m(skyrl_entrypoint pid=391889)[0m override_cpus: 1
|
||
[36m(skyrl_entrypoint pid=391889)[0m override_memory_mb: 2048
|
||
[36m(skyrl_entrypoint pid=391889)[0m override_storage_mb: 2048
|
||
[36m(skyrl_entrypoint pid=391889)[0m auto_snapshot: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m verifier_override_timeout_sec: 120
|
||
[36m(skyrl_entrypoint pid=391889)[0m max_retries: 3
|
||
[36m(skyrl_entrypoint pid=391889)[0m min_wait_sec: 60.0
|
||
[36m(skyrl_entrypoint pid=391889)[0m max_wait_sec: 600.0
|
||
[36m(skyrl_entrypoint pid=391889)[0m wait_multiplier: 2.0
|
||
[36m(skyrl_entrypoint pid=391889)[0m exclude_exceptions:
|
||
[36m(skyrl_entrypoint pid=391889)[0m - VerifierTimeoutError
|
||
[36m(skyrl_entrypoint pid=391889)[0m - VerifierRuntimeError
|
||
[36m(skyrl_entrypoint pid=391889)[0m - RewardFileNotFoundError
|
||
[36m(skyrl_entrypoint pid=391889)[0m - RewardFileEmptyError
|
||
[36m(skyrl_entrypoint pid=391889)[0m - VerifierOutputParseError
|
||
[36m(skyrl_entrypoint pid=391889)[0m n_concurrent_trials: 675
|
||
[36m(skyrl_entrypoint pid=391889)[0m log_level: INFO
|
||
[36m(skyrl_entrypoint pid=391889)[0m enable_reward_shaping: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m enable_error_classification: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m mask_exceptions:
|
||
[36m(skyrl_entrypoint pid=391889)[0m - DaytonaError
|
||
[36m(skyrl_entrypoint pid=391889)[0m - EnvironmentStartTimeoutError
|
||
[36m(skyrl_entrypoint pid=391889)[0m - NetworkError
|
||
[36m(skyrl_entrypoint pid=391889)[0m - ConnectionError
|
||
[36m(skyrl_entrypoint pid=391889)[0m - RewardFileNotFoundError
|
||
[36m(skyrl_entrypoint pid=391889)[0m - RewardFileEmptyError
|
||
[36m(skyrl_entrypoint pid=391889)[0m - AgentEnvironmentTimeoutError
|
||
[36m(skyrl_entrypoint pid=391889)[0m - ContextLengthExceededError
|
||
[36m(skyrl_entrypoint pid=391889)[0m default_error_treatment: zero
|
||
[36m(skyrl_entrypoint pid=391889)[0m passthrough_exceptions:
|
||
[36m(skyrl_entrypoint pid=391889)[0m - AgentTimeoutError
|
||
[36m(skyrl_entrypoint pid=391889)[0m zero_exceptions: []
|
||
[36m(skyrl_entrypoint pid=391889)[0m model_info:
|
||
[36m(skyrl_entrypoint pid=391889)[0m max_input_tokens: 32000
|
||
[36m(skyrl_entrypoint pid=391889)[0m max_output_tokens: 4096
|
||
[36m(skyrl_entrypoint pid=391889)[0m archiving:
|
||
[36m(skyrl_entrypoint pid=391889)[0m enabled: false
|
||
[36m(skyrl_entrypoint pid=391889)[0m trace_upload:
|
||
[36m(skyrl_entrypoint pid=391889)[0m enabled: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m repo_org: DCAgent
|
||
[36m(skyrl_entrypoint pid=391889)[0m episodes: last
|
||
[36m(skyrl_entrypoint pid=391889)[0m dataset_type: SFT
|
||
[36m(skyrl_entrypoint pid=391889)[0m cleanup: true
|
||
[36m(skyrl_entrypoint pid=391889)[0m [0m
|
||
[36m(skyrl_entrypoint pid=391889)[0m [2026-06-07 16:03:03,420 E 391889 391932] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 2x across cluster][0m
|
||
[36m(skyrl_entrypoint pid=391889)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:
|
||
[36m(skyrl_entrypoint pid=391889)[0m No module named 'vllm._version'
|
||
[36m(skyrl_entrypoint pid=391889)[0m from .version import __version__, __version_tuple__ # isort:skip
|
||
[36m(skyrl_entrypoint pid=391889)[0m W0607 16:06:28.744000 391889 envs/rl/lib/python3.12/site-packages/torch/utils/cpp_extension.py:117] No CUDA runtime is found, using CUDA_HOME='/e/software/default/stages/2026/software/CUDA/13'
|
||
[33m(raylet)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e
|
||
[36m(get_all_env_variables pid=392160)[0m [2026-06-07 16:07:18,837 E 392160 392200] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14
|
||
[33m(raylet, ip=10.128.32.42)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 29x across cluster][0m
|
||
[36m(pid=3113603, ip=10.128.32.42)[0m [2026-06-07 16:07:21,100 E 3113603 3113706] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 4x across cluster][0m
|
||
[33m(raylet, ip=10.128.32.37)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 2x across cluster][0m
|
||
[36m(pid=3113603, ip=10.128.32.42)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:
|
||
[36m(pid=3113603, ip=10.128.32.42)[0m No module named 'vllm._version'
|
||
[36m(pid=3113603, ip=10.128.32.42)[0m from .version import __version__, __version_tuple__ # isort:skip
|
||
[33m(raylet, ip=10.128.32.37)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 10x across cluster][0m
|
||
[36m(get_addr_port pid=1022558, ip=10.128.32.37)[0m [2026-06-07 16:08:25,572 E 1022558 1022633] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14
|
||
[33m(raylet, ip=10.128.32.35)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e
|
||
[36m(pid=1022559, ip=10.128.32.37)[0m [2026-06-07 16:08:25,619 E 1022559 1022661] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14
|
||
[36m(get_addr_port pid=899452, ip=10.128.32.35)[0m [2026-06-07 16:13:12,407 E 899452 899534] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14
|
||
[33m(raylet, ip=10.128.32.37)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 23x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m 2026-06-07 16:16:18.252 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:152 - setup_envvars_for_vllm: distributed_executor_backend=uni, SKYRL_ENABLE_NUMA_AFFINITY=1, CUDA_VISIBLE_DEVICES=0, VLLM_ENABLE_V1_MULTIPROCESSING=<unset>
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m 2026-06-07 16:16:18.253 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:162 - setup_envvars_for_vllm: numa_enabled=True
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m 2026-06-07 16:16:18.253 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:165 - setup_envvars_for_vllm: set VLLM_ENABLE_V1_MULTIPROCESSING=0 for NUMA affinity
|
||
[36m(pid=1022866, ip=10.128.32.37)[0m [2026-06-07 16:13:12,454 E 1022866 1022968] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m 2026-06-07 16:16:18.863 | INFO | skyrl_train.utils.numa:set_numa_affinity_for_gpu:340 - NUMA affinity: bound process to CPUs 0-71 (NUMA node 0) for GPU 0
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m 2026-06-07 16:16:19.203 | DEBUG | skyrl_train.utils.numa:_set_membind_via_libnuma:375 - NUMA affinity: set memory preferred to NUMA node 0
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m 2026-06-07 16:16:19.203 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:613 - BaseVLLMInferenceEngine: vllm_v1_disable_multiproc=True, vllm.__version__=dev, VLLM_ENABLE_V1_MULTIPROCESSING=0
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m 2026-06-07 16:16:19.203 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:621 - BaseVLLMInferenceEngine: set VLLM_ENABLE_V1_MULTIPROCESSING=0
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m 2026-06-07 16:16:19.203 | WARNING | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1187 - OpenAI API sampling params overridden: temperature=0.7, top_p=0.95, top_k=20
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m 2026-06-07 16:16:19.249 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1239 - Engine startup stagger: sleeping 2.96s (attempt 1/5) to avoid port collisions
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m No module named 'vllm._version'
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m from .version import __version__, __version_tuple__ # isort:skip
|
||
[33m(raylet, ip=10.128.32.35)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 3x across cluster][0m
|
||
[36m(pid=2697284, ip=10.128.32.43)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:
|
||
[36m(pid=2697284, ip=10.128.32.43)[0m No module named 'vllm._version'
|
||
[36m(pid=2697284, ip=10.128.32.43)[0m from .version import __version__, __version_tuple__ # isort:skip
|
||
[33m(raylet, ip=10.128.32.42)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 81x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114084) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/__init__.py:1617: UserWarning: Please use the new API settings to control TF32 behavior, such as torch.backends.cudnn.conv.fp32_precision = 'tf32' or torch.backends.cuda.matmul.fp32_precision = 'ieee'. Old settings, e.g, torch.backends.cuda.matmul.allow_tf32 = True, torch.backends.cudnn.allow_tf32 = True, allowTF32CuDNN() and allowTF32CuBLAS() will be deprecated after Pytorch 2.9. Please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:80.)
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114084) _C._set_float32_matmul_precision(precision)
|
||
[36m(pid=3114114, ip=10.128.32.42)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 6x across cluster][0m
|
||
[36m(pid=3114114, ip=10.128.32.42)[0m No module named 'vllm._version'[32m [repeated 6x across cluster][0m
|
||
[36m(pid=3114114, ip=10.128.32.42)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 6x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m [W607 16:18:19.012223938 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-041-42-interconnect-1.jupiter.internal]:48887 (errno: 97 - Address family not supported by protocol).
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m [W607 16:18:19.013894942 Utils.hpp:166] Warning: Environment variable NCCL_BLOCKING_WAIT is deprecated; use TORCH_NCCL_BLOCKING_WAIT instead (function operator())
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m [rank0]:[W607 16:18:19.016313609 Utils.hpp:137] Warning: Environment variable TORCH_NCCL_TRACE_BUFFER_SIZE is deprecated; use TORCH_FR_BUFFER_SIZE instead (function operator())
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114084)
|
||
Loading safetensors checkpoint shards: 0% Completed | 0/4 [00:00<?, ?it/s]
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114084)
|
||
Loading safetensors checkpoint shards: 25% Completed | 1/4 [00:03<00:11, 3.70s/it]
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114084)
|
||
Loading safetensors checkpoint shards: 50% Completed | 2/4 [00:05<00:05, 2.63s/it]
|
||
[36m(get_addr_port pid=2697285, ip=10.128.32.43)[0m [2026-06-07 16:18:27,672 E 2697285 2697359] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114084)
|
||
Loading safetensors checkpoint shards: 75% Completed | 3/4 [00:07<00:02, 2.29s/it]
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114084)
|
||
Loading safetensors checkpoint shards: 100% Completed | 4/4 [00:08<00:00, 1.67s/it]
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114084)
|
||
Loading safetensors checkpoint shards: 100% Completed | 4/4 [00:08<00:00, 2.05s/it]
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114084)
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114084) The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m 2026-06-07 16:18:32.160 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1296 - Initializing OpenAIServingChat with custom_chat_template read from: chat_templates/qwen3_thinking_acc.jinja2
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m 2026-06-07 16:18:37.079 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:152 - setup_envvars_for_vllm: distributed_executor_backend=uni, SKYRL_ENABLE_NUMA_AFFINITY=1, CUDA_VISIBLE_DEVICES=2, VLLM_ENABLE_V1_MULTIPROCESSING=<unset>
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m 2026-06-07 16:18:37.080 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:162 - setup_envvars_for_vllm: numa_enabled=True
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m 2026-06-07 16:18:37.080 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:165 - setup_envvars_for_vllm: set VLLM_ENABLE_V1_MULTIPROCESSING=0 for NUMA affinity
|
||
[36m(pid=3114114, ip=10.128.32.42)[0m [2026-06-07 16:18:30,329 E 3114114 3114278] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 13x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m 2026-06-07 16:18:38.033 | INFO | skyrl_train.utils.numa:set_numa_affinity_for_gpu:340 - NUMA affinity: bound process to CPUs 144-215 (NUMA node 2) for GPU 2
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m 2026-06-07 16:18:38.037 | DEBUG | skyrl_train.utils.numa:_set_membind_via_libnuma:375 - NUMA affinity: set memory preferred to NUMA node 2
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m 2026-06-07 16:18:38.037 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:613 - BaseVLLMInferenceEngine: vllm_v1_disable_multiproc=True, vllm.__version__=dev, VLLM_ENABLE_V1_MULTIPROCESSING=0
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m 2026-06-07 16:18:38.037 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:621 - BaseVLLMInferenceEngine: set VLLM_ENABLE_V1_MULTIPROCESSING=0
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m 2026-06-07 16:18:38.038 | WARNING | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1187 - OpenAI API sampling params overridden: temperature=0.7, top_p=0.95, top_k=20
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m 2026-06-07 16:18:38.082 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1239 - Engine startup stagger: sleeping 1.57s (attempt 1/5) to avoid port collisions
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m No module named 'vllm._version'
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m from .version import __version__, __version_tuple__ # isort:skip
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114382) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/__init__.py:1617: UserWarning: Please use the new API settings to control TF32 behavior, such as torch.backends.cudnn.conv.fp32_precision = 'tf32' or torch.backends.cuda.matmul.fp32_precision = 'ieee'. Old settings, e.g, torch.backends.cuda.matmul.allow_tf32 = True, torch.backends.cudnn.allow_tf32 = True, allowTF32CuDNN() and allowTF32CuBLAS() will be deprecated after Pytorch 2.9. Please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:80.)
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114382) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m 2026-06-07 16:18:37.079 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:152 - setup_envvars_for_vllm: distributed_executor_backend=uni, SKYRL_ENABLE_NUMA_AFFINITY=1, CUDA_VISIBLE_DEVICES=3, VLLM_ENABLE_V1_MULTIPROCESSING=<unset>[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m 2026-06-07 16:18:37.080 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:162 - setup_envvars_for_vllm: numa_enabled=True[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m 2026-06-07 16:18:37.080 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:165 - setup_envvars_for_vllm: set VLLM_ENABLE_V1_MULTIPROCESSING=0 for NUMA affinity[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m 2026-06-07 16:18:38.037 | INFO | skyrl_train.utils.numa:set_numa_affinity_for_gpu:340 - NUMA affinity: bound process to CPUs 216-287 (NUMA node 3) for GPU 3[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m 2026-06-07 16:18:38.040 | DEBUG | skyrl_train.utils.numa:_set_membind_via_libnuma:375 - NUMA affinity: set memory preferred to NUMA node 3[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m 2026-06-07 16:18:38.040 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:613 - BaseVLLMInferenceEngine: vllm_v1_disable_multiproc=True, vllm.__version__=dev, VLLM_ENABLE_V1_MULTIPROCESSING=0[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m 2026-06-07 16:18:38.040 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:621 - BaseVLLMInferenceEngine: set VLLM_ENABLE_V1_MULTIPROCESSING=0[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m 2026-06-07 16:18:38.040 | WARNING | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1187 - OpenAI API sampling params overridden: temperature=0.7, top_p=0.95, top_k=20[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m 2026-06-07 16:18:38.082 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1239 - Engine startup stagger: sleeping 2.51s (attempt 1/5) to avoid port collisions[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.[32m [repeated 4x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 4x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m No module named 'vllm._version'[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114394) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m [W607 16:18:52.506534225 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-041-42.jupiter.internal]:44549 (errno: 97 - Address family not supported by protocol).
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m [W607 16:18:52.507043782 Utils.hpp:166] Warning: Environment variable NCCL_BLOCKING_WAIT is deprecated; use TORCH_NCCL_BLOCKING_WAIT instead (function operator())
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m [rank0]:[W607 16:18:52.509417811 Utils.hpp:137] Warning: Environment variable TORCH_NCCL_TRACE_BUFFER_SIZE is deprecated; use TORCH_FR_BUFFER_SIZE instead (function operator())
|
||
[36m(AsyncVLLMInferenceEngine pid=3114112, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114390) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114394)
|
||
Loading safetensors checkpoint shards: 0% Completed | 0/4 [00:00<?, ?it/s]
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114382)
|
||
Loading safetensors checkpoint shards: 25% Completed | 1/4 [00:01<00:03, 1.24s/it]
|
||
[36m(AsyncVLLMInferenceEngine pid=3114112, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114390) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/__init__.py:1617: UserWarning: Please use the new API settings to control TF32 behavior, such as torch.backends.cudnn.conv.fp32_precision = 'tf32' or torch.backends.cuda.matmul.fp32_precision = 'ieee'. Old settings, e.g, torch.backends.cuda.matmul.allow_tf32 = True, torch.backends.cudnn.allow_tf32 = True, allowTF32CuDNN() and allowTF32CuBLAS() will be deprecated after Pytorch 2.9. Please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:80.)[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114112, ip=10.128.32.42)[0m [W607 16:18:53.056472543 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-041-42.jupiter.internal]:51575 (errno: 97 - Address family not supported by protocol).[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114112, ip=10.128.32.42)[0m [W607 16:18:53.057000595 Utils.hpp:166] Warning: Environment variable NCCL_BLOCKING_WAIT is deprecated; use TORCH_NCCL_BLOCKING_WAIT instead (function operator())[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114112, ip=10.128.32.42)[0m [rank0]:[W607 16:18:53.059114149 Utils.hpp:137] Warning: Environment variable TORCH_NCCL_TRACE_BUFFER_SIZE is deprecated; use TORCH_FR_BUFFER_SIZE instead (function operator())[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114382)
|
||
[36m(AsyncVLLMInferenceEngine pid=3114112, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114390)
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114394)
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114382) The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.
|
||
[36m(AsyncVLLMInferenceEngine pid=3114112, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114390)
|
||
Loading safetensors checkpoint shards: 0% Completed | 0/4 [00:00<?, ?it/s][32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114394)
|
||
Loading safetensors checkpoint shards: 100% Completed | 4/4 [00:04<00:00, 1.09s/it][32m [repeated 14x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m 2026-06-07 16:19:01.699 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1296 - Initializing OpenAIServingChat with custom_chat_template read from: chat_templates/qwen3_thinking_acc.jinja2
|
||
[33m(raylet, ip=10.128.32.40)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114394) The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m 2026-06-07 16:19:01.845 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1296 - Initializing OpenAIServingChat with custom_chat_template read from: chat_templates/qwen3_thinking_acc.jinja2[32m [repeated 2x across cluster][0m
|
||
[36m(pid=2739368, ip=10.128.32.34)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:
|
||
[36m(pid=2739368, ip=10.128.32.34)[0m No module named 'vllm._version'
|
||
[36m(pid=2739368, ip=10.128.32.34)[0m from .version import __version__, __version_tuple__ # isort:skip
|
||
[36m(get_addr_port pid=2576534, ip=10.128.32.40)[0m [2026-06-07 16:21:33,156 E 2576534 2576609] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14
|
||
[33m(raylet, ip=10.128.32.40)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 29x across cluster][0m
|
||
[36m(pid=2739604, ip=10.128.32.34)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 3x across cluster][0m
|
||
[36m(pid=2739604, ip=10.128.32.34)[0m No module named 'vllm._version'[32m [repeated 3x across cluster][0m
|
||
[36m(pid=2739604, ip=10.128.32.34)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 3x across cluster][0m
|
||
[36m(pid=2739604, ip=10.128.32.34)[0m [2026-06-07 16:21:33,411 E 2739604 2739772] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 4x across cluster][0m
|
||
[33m(raylet, ip=10.128.32.40)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 3x across cluster][0m
|
||
[36m(pid=2576535, ip=10.128.32.40)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:
|
||
[36m(pid=2576535, ip=10.128.32.40)[0m No module named 'vllm._version'
|
||
[36m(pid=2576535, ip=10.128.32.40)[0m from .version import __version__, __version_tuple__ # isort:skip
|
||
[36m(get_addr_port pid=2631426, ip=10.128.32.36)[0m [2026-06-07 16:25:01,686 E 2631426 2631501] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14
|
||
[33m(raylet, ip=10.128.32.40)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 27x across cluster][0m
|
||
[36m(pid=2576838, ip=10.128.32.40)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 3x across cluster][0m
|
||
[36m(pid=2576838, ip=10.128.32.40)[0m No module named 'vllm._version'[32m [repeated 3x across cluster][0m
|
||
[36m(pid=2576838, ip=10.128.32.40)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 3x across cluster][0m
|
||
[36m(pid=2631427, ip=10.128.32.36)[0m [2026-06-07 16:25:01,733 E 2631427 2631529] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14
|
||
[36m(pid=2631427, ip=10.128.32.36)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:
|
||
[36m(pid=2631427, ip=10.128.32.36)[0m No module named 'vllm._version'
|
||
[36m(pid=2631427, ip=10.128.32.36)[0m from .version import __version__, __version_tuple__ # isort:skip
|
||
[36m(pid=2576838, ip=10.128.32.40)[0m [2026-06-07 16:25:01,997 E 2576838 2577002] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 3x across cluster][0m
|
||
[36m(pid=899451, ip=10.128.32.35)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:
|
||
[36m(pid=899451, ip=10.128.32.35)[0m No module named 'vllm._version'
|
||
[36m(pid=899451, ip=10.128.32.35)[0m from .version import __version__, __version_tuple__ # isort:skip
|
||
[36m(pid=1022559, ip=10.128.32.37)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 4x across cluster][0m
|
||
[36m(pid=1022559, ip=10.128.32.37)[0m No module named 'vllm._version'[32m [repeated 4x across cluster][0m
|
||
[36m(pid=1022559, ip=10.128.32.37)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 4x across cluster][0m
|
||
[33m(raylet, ip=10.128.32.38)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e
|
||
[36m(pid=1023166, ip=10.128.32.37)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 3x across cluster][0m
|
||
[36m(pid=1023166, ip=10.128.32.37)[0m No module named 'vllm._version'[32m [repeated 3x across cluster][0m
|
||
[36m(pid=1023166, ip=10.128.32.37)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m 2026-06-07 16:28:40.941 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:152 - setup_envvars_for_vllm: distributed_executor_backend=uni, SKYRL_ENABLE_NUMA_AFFINITY=1, CUDA_VISIBLE_DEVICES=0, VLLM_ENABLE_V1_MULTIPROCESSING=<unset>
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m 2026-06-07 16:28:40.952 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:162 - setup_envvars_for_vllm: numa_enabled=True
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m 2026-06-07 16:28:40.952 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:165 - setup_envvars_for_vllm: set VLLM_ENABLE_V1_MULTIPROCESSING=0 for NUMA affinity
|
||
[33m(raylet, ip=10.128.32.44)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 23x across cluster][0m
|
||
[36m(pid=3011432, ip=10.128.32.44)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 3x across cluster][0m
|
||
[36m(pid=3011432, ip=10.128.32.44)[0m No module named 'vllm._version'[32m [repeated 3x across cluster][0m
|
||
[36m(pid=3011432, ip=10.128.32.44)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m 2026-06-07 16:28:42.189 | INFO | skyrl_train.utils.numa:set_numa_affinity_for_gpu:340 - NUMA affinity: bound process to CPUs 0-71 (NUMA node 0) for GPU 0
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m 2026-06-07 16:28:42.404 | DEBUG | skyrl_train.utils.numa:_set_membind_via_libnuma:375 - NUMA affinity: set memory preferred to NUMA node 0
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m 2026-06-07 16:28:42.404 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:613 - BaseVLLMInferenceEngine: vllm_v1_disable_multiproc=True, vllm.__version__=dev, VLLM_ENABLE_V1_MULTIPROCESSING=0
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m 2026-06-07 16:28:42.404 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:621 - BaseVLLMInferenceEngine: set VLLM_ENABLE_V1_MULTIPROCESSING=0
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m 2026-06-07 16:28:42.404 | WARNING | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1187 - OpenAI API sampling params overridden: temperature=0.7, top_p=0.95, top_k=20
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m 2026-06-07 16:28:42.452 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1239 - Engine startup stagger: sleeping 2.34s (attempt 1/5) to avoid port collisions
|
||
[36m(get_addr_port pid=2705576, ip=10.128.32.38)[0m [2026-06-07 16:28:59,577 E 2705576 2705650] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m 2026-06-07 16:28:40.941 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:152 - setup_envvars_for_vllm: distributed_executor_backend=uni, SKYRL_ENABLE_NUMA_AFFINITY=1, CUDA_VISIBLE_DEVICES=3, VLLM_ENABLE_V1_MULTIPROCESSING=<unset>[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m 2026-06-07 16:28:40.952 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:162 - setup_envvars_for_vllm: numa_enabled=True[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m 2026-06-07 16:28:40.952 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:165 - setup_envvars_for_vllm: set VLLM_ENABLE_V1_MULTIPROCESSING=0 for NUMA affinity[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m 2026-06-07 16:28:42.192 | INFO | skyrl_train.utils.numa:set_numa_affinity_for_gpu:340 - NUMA affinity: bound process to CPUs 216-287 (NUMA node 3) for GPU 3[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m 2026-06-07 16:28:42.404 | DEBUG | skyrl_train.utils.numa:_set_membind_via_libnuma:375 - NUMA affinity: set memory preferred to NUMA node 3[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m 2026-06-07 16:28:42.404 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:613 - BaseVLLMInferenceEngine: vllm_v1_disable_multiproc=True, vllm.__version__=dev, VLLM_ENABLE_V1_MULTIPROCESSING=0[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m 2026-06-07 16:28:42.404 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:621 - BaseVLLMInferenceEngine: set VLLM_ENABLE_V1_MULTIPROCESSING=0[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m 2026-06-07 16:28:42.404 | WARNING | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1187 - OpenAI API sampling params overridden: temperature=0.7, top_p=0.95, top_k=20[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m 2026-06-07 16:28:42.452 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1239 - Engine startup stagger: sleeping 2.44s (attempt 1/5) to avoid port collisions[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.
|
||
[36m(pid=3011432, ip=10.128.32.44)[0m [2026-06-07 16:28:59,741 E 3011432 3011535] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m No module named 'vllm._version'
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m from .version import __version__, __version_tuple__ # isort:skip
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.[32m [repeated 7x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.[32m [repeated 3x across cluster][0m
|
||
[33m(raylet, ip=10.128.32.41)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 8x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m No module named 'vllm._version'[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 3x across cluster][0m
|
||
[36m(get_addr_port pid=2614264, ip=10.128.32.41)[0m [2026-06-07 16:31:12,398 E 2614264 2614340] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14
|
||
[33m(raylet, ip=10.128.32.38)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 28x across cluster][0m
|
||
[36m(pid=2705803, ip=10.128.32.38)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 4x across cluster][0m
|
||
[36m(pid=2705803, ip=10.128.32.38)[0m No module named 'vllm._version'[32m [repeated 4x across cluster][0m
|
||
[36m(pid=2705803, ip=10.128.32.38)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 4x across cluster][0m
|
||
[36m(pid=2705803, ip=10.128.32.38)[0m [2026-06-07 16:31:12,700 E 2705803 2705968] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 4x across cluster][0m
|
||
[33m(raylet, ip=10.128.32.41)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 3x across cluster][0m
|
||
[36m(pid=3011671, ip=10.128.32.44)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:
|
||
[36m(pid=3011671, ip=10.128.32.44)[0m No module named 'vllm._version'
|
||
[36m(pid=3011671, ip=10.128.32.44)[0m from .version import __version__, __version_tuple__ # isort:skip
|
||
[33m(raylet, ip=10.128.32.39)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 53x across cluster][0m
|
||
[36m(pid=2631820, ip=10.128.32.36)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 3x across cluster][0m
|
||
[36m(pid=2631820, ip=10.128.32.36)[0m No module named 'vllm._version'[32m [repeated 3x across cluster][0m
|
||
[36m(pid=2631820, ip=10.128.32.36)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 3x across cluster][0m
|
||
[36m(get_addr_port pid=2831549, ip=10.128.32.45)[0m [2026-06-07 16:32:34,010 E 2831549 2831630] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14
|
||
[33m(raylet, ip=10.128.32.39)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 28x across cluster][0m
|
||
[36m(pid=2831686, ip=10.128.32.45)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 4x across cluster][0m
|
||
[36m(pid=2831686, ip=10.128.32.45)[0m No module named 'vllm._version'[32m [repeated 4x across cluster][0m
|
||
[36m(pid=2831686, ip=10.128.32.45)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 4x across cluster][0m
|
||
[36m(pid=2831684, ip=10.128.32.45)[0m [2026-06-07 16:33:02,327 E 2831684 2831794] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 9x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m 2026-06-07 16:37:47.395 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:152 - setup_envvars_for_vllm: distributed_executor_backend=uni, SKYRL_ENABLE_NUMA_AFFINITY=1, CUDA_VISIBLE_DEVICES=0, VLLM_ENABLE_V1_MULTIPROCESSING=<unset>
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m 2026-06-07 16:37:47.396 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:162 - setup_envvars_for_vllm: numa_enabled=True
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m 2026-06-07 16:37:47.396 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:165 - setup_envvars_for_vllm: set VLLM_ENABLE_V1_MULTIPROCESSING=0 for NUMA affinity
|
||
[36m(pid=2994914, ip=10.128.32.39)[0m [2026-06-07 16:33:02,705 E 2994914 2995017] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 4x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m 2026-06-07 16:37:48.608 | INFO | skyrl_train.utils.numa:set_numa_affinity_for_gpu:340 - NUMA affinity: bound process to CPUs 0-71 (NUMA node 0) for GPU 0
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m 2026-06-07 16:37:48.637 | DEBUG | skyrl_train.utils.numa:_set_membind_via_libnuma:375 - NUMA affinity: set memory preferred to NUMA node 0
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m 2026-06-07 16:37:48.637 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:613 - BaseVLLMInferenceEngine: vllm_v1_disable_multiproc=True, vllm.__version__=dev, VLLM_ENABLE_V1_MULTIPROCESSING=0
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m 2026-06-07 16:37:48.637 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:621 - BaseVLLMInferenceEngine: set VLLM_ENABLE_V1_MULTIPROCESSING=0
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m 2026-06-07 16:37:48.637 | WARNING | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1187 - OpenAI API sampling params overridden: temperature=0.7, top_p=0.95, top_k=20
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m 2026-06-07 16:37:48.667 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1239 - Engine startup stagger: sleeping 2.38s (attempt 1/5) to avoid port collisions
|
||
[36m(AsyncVLLMInferenceEngine pid=2576836, ip=10.128.32.40)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.
|
||
[36m(AsyncVLLMInferenceEngine pid=2576836, ip=10.128.32.40)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.
|
||
[36m(AsyncVLLMInferenceEngine pid=2576836, ip=10.128.32.40)[0m The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.
|
||
[36m(AsyncVLLMInferenceEngine pid=2576836, ip=10.128.32.40)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e
|
||
[36m(AsyncVLLMInferenceEngine pid=2576836, ip=10.128.32.40)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e
|
||
[36m(AsyncVLLMInferenceEngine pid=2576836, ip=10.128.32.40)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:
|
||
[36m(AsyncVLLMInferenceEngine pid=2576836, ip=10.128.32.40)[0m No module named 'vllm._version'
|
||
[36m(AsyncVLLMInferenceEngine pid=2576836, ip=10.128.32.40)[0m from .version import __version__, __version_tuple__ # isort:skip
|
||
[36m(skyrl_entrypoint pid=391889)[0m [2026-06-07 16:38:35] INFO inference_engine_client_http_endpoint.py:350: Starting server on 0.0.0.0:8000
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m 2026-06-07 16:37:47.395 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:152 - setup_envvars_for_vllm: distributed_executor_backend=uni, SKYRL_ENABLE_NUMA_AFFINITY=1, CUDA_VISIBLE_DEVICES=3, VLLM_ENABLE_V1_MULTIPROCESSING=<unset>[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m 2026-06-07 16:37:47.396 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:162 - setup_envvars_for_vllm: numa_enabled=True[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m 2026-06-07 16:37:47.396 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:165 - setup_envvars_for_vllm: set VLLM_ENABLE_V1_MULTIPROCESSING=0 for NUMA affinity[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m 2026-06-07 16:37:48.609 | INFO | skyrl_train.utils.numa:set_numa_affinity_for_gpu:340 - NUMA affinity: bound process to CPUs 216-287 (NUMA node 3) for GPU 3[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m 2026-06-07 16:37:48.636 | DEBUG | skyrl_train.utils.numa:_set_membind_via_libnuma:375 - NUMA affinity: set memory preferred to NUMA node 3[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m 2026-06-07 16:37:48.637 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:613 - BaseVLLMInferenceEngine: vllm_v1_disable_multiproc=True, vllm.__version__=dev, VLLM_ENABLE_V1_MULTIPROCESSING=0[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m 2026-06-07 16:37:48.637 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:621 - BaseVLLMInferenceEngine: set VLLM_ENABLE_V1_MULTIPROCESSING=0[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m 2026-06-07 16:37:48.637 | WARNING | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1187 - OpenAI API sampling params overridden: temperature=0.7, top_p=0.95, top_k=20[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m 2026-06-07 16:37:48.668 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1239 - Engine startup stagger: sleeping 2.17s (attempt 1/5) to avoid port collisions[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.[32m [repeated 6x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 6x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m No module named 'vllm._version'[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 3x across cluster][0m
|
||
[36m(skyrl_entrypoint pid=391889)[0m [2026-06-07 16:38:35] INFO inference_engine_client_http_endpoint.py:242: Starting inference HTTP endpoint...
|
||
[36m(skyrl_entrypoint pid=391889)[0m [2026-06-07 16:38:36] INFO inference_engine_client_http_endpoint.py:229: Server ready after 2 attempts (2 seconds)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [32m2026-06-07 16:38:36.246[0m | [1mINFO [0m | [36mskyrl_train.inference_engines.inference_engine_client[0m:[36m_spin_up_http_endpoint[0m:[36m969[0m - [1mInferenceEngineClient HTTP endpoint started on 127.0.0.1:8000[0m
|
||
[36m(skyrl_entrypoint pid=391889)[0m [32m2026-06-07 16:38:36.246[0m | [1mINFO [0m | [36mskyrl_train.inference_engines.inference_engine_client[0m:[36m__init__[0m:[36m61[0m - [1mInferenceEngineClient initialized with 48 engines.[0m
|
||
[36m(skyrl_entrypoint pid=391889)[0m [32m2026-06-07 16:38:36.249[0m | [1mINFO [0m | [36mexamples.terminal_bench.terminal_bench_generator[0m:[36m_configure_harbor_logging[0m:[36m242[0m - [1mHarbor logging level set to INFO[0m
|
||
[36m(skyrl_entrypoint pid=391889)[0m [32m2026-06-07 16:38:36.250[0m | [1mINFO [0m | [36mexamples.terminal_bench.terminal_bench_generator[0m:[36m__init__[0m:[36m142[0m - [1mTerminalBenchGenerator initialized with HarborConfigBuilder. Exposed fields: ['name', 'max_episodes', 'enable_summarize', 'store_all_messages', 'trajectory_config', 'enable_episode_logging', 'record_terminal_session', 'enable_pane_logging', 'strict_json_parser', 'interleaved_thinking', 'extra_body', 'override_timeout_sec', 'override_cpus', 'override_memory_mb', 'override_storage_mb', 'auto_snapshot', 'verifier_override_timeout_sec', 'max_retries', 'min_wait_sec', 'max_wait_sec', 'wait_multiplier', 'exclude_exceptions', 'n_concurrent_trials', 'log_level', 'enable_reward_shaping', 'enable_error_classification', 'mask_exceptions', 'default_error_treatment', 'passthrough_exceptions', 'zero_exceptions']. Retry config: max_retries=3, backoff=60.0-600.0s. Concurrent trials: 675. Reward shaping: enabled=False, shaper=pass_ratio. Error classification: enabled=True[0m
|
||
[36m(skyrl_entrypoint pid=391889)[0m [32m2026-06-07 16:38:36.251[0m | [1mINFO [0m | [36mexamples.terminal_bench.terminal_bench_generator[0m:[36m__init__[0m:[36m158[0m - [1mTerminalBenchGenerator initialized with custom chat template read from: chat_templates/qwen3_thinking_acc.jinja2[0m
|
||
[36m(skyrl_entrypoint pid=391889)[0m [32m2026-06-07 16:38:36.252[0m | [1mINFO [0m | [36mskyrl_train.utils.trainer_utils[0m:[36mbuild_dataloader[0m:[36m656[0m - [1mTotal steps: 156[0m
|
||
[36m(skyrl_entrypoint pid=391889)[0m [32m2026-06-07 16:38:36.252[0m | [1mINFO [0m | [36mskyrl_train.fully_async_trainer[0m:[36m_build_train_dataloader_and_compute_training_steps[0m:[36m357[0m - [1mLength of train_dataloader: 5000[0m
|
||
[36m(skyrl_entrypoint pid=391889)[0m [32m2026-06-07 16:38:36.252[0m | [1mINFO [0m | [36mskyrl_train.fully_async_trainer[0m:[36m_build_train_dataloader_and_compute_training_steps[0m:[36m358[0m - [1mNumber of steps per epoch: 78[0m
|
||
[36m(skyrl_entrypoint pid=391889)[0m [32m2026-06-07 16:38:36.252[0m | [1mINFO [0m | [36mskyrl_train.fully_async_trainer[0m:[36m_build_train_dataloader_and_compute_training_steps[0m:[36m359[0m - [1mTotal training steps: 80[0m
|
||
[36m(pid=2995275, ip=10.128.32.39)[0m [2026-06-07 16:39:06,386 E 2995275 2995386] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14
|
||
[33m(raylet, ip=10.128.32.46)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 25x across cluster][0m
|
||
[36m(pid=2662625, ip=10.128.32.46)[0m Using blocking ray.get inside async actor. This blocks the event loop. Please use `await` on object ref with asyncio.gather if you want to yield execution to the event loop instead.
|
||
[36m(pid=2662625, ip=10.128.32.46)[0m [2026-06-07 16:39:07,218 E 2662625 2662665] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m 2026-06-07 16:40:20.342 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:152 - setup_envvars_for_vllm: distributed_executor_backend=uni, SKYRL_ENABLE_NUMA_AFFINITY=1, CUDA_VISIBLE_DEVICES=0, VLLM_ENABLE_V1_MULTIPROCESSING=<unset>
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m 2026-06-07 16:40:20.343 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:162 - setup_envvars_for_vllm: numa_enabled=True
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m 2026-06-07 16:40:20.344 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:165 - setup_envvars_for_vllm: set VLLM_ENABLE_V1_MULTIPROCESSING=0 for NUMA affinity
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m 2026-06-07 16:40:21.623 | INFO | skyrl_train.utils.numa:set_numa_affinity_for_gpu:340 - NUMA affinity: bound process to CPUs 0-71 (NUMA node 0) for GPU 0
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m 2026-06-07 16:40:21.875 | DEBUG | skyrl_train.utils.numa:_set_membind_via_libnuma:375 - NUMA affinity: set memory preferred to NUMA node 0
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m 2026-06-07 16:40:21.875 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:613 - BaseVLLMInferenceEngine: vllm_v1_disable_multiproc=True, vllm.__version__=dev, VLLM_ENABLE_V1_MULTIPROCESSING=0
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m 2026-06-07 16:40:21.875 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:621 - BaseVLLMInferenceEngine: set VLLM_ENABLE_V1_MULTIPROCESSING=0
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m 2026-06-07 16:40:21.875 | WARNING | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1187 - OpenAI API sampling params overridden: temperature=0.7, top_p=0.95, top_k=20
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m 2026-06-07 16:40:21.925 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1239 - Engine startup stagger: sleeping 2.28s (attempt 1/5) to avoid port collisions
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698000) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/__init__.py:1617: UserWarning: Please use the new API settings to control TF32 behavior, such as torch.backends.cudnn.conv.fp32_precision = 'tf32' or torch.backends.cuda.matmul.fp32_precision = 'ieee'. Old settings, e.g, torch.backends.cuda.matmul.allow_tf32 = True, torch.backends.cudnn.allow_tf32 = True, allowTF32CuDNN() and allowTF32CuBLAS() will be deprecated after Pytorch 2.9. Please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:80.)
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698000) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=2697412, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698009) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=2697414, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698001) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698005) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m [W607 16:40:24.572522745 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-041-43-interconnect-1.jupiter.internal]:60183 (errno: 97 - Address family not supported by protocol).
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m [W607 16:40:24.576008342 Utils.hpp:166] Warning: Environment variable NCCL_BLOCKING_WAIT is deprecated; use TORCH_NCCL_BLOCKING_WAIT instead (function operator())
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m [rank0]:[W607 16:40:24.582010146 Utils.hpp:137] Warning: Environment variable TORCH_NCCL_TRACE_BUFFER_SIZE is deprecated; use TORCH_FR_BUFFER_SIZE instead (function operator())
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m No module named 'vllm._version'
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m from .version import __version__, __version_tuple__ # isort:skip
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m 2026-06-07 16:40:20.342 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:152 - setup_envvars_for_vllm: distributed_executor_backend=uni, SKYRL_ENABLE_NUMA_AFFINITY=1, CUDA_VISIBLE_DEVICES=3, VLLM_ENABLE_V1_MULTIPROCESSING=<unset>[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m 2026-06-07 16:40:20.343 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:162 - setup_envvars_for_vllm: numa_enabled=True[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m 2026-06-07 16:40:20.344 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:165 - setup_envvars_for_vllm: set VLLM_ENABLE_V1_MULTIPROCESSING=0 for NUMA affinity[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698000)
|
||
Loading safetensors checkpoint shards: 0% Completed | 0/4 [00:00<?, ?it/s]
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m 2026-06-07 16:40:21.623 | INFO | skyrl_train.utils.numa:set_numa_affinity_for_gpu:340 - NUMA affinity: bound process to CPUs 216-287 (NUMA node 3) for GPU 3[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m 2026-06-07 16:40:21.875 | DEBUG | skyrl_train.utils.numa:_set_membind_via_libnuma:375 - NUMA affinity: set memory preferred to NUMA node 3[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m 2026-06-07 16:40:21.875 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:613 - BaseVLLMInferenceEngine: vllm_v1_disable_multiproc=True, vllm.__version__=dev, VLLM_ENABLE_V1_MULTIPROCESSING=0[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m 2026-06-07 16:40:21.875 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:621 - BaseVLLMInferenceEngine: set VLLM_ENABLE_V1_MULTIPROCESSING=0[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m 2026-06-07 16:40:21.875 | WARNING | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1187 - OpenAI API sampling params overridden: temperature=0.7, top_p=0.95, top_k=20[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m 2026-06-07 16:40:21.925 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1239 - Engine startup stagger: sleeping 2.51s (attempt 1/5) to avoid port collisions[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698000)
|
||
Loading safetensors checkpoint shards: 25% Completed | 1/4 [00:04<00:12, 4.17s/it]
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698005) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/__init__.py:1617: UserWarning: Please use the new API settings to control TF32 behavior, such as torch.backends.cudnn.conv.fp32_precision = 'tf32' or torch.backends.cuda.matmul.fp32_precision = 'ieee'. Old settings, e.g, torch.backends.cuda.matmul.allow_tf32 = True, torch.backends.cudnn.allow_tf32 = True, allowTF32CuDNN() and allowTF32CuBLAS() will be deprecated after Pytorch 2.9. Please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:80.)[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739603, ip=10.128.32.34)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.[32m [repeated 6x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m [W607 16:40:24.572461210 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-041-43-interconnect-1.jupiter.internal]:46619 (errno: 97 - Address family not supported by protocol).[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m [W607 16:40:24.576002902 Utils.hpp:166] Warning: Environment variable NCCL_BLOCKING_WAIT is deprecated; use TORCH_NCCL_BLOCKING_WAIT instead (function operator())[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m [rank0]:[W607 16:40:24.582556632 Utils.hpp:137] Warning: Environment variable TORCH_NCCL_TRACE_BUFFER_SIZE is deprecated; use TORCH_FR_BUFFER_SIZE instead (function operator())[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739603, ip=10.128.32.34)[0m The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739603, ip=10.128.32.34)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 6x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739603, ip=10.128.32.34)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739603, ip=10.128.32.34)[0m No module named 'vllm._version'[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739603, ip=10.128.32.34)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698005)
|
||
Loading safetensors checkpoint shards: 0% Completed | 0/4 [00:00<?, ?it/s][32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698000)
|
||
Loading safetensors checkpoint shards: 75% Completed | 3/4 [00:12<00:04, 4.12s/it][32m [repeated 8x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698000)
|
||
[36m(AsyncVLLMInferenceEngine pid=2697412, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698009)
|
||
[36m(AsyncVLLMInferenceEngine pid=2697414, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698001)
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698005)
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698000) The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m 2026-06-07 16:40:45.443 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1296 - Initializing OpenAIServingChat with custom_chat_template read from: chat_templates/qwen3_thinking_acc.jinja2
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698005)
|
||
Loading safetensors checkpoint shards: 100% Completed | 4/4 [00:13<00:00, 3.44s/it][32m [repeated 11x across cluster][0m
|
||
[36m(pid=2614265, ip=10.128.32.41)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:
|
||
[36m(pid=2614265, ip=10.128.32.41)[0m No module named 'vllm._version'
|
||
[36m(pid=2614265, ip=10.128.32.41)[0m from .version import __version__, __version_tuple__ # isort:skip
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698005) The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m 2026-06-07 16:40:45.422 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1296 - Initializing OpenAIServingChat with custom_chat_template read from: chat_templates/qwen3_thinking_acc.jinja2[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m 2026-06-07 16:43:50.391 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:152 - setup_envvars_for_vllm: distributed_executor_backend=uni, SKYRL_ENABLE_NUMA_AFFINITY=1, CUDA_VISIBLE_DEVICES=0, VLLM_ENABLE_V1_MULTIPROCESSING=<unset>
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m 2026-06-07 16:43:50.392 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:162 - setup_envvars_for_vllm: numa_enabled=True
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m 2026-06-07 16:43:50.392 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:165 - setup_envvars_for_vllm: set VLLM_ENABLE_V1_MULTIPROCESSING=0 for NUMA affinity
|
||
[36m(pid=2614410, ip=10.128.32.41)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 3x across cluster][0m
|
||
[36m(pid=2614410, ip=10.128.32.41)[0m No module named 'vllm._version'[32m [repeated 3x across cluster][0m
|
||
[36m(pid=2614410, ip=10.128.32.41)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m 2026-06-07 16:43:51.752 | INFO | skyrl_train.utils.numa:set_numa_affinity_for_gpu:340 - NUMA affinity: bound process to CPUs 0-71 (NUMA node 0) for GPU 0
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m 2026-06-07 16:43:51.780 | DEBUG | skyrl_train.utils.numa:_set_membind_via_libnuma:375 - NUMA affinity: set memory preferred to NUMA node 0
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m 2026-06-07 16:43:51.780 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:613 - BaseVLLMInferenceEngine: vllm_v1_disable_multiproc=True, vllm.__version__=dev, VLLM_ENABLE_V1_MULTIPROCESSING=0
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m 2026-06-07 16:43:51.780 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:621 - BaseVLLMInferenceEngine: set VLLM_ENABLE_V1_MULTIPROCESSING=0
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m 2026-06-07 16:43:51.780 | WARNING | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1187 - OpenAI API sampling params overridden: temperature=0.7, top_p=0.95, top_k=20
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m 2026-06-07 16:43:51.824 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1239 - Engine startup stagger: sleeping 2.15s (attempt 1/5) to avoid port collisions
|
||
[36m(AsyncVLLMInferenceEngine pid=2631819, ip=10.128.32.36)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.
|
||
[36m(AsyncVLLMInferenceEngine pid=2631819, ip=10.128.32.36)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.
|
||
[36m(AsyncVLLMInferenceEngine pid=2631819, ip=10.128.32.36)[0m The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.
|
||
[36m(AsyncVLLMInferenceEngine pid=2631819, ip=10.128.32.36)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e
|
||
[36m(AsyncVLLMInferenceEngine pid=2631819, ip=10.128.32.36)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m 2026-06-07 16:43:50.391 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:152 - setup_envvars_for_vllm: distributed_executor_backend=uni, SKYRL_ENABLE_NUMA_AFFINITY=1, CUDA_VISIBLE_DEVICES=3, VLLM_ENABLE_V1_MULTIPROCESSING=<unset>[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m 2026-06-07 16:43:50.392 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:162 - setup_envvars_for_vllm: numa_enabled=True[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m 2026-06-07 16:43:50.392 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:165 - setup_envvars_for_vllm: set VLLM_ENABLE_V1_MULTIPROCESSING=0 for NUMA affinity[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 4x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m No module named 'vllm._version'[32m [repeated 4x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 4x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577500) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/__init__.py:1617: UserWarning: Please use the new API settings to control TF32 behavior, such as torch.backends.cudnn.conv.fp32_precision = 'tf32' or torch.backends.cuda.matmul.fp32_precision = 'ieee'. Old settings, e.g, torch.backends.cuda.matmul.allow_tf32 = True, torch.backends.cudnn.allow_tf32 = True, allowTF32CuDNN() and allowTF32CuBLAS() will be deprecated after Pytorch 2.9. Please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:80.)
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577500) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m 2026-06-07 16:43:51.752 | INFO | skyrl_train.utils.numa:set_numa_affinity_for_gpu:340 - NUMA affinity: bound process to CPUs 216-287 (NUMA node 3) for GPU 3[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m 2026-06-07 16:43:51.780 | DEBUG | skyrl_train.utils.numa:_set_membind_via_libnuma:375 - NUMA affinity: set memory preferred to NUMA node 3[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m 2026-06-07 16:43:51.780 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:613 - BaseVLLMInferenceEngine: vllm_v1_disable_multiproc=True, vllm.__version__=dev, VLLM_ENABLE_V1_MULTIPROCESSING=0[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m 2026-06-07 16:43:51.780 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:621 - BaseVLLMInferenceEngine: set VLLM_ENABLE_V1_MULTIPROCESSING=0[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m 2026-06-07 16:43:51.780 | WARNING | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1187 - OpenAI API sampling params overridden: temperature=0.7, top_p=0.95, top_k=20[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m 2026-06-07 16:43:51.824 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1239 - Engine startup stagger: sleeping 2.92s (attempt 1/5) to avoid port collisions[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.[32m [repeated 6x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 6x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576836, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577485) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=2576837, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577489) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577495) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m [W607 16:46:26.579304849 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-041-40-interconnect-1.jupiter.internal]:39565 (errno: 97 - Address family not supported by protocol).
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m [W607 16:46:26.580154880 Utils.hpp:166] Warning: Environment variable NCCL_BLOCKING_WAIT is deprecated; use TORCH_NCCL_BLOCKING_WAIT instead (function operator())
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m [rank0]:[W607 16:46:26.586558911 Utils.hpp:137] Warning: Environment variable TORCH_NCCL_TRACE_BUFFER_SIZE is deprecated; use TORCH_FR_BUFFER_SIZE instead (function operator())
|
||
[36m(AsyncVLLMInferenceEngine pid=2576837, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577489)
|
||
Loading safetensors checkpoint shards: 0% Completed | 0/4 [00:00<?, ?it/s]
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577495) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/__init__.py:1617: UserWarning: Please use the new API settings to control TF32 behavior, such as torch.backends.cudnn.conv.fp32_precision = 'tf32' or torch.backends.cuda.matmul.fp32_precision = 'ieee'. Old settings, e.g, torch.backends.cuda.matmul.allow_tf32 = True, torch.backends.cudnn.allow_tf32 = True, allowTF32CuDNN() and allowTF32CuBLAS() will be deprecated after Pytorch 2.9. Please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:80.)[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m [W607 16:46:26.579359856 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-041-40.jupiter.internal]:52891 (errno: 97 - Address family not supported by protocol).[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m [W607 16:46:26.580154720 Utils.hpp:166] Warning: Environment variable NCCL_BLOCKING_WAIT is deprecated; use TORCH_NCCL_BLOCKING_WAIT instead (function operator())[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m [rank0]:[W607 16:46:26.586350275 Utils.hpp:137] Warning: Environment variable TORCH_NCCL_TRACE_BUFFER_SIZE is deprecated; use TORCH_FR_BUFFER_SIZE instead (function operator())[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577500)
|
||
Loading safetensors checkpoint shards: 25% Completed | 1/4 [00:03<00:10, 3.33s/it]
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577495)
|
||
Loading safetensors checkpoint shards: 0% Completed | 0/4 [00:00<?, ?it/s][32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577500)
|
||
Loading safetensors checkpoint shards: 75% Completed | 3/4 [00:11<00:04, 4.01s/it][32m [repeated 8x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577500)
|
||
[36m(AsyncVLLMInferenceEngine pid=2576836, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577485)
|
||
[36m(AsyncVLLMInferenceEngine pid=2576837, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577489)
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577495)
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577500) The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.
|
||
[36m(AsyncVLLMInferenceEngine pid=2576836, ip=10.128.32.40)[0m 2026-06-07 16:47:11.175 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1296 - Initializing OpenAIServingChat with custom_chat_template read from: chat_templates/qwen3_thinking_acc.jinja2
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577495)
|
||
Loading safetensors checkpoint shards: 100% Completed | 4/4 [00:13<00:00, 3.27s/it][32m [repeated 11x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m 2026-06-07 16:48:36.432 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:152 - setup_envvars_for_vllm: distributed_executor_backend=uni, SKYRL_ENABLE_NUMA_AFFINITY=1, CUDA_VISIBLE_DEVICES=0, VLLM_ENABLE_V1_MULTIPROCESSING=<unset>
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m 2026-06-07 16:48:36.433 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:162 - setup_envvars_for_vllm: numa_enabled=True
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m 2026-06-07 16:48:36.433 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:165 - setup_envvars_for_vllm: set VLLM_ENABLE_V1_MULTIPROCESSING=0 for NUMA affinity
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577495) The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m 2026-06-07 16:47:11.226 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1296 - Initializing OpenAIServingChat with custom_chat_template read from: chat_templates/qwen3_thinking_acc.jinja2[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m 2026-06-07 16:48:37.639 | INFO | skyrl_train.utils.numa:set_numa_affinity_for_gpu:340 - NUMA affinity: bound process to CPUs 0-71 (NUMA node 0) for GPU 0
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m 2026-06-07 16:48:37.989 | DEBUG | skyrl_train.utils.numa:_set_membind_via_libnuma:375 - NUMA affinity: set memory preferred to NUMA node 0
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m 2026-06-07 16:48:37.989 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:613 - BaseVLLMInferenceEngine: vllm_v1_disable_multiproc=True, vllm.__version__=dev, VLLM_ENABLE_V1_MULTIPROCESSING=0
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m 2026-06-07 16:48:37.989 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:621 - BaseVLLMInferenceEngine: set VLLM_ENABLE_V1_MULTIPROCESSING=0
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m 2026-06-07 16:48:37.989 | WARNING | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1187 - OpenAI API sampling params overridden: temperature=0.7, top_p=0.95, top_k=20
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m 2026-06-07 16:48:38.038 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1239 - Engine startup stagger: sleeping 2.40s (attempt 1/5) to avoid port collisions
|
||
[36m(AsyncVLLMInferenceEngine pid=3011433, ip=10.128.32.44)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m 2026-06-07 16:48:36.432 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:152 - setup_envvars_for_vllm: distributed_executor_backend=uni, SKYRL_ENABLE_NUMA_AFFINITY=1, CUDA_VISIBLE_DEVICES=3, VLLM_ENABLE_V1_MULTIPROCESSING=<unset>[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m 2026-06-07 16:48:36.433 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:162 - setup_envvars_for_vllm: numa_enabled=True[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m 2026-06-07 16:48:36.433 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:165 - setup_envvars_for_vllm: set VLLM_ENABLE_V1_MULTIPROCESSING=0 for NUMA affinity[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011432, ip=10.128.32.44)[0m The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.[32m [repeated 4x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m 2026-06-07 16:48:37.639 | INFO | skyrl_train.utils.numa:set_numa_affinity_for_gpu:340 - NUMA affinity: bound process to CPUs 216-287 (NUMA node 3) for GPU 3[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m 2026-06-07 16:48:37.989 | DEBUG | skyrl_train.utils.numa:_set_membind_via_libnuma:375 - NUMA affinity: set memory preferred to NUMA node 3[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m 2026-06-07 16:48:37.989 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:613 - BaseVLLMInferenceEngine: vllm_v1_disable_multiproc=True, vllm.__version__=dev, VLLM_ENABLE_V1_MULTIPROCESSING=0[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m 2026-06-07 16:48:37.989 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:621 - BaseVLLMInferenceEngine: set VLLM_ENABLE_V1_MULTIPROCESSING=0[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m 2026-06-07 16:48:37.989 | WARNING | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1187 - OpenAI API sampling params overridden: temperature=0.7, top_p=0.95, top_k=20[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m 2026-06-07 16:48:38.038 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1239 - Engine startup stagger: sleeping 2.04s (attempt 1/5) to avoid port collisions[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011432, ip=10.128.32.44)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.[32m [repeated 7x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m No module named 'vllm._version'
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m from .version import __version__, __version_tuple__ # isort:skip
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740424) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/__init__.py:1617: UserWarning: Please use the new API settings to control TF32 behavior, such as torch.backends.cudnn.conv.fp32_precision = 'tf32' or torch.backends.cuda.matmul.fp32_precision = 'ieee'. Old settings, e.g, torch.backends.cuda.matmul.allow_tf32 = True, torch.backends.cudnn.allow_tf32 = True, allowTF32CuDNN() and allowTF32CuBLAS() will be deprecated after Pytorch 2.9. Please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:80.)
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740424) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 6x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m No module named 'vllm._version'[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739605, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740434) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=2739603, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740438) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740428) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m [W607 16:49:22.529996202 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-041-34.jupiter.internal]:58359 (errno: 97 - Address family not supported by protocol).
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m [W607 16:49:22.533237515 Utils.hpp:166] Warning: Environment variable NCCL_BLOCKING_WAIT is deprecated; use TORCH_NCCL_BLOCKING_WAIT instead (function operator())
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m [rank0]:[W607 16:49:22.539139353 Utils.hpp:137] Warning: Environment variable TORCH_NCCL_TRACE_BUFFER_SIZE is deprecated; use TORCH_FR_BUFFER_SIZE instead (function operator())
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740424)
|
||
Loading safetensors checkpoint shards: 0% Completed | 0/4 [00:00<?, ?it/s]
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740428) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/__init__.py:1617: UserWarning: Please use the new API settings to control TF32 behavior, such as torch.backends.cudnn.conv.fp32_precision = 'tf32' or torch.backends.cuda.matmul.fp32_precision = 'ieee'. Old settings, e.g, torch.backends.cuda.matmul.allow_tf32 = True, torch.backends.cudnn.allow_tf32 = True, allowTF32CuDNN() and allowTF32CuBLAS() will be deprecated after Pytorch 2.9. Please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:80.)[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m [W607 16:49:22.530487424 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-041-34-interconnect-1.jupiter.internal]:58341 (errno: 97 - Address family not supported by protocol).[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m [W607 16:49:22.533253643 Utils.hpp:166] Warning: Environment variable NCCL_BLOCKING_WAIT is deprecated; use TORCH_NCCL_BLOCKING_WAIT instead (function operator())[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m [rank0]:[W607 16:49:22.539392180 Utils.hpp:137] Warning: Environment variable TORCH_NCCL_TRACE_BUFFER_SIZE is deprecated; use TORCH_FR_BUFFER_SIZE instead (function operator())[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740424)
|
||
Loading safetensors checkpoint shards: 25% Completed | 1/4 [00:03<00:10, 3.55s/it]
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740428)
|
||
Loading safetensors checkpoint shards: 0% Completed | 0/4 [00:00<?, ?it/s][32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740424)
|
||
Loading safetensors checkpoint shards: 75% Completed | 3/4 [00:11<00:03, 3.94s/it][32m [repeated 8x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740424)
|
||
[36m(AsyncVLLMInferenceEngine pid=2739605, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740434)
|
||
[36m(AsyncVLLMInferenceEngine pid=2739603, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740438)
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740428)
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740424) The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m 2026-06-07 16:50:33.345 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1296 - Initializing OpenAIServingChat with custom_chat_template read from: chat_templates/qwen3_thinking_acc.jinja2
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740428)
|
||
Loading safetensors checkpoint shards: 100% Completed | 4/4 [00:12<00:00, 3.22s/it][32m [repeated 11x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631821, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632466) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/__init__.py:1617: UserWarning: Please use the new API settings to control TF32 behavior, such as torch.backends.cudnn.conv.fp32_precision = 'tf32' or torch.backends.cuda.matmul.fp32_precision = 'ieee'. Old settings, e.g, torch.backends.cuda.matmul.allow_tf32 = True, torch.backends.cudnn.allow_tf32 = True, allowTF32CuDNN() and allowTF32CuBLAS() will be deprecated after Pytorch 2.9. Please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:80.)
|
||
[36m(AsyncVLLMInferenceEngine pid=2631821, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632466) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=2739603, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740438) The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m 2026-06-07 16:50:33.384 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1296 - Initializing OpenAIServingChat with custom_chat_template read from: chat_templates/qwen3_thinking_acc.jinja2[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631819, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632450) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632471) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=2631821, ip=10.128.32.36)[0m [W607 16:50:43.765319760 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-041-36.jupiter.internal]:48845 (errno: 97 - Address family not supported by protocol).
|
||
[36m(AsyncVLLMInferenceEngine pid=2631821, ip=10.128.32.36)[0m [W607 16:50:43.767120941 Utils.hpp:166] Warning: Environment variable NCCL_BLOCKING_WAIT is deprecated; use TORCH_NCCL_BLOCKING_WAIT instead (function operator())
|
||
[36m(AsyncVLLMInferenceEngine pid=2631821, ip=10.128.32.36)[0m [rank0]:[W607 16:50:43.772874492 Utils.hpp:137] Warning: Environment variable TORCH_NCCL_TRACE_BUFFER_SIZE is deprecated; use TORCH_FR_BUFFER_SIZE instead (function operator())
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632455) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632455)
|
||
Loading safetensors checkpoint shards: 0% Completed | 0/4 [00:00<?, ?it/s]
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632455) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/__init__.py:1617: UserWarning: Please use the new API settings to control TF32 behavior, such as torch.backends.cudnn.conv.fp32_precision = 'tf32' or torch.backends.cuda.matmul.fp32_precision = 'ieee'. Old settings, e.g, torch.backends.cuda.matmul.allow_tf32 = True, torch.backends.cudnn.allow_tf32 = True, allowTF32CuDNN() and allowTF32CuBLAS() will be deprecated after Pytorch 2.9. Please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:80.)[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m [W607 16:50:43.107187519 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-041-36-interconnect-1.jupiter.internal]:42177 (errno: 97 - Address family not supported by protocol).[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m [W607 16:50:43.107636054 Utils.hpp:166] Warning: Environment variable NCCL_BLOCKING_WAIT is deprecated; use TORCH_NCCL_BLOCKING_WAIT instead (function operator())[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m [rank0]:[W607 16:50:43.109791084 Utils.hpp:137] Warning: Environment variable TORCH_NCCL_TRACE_BUFFER_SIZE is deprecated; use TORCH_FR_BUFFER_SIZE instead (function operator())[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m 2026-06-07 16:52:06.305 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:152 - setup_envvars_for_vllm: distributed_executor_backend=uni, SKYRL_ENABLE_NUMA_AFFINITY=1, CUDA_VISIBLE_DEVICES=0, VLLM_ENABLE_V1_MULTIPROCESSING=<unset>
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m 2026-06-07 16:52:06.306 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:162 - setup_envvars_for_vllm: numa_enabled=True
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m 2026-06-07 16:52:06.306 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:165 - setup_envvars_for_vllm: set VLLM_ENABLE_V1_MULTIPROCESSING=0 for NUMA affinity
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632455)
|
||
Loading safetensors checkpoint shards: 25% Completed | 1/4 [00:04<00:12, 4.01s/it]
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m 2026-06-07 16:52:07.518 | INFO | skyrl_train.utils.numa:set_numa_affinity_for_gpu:340 - NUMA affinity: bound process to CPUs 0-71 (NUMA node 0) for GPU 0
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632471)
|
||
Loading safetensors checkpoint shards: 0% Completed | 0/4 [00:00<?, ?it/s][32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m 2026-06-07 16:52:07.792 | DEBUG | skyrl_train.utils.numa:_set_membind_via_libnuma:375 - NUMA affinity: set memory preferred to NUMA node 0
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m 2026-06-07 16:52:07.792 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:613 - BaseVLLMInferenceEngine: vllm_v1_disable_multiproc=True, vllm.__version__=dev, VLLM_ENABLE_V1_MULTIPROCESSING=0
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m 2026-06-07 16:52:07.792 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:621 - BaseVLLMInferenceEngine: set VLLM_ENABLE_V1_MULTIPROCESSING=0
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m 2026-06-07 16:52:07.792 | WARNING | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1187 - OpenAI API sampling params overridden: temperature=0.7, top_p=0.95, top_k=20
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m 2026-06-07 16:52:07.835 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1239 - Engine startup stagger: sleeping 2.77s (attempt 1/5) to avoid port collisions
|
||
[36m(AsyncVLLMInferenceEngine pid=2705804, ip=10.128.32.38)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.
|
||
[36m(AsyncVLLMInferenceEngine pid=2705804, ip=10.128.32.38)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.
|
||
[36m(AsyncVLLMInferenceEngine pid=2705804, ip=10.128.32.38)[0m The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.
|
||
[36m(AsyncVLLMInferenceEngine pid=2705804, ip=10.128.32.38)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e
|
||
[36m(AsyncVLLMInferenceEngine pid=2705804, ip=10.128.32.38)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m No module named 'vllm._version'
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m from .version import __version__, __version_tuple__ # isort:skip
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m 2026-06-07 16:52:06.305 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:152 - setup_envvars_for_vllm: distributed_executor_backend=uni, SKYRL_ENABLE_NUMA_AFFINITY=1, CUDA_VISIBLE_DEVICES=3, VLLM_ENABLE_V1_MULTIPROCESSING=<unset>[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m 2026-06-07 16:52:06.306 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:162 - setup_envvars_for_vllm: numa_enabled=True[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m 2026-06-07 16:52:06.306 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:165 - setup_envvars_for_vllm: set VLLM_ENABLE_V1_MULTIPROCESSING=0 for NUMA affinity[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632471)
|
||
Loading safetensors checkpoint shards: 50% Completed | 2/4 [00:08<00:08, 4.06s/it][32m [repeated 7x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m 2026-06-07 16:52:07.518 | INFO | skyrl_train.utils.numa:set_numa_affinity_for_gpu:340 - NUMA affinity: bound process to CPUs 216-287 (NUMA node 3) for GPU 3[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m 2026-06-07 16:52:07.791 | DEBUG | skyrl_train.utils.numa:_set_membind_via_libnuma:375 - NUMA affinity: set memory preferred to NUMA node 3[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m 2026-06-07 16:52:07.791 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:613 - BaseVLLMInferenceEngine: vllm_v1_disable_multiproc=True, vllm.__version__=dev, VLLM_ENABLE_V1_MULTIPROCESSING=0[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m 2026-06-07 16:52:07.791 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:621 - BaseVLLMInferenceEngine: set VLLM_ENABLE_V1_MULTIPROCESSING=0[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m 2026-06-07 16:52:07.791 | WARNING | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1187 - OpenAI API sampling params overridden: temperature=0.7, top_p=0.95, top_k=20[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m 2026-06-07 16:52:07.835 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1239 - Engine startup stagger: sleeping 1.87s (attempt 1/5) to avoid port collisions[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.[32m [repeated 6x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632455)
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 6x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m No module named 'vllm._version'[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631821, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632466)
|
||
[36m(AsyncVLLMInferenceEngine pid=2631819, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632450)
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632471)
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632471)
|
||
Loading safetensors checkpoint shards: 100% Completed | 4/4 [00:13<00:00, 3.34s/it][32m [repeated 12x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m 2026-06-07 16:52:19.686 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1296 - Initializing OpenAIServingChat with custom_chat_template read from: chat_templates/qwen3_thinking_acc.jinja2
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632471) The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.[32m [repeated 4x across cluster][0m
|
||
[36m(pid=2994914, ip=10.128.32.39)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:
|
||
[36m(pid=2994914, ip=10.128.32.39)[0m No module named 'vllm._version'
|
||
[36m(pid=2994914, ip=10.128.32.39)[0m from .version import __version__, __version_tuple__ # isort:skip
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m 2026-06-07 16:52:19.639 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1296 - Initializing OpenAIServingChat with custom_chat_template read from: chat_templates/qwen3_thinking_acc.jinja2[32m [repeated 3x across cluster][0m
|
||
[36m(pid=2995275, ip=10.128.32.39)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:
|
||
[36m(pid=2995275, ip=10.128.32.39)[0m No module named 'vllm._version'
|
||
[36m(pid=2995275, ip=10.128.32.39)[0m from .version import __version__, __version_tuple__ # isort:skip
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012340) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/__init__.py:1617: UserWarning: Please use the new API settings to control TF32 behavior, such as torch.backends.cudnn.conv.fp32_precision = 'tf32' or torch.backends.cuda.matmul.fp32_precision = 'ieee'. Old settings, e.g, torch.backends.cuda.matmul.allow_tf32 = True, torch.backends.cudnn.allow_tf32 = True, allowTF32CuDNN() and allowTF32CuBLAS() will be deprecated after Pytorch 2.9. Please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:80.)
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012340) _C._set_float32_matmul_precision(precision)
|
||
[36m(pid=2995274, ip=10.128.32.39)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 2x across cluster][0m
|
||
[36m(pid=2995274, ip=10.128.32.39)[0m No module named 'vllm._version'[32m [repeated 2x across cluster][0m
|
||
[36m(pid=2995274, ip=10.128.32.39)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011433, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012342) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=3011432, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012343) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012341) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m [W607 16:54:12.735426276 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-041-44.jupiter.internal]:48903 (errno: 97 - Address family not supported by protocol).
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m [W607 16:54:12.739040171 Utils.hpp:166] Warning: Environment variable NCCL_BLOCKING_WAIT is deprecated; use TORCH_NCCL_BLOCKING_WAIT instead (function operator())
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m [rank0]:[W607 16:54:12.745053709 Utils.hpp:137] Warning: Environment variable TORCH_NCCL_TRACE_BUFFER_SIZE is deprecated; use TORCH_FR_BUFFER_SIZE instead (function operator())
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012340)
|
||
Loading safetensors checkpoint shards: 0% Completed | 0/4 [00:00<?, ?it/s]
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012341) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/__init__.py:1617: UserWarning: Please use the new API settings to control TF32 behavior, such as torch.backends.cudnn.conv.fp32_precision = 'tf32' or torch.backends.cuda.matmul.fp32_precision = 'ieee'. Old settings, e.g, torch.backends.cuda.matmul.allow_tf32 = True, torch.backends.cudnn.allow_tf32 = True, allowTF32CuDNN() and allowTF32CuBLAS() will be deprecated after Pytorch 2.9. Please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:80.)[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m [W607 16:54:12.735348261 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-041-44.jupiter.internal]:53241 (errno: 97 - Address family not supported by protocol).[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m [W607 16:54:12.739023275 Utils.hpp:166] Warning: Environment variable NCCL_BLOCKING_WAIT is deprecated; use TORCH_NCCL_BLOCKING_WAIT instead (function operator())[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m [rank0]:[W607 16:54:12.745099244 Utils.hpp:137] Warning: Environment variable TORCH_NCCL_TRACE_BUFFER_SIZE is deprecated; use TORCH_FR_BUFFER_SIZE instead (function operator())[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012340)
|
||
Loading safetensors checkpoint shards: 25% Completed | 1/4 [00:03<00:11, 3.83s/it]
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012341)
|
||
Loading safetensors checkpoint shards: 0% Completed | 0/4 [00:00<?, ?it/s][32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012340)
|
||
Loading safetensors checkpoint shards: 75% Completed | 3/4 [00:12<00:04, 4.15s/it][32m [repeated 8x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012340)
|
||
[36m(AsyncVLLMInferenceEngine pid=3011433, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012342)
|
||
[36m(AsyncVLLMInferenceEngine pid=3011432, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012343)
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012341)
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012340) The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m 2026-06-07 16:55:23.546 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1296 - Initializing OpenAIServingChat with custom_chat_template read from: chat_templates/qwen3_thinking_acc.jinja2
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012341)
|
||
Loading safetensors checkpoint shards: 100% Completed | 4/4 [00:13<00:00, 3.35s/it][32m [repeated 11x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706759) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/__init__.py:1617: UserWarning: Please use the new API settings to control TF32 behavior, such as torch.backends.cudnn.conv.fp32_precision = 'tf32' or torch.backends.cuda.matmul.fp32_precision = 'ieee'. Old settings, e.g, torch.backends.cuda.matmul.allow_tf32 = True, torch.backends.cudnn.allow_tf32 = True, allowTF32CuDNN() and allowTF32CuBLAS() will be deprecated after Pytorch 2.9. Please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:80.)
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706759) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012341) The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m 2026-06-07 16:55:23.611 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1296 - Initializing OpenAIServingChat with custom_chat_template read from: chat_templates/qwen3_thinking_acc.jinja2[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705804, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706735) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=2705802, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706743) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706739) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m [W607 16:55:33.477299204 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-041-38.jupiter.internal]:39711 (errno: 97 - Address family not supported by protocol).
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m [W607 16:55:33.479730539 Utils.hpp:166] Warning: Environment variable NCCL_BLOCKING_WAIT is deprecated; use TORCH_NCCL_BLOCKING_WAIT instead (function operator())
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m [rank0]:[W607 16:55:33.486001944 Utils.hpp:137] Warning: Environment variable TORCH_NCCL_TRACE_BUFFER_SIZE is deprecated; use TORCH_FR_BUFFER_SIZE instead (function operator())
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m 2026-06-07 16:55:38.696 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:152 - setup_envvars_for_vllm: distributed_executor_backend=uni, SKYRL_ENABLE_NUMA_AFFINITY=1, CUDA_VISIBLE_DEVICES=0, VLLM_ENABLE_V1_MULTIPROCESSING=<unset>
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m 2026-06-07 16:55:38.697 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:162 - setup_envvars_for_vllm: numa_enabled=True
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m 2026-06-07 16:55:38.697 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:165 - setup_envvars_for_vllm: set VLLM_ENABLE_V1_MULTIPROCESSING=0 for NUMA affinity
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706739) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/__init__.py:1617: UserWarning: Please use the new API settings to control TF32 behavior, such as torch.backends.cudnn.conv.fp32_precision = 'tf32' or torch.backends.cuda.matmul.fp32_precision = 'ieee'. Old settings, e.g, torch.backends.cuda.matmul.allow_tf32 = True, torch.backends.cudnn.allow_tf32 = True, allowTF32CuDNN() and allowTF32CuBLAS() will be deprecated after Pytorch 2.9. Please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:80.)[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m [W607 16:55:33.476236797 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-041-38-interconnect-1.jupiter.internal]:41289 (errno: 97 - Address family not supported by protocol).[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m [W607 16:55:33.479705003 Utils.hpp:166] Warning: Environment variable NCCL_BLOCKING_WAIT is deprecated; use TORCH_NCCL_BLOCKING_WAIT instead (function operator())[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m [rank0]:[W607 16:55:33.484923505 Utils.hpp:137] Warning: Environment variable TORCH_NCCL_TRACE_BUFFER_SIZE is deprecated; use TORCH_FR_BUFFER_SIZE instead (function operator())[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m 2026-06-07 16:55:40.047 | INFO | skyrl_train.utils.numa:set_numa_affinity_for_gpu:340 - NUMA affinity: bound process to CPUs 0-71 (NUMA node 0) for GPU 0
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m 2026-06-07 16:55:40.076 | DEBUG | skyrl_train.utils.numa:_set_membind_via_libnuma:375 - NUMA affinity: set memory preferred to NUMA node 0
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m 2026-06-07 16:55:40.077 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:613 - BaseVLLMInferenceEngine: vllm_v1_disable_multiproc=True, vllm.__version__=dev, VLLM_ENABLE_V1_MULTIPROCESSING=0
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m 2026-06-07 16:55:40.077 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:621 - BaseVLLMInferenceEngine: set VLLM_ENABLE_V1_MULTIPROCESSING=0
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m 2026-06-07 16:55:40.077 | WARNING | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1187 - OpenAI API sampling params overridden: temperature=0.7, top_p=0.95, top_k=20
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m 2026-06-07 16:55:40.144 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1239 - Engine startup stagger: sleeping 2.39s (attempt 1/5) to avoid port collisions
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m 2026-06-07 16:55:38.696 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:152 - setup_envvars_for_vllm: distributed_executor_backend=uni, SKYRL_ENABLE_NUMA_AFFINITY=1, CUDA_VISIBLE_DEVICES=3, VLLM_ENABLE_V1_MULTIPROCESSING=<unset>[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m 2026-06-07 16:55:38.697 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:162 - setup_envvars_for_vllm: numa_enabled=True[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m 2026-06-07 16:55:38.697 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:165 - setup_envvars_for_vllm: set VLLM_ENABLE_V1_MULTIPROCESSING=0 for NUMA affinity[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m 2026-06-07 16:55:40.048 | INFO | skyrl_train.utils.numa:set_numa_affinity_for_gpu:340 - NUMA affinity: bound process to CPUs 216-287 (NUMA node 3) for GPU 3[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m 2026-06-07 16:55:40.076 | DEBUG | skyrl_train.utils.numa:_set_membind_via_libnuma:375 - NUMA affinity: set memory preferred to NUMA node 3[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m 2026-06-07 16:55:40.076 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:613 - BaseVLLMInferenceEngine: vllm_v1_disable_multiproc=True, vllm.__version__=dev, VLLM_ENABLE_V1_MULTIPROCESSING=0[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m 2026-06-07 16:55:40.076 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:621 - BaseVLLMInferenceEngine: set VLLM_ENABLE_V1_MULTIPROCESSING=0[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m 2026-06-07 16:55:40.076 | WARNING | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1187 - OpenAI API sampling params overridden: temperature=0.7, top_p=0.95, top_k=20[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m 2026-06-07 16:55:40.144 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1239 - Engine startup stagger: sleeping 2.33s (attempt 1/5) to avoid port collisions[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.[32m [repeated 4x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e
|
||
[36m(AsyncVLLMInferenceEngine pid=2831685, ip=10.128.32.45)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:
|
||
[36m(AsyncVLLMInferenceEngine pid=2831685, ip=10.128.32.45)[0m No module named 'vllm._version'
|
||
[36m(AsyncVLLMInferenceEngine pid=2831685, ip=10.128.32.45)[0m from .version import __version__, __version_tuple__ # isort:skip
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706759)
|
||
Loading safetensors checkpoint shards: 0% Completed | 0/4 [00:00<?, ?it/s]
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 6x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m No module named 'vllm._version'[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706759)
|
||
Loading safetensors checkpoint shards: 25% Completed | 1/4 [00:03<00:11, 3.97s/it]
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706739)
|
||
Loading safetensors checkpoint shards: 0% Completed | 0/4 [00:00<?, ?it/s][32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706759)
|
||
Loading safetensors checkpoint shards: 75% Completed | 3/4 [00:12<00:04, 4.03s/it][32m [repeated 8x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706759)
|
||
[36m(AsyncVLLMInferenceEngine pid=2705804, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706735)
|
||
[36m(AsyncVLLMInferenceEngine pid=2705802, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706743)
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706739)
|
||
[36m(AsyncVLLMInferenceEngine pid=2705804, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706735) The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m 2026-06-07 16:57:12.397 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1296 - Initializing OpenAIServingChat with custom_chat_template read from: chat_templates/qwen3_thinking_acc.jinja2
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706739)
|
||
Loading safetensors checkpoint shards: 100% Completed | 4/4 [00:13<00:00, 3.35s/it][32m [repeated 11x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832747) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/__init__.py:1617: UserWarning: Please use the new API settings to control TF32 behavior, such as torch.backends.cudnn.conv.fp32_precision = 'tf32' or torch.backends.cuda.matmul.fp32_precision = 'ieee'. Old settings, e.g, torch.backends.cuda.matmul.allow_tf32 = True, torch.backends.cudnn.allow_tf32 = True, allowTF32CuDNN() and allowTF32CuBLAS() will be deprecated after Pytorch 2.9. Please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:80.)
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832747) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=2705802, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706743) The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m 2026-06-07 16:57:12.355 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1296 - Initializing OpenAIServingChat with custom_chat_template read from: chat_templates/qwen3_thinking_acc.jinja2[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831684, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832749) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=2831685, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832737) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832741) _C._set_float32_matmul_precision(precision)
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m [W607 17:00:00.617336904 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-041-45.jupiter.internal]:41873 (errno: 97 - Address family not supported by protocol).
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m [W607 17:00:00.621348216 Utils.hpp:166] Warning: Environment variable NCCL_BLOCKING_WAIT is deprecated; use TORCH_NCCL_BLOCKING_WAIT instead (function operator())
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m [rank0]:[W607 17:00:00.627687930 Utils.hpp:137] Warning: Environment variable TORCH_NCCL_TRACE_BUFFER_SIZE is deprecated; use TORCH_FR_BUFFER_SIZE instead (function operator())
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832747)
|
||
Loading safetensors checkpoint shards: 0% Completed | 0/4 [00:00<?, ?it/s]
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832741) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/__init__.py:1617: UserWarning: Please use the new API settings to control TF32 behavior, such as torch.backends.cudnn.conv.fp32_precision = 'tf32' or torch.backends.cuda.matmul.fp32_precision = 'ieee'. Old settings, e.g, torch.backends.cuda.matmul.allow_tf32 = True, torch.backends.cudnn.allow_tf32 = True, allowTF32CuDNN() and allowTF32CuBLAS() will be deprecated after Pytorch 2.9. Please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:80.)[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m [W607 17:00:00.617384583 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-041-45.jupiter.internal]:60113 (errno: 97 - Address family not supported by protocol).[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m [W607 17:00:00.621367768 Utils.hpp:166] Warning: Environment variable NCCL_BLOCKING_WAIT is deprecated; use TORCH_NCCL_BLOCKING_WAIT instead (function operator())[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m [rank0]:[W607 17:00:00.627463967 Utils.hpp:137] Warning: Environment variable TORCH_NCCL_TRACE_BUFFER_SIZE is deprecated; use TORCH_FR_BUFFER_SIZE instead (function operator())[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832747)
|
||
Loading safetensors checkpoint shards: 25% Completed | 1/4 [00:03<00:11, 3.98s/it]
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832741)
|
||
Loading safetensors checkpoint shards: 0% Completed | 0/4 [00:00<?, ?it/s][32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832747)
|
||
Loading safetensors checkpoint shards: 75% Completed | 3/4 [00:12<00:04, 4.08s/it][32m [repeated 8x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832747)
|
||
[36m(AsyncVLLMInferenceEngine pid=2831684, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832749)
|
||
[36m(AsyncVLLMInferenceEngine pid=2831685, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832737)
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832741)
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832747) The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.
|
||
[36m(AsyncVLLMInferenceEngine pid=2831684, ip=10.128.32.45)[0m 2026-06-07 17:01:11.571 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1296 - Initializing OpenAIServingChat with custom_chat_template read from: chat_templates/qwen3_thinking_acc.jinja2
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832741)
|
||
Loading safetensors checkpoint shards: 100% Completed | 4/4 [00:13<00:00, 3.39s/it][32m [repeated 11x across cluster][0m
|
||
[36m(FSDPPolicyWorkerBase pid=2662625, ip=10.128.32.46)[0m 2026-06-07 17:02:34 INFO [ipv4-debug] hostname=jpbo-041-46.jupiter.internal
|
||
[36m(FSDPPolicyWorkerBase pid=2662625, ip=10.128.32.46)[0m 2026-06-07 17:02:34 INFO [ipv4-debug] _global_node.node_ip_address=10.128.32.46
|
||
[36m(FSDPPolicyWorkerBase pid=2662625, ip=10.128.32.46)[0m 2026-06-07 17:02:34 INFO [ipv4-debug] get_node_ip_address()=10.128.32.46
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832741) The tokenizer you are loading from '/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m 2026-06-07 17:01:11.656 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1296 - Initializing OpenAIServingChat with custom_chat_template read from: chat_templates/qwen3_thinking_acc.jinja2[32m [repeated 3x across cluster][0m
|
||
[33m(raylet, ip=10.128.32.46)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e
|
||
[36m(FSDPPolicyWorkerBase pid=2662625, ip=10.128.32.46)[0m [W607 17:02:34.731583451 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-041-46-interconnect-1.jupiter.internal]:41111 (errno: 97 - Address family not supported by protocol).
|
||
[36m(skyrl_entrypoint pid=391889)[0m [32m2026-06-07 17:02:34.937[0m | [1m[32mINFO [0m | [36mskyrl_train.workers.worker[0m:[36m_initiate_actors[0m:[36m636[0m - [1m[32mInitializing process group for RayActorGroup[0m
|
||
[36m(pid=2663493, ip=10.128.32.46)[0m Using blocking ray.get inside async actor. This blocks the event loop. Please use `await` on object ref with asyncio.gather if you want to yield execution to the event loop instead.
|
||
[33m(raylet)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 41x across cluster][0m
|
||
[36m(pid=394709)[0m Using blocking ray.get inside async actor. This blocks the event loop. Please use `await` on object ref with asyncio.gather if you want to yield execution to the event loop instead.[32m [repeated 3x across cluster][0m
|
||
[36m(pid=2663493, ip=10.128.32.46)[0m [2026-06-07 17:03:05,886 E 2663493 2663634] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14
|
||
[36m(FSDPPolicyWorkerBase pid=2663493, ip=10.128.32.46)[0m [W607 17:03:10.567912961 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-041-46-interconnect-1.jupiter.internal]:41111 (errno: 97 - Address family not supported by protocol).
|
||
[36m(FSDPPolicyWorkerBase pid=2663493, ip=10.128.32.46)[0m [W607 17:03:10.570281892 Utils.hpp:166] Warning: Environment variable NCCL_BLOCKING_WAIT is deprecated; use TORCH_NCCL_BLOCKING_WAIT instead (function operator())
|
||
[36m(pid=394710)[0m Using blocking ray.get inside async actor. This blocks the event loop. Please use `await` on object ref with asyncio.gather if you want to yield execution to the event loop instead.[32m [repeated 3x across cluster][0m
|
||
2026-06-07 17:12:36.032 | ERROR | __main__:main:134 - Training failed: [36mray::skyrl_entrypoint()[39m (pid=391889, ip=10.128.32.33)
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train/examples/terminal_bench/entrypoints/main_tbench.py", line 105, in skyrl_entrypoint
|
||
exp.run()
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train/skyrl_train/entrypoints/main_base.py", line 481, in run
|
||
trainer = self._setup_trainer()
|
||
^^^^^^^^^^^^^^^^^^^^^
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train/skyrl_train/entrypoints/main_base.py", line 450, in _setup_trainer
|
||
trainer.build_models(PolicyWorker, CriticWorker, RefWorker, policy_pg=self.policy_pg)
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train/skyrl_train/trainer.py", line 768, in build_models
|
||
policy_model = PPORayActorGroup(
|
||
^^^^^^^^^^^^^^^^^
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train/skyrl_train/workers/worker.py", line 511, in __init__
|
||
self._initiate_actors(pg, num_gpus_per_actor)
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train/skyrl_train/workers/worker.py", line 637, in _initiate_actors
|
||
ray.get([actor.init_worker_process_group.remote() for actor in self._actor_handlers])
|
||
^^^^^^^^^^^^^^^^^^^
|
||
^^^^^^^^^^^^^^^^^^^^^
|
||
^^^^^^^^^^^^^^^^^^^
|
||
ray.exceptions.RayTaskError(DistStoreError): [36mray::FSDPPolicyWorkerBase.init_worker_process_group()[39m (pid=2662625, ip=10.128.32.46, actor_id=b5939e2b137362308272f3e202000000, repr=<skyrl_train.workers.fsdp.fsdp_worker.FSDPPolicyWorkerBase object at 0x400df5e2cf50>)
|
||
File "/e/scratch/jureap59/feuer1/miniforge3/lib/python3.12/concurrent/futures/_base.py", line 449, in result
|
||
return self.__get_result()
|
||
^^^^^^^^^^^^^^^^^^^
|
||
File "/e/scratch/jureap59/feuer1/miniforge3/lib/python3.12/concurrent/futures/_base.py", line 401, in __get_result
|
||
raise self._exception
|
||
^^^^^^^^^^^^^^^^^^^^^
|
||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train/skyrl_train/workers/worker.py", line 149, in init_worker_process_group
|
||
torch.distributed.init_process_group(
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/distributed/c10d_logger.py", line 81, in wrapper
|
||
return func(*args, **kwargs)
|
||
^^^^^^^^^^^^^^^^^^^^^
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/distributed/c10d_logger.py", line 95, in wrapper
|
||
func_return = func(*args, **kwargs)
|
||
^^^^^^^^^^^^^^^^^^^^^
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/distributed/distributed_c10d.py", line 1762, in init_process_group
|
||
store, rank, world_size = next(rendezvous_iterator)
|
||
^^^^^^^^^^^^^^^^^^^^^^^^^
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/distributed/rendezvous.py", line 278, in _env_rendezvous_handler
|
||
store = _create_c10d_store(
|
||
^^^^^^^^^^^^^^^^^^^
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/distributed/rendezvous.py", line 198, in _create_c10d_store
|
||
return TCPStore(
|
||
^^^^^^^^^
|
||
torch.distributed.DistStoreError: Timed out after 601 seconds waiting for clients. 4/8 clients joined.
|
||
Exception raised from waitForWorkers at /pytorch/torch/csrc/distributed/c10d/TCPStore.cpp:396 (most recent call first):
|
||
frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0xb0 (0x4000bf09c700 in /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/lib/libc10.so)
|
||
frame #1: <unknown function> + 0x5e9c9c0 (0x400ae12fc9c0 in /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/lib/libtorch_cpu.so)
|
||
frame #2: c10d::TCPStore::waitForWorkers() + 0x350 (0x400ae1391410 in /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/lib/libtorch_cpu.so)
|
||
frame #3: c10d::TCPStore::TCPStore(std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, c10d::TCPStoreOptions const&) + 0x468 (0x400ae13918c8 in /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/lib/libtorch_cpu.so)
|
||
frame #4: <unknown function> + 0x109a094 (0x4000be86a094 in /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/lib/libtorch_python.so)
|
||
frame #5: <unknown function> + 0x113236c (0x4000be90236c in /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/lib/libtorch_python.so)
|
||
frame #6: <unknown function> + 0x5d6d60 (0x4000bdda6d60 in /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/lib/libtorch_python.so)
|
||
frame #7: <unknown function> + 0x1b7a38 (0xaaaacb007a38 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #8: _PyObject_MakeTpCall + 0x98 (0xaaaacafb5db8 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #9: <unknown function> + 0x169f50 (0xaaaacafb9f50 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #10: <unknown function> + 0x1682e4 (0xaaaacafb82e4 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #11: <unknown function> + 0x1e0ce8 (0xaaaacb030ce8 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #12: <unknown function> + 0x1d7ddc (0xaaaacb027ddc in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #13: <unknown function> + 0x646b0c (0x4000bde16b0c in /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/lib/libtorch_python.so)
|
||
frame #14: _PyObject_MakeTpCall + 0x98 (0xaaaacafb5db8 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #15: _PyEval_EvalFrameDefault + 0x280c (0xaaaacb0bae54 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #16: <unknown function> + 0x1808c0 (0xaaaacafd08c0 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #17: <unknown function> + 0x182bf8 (0xaaaacafd2bf8 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #18: <unknown function> + 0x25fd30 (0xaaaacb0afd30 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #19: <unknown function> + 0x1b7d20 (0xaaaacb007d20 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #20: PyObject_Vectorcall + 0x54 (0xaaaacafb60e4 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #21: _PyEval_EvalFrameDefault + 0x280c (0xaaaacb0bae54 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #22: <unknown function> + 0x1808c0 (0xaaaacafd08c0 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #23: <unknown function> + 0x1835c8 (0xaaaacafd35c8 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #24: <unknown function> + 0x84b708 (0x40000a80b708 in /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/ray/_raylet.so)
|
||
frame #25: <unknown function> + 0x865fe4 (0x40000a825fe4 in /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/ray/_raylet.so)
|
||
frame #26: <unknown function> + 0x823470 (0x40000a7e3470 in /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/ray/_raylet.so)
|
||
frame #27: <unknown function> + 0x830110 (0x40000a7f0110 in /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/ray/_raylet.so)
|
||
frame #28: <unknown function> + 0x174a60 (0xaaaacafc4a60 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #29: PyObject_VectorcallMethod + 0xa4 (0xaaaacafb6270 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #30: PyIter_Send + 0xbc (0xaaaacaf9b62c in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #31: <unknown function> + 0xaa04 (0x40000ca7aa04 in /e/scratch/jureap59/feuer1/miniforge3/lib/python3.12/lib-dynload/_asyncio.cpython-312-aarch64-linux-gnu.so)
|
||
frame #32: <unknown function> + 0xbcc4 (0x40000ca7bcc4 in /e/scratch/jureap59/feuer1/miniforge3/lib/python3.12/lib-dynload/_asyncio.cpython-312-aarch64-linux-gnu.so)
|
||
frame #33: _PyObject_MakeTpCall + 0x98 (0xaaaacafb5db8 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #34: <unknown function> + 0x28cfac (0xaaaacb0dcfac in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #35: <unknown function> + 0x1b7b68 (0xaaaacb007b68 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #36: _PyEval_EvalFrameDefault + 0x52ac (0xaaaacb0bd8f4 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #37: <unknown function> + 0x83c564 (0x40000a7fc564 in /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/ray/_raylet.so)
|
||
frame #38: _PyEval_EvalFrameDefault + 0x52ac (0xaaaacb0bd8f4 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #39: <unknown function> + 0x169f88 (0xaaaacafb9f88 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #40: <unknown function> + 0x36002c (0xaaaacb1b002c in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #41: <unknown function> + 0x2e4024 (0xaaaacb134024 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #42: <unknown function> + 0x80e00 (0x400009730e00 in /lib64/libc.so.6)
|
||
frame #43: <unknown function> + 0xeb49c (0x40000979b49c in /lib64/libc.so.6)
|
||
2026-06-07 17:12:36.033 | INFO | __main__:main:137 - Shutting down Ray on head node...
|
||
[36m(pid=394710)[0m [2026-06-07 17:03:06,091 E 394710 394918] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 6x across cluster][0m
|
||
[36m(FSDPPolicyWorkerBase pid=2663495, ip=10.128.32.46)[0m [W607 17:03:10.575118569 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-041-46.jupiter.internal]:41111 (errno: 97 - Address family not supported by protocol).[32m [repeated 2x across cluster][0m
|
||
[36m(FSDPPolicyWorkerBase pid=2663495, ip=10.128.32.46)[0m [W607 17:03:10.575405606 Utils.hpp:166] Warning: Environment variable NCCL_BLOCKING_WAIT is deprecated; use TORCH_NCCL_BLOCKING_WAIT instead (function operator())[32m [repeated 2x across cluster][0m
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] Started monitoring (every 120s)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:02:41] OK: 50 / 131,072 FDs open (0.0% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:02:41] OK: RSS 1.38 GiB | node mem 161.1/858.0 GiB used (18.8%), avail 696.9 GiB
|
||
[36m(skyrl_entrypoint pid=391889)[0m ⚙️ Running in WANDB offline mode
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:04:41] OK: 50 / 131,072 FDs open (0.0% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:04:41] OK: RSS 1.50 GiB | node mem 161.4/858.0 GiB used (18.8%), avail 696.6 GiB
|
||
[36m(skyrl_entrypoint pid=391889)[0m INFO 06-07 16:06:28 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:06:41] OK: 51 / 131,072 FDs open (0.0% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:06:41] OK: RSS 1.59 GiB | node mem 161.5/858.0 GiB used (18.8%), avail 696.5 GiB
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:08:41] OK: 62 / 131,072 FDs open (0.0% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:08:41] OK: RSS 1.63 GiB | node mem 161.0/858.0 GiB used (18.8%), avail 696.9 GiB
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:10:41] OK: 62 / 131,072 FDs open (0.0% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:10:41] OK: RSS 1.63 GiB | node mem 161.0/858.0 GiB used (18.8%), avail 696.9 GiB
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:12:41] OK: 63 / 131,072 FDs open (0.0% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:12:41] OK: RSS 1.63 GiB | node mem 161.1/858.0 GiB used (18.8%), avail 696.8 GiB
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:14:41] OK: 64 / 131,072 FDs open (0.0% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:14:41] OK: RSS 1.63 GiB | node mem 161.1/858.0 GiB used (18.8%), avail 696.9 GiB
|
||
[36m(pid=3113603, ip=10.128.32.42)[0m INFO 06-07 16:16:16 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:16:41] OK: 64 / 131,072 FDs open (0.0% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:16:41] OK: RSS 1.63 GiB | node mem 161.1/858.0 GiB used (18.8%), avail 696.9 GiB
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m WARNING 06-07 16:16:48 [arg_utils.py:1256] The global random seed is set to 42. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM.
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m INFO 06-07 16:16:48 [model.py:529] Resolved architecture: Qwen3ForCausalLM
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m INFO 06-07 16:16:48 [model.py:1549] Using max model len 32768
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m INFO 06-07 16:16:48 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'NCCL_DEBUG': '***', 'NCCL_SOCKET_FAMILY': '***', 'NCCL_SOCKET_IFNAME': '***', 'RAY_ADDRESS': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}}
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m INFO 06-07 16:16:48 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=65536.
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m INFO 06-07 16:16:48 [vllm.py:690] Asynchronous scheduling is enabled.
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m WARNING 06-07 16:16:48 [vllm.py:728] Enforce eager set, overriding optimization level to -O0
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m INFO 06-07 16:16:48 [vllm.py:846] Cudagraph is disabled under eager mode
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m WARNING 06-07 16:16:48 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m WARNING 06-07 16:16:48 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m INFO 06-07 16:17:22 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114084) INFO 06-07 16:17:23 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', speculative_config=None, tokenizer='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=42, served_model_name=0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [65536], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114084) INFO 06-07 16:18:18 [worker_base.py:289] Injected <class 'skyrl_train.inference_engines.vllm.vllm_engine.WorkerWrap'> into <class 'vllm.v1.worker.gpu_worker.Worker'> for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'read_named_weights', 'set_numa_affinity', 'test_rpc']
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114084) INFO 06-07 16:18:19 [parallel_state.py:1234] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.128.32.42:48887 backend=nccl
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114084) INFO 06-07 16:18:19 [parallel_state.py:1445] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114084) INFO 06-07 16:18:19 [gpu_model_runner.py:4125] Starting to load model /e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6...
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114084) INFO 06-07 16:18:20 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114084) INFO 06-07 16:18:28 [default_loader.py:293] Loading weights took 8.20 seconds
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114084) INFO 06-07 16:18:28 [gpu_model_runner.py:4222] Model loading took 15.27 GiB memory and 8.778860 seconds
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114084) INFO 06-07 16:18:31 [gpu_worker.py:373] Available KV cache memory: 49.32 GiB
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114084) INFO 06-07 16:18:31 [kv_cache_utils.py:1307] GPU KV cache size: 359,104 tokens
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114084) INFO 06-07 16:18:31 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 10.96x
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114084) INFO 06-07 16:18:31 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled.
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114084) INFO 06-07 16:18:31 [core.py:278] init engine (profile, create kv cache, warmup model) took 2.54 seconds
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114084) WARNING 06-07 16:18:32 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114084) INFO 06-07 16:18:32 [vllm.py:690] Asynchronous scheduling is enabled.
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114084) WARNING 06-07 16:18:32 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114084) INFO 06-07 16:18:32 [vllm.py:846] Cudagraph is disabled under eager mode
|
||
[36m(AsyncVLLMInferenceEngine pid=3113603, ip=10.128.32.42)[0m WARNING 06-07 16:18:32 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.
|
||
[36m(pid=3114113, ip=10.128.32.42)[0m INFO 06-07 16:18:35 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m WARNING 06-07 16:18:39 [arg_utils.py:1256] The global random seed is set to 56. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM.
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m INFO 06-07 16:18:39 [model.py:529] Resolved architecture: Qwen3ForCausalLM
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m INFO 06-07 16:18:39 [model.py:1549] Using max model len 32768
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m INFO 06-07 16:18:39 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'NCCL_DEBUG': '***', 'NCCL_SOCKET_FAMILY': '***', 'NCCL_SOCKET_IFNAME': '***', 'RAY_ADDRESS': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}}
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m INFO 06-07 16:18:39 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=65536.
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m WARNING 06-07 16:18:39 [vllm.py:728] Enforce eager set, overriding optimization level to -O0
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m INFO 06-07 16:18:39 [vllm.py:690] Asynchronous scheduling is enabled.
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m INFO 06-07 16:18:39 [vllm.py:846] Cudagraph is disabled under eager mode
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m WARNING 06-07 16:18:40 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m WARNING 06-07 16:18:40 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned
|
||
[36m(AsyncVLLMInferenceEngine pid=3114112, ip=10.128.32.42)[0m INFO 06-07 16:18:40 [vllm.py:690] Asynchronous scheduling is enabled.
|
||
[36m(AsyncVLLMInferenceEngine pid=3114112, ip=10.128.32.42)[0m INFO 06-07 16:18:40 [vllm.py:846] Cudagraph is disabled under eager mode
|
||
[36m(pid=3114114, ip=10.128.32.42)[0m INFO 06-07 16:18:35 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s[32m [repeated 2x across cluster][0m
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:18:41] OK: 68 / 131,072 FDs open (0.1% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:18:41] OK: RSS 1.63 GiB | node mem 161.1/858.0 GiB used (18.8%), avail 696.9 GiB
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m WARNING 06-07 16:18:40 [arg_utils.py:1256] The global random seed is set to 57. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM.[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m INFO 06-07 16:18:40 [model.py:529] Resolved architecture: Qwen3ForCausalLM[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m INFO 06-07 16:18:40 [model.py:1549] Using max model len 32768[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m INFO 06-07 16:18:40 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'NCCL_DEBUG': '***', 'NCCL_SOCKET_FAMILY': '***', 'NCCL_SOCKET_IFNAME': '***', 'RAY_ADDRESS': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}}[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m INFO 06-07 16:18:40 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=65536.[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m WARNING 06-07 16:18:40 [vllm.py:728] Enforce eager set, overriding optimization level to -O0[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m WARNING 06-07 16:18:40 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m WARNING 06-07 16:18:40 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m INFO 06-07 16:18:40 [vllm.py:690] Asynchronous scheduling is enabled.
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m INFO 06-07 16:18:40 [vllm.py:846] Cudagraph is disabled under eager mode
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m INFO 06-07 16:18:48 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m INFO 06-07 16:18:48 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114382) INFO 06-07 16:18:48 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', speculative_config=None, tokenizer='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=56, served_model_name=0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [65536], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114382) INFO 06-07 16:18:51 [worker_base.py:289] Injected <class 'skyrl_train.inference_engines.vllm.vllm_engine.WorkerWrap'> into <class 'vllm.v1.worker.gpu_worker.Worker'> for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'read_named_weights', 'set_numa_affinity', 'test_rpc']
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114382) INFO 06-07 16:18:52 [parallel_state.py:1234] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.128.32.42:44549 backend=nccl
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114382) INFO 06-07 16:18:52 [parallel_state.py:1445] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114382) INFO 06-07 16:18:53 [gpu_model_runner.py:4125] Starting to load model /e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6...
|
||
[36m(AsyncVLLMInferenceEngine pid=3114112, ip=10.128.32.42)[0m INFO 06-07 16:18:48 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114382) INFO 06-07 16:18:53 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
||
[36m(AsyncVLLMInferenceEngine pid=3114112, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114390) INFO 06-07 16:18:49 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', speculative_config=None, tokenizer='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=55, served_model_name=0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [65536], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114382) INFO 06-07 16:18:58 [default_loader.py:293] Loading weights took 4.13 seconds
|
||
[36m(AsyncVLLMInferenceEngine pid=3114112, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114390) INFO 06-07 16:18:52 [worker_base.py:289] Injected <class 'skyrl_train.inference_engines.vllm.vllm_engine.WorkerWrap'> into <class 'vllm.v1.worker.gpu_worker.Worker'> for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'read_named_weights', 'set_numa_affinity', 'test_rpc'][32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114112, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114390) INFO 06-07 16:18:53 [parallel_state.py:1234] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.128.32.42:51575 backend=nccl[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114112, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114390) INFO 06-07 16:18:53 [parallel_state.py:1445] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114112, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114390) INFO 06-07 16:18:53 [gpu_model_runner.py:4125] Starting to load model /e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6...[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114382) INFO 06-07 16:18:58 [gpu_model_runner.py:4222] Model loading took 15.27 GiB memory and 4.877260 seconds
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114382) INFO 06-07 16:19:00 [gpu_worker.py:373] Available KV cache memory: 49.32 GiB
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114382) INFO 06-07 16:19:00 [kv_cache_utils.py:1307] GPU KV cache size: 359,104 tokens
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114382) INFO 06-07 16:19:00 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 10.96x
|
||
[36m(AsyncVLLMInferenceEngine pid=3114112, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114390) INFO 06-07 16:18:53 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114382) INFO 06-07 16:19:00 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled.
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114382) INFO 06-07 16:19:00 [core.py:278] init engine (profile, create kv cache, warmup model) took 2.44 seconds
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114382) WARNING 06-07 16:19:01 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114382) INFO 06-07 16:19:01 [vllm.py:690] Asynchronous scheduling is enabled.
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114382) WARNING 06-07 16:19:01 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114382) INFO 06-07 16:19:01 [vllm.py:846] Cudagraph is disabled under eager mode
|
||
[36m(AsyncVLLMInferenceEngine pid=3114113, ip=10.128.32.42)[0m WARNING 06-07 16:19:01 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:20:41] OK: 68 / 131,072 FDs open (0.1% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:20:41] OK: RSS 1.63 GiB | node mem 161.1/858.0 GiB used (18.8%), avail 696.9 GiB
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114394) INFO 06-07 16:18:58 [default_loader.py:293] Loading weights took 4.38 seconds[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114394) INFO 06-07 16:18:58 [gpu_model_runner.py:4222] Model loading took 15.27 GiB memory and 5.087965 seconds[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114394) INFO 06-07 16:19:01 [gpu_worker.py:373] Available KV cache memory: 49.32 GiB[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114394) INFO 06-07 16:19:01 [kv_cache_utils.py:1307] GPU KV cache size: 359,104 tokens[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114394) INFO 06-07 16:19:01 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 10.96x[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114394) INFO 06-07 16:19:01 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled.[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114394) INFO 06-07 16:19:01 [core.py:278] init engine (profile, create kv cache, warmup model) took 2.46 seconds[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114394) WARNING 06-07 16:19:01 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114394) INFO 06-07 16:19:01 [vllm.py:690] Asynchronous scheduling is enabled.[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114394) WARNING 06-07 16:19:01 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m (EngineCore_DP0 pid=3114394) INFO 06-07 16:19:01 [vllm.py:846] Cudagraph is disabled under eager mode[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3114114, ip=10.128.32.42)[0m WARNING 06-07 16:19:01 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.[32m [repeated 2x across cluster][0m
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:22:41] OK: 70 / 131,072 FDs open (0.1% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:22:41] OK: RSS 1.63 GiB | node mem 161.1/858.0 GiB used (18.8%), avail 696.8 GiB
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:24:41] OK: 72 / 131,072 FDs open (0.1% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:24:41] OK: RSS 1.63 GiB | node mem 161.1/858.0 GiB used (18.8%), avail 696.9 GiB
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:26:41] OK: 72 / 131,072 FDs open (0.1% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:26:41] OK: RSS 1.63 GiB | node mem 161.0/858.0 GiB used (18.8%), avail 696.9 GiB
|
||
[36m(pid=2697284, ip=10.128.32.43)[0m INFO 06-07 16:28:12 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:28:41] OK: 74 / 131,072 FDs open (0.1% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:28:41] OK: RSS 1.63 GiB | node mem 161.1/858.0 GiB used (18.8%), avail 696.9 GiB
|
||
[36m(pid=2697413, ip=10.128.32.43)[0m INFO 06-07 16:28:12 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m WARNING 06-07 16:29:10 [arg_utils.py:1256] The global random seed is set to 50. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM.
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m INFO 06-07 16:29:11 [model.py:529] Resolved architecture: Qwen3ForCausalLM
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m INFO 06-07 16:29:11 [model.py:1549] Using max model len 32768
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m INFO 06-07 16:29:11 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'NCCL_DEBUG': '***', 'NCCL_SOCKET_FAMILY': '***', 'NCCL_SOCKET_IFNAME': '***', 'RAY_ADDRESS': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}}
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m INFO 06-07 16:29:11 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=65536.
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m INFO 06-07 16:29:11 [vllm.py:690] Asynchronous scheduling is enabled.
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m WARNING 06-07 16:29:11 [vllm.py:728] Enforce eager set, overriding optimization level to -O0
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m INFO 06-07 16:29:11 [vllm.py:846] Cudagraph is disabled under eager mode
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m WARNING 06-07 16:29:11 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m WARNING 06-07 16:29:11 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:30:41] OK: 75 / 131,072 FDs open (0.1% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:30:41] OK: RSS 1.63 GiB | node mem 161.0/858.0 GiB used (18.8%), avail 696.9 GiB
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m WARNING 06-07 16:29:10 [arg_utils.py:1256] The global random seed is set to 53. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m INFO 06-07 16:29:11 [model.py:529] Resolved architecture: Qwen3ForCausalLM[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m INFO 06-07 16:29:11 [model.py:1549] Using max model len 32768[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m INFO 06-07 16:29:11 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'NCCL_DEBUG': '***', 'NCCL_SOCKET_FAMILY': '***', 'NCCL_SOCKET_IFNAME': '***', 'RAY_ADDRESS': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}}[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m INFO 06-07 16:29:11 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=65536.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m INFO 06-07 16:29:11 [vllm.py:690] Asynchronous scheduling is enabled.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m WARNING 06-07 16:29:11 [vllm.py:728] Enforce eager set, overriding optimization level to -O0[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m INFO 06-07 16:29:11 [vllm.py:846] Cudagraph is disabled under eager mode[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m WARNING 06-07 16:29:11 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m WARNING 06-07 16:29:11 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned[32m [repeated 3x across cluster][0m
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:32:41] OK: 80 / 131,072 FDs open (0.1% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:32:41] OK: RSS 1.63 GiB | node mem 161.0/858.0 GiB used (18.8%), avail 697.0 GiB
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:34:41] OK: 80 / 131,072 FDs open (0.1% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:34:41] OK: RSS 1.63 GiB | node mem 161.1/858.0 GiB used (18.8%), avail 696.9 GiB
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m INFO 06-07 16:35:54 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698000) INFO 06-07 16:35:55 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', speculative_config=None, tokenizer='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=50, served_model_name=0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [65536], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:36:41] OK: 80 / 131,072 FDs open (0.1% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:36:41] OK: RSS 1.63 GiB | node mem 161.1/858.0 GiB used (18.8%), avail 696.9 GiB
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m INFO 06-07 16:35:54 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698005) INFO 06-07 16:35:55 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', speculative_config=None, tokenizer='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=53, served_model_name=0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [65536], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}[32m [repeated 3x across cluster][0m
|
||
[36m(pid=2576535, ip=10.128.32.40)[0m INFO 06-07 16:37:45 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s
|
||
[36m(pid=2576836, ip=10.128.32.40)[0m INFO 06-07 16:37:45 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s
|
||
[36m(AsyncVLLMInferenceEngine pid=2576836, ip=10.128.32.40)[0m WARNING 06-07 16:37:50 [arg_utils.py:1256] The global random seed is set to 63. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM.
|
||
[36m(pid=2576838, ip=10.128.32.40)[0m INFO 06-07 16:37:45 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576836, ip=10.128.32.40)[0m INFO 06-07 16:37:50 [model.py:529] Resolved architecture: Qwen3ForCausalLM
|
||
[36m(AsyncVLLMInferenceEngine pid=2576836, ip=10.128.32.40)[0m INFO 06-07 16:37:50 [model.py:1549] Using max model len 32768
|
||
[36m(AsyncVLLMInferenceEngine pid=2576836, ip=10.128.32.40)[0m INFO 06-07 16:37:50 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'NCCL_DEBUG': '***', 'NCCL_SOCKET_FAMILY': '***', 'NCCL_SOCKET_IFNAME': '***', 'RAY_ADDRESS': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}}
|
||
[36m(AsyncVLLMInferenceEngine pid=2576836, ip=10.128.32.40)[0m INFO 06-07 16:37:50 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=65536.
|
||
[36m(AsyncVLLMInferenceEngine pid=2576836, ip=10.128.32.40)[0m INFO 06-07 16:37:50 [vllm.py:690] Asynchronous scheduling is enabled.
|
||
[36m(AsyncVLLMInferenceEngine pid=2576836, ip=10.128.32.40)[0m WARNING 06-07 16:37:50 [vllm.py:728] Enforce eager set, overriding optimization level to -O0
|
||
[36m(AsyncVLLMInferenceEngine pid=2576836, ip=10.128.32.40)[0m INFO 06-07 16:37:50 [vllm.py:846] Cudagraph is disabled under eager mode
|
||
[36m(AsyncVLLMInferenceEngine pid=2576836, ip=10.128.32.40)[0m WARNING 06-07 16:37:50 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1
|
||
[36m(AsyncVLLMInferenceEngine pid=2576836, ip=10.128.32.40)[0m WARNING 06-07 16:37:50 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:38:41] OK: 88 / 131,072 FDs open (0.1% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:38:41] OK: RSS 1.64 GiB | node mem 161.1/858.0 GiB used (18.8%), avail 696.9 GiB
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m WARNING 06-07 16:37:51 [arg_utils.py:1256] The global random seed is set to 62. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m INFO 06-07 16:37:51 [model.py:529] Resolved architecture: Qwen3ForCausalLM[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m INFO 06-07 16:37:51 [model.py:1549] Using max model len 32768[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m INFO 06-07 16:37:51 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'NCCL_DEBUG': '***', 'NCCL_SOCKET_FAMILY': '***', 'NCCL_SOCKET_IFNAME': '***', 'RAY_ADDRESS': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}}[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m INFO 06-07 16:37:51 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=65536.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m INFO 06-07 16:37:51 [vllm.py:690] Asynchronous scheduling is enabled.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m WARNING 06-07 16:37:51 [vllm.py:728] Enforce eager set, overriding optimization level to -O0[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m INFO 06-07 16:37:51 [vllm.py:846] Cudagraph is disabled under eager mode[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m WARNING 06-07 16:37:51 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m WARNING 06-07 16:37:51 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned[32m [repeated 3x across cluster][0m
|
||
[36m(pid=2739368, ip=10.128.32.34)[0m INFO 06-07 16:39:00 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698000) INFO 06-07 16:40:23 [worker_base.py:289] Injected <class 'skyrl_train.inference_engines.vllm.vllm_engine.WorkerWrap'> into <class 'vllm.v1.worker.gpu_worker.Worker'> for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'read_named_weights', 'set_numa_affinity', 'test_rpc']
|
||
[36m(pid=2739604, ip=10.128.32.34)[0m INFO 06-07 16:39:00 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m WARNING 06-07 16:40:24 [arg_utils.py:1256] The global random seed is set to 58. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM.
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698000) INFO 06-07 16:40:24 [parallel_state.py:1234] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.128.32.43:60183 backend=nccl
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m INFO 06-07 16:40:24 [model.py:529] Resolved architecture: Qwen3ForCausalLM
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m INFO 06-07 16:40:24 [model.py:1549] Using max model len 32768
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698000) INFO 06-07 16:40:24 [parallel_state.py:1445] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m INFO 06-07 16:40:24 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'NCCL_DEBUG': '***', 'NCCL_SOCKET_FAMILY': '***', 'NCCL_SOCKET_IFNAME': '***', 'RAY_ADDRESS': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}}
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m INFO 06-07 16:40:24 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=65536.
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m INFO 06-07 16:40:24 [vllm.py:690] Asynchronous scheduling is enabled.
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m WARNING 06-07 16:40:24 [vllm.py:728] Enforce eager set, overriding optimization level to -O0
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m INFO 06-07 16:40:24 [vllm.py:846] Cudagraph is disabled under eager mode
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m WARNING 06-07 16:40:24 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m WARNING 06-07 16:40:24 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698000) INFO 06-07 16:40:25 [gpu_model_runner.py:4125] Starting to load model /e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6...
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698000) INFO 06-07 16:40:27 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698000) INFO 06-07 16:40:41 [default_loader.py:293] Loading weights took 13.76 seconds
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698005) INFO 06-07 16:40:23 [worker_base.py:289] Injected <class 'skyrl_train.inference_engines.vllm.vllm_engine.WorkerWrap'> into <class 'vllm.v1.worker.gpu_worker.Worker'> for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'read_named_weights', 'set_numa_affinity', 'test_rpc'][32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739603, ip=10.128.32.34)[0m WARNING 06-07 16:40:24 [arg_utils.py:1256] The global random seed is set to 60. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698005) INFO 06-07 16:40:24 [parallel_state.py:1234] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.128.32.43:46619 backend=nccl[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739603, ip=10.128.32.34)[0m INFO 06-07 16:40:24 [model.py:529] Resolved architecture: Qwen3ForCausalLM[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739603, ip=10.128.32.34)[0m INFO 06-07 16:40:24 [model.py:1549] Using max model len 32768[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698005) INFO 06-07 16:40:24 [parallel_state.py:1445] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739603, ip=10.128.32.34)[0m INFO 06-07 16:40:24 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'NCCL_DEBUG': '***', 'NCCL_SOCKET_FAMILY': '***', 'NCCL_SOCKET_IFNAME': '***', 'RAY_ADDRESS': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}}[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739603, ip=10.128.32.34)[0m INFO 06-07 16:40:24 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=65536.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739603, ip=10.128.32.34)[0m INFO 06-07 16:40:24 [vllm.py:690] Asynchronous scheduling is enabled.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739603, ip=10.128.32.34)[0m WARNING 06-07 16:40:24 [vllm.py:728] Enforce eager set, overriding optimization level to -O0[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739603, ip=10.128.32.34)[0m INFO 06-07 16:40:24 [vllm.py:846] Cudagraph is disabled under eager mode[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739603, ip=10.128.32.34)[0m WARNING 06-07 16:40:25 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739603, ip=10.128.32.34)[0m WARNING 06-07 16:40:25 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698005) INFO 06-07 16:40:25 [gpu_model_runner.py:4125] Starting to load model /e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6...[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698005) INFO 06-07 16:40:27 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].[32m [repeated 3x across cluster][0m
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:40:41] OK: 88 / 131,072 FDs open (0.1% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:40:41] OK: RSS 1.64 GiB | node mem 161.0/858.0 GiB used (18.8%), avail 696.9 GiB
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698000) INFO 06-07 16:40:41 [gpu_model_runner.py:4222] Model loading took 15.27 GiB memory and 15.544833 seconds
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698000) INFO 06-07 16:40:44 [gpu_worker.py:373] Available KV cache memory: 49.32 GiB
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698000) INFO 06-07 16:40:44 [kv_cache_utils.py:1307] GPU KV cache size: 359,104 tokens
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698000) INFO 06-07 16:40:44 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 10.96x
|
||
[36m(AsyncVLLMInferenceEngine pid=2697412, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698009) INFO 06-07 16:40:44 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled.
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698000) INFO 06-07 16:40:44 [core.py:278] init engine (profile, create kv cache, warmup model) took 2.82 seconds
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698000) WARNING 06-07 16:40:45 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
||
[36m(AsyncVLLMInferenceEngine pid=2697284, ip=10.128.32.43)[0m WARNING 06-07 16:40:45 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:42:41] OK: 88 / 131,072 FDs open (0.1% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:42:41] OK: RSS 1.64 GiB | node mem 161.1/858.0 GiB used (18.8%), avail 696.9 GiB
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698005) INFO 06-07 16:40:41 [default_loader.py:293] Loading weights took 13.76 seconds[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698005) INFO 06-07 16:40:45 [vllm.py:690] Asynchronous scheduling is enabled.[32m [repeated 4x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698005) INFO 06-07 16:40:45 [vllm.py:846] Cudagraph is disabled under eager mode[32m [repeated 4x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698005) WARNING 06-07 16:40:45 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1[32m [repeated 4x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698005) INFO 06-07 16:40:41 [gpu_model_runner.py:4222] Model loading took 15.27 GiB memory and 15.540163 seconds[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698005) INFO 06-07 16:40:44 [gpu_worker.py:373] Available KV cache memory: 49.32 GiB[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698005) INFO 06-07 16:40:44 [kv_cache_utils.py:1307] GPU KV cache size: 359,104 tokens[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698005) INFO 06-07 16:40:44 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 10.96x[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698005) INFO 06-07 16:40:44 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698005) INFO 06-07 16:40:44 [core.py:278] init engine (profile, create kv cache, warmup model) took 2.85 seconds[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m (EngineCore_DP0 pid=2698005) WARNING 06-07 16:40:45 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2697413, ip=10.128.32.43)[0m WARNING 06-07 16:40:45 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.[32m [repeated 3x across cluster][0m
|
||
[36m(pid=2631427, ip=10.128.32.36)[0m INFO 06-07 16:43:22 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s
|
||
[36m(AsyncVLLMInferenceEngine pid=2631819, ip=10.128.32.36)[0m WARNING 06-07 16:43:53 [arg_utils.py:1256] The global random seed is set to 79. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM.
|
||
[36m(pid=2631820, ip=10.128.32.36)[0m INFO 06-07 16:43:22 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631819, ip=10.128.32.36)[0m INFO 06-07 16:43:53 [model.py:529] Resolved architecture: Qwen3ForCausalLM
|
||
[36m(AsyncVLLMInferenceEngine pid=2631819, ip=10.128.32.36)[0m INFO 06-07 16:43:53 [model.py:1549] Using max model len 32768
|
||
[36m(AsyncVLLMInferenceEngine pid=2631819, ip=10.128.32.36)[0m INFO 06-07 16:43:53 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'NCCL_DEBUG': '***', 'NCCL_SOCKET_FAMILY': '***', 'NCCL_SOCKET_IFNAME': '***', 'RAY_ADDRESS': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}}
|
||
[36m(AsyncVLLMInferenceEngine pid=2631819, ip=10.128.32.36)[0m INFO 06-07 16:43:53 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=65536.
|
||
[36m(AsyncVLLMInferenceEngine pid=2631819, ip=10.128.32.36)[0m INFO 06-07 16:43:53 [vllm.py:690] Asynchronous scheduling is enabled.
|
||
[36m(AsyncVLLMInferenceEngine pid=2631819, ip=10.128.32.36)[0m WARNING 06-07 16:43:53 [vllm.py:728] Enforce eager set, overriding optimization level to -O0
|
||
[36m(AsyncVLLMInferenceEngine pid=2631819, ip=10.128.32.36)[0m INFO 06-07 16:43:53 [vllm.py:846] Cudagraph is disabled under eager mode
|
||
[36m(AsyncVLLMInferenceEngine pid=2631819, ip=10.128.32.36)[0m WARNING 06-07 16:43:54 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1
|
||
[36m(AsyncVLLMInferenceEngine pid=2631819, ip=10.128.32.36)[0m WARNING 06-07 16:43:54 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m WARNING 06-07 16:43:54 [arg_utils.py:1256] The global random seed is set to 81. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m INFO 06-07 16:44:08 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m INFO 06-07 16:43:54 [model.py:529] Resolved architecture: Qwen3ForCausalLM[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m INFO 06-07 16:43:54 [model.py:1549] Using max model len 32768[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m INFO 06-07 16:43:54 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'NCCL_DEBUG': '***', 'NCCL_SOCKET_FAMILY': '***', 'NCCL_SOCKET_IFNAME': '***', 'RAY_ADDRESS': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}}[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m INFO 06-07 16:43:54 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=65536.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m INFO 06-07 16:43:54 [vllm.py:690] Asynchronous scheduling is enabled.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m WARNING 06-07 16:43:54 [vllm.py:728] Enforce eager set, overriding optimization level to -O0[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m INFO 06-07 16:43:54 [vllm.py:846] Cudagraph is disabled under eager mode[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m WARNING 06-07 16:43:55 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m WARNING 06-07 16:43:55 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576836, ip=10.128.32.40)[0m INFO 06-07 16:44:08 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577500) INFO 06-07 16:44:09 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', speculative_config=None, tokenizer='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=62, served_model_name=0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [65536], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:44:41] OK: 88 / 131,072 FDs open (0.1% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:44:41] OK: RSS 1.64 GiB | node mem 161.1/858.0 GiB used (18.8%), avail 696.9 GiB
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m INFO 06-07 16:44:08 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577495) INFO 06-07 16:44:09 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', speculative_config=None, tokenizer='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=65, served_model_name=0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [65536], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577500) INFO 06-07 16:46:25 [worker_base.py:289] Injected <class 'skyrl_train.inference_engines.vllm.vllm_engine.WorkerWrap'> into <class 'vllm.v1.worker.gpu_worker.Worker'> for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'read_named_weights', 'set_numa_affinity', 'test_rpc']
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577500) INFO 06-07 16:46:26 [parallel_state.py:1234] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.128.32.40:39565 backend=nccl
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577500) INFO 06-07 16:46:26 [parallel_state.py:1445] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577500) INFO 06-07 16:46:26 [gpu_model_runner.py:4125] Starting to load model /e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6...
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:46:41] OK: 88 / 131,072 FDs open (0.1% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:46:41] OK: RSS 1.64 GiB | node mem 161.1/858.0 GiB used (18.8%), avail 696.9 GiB
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577495) INFO 06-07 16:46:25 [worker_base.py:289] Injected <class 'skyrl_train.inference_engines.vllm.vllm_engine.WorkerWrap'> into <class 'vllm.v1.worker.gpu_worker.Worker'> for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'read_named_weights', 'set_numa_affinity', 'test_rpc'][32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577495) INFO 06-07 16:46:26 [parallel_state.py:1234] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.128.32.40:52891 backend=nccl[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577495) INFO 06-07 16:46:26 [parallel_state.py:1445] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577495) INFO 06-07 16:46:26 [gpu_model_runner.py:4125] Starting to load model /e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6...[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576837, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577489) INFO 06-07 16:46:53 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577500) INFO 06-07 16:47:07 [default_loader.py:293] Loading weights took 13.09 seconds
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577495) INFO 06-07 16:46:54 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577500) INFO 06-07 16:47:07 [gpu_model_runner.py:4222] Model loading took 15.27 GiB memory and 40.482660 seconds
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577500) INFO 06-07 16:47:10 [gpu_worker.py:373] Available KV cache memory: 49.32 GiB
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577500) INFO 06-07 16:47:10 [kv_cache_utils.py:1307] GPU KV cache size: 359,104 tokens
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577500) INFO 06-07 16:47:10 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 10.96x
|
||
[36m(AsyncVLLMInferenceEngine pid=2576837, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577489) INFO 06-07 16:47:10 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled.
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577500) INFO 06-07 16:47:10 [core.py:278] init engine (profile, create kv cache, warmup model) took 2.84 seconds
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577500) WARNING 06-07 16:47:11 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577500) INFO 06-07 16:47:11 [vllm.py:690] Asynchronous scheduling is enabled.
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577500) WARNING 06-07 16:47:11 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577500) INFO 06-07 16:47:11 [vllm.py:846] Cudagraph is disabled under eager mode
|
||
[36m(AsyncVLLMInferenceEngine pid=2576836, ip=10.128.32.40)[0m WARNING 06-07 16:47:11 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m INFO 06-07 16:47:33 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577495) INFO 06-07 16:47:07 [default_loader.py:293] Loading weights took 13.09 seconds[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577495) INFO 06-07 16:47:07 [gpu_model_runner.py:4222] Model loading took 15.27 GiB memory and 40.481976 seconds[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577495) INFO 06-07 16:47:10 [gpu_worker.py:373] Available KV cache memory: 49.32 GiB[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577495) INFO 06-07 16:47:10 [kv_cache_utils.py:1307] GPU KV cache size: 359,104 tokens[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577495) INFO 06-07 16:47:10 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 10.96x[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577495) INFO 06-07 16:47:10 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577495) INFO 06-07 16:47:10 [core.py:278] init engine (profile, create kv cache, warmup model) took 2.86 seconds[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577495) WARNING 06-07 16:47:11 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577495) INFO 06-07 16:47:11 [vllm.py:690] Asynchronous scheduling is enabled.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577495) WARNING 06-07 16:47:11 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576838, ip=10.128.32.40)[0m (EngineCore_DP0 pid=2577495) INFO 06-07 16:47:11 [vllm.py:846] Cudagraph is disabled under eager mode[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2576535, ip=10.128.32.40)[0m WARNING 06-07 16:47:11 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740424) INFO 06-07 16:47:34 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', speculative_config=None, tokenizer='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=58, served_model_name=0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [65536], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}
|
||
[36m(pid=3010613, ip=10.128.32.44)[0m INFO 06-07 16:48:08 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s[32m [repeated 4x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740428) INFO 06-07 16:47:34 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', speculative_config=None, tokenizer='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=61, served_model_name=0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [65536], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m INFO 06-07 16:48:26 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s[32m [repeated 4x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632455) INFO 06-07 16:48:27 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', speculative_config=None, tokenizer='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=66, served_model_name=0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [65536], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}
|
||
[36m(AsyncVLLMInferenceEngine pid=3011433, ip=10.128.32.44)[0m WARNING 06-07 16:48:40 [arg_utils.py:1256] The global random seed is set to 68. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM.
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m INFO 06-07 16:48:26 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632471) INFO 06-07 16:48:27 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', speculative_config=None, tokenizer='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=81, served_model_name=0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [65536], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011433, ip=10.128.32.44)[0m INFO 06-07 16:48:40 [model.py:529] Resolved architecture: Qwen3ForCausalLM
|
||
[36m(AsyncVLLMInferenceEngine pid=3011433, ip=10.128.32.44)[0m INFO 06-07 16:48:40 [model.py:1549] Using max model len 32768
|
||
[36m(AsyncVLLMInferenceEngine pid=3011433, ip=10.128.32.44)[0m INFO 06-07 16:48:40 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'NCCL_DEBUG': '***', 'NCCL_SOCKET_FAMILY': '***', 'NCCL_SOCKET_IFNAME': '***', 'RAY_ADDRESS': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}}
|
||
[36m(AsyncVLLMInferenceEngine pid=3011433, ip=10.128.32.44)[0m INFO 06-07 16:48:40 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=65536.
|
||
[36m(AsyncVLLMInferenceEngine pid=3011433, ip=10.128.32.44)[0m INFO 06-07 16:48:40 [vllm.py:690] Asynchronous scheduling is enabled.
|
||
[36m(AsyncVLLMInferenceEngine pid=3011433, ip=10.128.32.44)[0m WARNING 06-07 16:48:40 [vllm.py:728] Enforce eager set, overriding optimization level to -O0
|
||
[36m(AsyncVLLMInferenceEngine pid=3011433, ip=10.128.32.44)[0m INFO 06-07 16:48:40 [vllm.py:846] Cudagraph is disabled under eager mode
|
||
[36m(AsyncVLLMInferenceEngine pid=3011433, ip=10.128.32.44)[0m WARNING 06-07 16:48:40 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1
|
||
[36m(AsyncVLLMInferenceEngine pid=3011433, ip=10.128.32.44)[0m WARNING 06-07 16:48:40 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:48:41] OK: 88 / 131,072 FDs open (0.1% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:48:41] OK: RSS 1.64 GiB | node mem 161.1/858.0 GiB used (18.8%), avail 696.9 GiB
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740424) INFO 06-07 16:49:22 [worker_base.py:289] Injected <class 'skyrl_train.inference_engines.vllm.vllm_engine.WorkerWrap'> into <class 'vllm.v1.worker.gpu_worker.Worker'> for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'read_named_weights', 'set_numa_affinity', 'test_rpc']
|
||
[36m(AsyncVLLMInferenceEngine pid=3011432, ip=10.128.32.44)[0m WARNING 06-07 16:48:40 [arg_utils.py:1256] The global random seed is set to 69. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011432, ip=10.128.32.44)[0m INFO 06-07 16:48:40 [model.py:529] Resolved architecture: Qwen3ForCausalLM[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011432, ip=10.128.32.44)[0m INFO 06-07 16:48:40 [model.py:1549] Using max model len 32768[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011432, ip=10.128.32.44)[0m INFO 06-07 16:48:40 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'NCCL_DEBUG': '***', 'NCCL_SOCKET_FAMILY': '***', 'NCCL_SOCKET_IFNAME': '***', 'RAY_ADDRESS': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}}[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011432, ip=10.128.32.44)[0m INFO 06-07 16:48:40 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=65536.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011432, ip=10.128.32.44)[0m INFO 06-07 16:48:40 [vllm.py:690] Asynchronous scheduling is enabled.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011432, ip=10.128.32.44)[0m WARNING 06-07 16:48:40 [vllm.py:728] Enforce eager set, overriding optimization level to -O0[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011432, ip=10.128.32.44)[0m INFO 06-07 16:48:40 [vllm.py:846] Cudagraph is disabled under eager mode[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011432, ip=10.128.32.44)[0m WARNING 06-07 16:48:40 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011432, ip=10.128.32.44)[0m WARNING 06-07 16:48:40 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740424) INFO 06-07 16:49:22 [parallel_state.py:1234] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.128.32.34:58359 backend=nccl
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740424) INFO 06-07 16:49:22 [parallel_state.py:1445] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740424) INFO 06-07 16:50:15 [gpu_model_runner.py:4125] Starting to load model /e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6...
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740428) INFO 06-07 16:49:22 [worker_base.py:289] Injected <class 'skyrl_train.inference_engines.vllm.vllm_engine.WorkerWrap'> into <class 'vllm.v1.worker.gpu_worker.Worker'> for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'read_named_weights', 'set_numa_affinity', 'test_rpc'][32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740428) INFO 06-07 16:49:22 [parallel_state.py:1234] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.128.32.34:58341 backend=nccl[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740428) INFO 06-07 16:49:22 [parallel_state.py:1445] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740424) INFO 06-07 16:50:16 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740424) INFO 06-07 16:50:29 [default_loader.py:293] Loading weights took 13.02 seconds
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740428) INFO 06-07 16:50:15 [gpu_model_runner.py:4125] Starting to load model /e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6...[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740428) INFO 06-07 16:50:16 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740424) INFO 06-07 16:50:29 [gpu_model_runner.py:4222] Model loading took 15.27 GiB memory and 14.171795 seconds
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740424) INFO 06-07 16:50:32 [gpu_worker.py:373] Available KV cache memory: 49.32 GiB
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740424) INFO 06-07 16:50:32 [kv_cache_utils.py:1307] GPU KV cache size: 359,104 tokens
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740424) INFO 06-07 16:50:32 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 10.96x
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740424) INFO 06-07 16:50:32 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled.
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740424) INFO 06-07 16:50:32 [core.py:278] init engine (profile, create kv cache, warmup model) took 2.81 seconds
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740424) WARNING 06-07 16:50:33 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740424) INFO 06-07 16:50:33 [vllm.py:690] Asynchronous scheduling is enabled.
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740424) WARNING 06-07 16:50:33 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740424) INFO 06-07 16:50:33 [vllm.py:846] Cudagraph is disabled under eager mode
|
||
[36m(AsyncVLLMInferenceEngine pid=2739368, ip=10.128.32.34)[0m WARNING 06-07 16:50:33 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:50:41] OK: 88 / 131,072 FDs open (0.1% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:50:41] OK: RSS 1.64 GiB | node mem 161.1/858.0 GiB used (18.8%), avail 696.9 GiB
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740428) INFO 06-07 16:50:29 [default_loader.py:293] Loading weights took 12.88 seconds[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740428) INFO 06-07 16:50:29 [gpu_model_runner.py:4222] Model loading took 15.27 GiB memory and 14.175118 seconds[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740428) INFO 06-07 16:50:32 [gpu_worker.py:373] Available KV cache memory: 49.32 GiB[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740428) INFO 06-07 16:50:32 [kv_cache_utils.py:1307] GPU KV cache size: 359,104 tokens[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740428) INFO 06-07 16:50:32 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 10.96x[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740428) INFO 06-07 16:50:32 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739603, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740438) INFO 06-07 16:50:32 [core.py:278] init engine (profile, create kv cache, warmup model) took 2.83 seconds[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739603, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740438) WARNING 06-07 16:50:33 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739603, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740438) INFO 06-07 16:50:33 [vllm.py:690] Asynchronous scheduling is enabled.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739603, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740438) WARNING 06-07 16:50:33 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739603, ip=10.128.32.34)[0m (EngineCore_DP0 pid=2740438) INFO 06-07 16:50:33 [vllm.py:846] Cudagraph is disabled under eager mode[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2739604, ip=10.128.32.34)[0m WARNING 06-07 16:50:33 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631821, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632466) INFO 06-07 16:50:42 [worker_base.py:289] Injected <class 'skyrl_train.inference_engines.vllm.vllm_engine.WorkerWrap'> into <class 'vllm.v1.worker.gpu_worker.Worker'> for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'read_named_weights', 'set_numa_affinity', 'test_rpc']
|
||
[36m(AsyncVLLMInferenceEngine pid=2631821, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632466) INFO 06-07 16:50:43 [parallel_state.py:1234] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.128.32.36:48845 backend=nccl
|
||
[36m(AsyncVLLMInferenceEngine pid=2631821, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632466) INFO 06-07 16:50:43 [parallel_state.py:1445] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632455) INFO 06-07 16:51:09 [gpu_model_runner.py:4125] Starting to load model /e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6...
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632455) INFO 06-07 16:50:43 [worker_base.py:289] Injected <class 'skyrl_train.inference_engines.vllm.vllm_engine.WorkerWrap'> into <class 'vllm.v1.worker.gpu_worker.Worker'> for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'read_named_weights', 'set_numa_affinity', 'test_rpc'][32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632455) INFO 06-07 16:50:43 [parallel_state.py:1234] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.128.32.36:42177 backend=nccl[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632455) INFO 06-07 16:50:43 [parallel_state.py:1445] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A[32m [repeated 3x across cluster][0m
|
||
[36m(pid=2705575, ip=10.128.32.38)[0m INFO 06-07 16:51:38 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632471) INFO 06-07 16:51:09 [gpu_model_runner.py:4125] Starting to load model /e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6...[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m INFO 06-07 16:51:54 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s[32m [repeated 4x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012340) INFO 06-07 16:51:56 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', speculative_config=None, tokenizer='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=67, served_model_name=0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [65536], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632455) INFO 06-07 16:52:02 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m INFO 06-07 16:51:54 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012341) INFO 06-07 16:51:56 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', speculative_config=None, tokenizer='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=78, served_model_name=0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [65536], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705804, ip=10.128.32.38)[0m WARNING 06-07 16:52:09 [arg_utils.py:1256] The global random seed is set to 72. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM.
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632471) INFO 06-07 16:52:02 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705804, ip=10.128.32.38)[0m INFO 06-07 16:52:09 [model.py:529] Resolved architecture: Qwen3ForCausalLM
|
||
[36m(AsyncVLLMInferenceEngine pid=2705804, ip=10.128.32.38)[0m INFO 06-07 16:52:09 [model.py:1549] Using max model len 32768
|
||
[36m(AsyncVLLMInferenceEngine pid=2705804, ip=10.128.32.38)[0m INFO 06-07 16:52:09 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'NCCL_DEBUG': '***', 'NCCL_SOCKET_FAMILY': '***', 'NCCL_SOCKET_IFNAME': '***', 'RAY_ADDRESS': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}}
|
||
[36m(AsyncVLLMInferenceEngine pid=2705804, ip=10.128.32.38)[0m INFO 06-07 16:52:09 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=65536.
|
||
[36m(AsyncVLLMInferenceEngine pid=2705804, ip=10.128.32.38)[0m INFO 06-07 16:52:09 [vllm.py:690] Asynchronous scheduling is enabled.
|
||
[36m(AsyncVLLMInferenceEngine pid=2705804, ip=10.128.32.38)[0m WARNING 06-07 16:52:09 [vllm.py:728] Enforce eager set, overriding optimization level to -O0
|
||
[36m(AsyncVLLMInferenceEngine pid=2705804, ip=10.128.32.38)[0m INFO 06-07 16:52:09 [vllm.py:846] Cudagraph is disabled under eager mode
|
||
[36m(AsyncVLLMInferenceEngine pid=2705804, ip=10.128.32.38)[0m WARNING 06-07 16:52:10 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1
|
||
[36m(AsyncVLLMInferenceEngine pid=2705804, ip=10.128.32.38)[0m WARNING 06-07 16:52:10 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632455) INFO 06-07 16:52:15 [default_loader.py:293] Loading weights took 13.37 seconds
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m WARNING 06-07 16:52:10 [arg_utils.py:1256] The global random seed is set to 70. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m INFO 06-07 16:52:10 [model.py:529] Resolved architecture: Qwen3ForCausalLM[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m INFO 06-07 16:52:10 [model.py:1549] Using max model len 32768[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m INFO 06-07 16:52:10 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'NCCL_DEBUG': '***', 'NCCL_SOCKET_FAMILY': '***', 'NCCL_SOCKET_IFNAME': '***', 'RAY_ADDRESS': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}}[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m INFO 06-07 16:52:10 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=65536.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m INFO 06-07 16:52:10 [vllm.py:690] Asynchronous scheduling is enabled.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m WARNING 06-07 16:52:10 [vllm.py:728] Enforce eager set, overriding optimization level to -O0[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m INFO 06-07 16:52:10 [vllm.py:846] Cudagraph is disabled under eager mode[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m WARNING 06-07 16:52:11 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m WARNING 06-07 16:52:11 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632455) INFO 06-07 16:52:16 [gpu_model_runner.py:4222] Model loading took 15.27 GiB memory and 65.938186 seconds
|
||
[36m(AsyncVLLMInferenceEngine pid=2631821, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632466) INFO 06-07 16:52:18 [gpu_worker.py:373] Available KV cache memory: 49.32 GiB
|
||
[36m(AsyncVLLMInferenceEngine pid=2631821, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632466) INFO 06-07 16:52:18 [kv_cache_utils.py:1307] GPU KV cache size: 359,104 tokens
|
||
[36m(AsyncVLLMInferenceEngine pid=2631821, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632466) INFO 06-07 16:52:18 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 10.96x
|
||
[36m(AsyncVLLMInferenceEngine pid=2631821, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632466) INFO 06-07 16:52:18 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled.
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632455) INFO 06-07 16:52:18 [core.py:278] init engine (profile, create kv cache, warmup model) took 2.81 seconds
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632455) WARNING 06-07 16:52:19 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m WARNING 06-07 16:52:19 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:52:41] OK: 88 / 131,072 FDs open (0.1% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:52:41] OK: RSS 1.64 GiB | node mem 161.0/858.0 GiB used (18.8%), avail 696.9 GiB
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632471) INFO 06-07 16:52:15 [default_loader.py:293] Loading weights took 13.37 seconds[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632471) INFO 06-07 16:52:19 [vllm.py:690] Asynchronous scheduling is enabled.[32m [repeated 4x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632471) INFO 06-07 16:52:19 [vllm.py:846] Cudagraph is disabled under eager mode[32m [repeated 4x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632471) WARNING 06-07 16:52:19 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1[32m [repeated 4x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632471) INFO 06-07 16:52:16 [gpu_model_runner.py:4222] Model loading took 15.27 GiB memory and 65.945630 seconds[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632455) INFO 06-07 16:52:18 [gpu_worker.py:373] Available KV cache memory: 49.32 GiB[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632455) INFO 06-07 16:52:18 [kv_cache_utils.py:1307] GPU KV cache size: 359,104 tokens[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631427, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632455) INFO 06-07 16:52:18 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 10.96x[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632471) INFO 06-07 16:52:18 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632471) INFO 06-07 16:52:18 [core.py:278] init engine (profile, create kv cache, warmup model) took 2.79 seconds[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m (EngineCore_DP0 pid=2632471) WARNING 06-07 16:52:19 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2631820, ip=10.128.32.36)[0m WARNING 06-07 16:52:19 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m INFO 06-07 16:53:15 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706759) INFO 06-07 16:53:17 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', speculative_config=None, tokenizer='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=70, served_model_name=0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [65536], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012340) INFO 06-07 16:54:12 [worker_base.py:289] Injected <class 'skyrl_train.inference_engines.vllm.vllm_engine.WorkerWrap'> into <class 'vllm.v1.worker.gpu_worker.Worker'> for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'read_named_weights', 'set_numa_affinity', 'test_rpc']
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m INFO 06-07 16:53:15 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706739) INFO 06-07 16:53:17 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', speculative_config=None, tokenizer='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=73, served_model_name=0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [65536], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012340) INFO 06-07 16:54:12 [parallel_state.py:1234] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.128.32.44:48903 backend=nccl
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012340) INFO 06-07 16:54:12 [parallel_state.py:1445] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012341) INFO 06-07 16:54:12 [worker_base.py:289] Injected <class 'skyrl_train.inference_engines.vllm.vllm_engine.WorkerWrap'> into <class 'vllm.v1.worker.gpu_worker.Worker'> for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'read_named_weights', 'set_numa_affinity', 'test_rpc'][32m [repeated 3x across cluster][0m
|
||
[36m(pid=2831550, ip=10.128.32.45)[0m INFO 06-07 16:54:18 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012341) INFO 06-07 16:54:12 [parallel_state.py:1234] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.128.32.44:53241 backend=nccl[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012341) INFO 06-07 16:54:12 [parallel_state.py:1445] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A[32m [repeated 3x across cluster][0m
|
||
[36m(pid=2831684, ip=10.128.32.45)[0m INFO 06-07 16:54:18 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:54:41] OK: 88 / 131,072 FDs open (0.1% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:54:41] OK: RSS 1.64 GiB | node mem 161.1/858.0 GiB used (18.8%), avail 696.9 GiB
|
||
[36m(pid=2831686, ip=10.128.32.45)[0m INFO 06-07 16:54:18 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s[32m [repeated 2x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012340) INFO 06-07 16:55:04 [gpu_model_runner.py:4125] Starting to load model /e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6...
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012340) INFO 06-07 16:55:06 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012340) INFO 06-07 16:55:19 [default_loader.py:293] Loading weights took 13.56 seconds
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012341) INFO 06-07 16:55:04 [gpu_model_runner.py:4125] Starting to load model /e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6...[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012341) INFO 06-07 16:55:06 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012340) INFO 06-07 16:55:20 [gpu_model_runner.py:4222] Model loading took 15.27 GiB memory and 14.714728 seconds
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012340) INFO 06-07 16:55:22 [gpu_worker.py:373] Available KV cache memory: 49.32 GiB
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012340) INFO 06-07 16:55:22 [kv_cache_utils.py:1307] GPU KV cache size: 359,104 tokens
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012340) INFO 06-07 16:55:22 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 10.96x
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012340) INFO 06-07 16:55:22 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled.
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012340) INFO 06-07 16:55:22 [core.py:278] init engine (profile, create kv cache, warmup model) took 2.82 seconds
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012340) WARNING 06-07 16:55:23 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012340) INFO 06-07 16:55:23 [vllm.py:690] Asynchronous scheduling is enabled.
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012340) WARNING 06-07 16:55:23 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012340) INFO 06-07 16:55:23 [vllm.py:846] Cudagraph is disabled under eager mode
|
||
[36m(AsyncVLLMInferenceEngine pid=3010613, ip=10.128.32.44)[0m WARNING 06-07 16:55:23 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706759) INFO 06-07 16:55:32 [worker_base.py:289] Injected <class 'skyrl_train.inference_engines.vllm.vllm_engine.WorkerWrap'> into <class 'vllm.v1.worker.gpu_worker.Worker'> for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'read_named_weights', 'set_numa_affinity', 'test_rpc']
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012341) INFO 06-07 16:55:19 [default_loader.py:293] Loading weights took 13.43 seconds[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012341) INFO 06-07 16:55:20 [gpu_model_runner.py:4222] Model loading took 15.27 GiB memory and 14.718624 seconds[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011433, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012342) INFO 06-07 16:55:22 [gpu_worker.py:373] Available KV cache memory: 49.32 GiB[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011433, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012342) INFO 06-07 16:55:22 [kv_cache_utils.py:1307] GPU KV cache size: 359,104 tokens[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011433, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012342) INFO 06-07 16:55:22 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 10.96x[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012341) INFO 06-07 16:55:22 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012341) INFO 06-07 16:55:22 [core.py:278] init engine (profile, create kv cache, warmup model) took 2.81 seconds[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012341) WARNING 06-07 16:55:23 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012341) INFO 06-07 16:55:23 [vllm.py:690] Asynchronous scheduling is enabled.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012341) WARNING 06-07 16:55:23 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m (EngineCore_DP0 pid=3012341) INFO 06-07 16:55:23 [vllm.py:846] Cudagraph is disabled under eager mode[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=3011671, ip=10.128.32.44)[0m WARNING 06-07 16:55:23 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706759) INFO 06-07 16:55:33 [parallel_state.py:1234] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.128.32.38:39711 backend=nccl
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706759) INFO 06-07 16:55:33 [parallel_state.py:1445] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m WARNING 06-07 16:55:42 [arg_utils.py:1256] The global random seed is set to 82. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM.
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706739) INFO 06-07 16:55:32 [worker_base.py:289] Injected <class 'skyrl_train.inference_engines.vllm.vllm_engine.WorkerWrap'> into <class 'vllm.v1.worker.gpu_worker.Worker'> for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'read_named_weights', 'set_numa_affinity', 'test_rpc'][32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706739) INFO 06-07 16:55:33 [parallel_state.py:1234] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.128.32.38:41289 backend=nccl[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706739) INFO 06-07 16:55:33 [parallel_state.py:1445] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706759) INFO 06-07 16:56:00 [gpu_model_runner.py:4125] Starting to load model /e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6...
|
||
[36m(AsyncVLLMInferenceEngine pid=2831685, ip=10.128.32.45)[0m WARNING 06-07 16:55:42 [arg_utils.py:1256] The global random seed is set to 84. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m INFO 06-07 16:56:09 [model.py:529] Resolved architecture: Qwen3ForCausalLM
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m INFO 06-07 16:56:09 [model.py:1549] Using max model len 32768
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m INFO 06-07 16:56:09 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'NCCL_DEBUG': '***', 'NCCL_SOCKET_FAMILY': '***', 'NCCL_SOCKET_IFNAME': '***', 'RAY_ADDRESS': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}}
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m INFO 06-07 16:56:09 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=65536.
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m INFO 06-07 16:56:09 [vllm.py:690] Asynchronous scheduling is enabled.
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m WARNING 06-07 16:56:09 [vllm.py:728] Enforce eager set, overriding optimization level to -O0
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m INFO 06-07 16:56:09 [vllm.py:846] Cudagraph is disabled under eager mode
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706739) INFO 06-07 16:56:00 [gpu_model_runner.py:4125] Starting to load model /e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6...[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m WARNING 06-07 16:56:09 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m WARNING 06-07 16:56:09 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:56:41] OK: 88 / 131,072 FDs open (0.1% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:56:41] OK: RSS 1.64 GiB | node mem 161.1/858.0 GiB used (18.8%), avail 696.9 GiB
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m INFO 06-07 16:56:09 [model.py:529] Resolved architecture: Qwen3ForCausalLM[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m INFO 06-07 16:56:09 [model.py:1549] Using max model len 32768[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m INFO 06-07 16:56:09 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'NCCL_DEBUG': '***', 'NCCL_SOCKET_FAMILY': '***', 'NCCL_SOCKET_IFNAME': '***', 'RAY_ADDRESS': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}}[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m INFO 06-07 16:56:09 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=65536.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m INFO 06-07 16:56:09 [vllm.py:690] Asynchronous scheduling is enabled.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m WARNING 06-07 16:56:09 [vllm.py:728] Enforce eager set, overriding optimization level to -O0[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m INFO 06-07 16:56:09 [vllm.py:846] Cudagraph is disabled under eager mode[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m WARNING 06-07 16:56:09 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m WARNING 06-07 16:56:09 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m INFO 06-07 16:56:50 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832747) INFO 06-07 16:56:52 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', speculative_config=None, tokenizer='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=82, served_model_name=0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [65536], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706759) INFO 06-07 16:56:54 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706759) INFO 06-07 16:57:08 [default_loader.py:293] Loading weights took 13.39 seconds
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m INFO 06-07 16:56:50 [pynccl.py:178] pynccl trace buffer enabled: size=10000 dump_dir=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps flush_interval=0 s[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832741) INFO 06-07 16:56:52 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', speculative_config=None, tokenizer='/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=85, served_model_name=0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [65536], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706739) INFO 06-07 16:56:54 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706759) INFO 06-07 16:57:08 [gpu_model_runner.py:4222] Model loading took 15.27 GiB memory and 67.984083 seconds
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706759) INFO 06-07 16:57:11 [gpu_worker.py:373] Available KV cache memory: 49.32 GiB
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706759) INFO 06-07 16:57:11 [kv_cache_utils.py:1307] GPU KV cache size: 359,104 tokens
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706759) INFO 06-07 16:57:11 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 10.96x
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706759) INFO 06-07 16:57:11 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled.
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706759) INFO 06-07 16:57:11 [core.py:278] init engine (profile, create kv cache, warmup model) took 2.83 seconds
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706759) WARNING 06-07 16:57:12 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706759) INFO 06-07 16:57:12 [vllm.py:690] Asynchronous scheduling is enabled.
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706759) WARNING 06-07 16:57:12 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706759) INFO 06-07 16:57:12 [vllm.py:846] Cudagraph is disabled under eager mode
|
||
[36m(AsyncVLLMInferenceEngine pid=2705575, ip=10.128.32.38)[0m WARNING 06-07 16:57:12 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:58:41] OK: 88 / 131,072 FDs open (0.1% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [16:58:41] OK: RSS 1.64 GiB | node mem 161.1/858.0 GiB used (18.8%), avail 696.9 GiB
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706739) INFO 06-07 16:57:08 [default_loader.py:293] Loading weights took 13.39 seconds[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706739) INFO 06-07 16:57:08 [gpu_model_runner.py:4222] Model loading took 15.27 GiB memory and 67.996617 seconds[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706739) INFO 06-07 16:57:11 [gpu_worker.py:373] Available KV cache memory: 49.32 GiB[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706739) INFO 06-07 16:57:11 [kv_cache_utils.py:1307] GPU KV cache size: 359,104 tokens[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706739) INFO 06-07 16:57:11 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 10.96x[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706739) INFO 06-07 16:57:11 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706739) INFO 06-07 16:57:11 [core.py:278] init engine (profile, create kv cache, warmup model) took 2.82 seconds[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706739) WARNING 06-07 16:57:12 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706739) INFO 06-07 16:57:12 [vllm.py:690] Asynchronous scheduling is enabled.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706739) WARNING 06-07 16:57:12 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m (EngineCore_DP0 pid=2706739) INFO 06-07 16:57:12 [vllm.py:846] Cudagraph is disabled under eager mode[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2705803, ip=10.128.32.38)[0m WARNING 06-07 16:57:12 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832747) INFO 06-07 16:59:59 [worker_base.py:289] Injected <class 'skyrl_train.inference_engines.vllm.vllm_engine.WorkerWrap'> into <class 'vllm.v1.worker.gpu_worker.Worker'> for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'read_named_weights', 'set_numa_affinity', 'test_rpc']
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832747) INFO 06-07 17:00:00 [parallel_state.py:1234] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.128.32.45:41873 backend=nccl
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832747) INFO 06-07 17:00:00 [parallel_state.py:1445] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [17:00:41] OK: 88 / 131,072 FDs open (0.1% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [17:00:41] OK: RSS 1.64 GiB | node mem 161.1/858.0 GiB used (18.8%), avail 696.9 GiB
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832741) INFO 06-07 16:59:59 [worker_base.py:289] Injected <class 'skyrl_train.inference_engines.vllm.vllm_engine.WorkerWrap'> into <class 'vllm.v1.worker.gpu_worker.Worker'> for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'read_named_weights', 'set_numa_affinity', 'test_rpc'][32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832741) INFO 06-07 17:00:00 [parallel_state.py:1234] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.128.32.45:60113 backend=nccl[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832741) INFO 06-07 17:00:00 [parallel_state.py:1445] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832747) INFO 06-07 17:00:52 [gpu_model_runner.py:4125] Starting to load model /e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6...
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832747) INFO 06-07 17:00:53 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832747) INFO 06-07 17:01:07 [default_loader.py:293] Loading weights took 13.57 seconds
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832741) INFO 06-07 17:00:52 [gpu_model_runner.py:4125] Starting to load model /e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6...[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832741) INFO 06-07 17:00:53 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832747) INFO 06-07 17:01:08 [gpu_model_runner.py:4222] Model loading took 15.27 GiB memory and 15.159333 seconds
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832747) INFO 06-07 17:01:10 [gpu_worker.py:373] Available KV cache memory: 49.32 GiB
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832747) INFO 06-07 17:01:10 [kv_cache_utils.py:1307] GPU KV cache size: 359,104 tokens
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832747) INFO 06-07 17:01:10 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 10.96x
|
||
[36m(AsyncVLLMInferenceEngine pid=2831684, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832749) INFO 06-07 17:01:10 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled.
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832747) INFO 06-07 17:01:10 [core.py:278] init engine (profile, create kv cache, warmup model) took 2.83 seconds
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832747) WARNING 06-07 17:01:11 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832747) INFO 06-07 17:01:11 [vllm.py:690] Asynchronous scheduling is enabled.
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832747) WARNING 06-07 17:01:11 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832747) INFO 06-07 17:01:11 [vllm.py:846] Cudagraph is disabled under eager mode
|
||
[36m(AsyncVLLMInferenceEngine pid=2831684, ip=10.128.32.45)[0m WARNING 06-07 17:01:11 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.
|
||
[36m(pid=2662625, ip=10.128.32.46)[0m ⚙️ Running in WANDB offline mode
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832741) INFO 06-07 17:01:07 [default_loader.py:293] Loading weights took 13.57 seconds[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832741) INFO 06-07 17:01:08 [gpu_model_runner.py:4222] Model loading took 15.27 GiB memory and 15.161752 seconds[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832741) INFO 06-07 17:01:10 [gpu_worker.py:373] Available KV cache memory: 49.32 GiB[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832741) INFO 06-07 17:01:10 [kv_cache_utils.py:1307] GPU KV cache size: 359,104 tokens[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832741) INFO 06-07 17:01:10 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 10.96x[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832741) INFO 06-07 17:01:10 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832741) INFO 06-07 17:01:10 [core.py:278] init engine (profile, create kv cache, warmup model) took 2.84 seconds[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832741) WARNING 06-07 17:01:11 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832741) INFO 06-07 17:01:11 [vllm.py:690] Asynchronous scheduling is enabled.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832741) WARNING 06-07 17:01:11 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831686, ip=10.128.32.45)[0m (EngineCore_DP0 pid=2832741) INFO 06-07 17:01:11 [vllm.py:846] Cudagraph is disabled under eager mode[32m [repeated 3x across cluster][0m
|
||
[36m(AsyncVLLMInferenceEngine pid=2831550, ip=10.128.32.45)[0m WARNING 06-07 17:01:11 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.[32m [repeated 3x across cluster][0m
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [17:02:41] OK: 86 / 131,072 FDs open (0.1% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [17:02:41] OK: RSS 1.64 GiB | node mem 162.3/858.0 GiB used (18.9%), avail 695.7 GiB
|
||
[36m(pid=2663493, ip=10.128.32.46)[0m ⚙️ Running in WANDB offline mode
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [17:04:41] OK: 89 / 131,072 FDs open (0.1% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [17:04:41] OK: RSS 1.64 GiB | node mem 162.4/858.0 GiB used (18.9%), avail 695.6 GiB
|
||
[36m(pid=2663495, ip=10.128.32.46)[0m ⚙️ Running in WANDB offline mode[32m [repeated 2x across cluster][0m
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [17:06:41] OK: 89 / 131,072 FDs open (0.1% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [17:06:41] OK: RSS 1.64 GiB | node mem 162.4/858.0 GiB used (18.9%), avail 695.5 GiB
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [17:08:41] OK: 83 / 131,072 FDs open (0.1% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [17:08:41] OK: RSS 1.64 GiB | node mem 162.4/858.0 GiB used (18.9%), avail 695.6 GiB
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [17:10:41] OK: 83 / 131,072 FDs open (0.1% of soft limit, hard limit: 131,072)
|
||
[36m(skyrl_entrypoint pid=391889)[0m [fd-monitor] [17:10:41] OK: RSS 1.64 GiB | node mem 162.5/858.0 GiB used (18.9%), avail 695.5 GiB
|
||
Error executing job with overrides: ['+terminal_bench_config=terminal_bench', 'trainer.strategy=fsdp2', 'trainer.algorithm.advantage_estimator=rloo_n', 'trainer.algorithm.use_kl_loss=false', 'trainer.algorithm.kl_loss_coef=0.0', 'trainer.algorithm.eps_clip_low=0.2', 'trainer.algorithm.eps_clip_high=0.05', 'trainer.algorithm.loss_reduction=sequence_mean', 'trainer.epochs=2', 'trainer.max_steps=80', 'trainer.update_epochs_per_batch=1', 'trainer.train_batch_size=64', 'trainer.policy_mini_batch_size=64', 'trainer.eval_batch_size=64', 'trainer.micro_forward_batch_size_per_gpu=4', 'trainer.micro_train_batch_size_per_gpu=1', 'trainer.max_prompt_length=999999', 'trainer.eval_interval=999999', 'trainer.eval_before_train=false', 'trainer.ckpt_interval=2', 'trainer.resume_mode=latest', 'trainer.hf_save_interval=5', '++trainer.hf_hub_repo_id=laion/ablation-pymethods2test-seqmean-arm0', '++trainer.hf_hub_private=false', '++trainer.hf_hub_revision=main', '++trainer.enable_db_registration=false', 'trainer.project_name=OpenThoughts-Agent', 'trainer.log_level=INFO', 'trainer.tracker_commit_each_step=true', 'trainer.logger=console', 'trainer.run_name=ablation-pymethods2test-seqmean-arm0', 'trainer.ckpt_path=/e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqmean-arm0/ablation-pymethods2test-seqmean-arm0/checkpoints', 'trainer.export_path=/e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqmean-arm0/ablation-pymethods2test-seqmean-arm0/exports', 'trainer.policy.optimizer_config.lr=8e-6', 'trainer.policy.optimizer_config.weight_decay=0.0', 'trainer.policy.optimizer_config.adam_betas=[0.9,0.999]', 'trainer.policy.optimizer_config.max_grad_norm=0.9', 'trainer.policy.fsdp_config.cpu_offload=false', 'trainer.policy.fsdp_config.reshard_after_forward=true', 'trainer.policy.fsdp_config.fsdp_size=4', 'trainer.policy.model.path=/e/home/jusers/feuer1/jupiter/.cache/huggingface/hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', 'trainer.ref.fsdp_config.cpu_offload=false', 'trainer.ref.fsdp_config.reshard_after_forward=true', 'trainer.ref.fsdp_config.fsdp_size=4', 'trainer.placement.colocate_all=false', 'trainer.placement.policy_num_nodes=2', 'trainer.placement.ref_num_nodes=2', 'trainer.placement.policy_num_gpus_per_node=4', 'trainer.placement.ref_num_gpus_per_node=4', 'trainer.fully_async.max_staleness_steps=16', 'trainer.fully_async.num_parallel_generation_workers=338', 'generator.backend=vllm', 'generator.timeout_multiplier=1.0', 'generator.model_dtype=bfloat16', 'generator.inference_engine_tensor_parallel_size=1', 'generator.num_inference_engines=48', 'generator.n_samples_per_prompt=8', 'generator.eval_n_samples_per_prompt=8', 'generator.gpu_memory_utilization=0.75', 'generator.max_num_seqs=24', 'generator.max_num_batched_tokens=65536', 'generator.enable_prefix_caching=true', 'generator.enable_chunked_prefill=true', 'generator.run_engines_locally=true', 'generator.weight_sync_backend=nccl', 'generator.async_engine=true', 'generator.batched=false', 'generator.enable_http_endpoint=true', 'generator.enable_ray_prometheus_stats=false', 'generator.vllm_stats_interval=1', 'generator.append_eos_token_after_stop_str_in_multi_turn=true', 'generator.max_turns=999999', 'generator.sampling_params.max_generate_length=4096', 'generator.sampling_params.temperature=0.7', 'generator.sampling_params.top_p=0.95', 'generator.sampling_params.top_k=20', '++generator.engine_init_kwargs.max_model_len=32768', '++generator.engine_init_kwargs.custom_chat_template_chat_completion_path=chat_templates/qwen3_thinking_acc.jinja2', '++generator.engine_init_kwargs.served_model_name=0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', 'data.train_data=["/e/scratch/jureap59/feuer1/tasks/exp_rpt_pymethods2test-large"]', 'data.val_data=[]', '+terminal_bench_config.trials_dir=/e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqmean-arm0/ablation-pymethods2test-seqmean-arm0/trace_jobs', '+terminal_bench_config.harbor.name=terminus-2', '+terminal_bench_config.harbor.max_episodes=999999', '+terminal_bench_config.harbor.enable_summarize=false', '+terminal_bench_config.harbor.store_all_messages=true', '+terminal_bench_config.harbor.trajectory_config.raw_content=true', '+terminal_bench_config.harbor.enable_episode_logging=false', '+terminal_bench_config.harbor.record_terminal_session=false', '+terminal_bench_config.harbor.enable_pane_logging=false', '+terminal_bench_config.harbor.strict_json_parser=true', '+terminal_bench_config.harbor.interleaved_thinking=true', '+terminal_bench_config.harbor.extra_body.chat_template_kwargs.enable_thinking=true', '+terminal_bench_config.harbor.override_timeout_sec=900', '+terminal_bench_config.harbor.override_cpus=1', '+terminal_bench_config.harbor.override_memory_mb=2048', '+terminal_bench_config.harbor.override_storage_mb=2048', '+terminal_bench_config.harbor.auto_snapshot=true', '+terminal_bench_config.harbor.verifier_override_timeout_sec=120', '+terminal_bench_config.harbor.max_retries=3', '+terminal_bench_config.harbor.min_wait_sec=60.0', '+terminal_bench_config.harbor.max_wait_sec=600.0', '+terminal_bench_config.harbor.wait_multiplier=2.0', '+terminal_bench_config.harbor.exclude_exceptions=["VerifierTimeoutError","VerifierRuntimeError","RewardFileNotFoundError","RewardFileEmptyError","VerifierOutputParseError"]', '+terminal_bench_config.harbor.n_concurrent_trials=675', '+terminal_bench_config.harbor.log_level=INFO', '+terminal_bench_config.harbor.enable_reward_shaping=false', '+terminal_bench_config.harbor.enable_error_classification=true', '+terminal_bench_config.harbor.mask_exceptions=["DaytonaError","EnvironmentStartTimeoutError","NetworkError","ConnectionError","RewardFileNotFoundError","RewardFileEmptyError","AgentEnvironmentTimeoutError","ContextLengthExceededError"]', '+terminal_bench_config.harbor.default_error_treatment=zero', '+terminal_bench_config.harbor.passthrough_exceptions=["AgentTimeoutError"]', '+terminal_bench_config.harbor.zero_exceptions=[]', '+terminal_bench_config.model_info.max_input_tokens=32000', '+terminal_bench_config.model_info.max_output_tokens=4096', '+terminal_bench_config.archiving.enabled=false', '+terminal_bench_config.trace_upload.enabled=true', '+terminal_bench_config.trace_upload.repo_org=DCAgent', '+terminal_bench_config.trace_upload.episodes=last', '+terminal_bench_config.trace_upload.dataset_type=SFT', '+terminal_bench_config.trace_upload.cleanup=true']
|
||
Traceback (most recent call last):
|
||
File "<frozen runpy>", line 198, in _run_module_as_main
|
||
File "<frozen runpy>", line 88, in _run_code
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train/examples/terminal_bench/entrypoints/main_tbench.py", line 142, in <module>
|
||
main()
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/hydra/main.py", line 94, in decorated_main
|
||
_run_hydra(
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/hydra/_internal/utils.py", line 394, in _run_hydra
|
||
_run_app(
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/hydra/_internal/utils.py", line 457, in _run_app
|
||
run_and_report(
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/hydra/_internal/utils.py", line 223, in run_and_report
|
||
raise ex
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/hydra/_internal/utils.py", line 220, in run_and_report
|
||
return func()
|
||
^^^^^^
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/hydra/_internal/utils.py", line 458, in <lambda>
|
||
lambda: hydra.run(
|
||
^^^^^^^^^^
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/hydra/_internal/hydra.py", line 132, in run
|
||
_ = ret.return_value
|
||
^^^^^^^^^^^^^^^^
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/hydra/core/utils.py", line 260, in return_value
|
||
raise self._return_value
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/hydra/core/utils.py", line 186, in run_job
|
||
ret.return_value = task_function(task_cfg)
|
||
^^^^^^^^^^^^^^^^^^^^^^^
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train/examples/terminal_bench/entrypoints/main_tbench.py", line 132, in main
|
||
ray.get(skyrl_entrypoint.remote(cfg))
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/ray/_private/auto_init_hook.py", line 22, in auto_init_wrapper
|
||
return fn(*args, **kwargs)
|
||
^^^^^^^^^^^^^^^^^^^
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/ray/_private/client_mode_hook.py", line 104, in wrapper
|
||
return func(*args, **kwargs)
|
||
^^^^^^^^^^^^^^^^^^^^^
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/ray/_private/worker.py", line 2967, in get
|
||
values, debugger_breakpoint = worker.get_objects(
|
||
^^^^^^^^^^^^^^^^^^^
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/ray/_private/worker.py", line 1015, in get_objects
|
||
raise value.as_instanceof_cause()
|
||
ray.exceptions.RayTaskError(DistStoreError): [36mray::skyrl_entrypoint()[39m (pid=391889, ip=10.128.32.33)
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train/examples/terminal_bench/entrypoints/main_tbench.py", line 105, in skyrl_entrypoint
|
||
exp.run()
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train/skyrl_train/entrypoints/main_base.py", line 481, in run
|
||
trainer = self._setup_trainer()
|
||
^^^^^^^^^^^^^^^^^^^^^
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train/skyrl_train/entrypoints/main_base.py", line 450, in _setup_trainer
|
||
trainer.build_models(PolicyWorker, CriticWorker, RefWorker, policy_pg=self.policy_pg)
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train/skyrl_train/trainer.py", line 768, in build_models
|
||
policy_model = PPORayActorGroup(
|
||
^^^^^^^^^^^^^^^^^
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train/skyrl_train/workers/worker.py", line 511, in __init__
|
||
self._initiate_actors(pg, num_gpus_per_actor)
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train/skyrl_train/workers/worker.py", line 637, in _initiate_actors
|
||
ray.get([actor.init_worker_process_group.remote() for actor in self._actor_handlers])
|
||
^^^^^^^^^^^^^^^^^^^
|
||
^^^^^^^^^^^^^^^^^^^^^
|
||
^^^^^^^^^^^^^^^^^^^
|
||
ray.exceptions.RayTaskError(DistStoreError): [36mray::FSDPPolicyWorkerBase.init_worker_process_group()[39m (pid=2662625, ip=10.128.32.46, actor_id=b5939e2b137362308272f3e202000000, repr=<skyrl_train.workers.fsdp.fsdp_worker.FSDPPolicyWorkerBase object at 0x400df5e2cf50>)
|
||
File "/e/scratch/jureap59/feuer1/miniforge3/lib/python3.12/concurrent/futures/_base.py", line 449, in result
|
||
return self.__get_result()
|
||
^^^^^^^^^^^^^^^^^^^
|
||
File "/e/scratch/jureap59/feuer1/miniforge3/lib/python3.12/concurrent/futures/_base.py", line 401, in __get_result
|
||
raise self._exception
|
||
^^^^^^^^^^^^^^^^^^^^^
|
||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train/skyrl_train/workers/worker.py", line 149, in init_worker_process_group
|
||
torch.distributed.init_process_group(
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/distributed/c10d_logger.py", line 81, in wrapper
|
||
return func(*args, **kwargs)
|
||
^^^^^^^^^^^^^^^^^^^^^
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/distributed/c10d_logger.py", line 95, in wrapper
|
||
func_return = func(*args, **kwargs)
|
||
^^^^^^^^^^^^^^^^^^^^^
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/distributed/distributed_c10d.py", line 1762, in init_process_group
|
||
store, rank, world_size = next(rendezvous_iterator)
|
||
^^^^^^^^^^^^^^^^^^^^^^^^^
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/distributed/rendezvous.py", line 278, in _env_rendezvous_handler
|
||
store = _create_c10d_store(
|
||
^^^^^^^^^^^^^^^^^^^
|
||
File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/distributed/rendezvous.py", line 198, in _create_c10d_store
|
||
return TCPStore(
|
||
^^^^^^^^^
|
||
torch.distributed.DistStoreError: Timed out after 601 seconds waiting for clients. 4/8 clients joined.
|
||
Exception raised from waitForWorkers at /pytorch/torch/csrc/distributed/c10d/TCPStore.cpp:396 (most recent call first):
|
||
frame #0: c10::Error::Error(c10::SourceLocation, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >) + 0xb0 (0x4000bf09c700 in /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/lib/libc10.so)
|
||
frame #1: <unknown function> + 0x5e9c9c0 (0x400ae12fc9c0 in /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/lib/libtorch_cpu.so)
|
||
frame #2: c10d::TCPStore::waitForWorkers() + 0x350 (0x400ae1391410 in /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/lib/libtorch_cpu.so)
|
||
frame #3: c10d::TCPStore::TCPStore(std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, c10d::TCPStoreOptions const&) + 0x468 (0x400ae13918c8 in /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/lib/libtorch_cpu.so)
|
||
frame #4: <unknown function> + 0x109a094 (0x4000be86a094 in /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/lib/libtorch_python.so)
|
||
frame #5: <unknown function> + 0x113236c (0x4000be90236c in /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/lib/libtorch_python.so)
|
||
frame #6: <unknown function> + 0x5d6d60 (0x4000bdda6d60 in /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/lib/libtorch_python.so)
|
||
frame #7: <unknown function> + 0x1b7a38 (0xaaaacb007a38 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #8: _PyObject_MakeTpCall + 0x98 (0xaaaacafb5db8 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #9: <unknown function> + 0x169f50 (0xaaaacafb9f50 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #10: <unknown function> + 0x1682e4 (0xaaaacafb82e4 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #11: <unknown function> + 0x1e0ce8 (0xaaaacb030ce8 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #12: <unknown function> + 0x1d7ddc (0xaaaacb027ddc in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #13: <unknown function> + 0x646b0c (0x4000bde16b0c in /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/lib/libtorch_python.so)
|
||
frame #14: _PyObject_MakeTpCall + 0x98 (0xaaaacafb5db8 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #15: _PyEval_EvalFrameDefault + 0x280c (0xaaaacb0bae54 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #16: <unknown function> + 0x1808c0 (0xaaaacafd08c0 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #17: <unknown function> + 0x182bf8 (0xaaaacafd2bf8 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #18: <unknown function> + 0x25fd30 (0xaaaacb0afd30 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #19: <unknown function> + 0x1b7d20 (0xaaaacb007d20 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #20: PyObject_Vectorcall + 0x54 (0xaaaacafb60e4 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #21: _PyEval_EvalFrameDefault + 0x280c (0xaaaacb0bae54 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #22: <unknown function> + 0x1808c0 (0xaaaacafd08c0 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #23: <unknown function> + 0x1835c8 (0xaaaacafd35c8 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #24: <unknown function> + 0x84b708 (0x40000a80b708 in /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/ray/_raylet.so)
|
||
frame #25: <unknown function> + 0x865fe4 (0x40000a825fe4 in /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/ray/_raylet.so)
|
||
frame #26: <unknown function> + 0x823470 (0x40000a7e3470 in /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/ray/_raylet.so)
|
||
frame #27: <unknown function> + 0x830110 (0x40000a7f0110 in /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/ray/_raylet.so)
|
||
frame #28: <unknown function> + 0x174a60 (0xaaaacafc4a60 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #29: PyObject_VectorcallMethod + 0xa4 (0xaaaacafb6270 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #30: PyIter_Send + 0xbc (0xaaaacaf9b62c in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #31: <unknown function> + 0xaa04 (0x40000ca7aa04 in /e/scratch/jureap59/feuer1/miniforge3/lib/python3.12/lib-dynload/_asyncio.cpython-312-aarch64-linux-gnu.so)
|
||
frame #32: <unknown function> + 0xbcc4 (0x40000ca7bcc4 in /e/scratch/jureap59/feuer1/miniforge3/lib/python3.12/lib-dynload/_asyncio.cpython-312-aarch64-linux-gnu.so)
|
||
frame #33: _PyObject_MakeTpCall + 0x98 (0xaaaacafb5db8 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #34: <unknown function> + 0x28cfac (0xaaaacb0dcfac in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #35: <unknown function> + 0x1b7b68 (0xaaaacb007b68 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #36: _PyEval_EvalFrameDefault + 0x52ac (0xaaaacb0bd8f4 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #37: <unknown function> + 0x83c564 (0x40000a7fc564 in /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/ray/_raylet.so)
|
||
frame #38: _PyEval_EvalFrameDefault + 0x52ac (0xaaaacb0bd8f4 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #39: <unknown function> + 0x169f88 (0xaaaacafb9f88 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #40: <unknown function> + 0x36002c (0xaaaacb1b002c in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #41: <unknown function> + 0x2e4024 (0xaaaacb134024 in ray::FSDPPolicyWorkerBase.init_worker_process_group)
|
||
frame #42: <unknown function> + 0x80e00 (0x400009730e00 in /lib64/libc.so.6)
|
||
frame #43: <unknown function> + 0xeb49c (0x40000979b49c in /lib64/libc.so.6)
|
||
Stopping Ray cluster...
|
||
Warning: Failed to stop Ray on jpbo-041-36: Command '['srun', '--export=ALL,WANDB_MODE=offline,GLOO_USE_IPV6=0,NCCL_SOCKET_FAMILY=AF_INET,VLLM_FORCE_IPV4=1,VLLM_SKIP_FLAG_DISCOVERY=1,SKYRL_ENABLE_NUMA_AFFINITY=1,DISABLE_AIOHTTP_TRANSPORT=True,VLLM_ALLREDUCE_USE_SYMM_MEM=0,TORCH_CUDNN_SDPA_ENABLED=0,PYTHONFAULTHANDLER=1,TORCH_NCCL_ASYNC_ERROR_HANDLING=1,TORCH_NCCL_HEARTBEAT_TIMEOUT_SEC=1800,TORCH_NCCL_BLOCKING_WAIT_TIMEOUT_MS=1800000,VLLM_MQ_MAX_CHUNKS=240,OT_AGENT_RAY_LOG_DIR=/e/data1/datasets/playground/ot-baf/experiments/_ray_logs,TORCH_NCCL_TRACE_BUFFER_SIZE=10000,TORCH_FR_BUFFER_SIZE=10000,TORCH_NCCL_DESYNC_DEBUG=1,TORCH_NCCL_DEBUG_INFO_TEMP_FILE=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps/nccl_trace,VLLM_PYNCCL_TRACE_BUFFER_SIZE=10000,VLLM_PYNCCL_TRACE_DUMP_DIR=/e/data1/datasets/playground/ot-baf/experiments/_nccl_dumps/nccl_trace,VLLM_RAY_EXTRA_ENV_VAR_PREFIXES_TO_COPY=TORCH_NCCL_,TORCH_FR_,VLLM_PYNCCL_', '--nodes=1', '--ntasks=1', '--overlap', '--cpu-bind=none', '-w', 'jpbo-041-36', 'bash', '-c', 'unset LD_PRELOAD PROXYCHAINS_CONF_FILE 2>/dev/null; ray stop --force']' timed out after 30 seconds
|
||
Ray cluster stopped
|
||
[RLJobRunner] Crash detected (exit!=0) — preserving Ray logs to /e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqmean-arm0/ablation-pymethods2test-seqmean-arm0/ray_logs BEFORE trace upload (so a wall-clock kill can't lose crash evidence)...
|
||
[RLJobRunner] Crash-time Ray log preservation timed out (600s); continuing.
|
||
[RLJobRunner] Launching trace upload (training exit code: 1):
|
||
repo_id: DCAgent/ablation-pymethods2test-seqmean-arm0
|
||
job_dir: /e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqmean-arm0/ablation-pymethods2test-seqmean-arm0
|
||
episodes: last
|
||
log: /e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqmean-arm0/logs/ablation-pymethods2test-seqmean-arm0_trace_upload.log
|
||
[RLJobRunner] Waiting for trace upload to complete...
|
||
[RLJobRunner] Trace upload failed with exit code 1.
|
||
Preserving Ray logs to /e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqmean-arm0/ray_logs/
|
||
Collecting Ray logs from worker jpbo-041-34...
|
||
Collecting Ray logs from worker jpbo-041-35...
|
||
Collecting Ray logs from worker jpbo-041-36...
|
||
Collecting Ray logs from worker jpbo-041-37...
|
||
Collecting Ray logs from worker jpbo-041-38...
|
||
Collecting Ray logs from worker jpbo-041-39...
|
||
Collecting Ray logs from worker jpbo-041-40...
|
||
Collecting Ray logs from worker jpbo-041-41...
|
||
Collecting Ray logs from worker jpbo-041-42...
|
||
Collecting Ray logs from worker jpbo-041-43...
|
||
Collecting Ray logs from worker jpbo-041-44...
|
||
Collecting Ray logs from worker jpbo-041-45...
|
||
Collecting Ray logs from worker jpbo-041-46...
|
||
Ray log preservation complete
|