The following have been reloaded with a version change: 1) GCCcore/.14.3.0 => GCCcore/14.3.0 Lmod is automatically replacing "GCC/14.3.0" with "nvidia-compilers/25.9-CUDA-13". Deactivating conda environment: /e/scratch/jureap59/feuer1/miniforge3/envs/otagent Activating RL environment: /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl Python executable: /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/bin/python Python path check: /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/bin/python [ray] RAY_TMPDIR=/tmp/ray/ray_736734 [triton_cache] Triton cache: /tmp/triton_cache_feuer1_736734 [triton_cache] TorchInductor cache: /tmp/torchinductor_cache_feuer1_736734 [proxy] ✓ Found proxychains binary at /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 [proxy] Setting up SSH tunnel to jpbl-s01-01 [proxy] SSH key: /e/home/jusers/feuer1/jupiter/.ssh/authorized_keys/id_ed25519_jsc [proxy] Tunnel port: 7003 [proxy] Node IP: 10.128.17.29 (workers will connect here) [proxy] ✓ SSH tunnel started successfully [proxy] ✓ Generated proxychains config at /e/home/jusers/feuer1/jupiter/.proxychains/proxychains_736734.conf [proxy] - Internal traffic (10.x.x.x, 172.x.x.x, 169.254.x.x) → DIRECT [proxy] - External traffic (internet) → PROXY via tunnel [proxy] ✓ Daytona timeout settings configured [proxy] Testing proxy connectivity... [proxychains] config file found: /e/home/jusers/feuer1/jupiter/.proxychains/proxychains_736734.conf [proxychains] preloading /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/lib/libproxychains4.so [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e [proxy] ✓ Proxy connectivity test passed (huggingface.co reachable via wrapped binary) [proxy] ✓ Tunnel accessible at 10.128.17.29:7003 (workers can connect) [proxy] ✓ Proxy setup complete (using wrapped binary for Ray workers) [container_runtime] Using cloud backend: daytona (no local container setup) === Universal RL Training Runner === Config: /e/data1/datasets/playground/ot-baf/explore-tis-minp/configs/explore-tis-minp_rl_config.json Working directory: /e/scratch/jureap59/feuer1/OpenThoughts-Agent Python: /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/bin/python Python version: Python 3.12.12 UV_USE_IO_URING: 0 Proxy: DISABLED (direct internet or not configured) ======================================== === RLJobRunner: explore-tis-minp === [wandb_utils] Fixing permissions on: /e/data1/datasets/playground/ot-baf/explore-tis-minp/wandb [wandb_utils] WandB directory ready: /e/data1/datasets/playground/ot-baf/explore-tis-minp/wandb HF_TOKEN=****pDbg HF_HUB_CACHE=/e/data1/datasets/playground/ot-baf/hf_hub SUPABASE_URL=https://rpzmyuapoqilpghynmza.s... (direct Supabase config) Environment configured: TENSOR_PARALLEL_SIZE=1 NUM_INFERENCE_ENGINES=56 POLICY_NUM_NODES=14 WANDB_DIR=/e/data1/datasets/playground/ot-baf/explore-tis-minp/wandb Starting Ray cluster with 14 nodes, 4 GPUs/node Cleaning up existing Ray instances... === Starting Ray Cluster === Nodes: 14 GPUs per node: 4 CPUs per node: 288 Head node: jpbo-006-45 (10.128.17.29) Ray port: 6379 ============================ Starting Ray head on jpbo-006-45 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_head_jpbo-006-45.log)... Command: srun --export=ALL,VLLM_HOST_IP=10.128.17.29 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-006-45 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --head --node-ip-address=10.128.17.29 --port=6379 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 Started Ray head on jpbo-006-45 Starting Ray worker on jpbo-007-02 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-007-02.log)... Command: srun --export=ALL,VLLM_HOST_IP=10.128.17.34 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-007-02 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.17.29:6379 --node-ip-address=10.128.17.34 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 Started Ray worker 1 on jpbo-007-02 Starting Ray worker on jpbo-007-11 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-007-11.log)... Command: srun --export=ALL,VLLM_HOST_IP=10.128.17.43 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-007-11 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.17.29:6379 --node-ip-address=10.128.17.43 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 Started Ray worker 2 on jpbo-007-11 Starting Ray worker on jpbo-007-21 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-007-21.log)... Command: srun --export=ALL,VLLM_HOST_IP=10.128.17.53 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-007-21 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.17.29:6379 --node-ip-address=10.128.17.53 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 Started Ray worker 3 on jpbo-007-21 Starting Ray worker on jpbo-007-23 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-007-23.log)... Command: srun --export=ALL,VLLM_HOST_IP=10.128.17.55 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-007-23 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.17.29:6379 --node-ip-address=10.128.17.55 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 Started Ray worker 4 on jpbo-007-23 Starting Ray worker on jpbo-007-31 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-007-31.log)... Command: srun --export=ALL,VLLM_HOST_IP=10.128.17.63 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-007-31 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.17.29:6379 --node-ip-address=10.128.17.63 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 Started Ray worker 5 on jpbo-007-31 Starting Ray worker on jpbo-010-46 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-010-46.log)... Command: srun --export=ALL,VLLM_HOST_IP=10.128.17.222 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-010-46 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.17.29:6379 --node-ip-address=10.128.17.222 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 Started Ray worker 6 on jpbo-010-46 Starting Ray worker on jpbo-010-47 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-010-47.log)... Command: srun --export=ALL,VLLM_HOST_IP=10.128.17.223 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-010-47 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.17.29:6379 --node-ip-address=10.128.17.223 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 Started Ray worker 7 on jpbo-010-47 Starting Ray worker on jpbo-010-48 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-010-48.log)... Command: srun --export=ALL,VLLM_HOST_IP=10.128.17.224 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-010-48 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.17.29:6379 --node-ip-address=10.128.17.224 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 Started Ray worker 8 on jpbo-010-48 Starting Ray worker on jpbo-051-33 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-051-33.log)... Command: srun --export=ALL,VLLM_HOST_IP=10.128.34.1 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-051-33 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.17.29:6379 --node-ip-address=10.128.34.1 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 Started Ray worker 9 on jpbo-051-33 Starting Ray worker on jpbo-051-34 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-051-34.log)... Command: srun --export=ALL,VLLM_HOST_IP=10.128.34.2 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-051-34 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.17.29:6379 --node-ip-address=10.128.34.2 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 Started Ray worker 10 on jpbo-051-34 Starting Ray worker on jpbo-051-37 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-051-37.log)... Command: srun --export=ALL,VLLM_HOST_IP=10.128.34.5 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-051-37 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.17.29:6379 --node-ip-address=10.128.34.5 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 Started Ray worker 11 on jpbo-051-37 Starting Ray worker on jpbo-051-38 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-051-38.log)... Command: srun --export=ALL,VLLM_HOST_IP=10.128.34.6 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-051-38 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.17.29:6379 --node-ip-address=10.128.34.6 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 Started Ray worker 12 on jpbo-051-38 Starting Ray worker on jpbo-051-40 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-051-40.log)... Command: srun --export=ALL,VLLM_HOST_IP=10.128.34.8 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-051-40 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.17.29:6379 --node-ip-address=10.128.34.8 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 Started Ray worker 13 on jpbo-051-40 Waiting for cluster (56 GPUs, 14 nodes)... Connecting to Ray at 10.128.17.29:6379 (expecting 14 nodes, 56.0 GPUs) Ray connection established, polling for resources... [Ray wait] nodes=14/14 GPUs=56.0/56.0 resources={'object_store_memory': 601295421440.0, 'memory': 10786241839104.0, 'node:10.128.34.5': 1.0, 'CPU': 4032.0, 'accelerator_type:GH200': 14.0, 'GPU': 56.0, 'node:10.128.17.34': 1.0, 'node:10.128.34.8': 1.0, 'node:10.128.17.223': 1.0, 'node:10.128.34.6': 1.0, 'node:10.128.17.53': 1.0, 'node:10.128.17.55': 1.0, 'node:10.128.17.224': 1.0, 'node:__internal_head__': 1.0, 'node:10.128.17.29': 1.0, 'node:10.128.34.1': 1.0, 'node:10.128.34.2': 1.0, 'node:10.128.17.63': 1.0, 'node:10.128.17.222': 1.0, 'node:10.128.17.43': 1.0} ✓ Ray cluster ready === Ray Cluster Ready === Address: 10.128.17.29:6379 Total GPUs: 56 ========================= Ray cluster ready at 10.128.17.29:6379 Total GPUs available: 56 [RLJobRunner] Pinggy check: url=False, token=False, needs_tunnel=False (agent=terminus-2, env=daytona) [RLJobRunner] No Pinggy tunnel needed, using local vLLM Running SkyRL: Python: /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/bin/python Entrypoint: examples.terminal_bench.entrypoints.main_tbench Args: 127 Hydra arguments Working dir: /e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train Using proxychains binary: /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 Executing command with srun: /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f $PROXYCHAINS_CONF_FILE /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/bin/python -m examples.terminal_bench.entrypoints.main_tbench +terminal_bench_config=terminal_bench trainer.strategy=fsdp2 trainer.algorithm.advantage_estimator=rloo_n trainer.algorithm.use_kl_loss=false trainer.algorithm.kl_loss_coef=0.0 trainer.algorithm.eps_clip_low=0.2 trainer.algorithm.eps_clip_high=0.05 trainer.algorithm.loss_reduction=seq_mean_token_sum_norm_global trainer.algorithm.use_tis=true trainer.algorithm.tis_imp_ratio_cap=2.0 trainer.epochs=2 trainer.max_steps=80 trainer.update_epochs_per_batch=1 trainer.train_batch_size=64 trainer.policy_mini_batch_size=64 trainer.eval_batch_size=64 trainer.micro_forward_batch_size_per_gpu=4 trainer.micro_train_batch_size_per_gpu=1 trainer.max_prompt_length=999999 trainer.eval_interval=999999 trainer.eval_before_train=false trainer.ckpt_interval=2 trainer.resume_mode=latest trainer.hf_save_interval=5 ++trainer.hf_hub_repo_id=laion/explore-tis-minp ++trainer.hf_hub_private=false ++trainer.hf_hub_revision=main ++trainer.enable_db_registration=false trainer.project_name=OpenThoughts-Agent trainer.log_level=INFO trainer.tracker_commit_each_step=true trainer.logger=console trainer.run_name=explore-tis-minp trainer.ckpt_path=/e/data1/datasets/playground/ot-baf/explore-tis-minp/explore-tis-minp/checkpoints trainer.export_path=/e/data1/datasets/playground/ot-baf/explore-tis-minp/explore-tis-minp/exports trainer.policy.optimizer_config.lr=8e-6 trainer.policy.optimizer_config.weight_decay=0.0 trainer.policy.optimizer_config.adam_betas=[0.9,0.999] trainer.policy.optimizer_config.max_grad_norm=0.9 trainer.policy.fsdp_config.cpu_offload=false trainer.policy.fsdp_config.reshard_after_forward=true trainer.policy.fsdp_config.fsdp_size=4 trainer.policy.model.path=/e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6 trainer.ref.fsdp_config.cpu_offload=false trainer.ref.fsdp_config.reshard_after_forward=true trainer.ref.fsdp_config.fsdp_size=4 trainer.placement.colocate_all=false trainer.placement.policy_num_nodes=2 trainer.placement.ref_num_nodes=2 trainer.placement.policy_num_gpus_per_node=4 trainer.placement.ref_num_gpus_per_node=4 trainer.fully_async.max_staleness_steps=16 trainer.fully_async.num_parallel_generation_workers=338 generator.backend=vllm generator.timeout_multiplier=1.0 generator.model_dtype=bfloat16 generator.inference_engine_tensor_parallel_size=1 generator.num_inference_engines=48 generator.n_samples_per_prompt=8 generator.eval_n_samples_per_prompt=8 generator.gpu_memory_utilization=0.75 generator.max_num_seqs=24 generator.max_num_batched_tokens=65536 generator.enable_prefix_caching=true generator.enable_chunked_prefill=true generator.run_engines_locally=true generator.weight_sync_backend=nccl generator.async_engine=true generator.batched=false generator.enable_http_endpoint=true generator.enable_ray_prometheus_stats=false generator.vllm_stats_interval=1 generator.append_eos_token_after_stop_str_in_multi_turn=true generator.max_turns=999999 generator.sampling_params.max_generate_length=4096 generator.sampling_params.temperature=0.7 generator.sampling_params.top_p=0.95 generator.sampling_params.top_k=20 ++generator.engine_init_kwargs.max_model_len=32768 ++generator.engine_init_kwargs.custom_chat_template_chat_completion_path=chat_templates/qwen3_thinking_acc.jinja2 ++generator.engine_init_kwargs.served_model_name=0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6 data.train_data=["/e/scratch/jureap59/feuer1/tasks/exp_rpt_pymethods2test-large"] data.val_data=[] +terminal_bench_config.trials_dir=/e/data1/datasets/playground/ot-baf/explore-tis-minp/explore-tis-minp/trace_jobs +terminal_bench_config.harbor.name=terminus-2 +terminal_bench_config.harbor.max_episodes=999999 +terminal_bench_config.harbor.enable_summarize=false +terminal_bench_config.harbor.store_all_messages=true +terminal_bench_config.harbor.trajectory_config.raw_content=true +terminal_bench_config.harbor.enable_episode_logging=false +terminal_bench_config.harbor.record_terminal_session=false +terminal_bench_config.harbor.enable_pane_logging=false +terminal_bench_config.harbor.strict_json_parser=true +terminal_bench_config.harbor.interleaved_thinking=true +terminal_bench_config.harbor.extra_body.chat_template_kwargs.enable_thinking=true +terminal_bench_config.harbor.override_timeout_sec=900 +terminal_bench_config.harbor.override_cpus=1 +terminal_bench_config.harbor.override_memory_mb=2048 +terminal_bench_config.harbor.override_storage_mb=2048 +terminal_bench_config.harbor.auto_snapshot=true +terminal_bench_config.harbor.verifier_override_timeout_sec=120 +terminal_bench_config.harbor.max_retries=3 +terminal_bench_config.harbor.min_wait_sec=60.0 +terminal_bench_config.harbor.max_wait_sec=600.0 +terminal_bench_config.harbor.wait_multiplier=2.0 +terminal_bench_config.harbor.exclude_exceptions=["VerifierTimeoutError","VerifierRuntimeError","RewardFileNotFoundError","RewardFileEmptyError","VerifierOutputParseError"] +terminal_bench_config.harbor.n_concurrent_trials=675 +terminal_bench_config.harbor.log_level=INFO +terminal_bench_config.harbor.enable_reward_shaping=false +terminal_bench_config.harbor.collect_rollout_details=true +terminal_bench_config.harbor.enable_error_classification=true +terminal_bench_config.harbor.mask_exceptions=["DaytonaError","EnvironmentStartTimeoutError","NetworkError","ConnectionError","RewardFileNotFoundError","RewardFileEmptyError","AgentEnvironmentTimeoutError","ContextLengthExceededError"] +terminal_bench_config.harbor.default_error_treatment=zero +terminal_bench_config.harbor.passthrough_exceptions=["AgentTimeoutError"] +terminal_bench_config.harbor.zero_exceptions=[] +terminal_bench_config.model_info.max_input_tokens=32000 +terminal_bench_config.model_info.max_output_tokens=4096 +terminal_bench_config.archiving.enabled=false +terminal_bench_config.trace_upload.enabled=true +terminal_bench_config.trace_upload.repo_org=DCAgent +terminal_bench_config.trace_upload.episodes=last +terminal_bench_config.trace_upload.dataset_type=SFT +terminal_bench_config.trace_upload.cleanup=true generator.sampling_params.temperature=1.0 generator.sampling_params.min_p=0.05 generator.sampling_params.top_k=-1 generator.sampling_params.top_p=1.0 [proxychains] config file found: /e/home/jusers/feuer1/jupiter/.proxychains/proxychains_736734.conf [proxychains] preloading /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/lib/libproxychains4.so [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e 2026-06-12 05:20:11.001 | WARNING | skyrl_train.utils.utils:validate_cfg:479 - `generator.sampling_params.logprobs` is `None` but `trainer.algorithm.use_tis` is `True`. Setting `logprobs` to `True`. 2026-06-12 05:20:11.170 | INFO | skyrl_train.utils.utils:prepare_runtime_environment:833 - Exporting wandb api key to ray runtime env 2026-06-12 05:20:11.170 | INFO | skyrl_train.utils.utils:prepare_runtime_environment:852 - Exporting RAY_ADDRESS to ray runtime env 2026-06-12 05:20:11.170 | INFO | skyrl_train.utils.utils:prepare_runtime_environment:877 - Exporting `NCCL_SOCKET_IFNAME` to ray runtime env: ib0 2026-06-12 05:20:11.170 | INFO | skyrl_train.utils.utils:prepare_runtime_environment:877 - Exporting `NCCL_SOCKET_FAMILY` to ray runtime env: AF_INET 2026-06-12 05:20:11.171 | INFO | skyrl_train.utils.utils:prepare_runtime_environment:877 - Exporting `NCCL_DEBUG` to ray runtime env: WARN 2026-06-12 05:20:11,171 INFO worker.py:1680 -- Using address 10.128.17.29:6379 set in the environment variable RAY_ADDRESS 2026-06-12 05:20:11,200 INFO worker.py:1821 -- Connecting to existing Ray cluster at address: 10.128.17.29:6379... 2026-06-12 05:20:11,207 WARNING node.py:1846 -- The object spilling config is specified from an unstable API - system config or environment variable. This is subject to change in the future. You can use the stable API - --object-spilling-directory in ray start or object_spilling_directory in ray.init() to specify the object spilling directory instead. If you need more advanced settings, please open a github issue with the Ray team. 2026-06-12 05:20:11,210 INFO worker.py:2007 -- Connected to Ray cluster. /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/ray/_private/worker.py:2046: FutureWarning: Tip: In future versions of Ray, Ray will no longer override accelerator visible devices env var if num_gpus=0 or num_gpus=None (default). To enable this behavior and turn off this error message, set RAY_ACCEL_ENV_VAR_OVERRIDE_ON_ZERO=0 warnings.warn( (raylet, ip=10.128.17.34) [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e (pid=471179) [2026-06-12 05:20:11,488 E 471179 472374] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14 (raylet, ip=10.128.17.34) 2026-06-12 05:20:12,333 WARNING node.py:1846 -- The object spilling config is specified from an unstable API - system config or environment variable. This is subject to change in the future. You can use the stable API - --object-spilling-directory in ray start or object_spilling_directory in ray.init() to specify the object spilling directory instead. If you need more advanced settings, please open a github issue with the Ray team. 2026-06-12 05:20:13.609 | INFO | skyrl_train.utils.ppo_utils:sync_registries:546 - Synced registries to ray actor (raylet) [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e [repeated 17x across cluster] (Ray deduplicates logs by default. Set RAY_DEDUP_LOGS=0 to disable log deduplication, or see https://docs.ray.io/en/master/ray-observability/user-guides/configure-logging.html#log-deduplication for more options.) (pid=471187) [2026-06-12 05:20:18,193 E 471187 472784] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14 [repeated 7x across cluster] (raylet) 2026-06-12 05:20:14,645 WARNING node.py:1846 -- The object spilling config is specified from an unstable API - system config or environment variable. This is subject to change in the future. You can use the stable API - --object-spilling-directory in ray start or object_spilling_directory in ray.init() to specify the object spilling directory instead. If you need more advanced settings, please open a github issue with the Ray team. [repeated 2x across cluster] (skyrl_entrypoint pid=487747) 2026-06-12 05:20:20.950 | INFO  | skyrl_train.entrypoints.main_base:_configure_log_level:214 - SkyRL log level set to: INFO (skyrl_entrypoint pid=487747) The tokenizer you are loading from '/e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue. (skyrl_entrypoint pid=487747) 2026-06-12 05:20:21.246 | INFO  | examples.terminal_bench.dataset:_load_data_files:40 - Loading data from: /e/scratch/jureap59/feuer1/tasks/exp_rpt_pymethods2test-large (skyrl_entrypoint pid=487747) 2026-06-12 05:20:23.315 | INFO  | examples.terminal_bench.dataset:_load_data_files:50 - Found 5000 valid task directories out of 5000 total directories (skyrl_entrypoint pid=487747) 2026-06-12 05:20:23.315 | INFO  | examples.terminal_bench.dataset:__init__:27 - TerminalBenchTaskDataset initialized with 5000 task paths (skyrl_entrypoint pid=487747) 2026-06-12 05:20:23.328 | INFO  | skyrl_train.entrypoints.main_base:_setup_trainer:407 - data: (skyrl_entrypoint pid=487747) train_data: (skyrl_entrypoint pid=487747) - /e/scratch/jureap59/feuer1/tasks/exp_rpt_pymethods2test-large (skyrl_entrypoint pid=487747) val_data: [] (skyrl_entrypoint pid=487747) trainer: (skyrl_entrypoint pid=487747) placement: (skyrl_entrypoint pid=487747) colocate_all: false (skyrl_entrypoint pid=487747) colocate_policy_ref: true (skyrl_entrypoint pid=487747) policy_num_nodes: 2 (skyrl_entrypoint pid=487747) policy_num_gpus_per_node: 4 (skyrl_entrypoint pid=487747) critic_num_nodes: 1 (skyrl_entrypoint pid=487747) critic_num_gpus_per_node: 4 (skyrl_entrypoint pid=487747) ref_num_nodes: 2 (skyrl_entrypoint pid=487747) ref_num_gpus_per_node: 4 (skyrl_entrypoint pid=487747) policy_strict_spread_pg: false (skyrl_entrypoint pid=487747) policy_per_gpu_bundles: false (skyrl_entrypoint pid=487747) policy_force_cvd_mask: false (skyrl_entrypoint pid=487747) sequence_parallel_backend: ulysses (skyrl_entrypoint pid=487747) strategy: fsdp2 (skyrl_entrypoint pid=487747) policy: (skyrl_entrypoint pid=487747) model: (skyrl_entrypoint pid=487747) path: /e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6 (skyrl_entrypoint pid=487747) lora: (skyrl_entrypoint pid=487747) rank: 0 (skyrl_entrypoint pid=487747) alpha: 16 (skyrl_entrypoint pid=487747) dropout: 0 (skyrl_entrypoint pid=487747) lora_sync_path: /tmp/skyrl_lora_sync (skyrl_entrypoint pid=487747) target_modules: all-linear (skyrl_entrypoint pid=487747) exclude_modules: null (skyrl_entrypoint pid=487747) deepspeed_config: ${deepspeed_config.train} (skyrl_entrypoint pid=487747) optimizer_config: (skyrl_entrypoint pid=487747) optimizer: AdamW (skyrl_entrypoint pid=487747) lr: 8.0e-06 (skyrl_entrypoint pid=487747) adam_betas: (skyrl_entrypoint pid=487747) - 0.9 (skyrl_entrypoint pid=487747) - 0.999 (skyrl_entrypoint pid=487747) weight_decay: 0.0 (skyrl_entrypoint pid=487747) max_grad_norm: 0.9 (skyrl_entrypoint pid=487747) offload_after_step: true (skyrl_entrypoint pid=487747) num_warmup_steps: 0 (skyrl_entrypoint pid=487747) scheduler: constant_with_warmup (skyrl_entrypoint pid=487747) optimizer_kwargs: {} (skyrl_entrypoint pid=487747) fsdp_config: (skyrl_entrypoint pid=487747) cpu_offload: false (skyrl_entrypoint pid=487747) reshard_after_forward: true (skyrl_entrypoint pid=487747) fsdp_size: 4 (skyrl_entrypoint pid=487747) expert_model_parallel_size: 1 (skyrl_entrypoint pid=487747) expert_tensor_parallel_size: 1 (skyrl_entrypoint pid=487747) moe_token_dispatcher_type: alltoall (skyrl_entrypoint pid=487747) moe_router_replay: false (skyrl_entrypoint pid=487747) moe_grouped_gemm: false (skyrl_entrypoint pid=487747) ep_comm_backend: torch (skyrl_entrypoint pid=487747) deepep_num_sms: 20 (skyrl_entrypoint pid=487747) deepep_token_chunk_size: null (skyrl_entrypoint pid=487747) sequence_parallel_size: 1 (skyrl_entrypoint pid=487747) use_torch_compile: false (skyrl_entrypoint pid=487747) record_memory: false (skyrl_entrypoint pid=487747) megatron_config: (skyrl_entrypoint pid=487747) tensor_model_parallel_size: 1 (skyrl_entrypoint pid=487747) pipeline_model_parallel_size: 1 (skyrl_entrypoint pid=487747) context_parallel_size: 1 (skyrl_entrypoint pid=487747) expert_model_parallel_size: 1 (skyrl_entrypoint pid=487747) expert_tensor_parallel_size: null (skyrl_entrypoint pid=487747) ddp_config: (skyrl_entrypoint pid=487747) grad_reduce_in_fp32: true (skyrl_entrypoint pid=487747) overlap_grad_reduce: false (skyrl_entrypoint pid=487747) overlap_param_gather: false (skyrl_entrypoint pid=487747) average_in_collective: true (skyrl_entrypoint pid=487747) model_config_kwargs: {} (skyrl_entrypoint pid=487747) torch_profiler_config: (skyrl_entrypoint pid=487747) enable: false (skyrl_entrypoint pid=487747) ranks: [] (skyrl_entrypoint pid=487747) save_path: null (skyrl_entrypoint pid=487747) optimizer_config_kwargs: (skyrl_entrypoint pid=487747) overlap_cpu_optimizer_d2h_h2d: false (skyrl_entrypoint pid=487747) use_precision_aware_optimizer: false (skyrl_entrypoint pid=487747) optimizer_cpu_offload: false (skyrl_entrypoint pid=487747) optimizer_offload_fraction: 0.0 (skyrl_entrypoint pid=487747) transformer_config_kwargs: (skyrl_entrypoint pid=487747) recompute_granularity: full (skyrl_entrypoint pid=487747) recompute_modules: (skyrl_entrypoint pid=487747) - core_attn (skyrl_entrypoint pid=487747) recompute_method: uniform (skyrl_entrypoint pid=487747) recompute_num_layers: 1 (skyrl_entrypoint pid=487747) empty_cuda_cache: true (skyrl_entrypoint pid=487747) ref: (skyrl_entrypoint pid=487747) model: (skyrl_entrypoint pid=487747) path: ${trainer.policy.model.path} (skyrl_entrypoint pid=487747) sequence_parallel_size: 1 (skyrl_entrypoint pid=487747) deepspeed_config: ${deepspeed_config.eval} (skyrl_entrypoint pid=487747) fsdp_config: (skyrl_entrypoint pid=487747) cpu_offload: false (skyrl_entrypoint pid=487747) reshard_after_forward: true (skyrl_entrypoint pid=487747) fsdp_size: 4 (skyrl_entrypoint pid=487747) expert_model_parallel_size: 1 (skyrl_entrypoint pid=487747) expert_tensor_parallel_size: 1 (skyrl_entrypoint pid=487747) moe_token_dispatcher_type: alltoall (skyrl_entrypoint pid=487747) moe_router_replay: false (skyrl_entrypoint pid=487747) moe_grouped_gemm: false (skyrl_entrypoint pid=487747) ep_comm_backend: torch (skyrl_entrypoint pid=487747) deepep_num_sms: 20 (skyrl_entrypoint pid=487747) deepep_token_chunk_size: null (skyrl_entrypoint pid=487747) megatron_config: (skyrl_entrypoint pid=487747) tensor_model_parallel_size: 1 (skyrl_entrypoint pid=487747) pipeline_model_parallel_size: 1 (skyrl_entrypoint pid=487747) context_parallel_size: 1 (skyrl_entrypoint pid=487747) expert_model_parallel_size: 1 (skyrl_entrypoint pid=487747) expert_tensor_parallel_size: 1 (skyrl_entrypoint pid=487747) model_config_kwargs: {} (skyrl_entrypoint pid=487747) transformer_config_kwargs: {} (skyrl_entrypoint pid=487747) critic: (skyrl_entrypoint pid=487747) model: (skyrl_entrypoint pid=487747) path: null (skyrl_entrypoint pid=487747) lora: (skyrl_entrypoint pid=487747) rank: 0 (skyrl_entrypoint pid=487747) alpha: 16 (skyrl_entrypoint pid=487747) dropout: 0 (skyrl_entrypoint pid=487747) target_modules: all-linear (skyrl_entrypoint pid=487747) exclude_modules: null (skyrl_entrypoint pid=487747) deepspeed_config: ${deepspeed_config.train} (skyrl_entrypoint pid=487747) optimizer_config: (skyrl_entrypoint pid=487747) optimizer: AdamW (skyrl_entrypoint pid=487747) lr: 5.0e-06 (skyrl_entrypoint pid=487747) adam_betas: (skyrl_entrypoint pid=487747) - 0.9 (skyrl_entrypoint pid=487747) - 0.999 (skyrl_entrypoint pid=487747) weight_decay: 0.01 (skyrl_entrypoint pid=487747) max_grad_norm: 1.0 (skyrl_entrypoint pid=487747) offload_after_step: true (skyrl_entrypoint pid=487747) num_warmup_steps: 0 (skyrl_entrypoint pid=487747) scheduler: constant_with_warmup (skyrl_entrypoint pid=487747) optimizer_kwargs: {} (skyrl_entrypoint pid=487747) fsdp_config: (skyrl_entrypoint pid=487747) cpu_offload: false (skyrl_entrypoint pid=487747) reshard_after_forward: true (skyrl_entrypoint pid=487747) fsdp_size: -1 (skyrl_entrypoint pid=487747) expert_model_parallel_size: 1 (skyrl_entrypoint pid=487747) expert_tensor_parallel_size: 1 (skyrl_entrypoint pid=487747) moe_token_dispatcher_type: alltoall (skyrl_entrypoint pid=487747) moe_router_replay: false (skyrl_entrypoint pid=487747) moe_grouped_gemm: false (skyrl_entrypoint pid=487747) ep_comm_backend: torch (skyrl_entrypoint pid=487747) deepep_num_sms: 20 (skyrl_entrypoint pid=487747) deepep_token_chunk_size: null (skyrl_entrypoint pid=487747) sequence_parallel_size: 1 (skyrl_entrypoint pid=487747) algorithm: (skyrl_entrypoint pid=487747) advantage_estimator: rloo_n (skyrl_entrypoint pid=487747) kl_ctrl: (skyrl_entrypoint pid=487747) type: fixed (skyrl_entrypoint pid=487747) kl_target: 0.1 (skyrl_entrypoint pid=487747) horizon: 10000 (skyrl_entrypoint pid=487747) kl_estimator_type: k3 (skyrl_entrypoint pid=487747) use_kl_estimator_k3: false (skyrl_entrypoint pid=487747) use_abs_kl: false (skyrl_entrypoint pid=487747) use_kl_in_reward: false (skyrl_entrypoint pid=487747) use_kl_loss: false (skyrl_entrypoint pid=487747) kl_loss_coef: 0.0 (skyrl_entrypoint pid=487747) use_entropy_loss: false (skyrl_entrypoint pid=487747) entropy_loss_coef: 0.01 (skyrl_entrypoint pid=487747) advantage_batch_normalize: false (skyrl_entrypoint pid=487747) value_head_prefix: value_head (skyrl_entrypoint pid=487747) policy_loss_type: regular (skyrl_entrypoint pid=487747) loss_reduction: seq_mean_token_sum_norm_global (skyrl_entrypoint pid=487747) global_loss_denom: null (skyrl_entrypoint pid=487747) grpo_norm_by_std: true (skyrl_entrypoint pid=487747) rloo_n_min_group_size: 4 (skyrl_entrypoint pid=487747) rloo_n_filter_zero_reward_groups: true (skyrl_entrypoint pid=487747) lambd: 1.0 (skyrl_entrypoint pid=487747) gamma: 1.0 (skyrl_entrypoint pid=487747) eps_clip_low: 0.2 (skyrl_entrypoint pid=487747) eps_clip_high: 0.05 (skyrl_entrypoint pid=487747) clip_ratio_c: 3.0 (skyrl_entrypoint pid=487747) tis_imp_ratio_cap: 2.0 (skyrl_entrypoint pid=487747) use_tis: true (skyrl_entrypoint pid=487747) sapo: (skyrl_entrypoint pid=487747) tau_pos: 1.0 (skyrl_entrypoint pid=487747) tau_neg: 1.05 (skyrl_entrypoint pid=487747) value_clip: 0.2 (skyrl_entrypoint pid=487747) dynamic_sampling: (skyrl_entrypoint pid=487747) type: null (skyrl_entrypoint pid=487747) max_sample_batches: 30 (skyrl_entrypoint pid=487747) min_replace_ratio: 0.3 (skyrl_entrypoint pid=487747) clip_cov: (skyrl_entrypoint pid=487747) clip_ratio: 0.0002 (skyrl_entrypoint pid=487747) clip_cov_lb: 1.0 (skyrl_entrypoint pid=487747) clip_cov_ub: 5.0 (skyrl_entrypoint pid=487747) kl_cov: (skyrl_entrypoint pid=487747) kl_cov_frac: 0.2 (skyrl_entrypoint pid=487747) ppo_kl_coef: 1.0 (skyrl_entrypoint pid=487747) cispo: (skyrl_entrypoint pid=487747) cispo_eps_clip_low: 0 (skyrl_entrypoint pid=487747) cispo_eps_clip_high: 5 (skyrl_entrypoint pid=487747) z_clip: (skyrl_entrypoint pid=487747) enabled: false (skyrl_entrypoint pid=487747) alpha: 0.97 (skyrl_entrypoint pid=487747) z_thresh: 2.5 (skyrl_entrypoint pid=487747) warmup_steps: 25 (skyrl_entrypoint pid=487747) mode: zscore (skyrl_entrypoint pid=487747) clip_option: adaptive_scaling (skyrl_entrypoint pid=487747) clip_factor: 1.0 (skyrl_entrypoint pid=487747) skip_update_on_spike: false (skyrl_entrypoint pid=487747) stale_clip: (skyrl_entrypoint pid=487747) enabled: false (skyrl_entrypoint pid=487747) alpha: 0.3 (skyrl_entrypoint pid=487747) entropy_threshold: 0.15 (skyrl_entrypoint pid=487747) entropy_window: 10 (skyrl_entrypoint pid=487747) min_lr_scale: 0.1 (skyrl_entrypoint pid=487747) max_seq_len: 1004095 (skyrl_entrypoint pid=487747) fully_async: (skyrl_entrypoint pid=487747) max_staleness_steps: 16 (skyrl_entrypoint pid=487747) num_parallel_generation_workers: 338 (skyrl_entrypoint pid=487747) gradient_checkpointing: true (skyrl_entrypoint pid=487747) gradient_checkpointing_use_reentrant: false (skyrl_entrypoint pid=487747) seed: 42 (skyrl_entrypoint pid=487747) resume_mode: latest (skyrl_entrypoint pid=487747) resume_path: null (skyrl_entrypoint pid=487747) ckpt_path: /e/data1/datasets/playground/ot-baf/explore-tis-minp/explore-tis-minp/checkpoints (skyrl_entrypoint pid=487747) max_ckpts_to_keep: -1 (skyrl_entrypoint pid=487747) ckpt_interval: 2 (skyrl_entrypoint pid=487747) hf_save_interval: 5 (skyrl_entrypoint pid=487747) hf_upload_mode: latest (skyrl_entrypoint pid=487747) export_path: /e/data1/datasets/playground/ot-baf/explore-tis-minp/explore-tis-minp/exports (skyrl_entrypoint pid=487747) bf16: true (skyrl_entrypoint pid=487747) epochs: 2 (skyrl_entrypoint pid=487747) max_steps: 80 (skyrl_entrypoint pid=487747) update_epochs_per_batch: 1 (skyrl_entrypoint pid=487747) train_batch_size: 64 (skyrl_entrypoint pid=487747) policy_mini_batch_size: 64 (skyrl_entrypoint pid=487747) critic_mini_batch_size: 256 (skyrl_entrypoint pid=487747) micro_train_batch_size_per_gpu: 1 (skyrl_entrypoint pid=487747) micro_forward_batch_size_per_gpu: 4 (skyrl_entrypoint pid=487747) update_ref_every_epoch: false (skyrl_entrypoint pid=487747) use_sample_packing: true (skyrl_entrypoint pid=487747) eval_batch_size: 64 (skyrl_entrypoint pid=487747) eval_before_train: false (skyrl_entrypoint pid=487747) eval_interval: 999999 (skyrl_entrypoint pid=487747) max_prompt_length: 999999 (skyrl_entrypoint pid=487747) flash_attn: true (skyrl_entrypoint pid=487747) disable_fast_tokenizer: false (skyrl_entrypoint pid=487747) target_modules: null (skyrl_entrypoint pid=487747) exclude_modules: null (skyrl_entrypoint pid=487747) project_name: OpenThoughts-Agent (skyrl_entrypoint pid=487747) run_name: explore-tis-minp (skyrl_entrypoint pid=487747) logger: console (skyrl_entrypoint pid=487747) tracker_commit_each_step: true (skyrl_entrypoint pid=487747) dump_data_batch: false (skyrl_entrypoint pid=487747) dump_eval_results: true (skyrl_entrypoint pid=487747) log_level: INFO (skyrl_entrypoint pid=487747) rope_scaling: null (skyrl_entrypoint pid=487747) rope_theta: null (skyrl_entrypoint pid=487747) step_wise_training: false (skyrl_entrypoint pid=487747) hf_hub_repo_id: laion/explore-tis-minp (skyrl_entrypoint pid=487747) hf_hub_private: false (skyrl_entrypoint pid=487747) hf_hub_revision: main (skyrl_entrypoint pid=487747) enable_db_registration: false (skyrl_entrypoint pid=487747) generator: (skyrl_entrypoint pid=487747) model_name: ${trainer.policy.model.path} (skyrl_entrypoint pid=487747) model_dtype: bfloat16 (skyrl_entrypoint pid=487747) timeout_multiplier: 1.0 (skyrl_entrypoint pid=487747) run_engines_locally: true (skyrl_entrypoint pid=487747) num_inference_engines: 48 (skyrl_entrypoint pid=487747) backend: vllm (skyrl_entrypoint pid=487747) weight_sync_backend: nccl (skyrl_entrypoint pid=487747) fuse_weights: false (skyrl_entrypoint pid=487747) weight_transfer_threshold_cuda_ipc_GB: 1.0 (skyrl_entrypoint pid=487747) inference_engine_tensor_parallel_size: 1 (skyrl_entrypoint pid=487747) inference_engine_pipeline_parallel_size: 1 (skyrl_entrypoint pid=487747) inference_engine_expert_parallel_size: 1 (skyrl_entrypoint pid=487747) inference_engine_data_parallel_size: 1 (skyrl_entrypoint pid=487747) inference_engine_mp_backend: false (skyrl_entrypoint pid=487747) n_samples_per_prompt: 8 (skyrl_entrypoint pid=487747) async_engine: true (skyrl_entrypoint pid=487747) batched: false (skyrl_entrypoint pid=487747) max_input_length: ${trainer.max_prompt_length} (skyrl_entrypoint pid=487747) vllm_v1_disable_multiproc: true (skyrl_entrypoint pid=487747) enable_prefix_caching: true (skyrl_entrypoint pid=487747) enable_chunked_prefill: true (skyrl_entrypoint pid=487747) max_num_batched_tokens: 65536 (skyrl_entrypoint pid=487747) enforce_eager: true (skyrl_entrypoint pid=487747) fully_sharded_loras: false (skyrl_entrypoint pid=487747) enable_ray_prometheus_stats: false (skyrl_entrypoint pid=487747) vllm_stats_interval: 1 (skyrl_entrypoint pid=487747) gpu_memory_utilization: 0.75 (skyrl_entrypoint pid=487747) max_num_seqs: 24 (skyrl_entrypoint pid=487747) remote_inference_engine_urls: (skyrl_entrypoint pid=487747) - 127.0.0.1:8001 (skyrl_entrypoint pid=487747) enable_http_endpoint: true (skyrl_entrypoint pid=487747) http_endpoint_host: 127.0.0.1 (skyrl_entrypoint pid=487747) http_endpoint_port: 8000 (skyrl_entrypoint pid=487747) max_turns: 999999 (skyrl_entrypoint pid=487747) chat_template: (skyrl_entrypoint pid=487747) source: name (skyrl_entrypoint pid=487747) name_or_path: null (skyrl_entrypoint pid=487747) chat_template_kwargs: {} (skyrl_entrypoint pid=487747) engine_init_kwargs: (skyrl_entrypoint pid=487747) max_model_len: 32768 (skyrl_entrypoint pid=487747) custom_chat_template_chat_completion_path: chat_templates/qwen3_thinking_acc.jinja2 (skyrl_entrypoint pid=487747) served_model_name: 0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6 (skyrl_entrypoint pid=487747) override_existing_update_group: disable (skyrl_entrypoint pid=487747) sampling_params: (skyrl_entrypoint pid=487747) max_generate_length: 4096 (skyrl_entrypoint pid=487747) repetition_penalty: 1.0 (skyrl_entrypoint pid=487747) temperature: 1.0 (skyrl_entrypoint pid=487747) top_p: 1.0 (skyrl_entrypoint pid=487747) min_p: 0.05 (skyrl_entrypoint pid=487747) top_k: -1 (skyrl_entrypoint pid=487747) logprobs: 0 (skyrl_entrypoint pid=487747) stop: null (skyrl_entrypoint pid=487747) use_conversation_multi_turn: true (skyrl_entrypoint pid=487747) append_eos_token_after_stop_str_in_multi_turn: true (skyrl_entrypoint pid=487747) eval_sampling_params: (skyrl_entrypoint pid=487747) max_generate_length: ${generator.sampling_params.max_generate_length} (skyrl_entrypoint pid=487747) repetition_penalty: 1.0 (skyrl_entrypoint pid=487747) temperature: 0.0 (skyrl_entrypoint pid=487747) top_p: 1.0 (skyrl_entrypoint pid=487747) min_p: 0.0 (skyrl_entrypoint pid=487747) top_k: -1 (skyrl_entrypoint pid=487747) logprobs: null (skyrl_entrypoint pid=487747) stop: null (skyrl_entrypoint pid=487747) eval_n_samples_per_prompt: 8 (skyrl_entrypoint pid=487747) zero_reward_on_non_stop: false (skyrl_entrypoint pid=487747) apply_overlong_filtering: false (skyrl_entrypoint pid=487747) rope_scaling: ${trainer.rope_scaling} (skyrl_entrypoint pid=487747) rope_theta: ${trainer.rope_theta} (skyrl_entrypoint pid=487747) teacher: (skyrl_entrypoint pid=487747) model_path: null (skyrl_entrypoint pid=487747) top_k_logprobs: 256 (skyrl_entrypoint pid=487747) num_inference_engines: 1 (skyrl_entrypoint pid=487747) inference_engine_tensor_parallel_size: 1 (skyrl_entrypoint pid=487747) inference_engine_pipeline_parallel_size: 1 (skyrl_entrypoint pid=487747) gpu_memory_utilization: 0.9 (skyrl_entrypoint pid=487747) enforce_eager: false (skyrl_entrypoint pid=487747) backend: vllm (skyrl_entrypoint pid=487747) engine_init_kwargs: {} (skyrl_entrypoint pid=487747) environment: (skyrl_entrypoint pid=487747) env_class: gsm8k (skyrl_entrypoint pid=487747) skyrl_gym: (skyrl_entrypoint pid=487747) max_env_workers: 32 (skyrl_entrypoint pid=487747) text2sql: (skyrl_entrypoint pid=487747) db_path: /home/ray/default/sql_data (skyrl_entrypoint pid=487747) llm_as_a_judge: (skyrl_entrypoint pid=487747) model: gpt-4o-mini (skyrl_entrypoint pid=487747) base_url: null (skyrl_entrypoint pid=487747) search: (skyrl_entrypoint pid=487747) log_requests: false (skyrl_entrypoint pid=487747) search_url: http://127.0.0.1:8000/retrieve (skyrl_entrypoint pid=487747) topk: 3 (skyrl_entrypoint pid=487747) timeout: 30 (skyrl_entrypoint pid=487747) rollout: (skyrl_entrypoint pid=487747) fanout: (skyrl_entrypoint pid=487747) enabled: true (skyrl_entrypoint pid=487747) num_coordinators: 4 (skyrl_entrypoint pid=487747) cpus_per_coordinator: 8 (skyrl_entrypoint pid=487747) deepspeed_config: (skyrl_entrypoint pid=487747) train: (skyrl_entrypoint pid=487747) zero_optimization: (skyrl_entrypoint pid=487747) stage: 3 (skyrl_entrypoint pid=487747) offload_param: (skyrl_entrypoint pid=487747) device: none (skyrl_entrypoint pid=487747) offload_optimizer: (skyrl_entrypoint pid=487747) device: none (skyrl_entrypoint pid=487747) pin_memory: true (skyrl_entrypoint pid=487747) sub_group_size: auto (skyrl_entrypoint pid=487747) reduce_bucket_size: auto (skyrl_entrypoint pid=487747) stage3_param_persistence_threshold: auto (skyrl_entrypoint pid=487747) stage3_prefetch_bucket_size: auto (skyrl_entrypoint pid=487747) stage3_max_live_parameters: auto (skyrl_entrypoint pid=487747) stage3_max_reuse_distance: auto (skyrl_entrypoint pid=487747) round_robin_gradients: true (skyrl_entrypoint pid=487747) zero_hpz_partition_size: 1 (skyrl_entrypoint pid=487747) zero_quantized_weights: false (skyrl_entrypoint pid=487747) zero_quantized_gradients: false (skyrl_entrypoint pid=487747) torch_autocast: (skyrl_entrypoint pid=487747) enabled: true (skyrl_entrypoint pid=487747) dtype: bfloat16 (skyrl_entrypoint pid=487747) disable_trace_cache: false (skyrl_entrypoint pid=487747) data_types: (skyrl_entrypoint pid=487747) grad_accum_dtype: fp32 (skyrl_entrypoint pid=487747) gradient_clipping: 1.0 (skyrl_entrypoint pid=487747) wall_clock_breakdown: false (skyrl_entrypoint pid=487747) prescale_gradient: false (skyrl_entrypoint pid=487747) eval: (skyrl_entrypoint pid=487747) zero_optimization: (skyrl_entrypoint pid=487747) stage: 3 (skyrl_entrypoint pid=487747) stage3_param_persistence_threshold: auto (skyrl_entrypoint pid=487747) offload_param: (skyrl_entrypoint pid=487747) device: cpu (skyrl_entrypoint pid=487747) pin_memory: true (skyrl_entrypoint pid=487747) torch_autocast: (skyrl_entrypoint pid=487747) enabled: true (skyrl_entrypoint pid=487747) dtype: bfloat16 (skyrl_entrypoint pid=487747) gradient_clipping: 1.0 (skyrl_entrypoint pid=487747) prescale_gradient: false (skyrl_entrypoint pid=487747) wall_clock_breakdown: false (skyrl_entrypoint pid=487747) terminal_bench_config: (skyrl_entrypoint pid=487747) trials_dir: /e/data1/datasets/playground/ot-baf/explore-tis-minp/explore-tis-minp/trace_jobs (skyrl_entrypoint pid=487747) harbor: (skyrl_entrypoint pid=487747) name: terminus-2 (skyrl_entrypoint pid=487747) max_episodes: 999999 (skyrl_entrypoint pid=487747) enable_summarize: false (skyrl_entrypoint pid=487747) store_all_messages: true (skyrl_entrypoint pid=487747) trajectory_config: (skyrl_entrypoint pid=487747) raw_content: true (skyrl_entrypoint pid=487747) enable_episode_logging: false (skyrl_entrypoint pid=487747) record_terminal_session: false (skyrl_entrypoint pid=487747) enable_pane_logging: false (skyrl_entrypoint pid=487747) strict_json_parser: true (skyrl_entrypoint pid=487747) interleaved_thinking: true (skyrl_entrypoint pid=487747) extra_body: (skyrl_entrypoint pid=487747) chat_template_kwargs: (skyrl_entrypoint pid=487747) enable_thinking: true (skyrl_entrypoint pid=487747) override_timeout_sec: 900 (skyrl_entrypoint pid=487747) override_cpus: 1 (skyrl_entrypoint pid=487747) override_memory_mb: 2048 (skyrl_entrypoint pid=487747) override_storage_mb: 2048 (skyrl_entrypoint pid=487747) auto_snapshot: true (skyrl_entrypoint pid=487747) verifier_override_timeout_sec: 120 (skyrl_entrypoint pid=487747) max_retries: 3 (skyrl_entrypoint pid=487747) min_wait_sec: 60.0 (skyrl_entrypoint pid=487747) max_wait_sec: 600.0 (skyrl_entrypoint pid=487747) wait_multiplier: 2.0 (skyrl_entrypoint pid=487747) exclude_exceptions: (skyrl_entrypoint pid=487747) - VerifierTimeoutError (skyrl_entrypoint pid=487747) - VerifierRuntimeError (skyrl_entrypoint pid=487747) - RewardFileNotFoundError (skyrl_entrypoint pid=487747) - RewardFileEmptyError (skyrl_entrypoint pid=487747) - VerifierOutputParseError (skyrl_entrypoint pid=487747) n_concurrent_trials: 675 (skyrl_entrypoint pid=487747) log_level: INFO (skyrl_entrypoint pid=487747) enable_reward_shaping: false (skyrl_entrypoint pid=487747) collect_rollout_details: true (skyrl_entrypoint pid=487747) enable_error_classification: true (skyrl_entrypoint pid=487747) mask_exceptions: (skyrl_entrypoint pid=487747) - DaytonaError (skyrl_entrypoint pid=487747) - EnvironmentStartTimeoutError (skyrl_entrypoint pid=487747) - NetworkError (skyrl_entrypoint pid=487747) - ConnectionError (skyrl_entrypoint pid=487747) - RewardFileNotFoundError (skyrl_entrypoint pid=487747) - RewardFileEmptyError (skyrl_entrypoint pid=487747) - AgentEnvironmentTimeoutError (skyrl_entrypoint pid=487747) - ContextLengthExceededError (skyrl_entrypoint pid=487747) default_error_treatment: zero (skyrl_entrypoint pid=487747) passthrough_exceptions: (skyrl_entrypoint pid=487747) - AgentTimeoutError (skyrl_entrypoint pid=487747) zero_exceptions: [] (skyrl_entrypoint pid=487747) model_info: (skyrl_entrypoint pid=487747) max_input_tokens: 32000 (skyrl_entrypoint pid=487747) max_output_tokens: 4096 (skyrl_entrypoint pid=487747) archiving: (skyrl_entrypoint pid=487747) enabled: false (skyrl_entrypoint pid=487747) trace_upload: (skyrl_entrypoint pid=487747) enabled: true (skyrl_entrypoint pid=487747) repo_org: DCAgent (skyrl_entrypoint pid=487747) episodes: last (skyrl_entrypoint pid=487747) dataset_type: SFT (skyrl_entrypoint pid=487747) cleanup: true (skyrl_entrypoint pid=487747)  (pid=471460) [2026-06-12 05:20:21,782 E 471460 487242] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14 [repeated 268x across cluster] (skyrl_entrypoint pid=487747) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash: (skyrl_entrypoint pid=487747) No module named 'vllm._version' (skyrl_entrypoint pid=487747) from .version import __version__, __version_tuple__ # isort:skip (skyrl_entrypoint pid=487747) W0612 05:20:28.196000 487747 envs/rl/lib/python3.12/site-packages/torch/utils/cpp_extension.py:117] No CUDA runtime is found, using CUDA_HOME='/e/software/default/stages/2026/software/CUDA/13' (raylet) [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e (raylet) 2026-06-12 05:20:31,137 WARNING node.py:1846 -- The object spilling config is specified from an unstable API - system config or environment variable. This is subject to change in the future. You can use the stable API - --object-spilling-directory in ray start or object_spilling_directory in ray.init() to specify the object spilling directory instead. If you need more advanced settings, please open a github issue with the Ray team. (raylet, ip=10.128.17.223) [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e [repeated 41x across cluster] (raylet, ip=10.128.17.223) 2026-06-12 05:20:35,364 WARNING node.py:1846 -- The object spilling config is specified from an unstable API - system config or environment variable. This is subject to change in the future. You can use the stable API - --object-spilling-directory in ray start or object_spilling_directory in ray.init() to specify the object spilling directory instead. If you need more advanced settings, please open a github issue with the Ray team. [repeated 6x across cluster] (pid=308714, ip=10.128.17.55) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash: (pid=308714, ip=10.128.17.55) No module named 'vllm._version' (pid=308714, ip=10.128.17.55) from .version import __version__, __version_tuple__ # isort:skip (raylet, ip=10.128.17.55) [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e [repeated 98x across cluster] [2026-06-12 05:20:41,365 E 487344 487721] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14 (raylet, ip=10.128.17.55) 2026-06-12 05:20:40,725 WARNING node.py:1846 -- The object spilling config is specified from an unstable API - system config or environment variable. This is subject to change in the future. You can use the stable API - --object-spilling-directory in ray start or object_spilling_directory in ray.init() to specify the object spilling directory instead. If you need more advanced settings, please open a github issue with the Ray team. [repeated 18x across cluster] (RegistryActor pid=631639, ip=10.128.17.34) [2026-06-12 05:20:42,366 E 631639 631679] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14 (pid=379669, ip=10.128.34.2) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash: [repeated 12x across cluster] (pid=379669, ip=10.128.34.2) No module named 'vllm._version' [repeated 12x across cluster] (pid=379669, ip=10.128.34.2) from .version import __version__, __version_tuple__ # isort:skip [repeated 12x across cluster] (raylet, ip=10.128.34.5) [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e [repeated 70x across cluster] (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) 2026-06-12 05:20:47.290 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:152 - setup_envvars_for_vllm: distributed_executor_backend=uni, SKYRL_ENABLE_NUMA_AFFINITY=1, CUDA_VISIBLE_DEVICES=0, VLLM_ENABLE_V1_MULTIPROCESSING= (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) 2026-06-12 05:20:47.291 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:167 - setup_envvars_for_vllm: numa_enabled=True (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) 2026-06-12 05:20:47.291 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:170 - setup_envvars_for_vllm: set VLLM_ENABLE_V1_MULTIPROCESSING=0 for NUMA affinity (raylet, ip=10.128.34.8) 2026-06-12 05:20:46,482 WARNING node.py:1846 -- The object spilling config is specified from an unstable API - system config or environment variable. This is subject to change in the future. You can use the stable API - --object-spilling-directory in ray start or object_spilling_directory in ray.init() to specify the object spilling directory instead. If you need more advanced settings, please open a github issue with the Ray team. [repeated 14x across cluster] (skyrl_entrypoint pid=487747) [2026-06-12 05:20:44,666 E 487747 487790] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14 [repeated 2x across cluster] (pid=592922, ip=10.128.34.1) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash: [repeated 12x across cluster] (pid=592922, ip=10.128.34.1) No module named 'vllm._version' [repeated 12x across cluster] (pid=592922, ip=10.128.34.1) from .version import __version__, __version_tuple__ # isort:skip [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) 2026-06-12 05:20:47.917 | INFO | skyrl_train.utils.numa:set_numa_affinity_for_gpu:340 - NUMA affinity: bound process to CPUs 0-71 (NUMA node 0) for GPU 0 (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) 2026-06-12 05:20:47.939 | DEBUG | skyrl_train.utils.numa:_set_membind_via_libnuma:375 - NUMA affinity: set memory preferred to NUMA node 0 (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) 2026-06-12 05:20:47.939 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:618 - BaseVLLMInferenceEngine: vllm_v1_disable_multiproc=True, vllm.__version__=dev, VLLM_ENABLE_V1_MULTIPROCESSING=0 (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) 2026-06-12 05:20:47.939 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:626 - BaseVLLMInferenceEngine: set VLLM_ENABLE_V1_MULTIPROCESSING=0 (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) 2026-06-12 05:20:47.939 | WARNING | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1192 - OpenAI API sampling params overridden: temperature=1.0, top_p=1.0, top_k=-1 (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) 2026-06-12 05:20:47.958 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1244 - Engine startup stagger: sleeping 2.90s (attempt 1/5) to avoid port collisions (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored. (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored. (raylet, ip=10.128.17.63) [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e [repeated 108x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) The tokenizer you are loading from '/e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue. (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) 2026-06-12 05:20:51.410 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:152 - setup_envvars_for_vllm: distributed_executor_backend=uni, SKYRL_ENABLE_NUMA_AFFINITY=1, CUDA_VISIBLE_DEVICES=3, VLLM_ENABLE_V1_MULTIPROCESSING= [repeated 15x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) 2026-06-12 05:20:51.411 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:167 - setup_envvars_for_vllm: numa_enabled=True [repeated 15x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) 2026-06-12 05:20:51.411 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:170 - setup_envvars_for_vllm: set VLLM_ENABLE_V1_MULTIPROCESSING=0 for NUMA affinity [repeated 15x across cluster] (raylet, ip=10.128.34.6) 2026-06-12 05:20:51,846 WARNING node.py:1846 -- The object spilling config is specified from an unstable API - system config or environment variable. This is subject to change in the future. You can use the stable API - --object-spilling-directory in ray start or object_spilling_directory in ray.init() to specify the object spilling directory instead. If you need more advanced settings, please open a github issue with the Ray team. [repeated 19x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash: [repeated 22x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) No module named 'vllm._version' [repeated 22x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) from .version import __version__, __version_tuple__ # isort:skip [repeated 22x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) 2026-06-12 05:20:52.568 | INFO | skyrl_train.utils.numa:set_numa_affinity_for_gpu:340 - NUMA affinity: bound process to CPUs 216-287 (NUMA node 3) for GPU 3 [repeated 15x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) 2026-06-12 05:20:52.589 | DEBUG | skyrl_train.utils.numa:_set_membind_via_libnuma:375 - NUMA affinity: set memory preferred to NUMA node 3 [repeated 15x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) 2026-06-12 05:20:52.589 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:618 - BaseVLLMInferenceEngine: vllm_v1_disable_multiproc=True, vllm.__version__=dev, VLLM_ENABLE_V1_MULTIPROCESSING=0 [repeated 15x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) 2026-06-12 05:20:52.589 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:626 - BaseVLLMInferenceEngine: set VLLM_ENABLE_V1_MULTIPROCESSING=0 [repeated 15x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) 2026-06-12 05:20:52.589 | WARNING | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1192 - OpenAI API sampling params overridden: temperature=1.0, top_p=1.0, top_k=-1 [repeated 15x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) 2026-06-12 05:20:52.622 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1244 - Engine startup stagger: sleeping 2.92s (attempt 1/5) to avoid port collisions [repeated 15x across cluster] (skyrl_entrypoint pid=487747) [2026-06-12 05:20:54] INFO inference_engine_client_http_endpoint.py:350: Starting server on 0.0.0.0:8000 (skyrl_entrypoint pid=487747) [2026-06-12 05:20:54] INFO inference_engine_client_http_endpoint.py:242: Starting inference HTTP endpoint... (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored. [repeated 26x across cluster] (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e [repeated 76x across cluster] (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) The tokenizer you are loading from '/e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue. [repeated 13x across cluster] (skyrl_entrypoint pid=487747) [2026-06-12 05:20:55] INFO inference_engine_client_http_endpoint.py:229: Server ready after 2 attempts (2 seconds) (skyrl_entrypoint pid=487747) 2026-06-12 05:20:55.781 | INFO  | skyrl_train.inference_engines.inference_engine_client:_spin_up_http_endpoint:969 - InferenceEngineClient HTTP endpoint started on 127.0.0.1:8000 (skyrl_entrypoint pid=487747) 2026-06-12 05:20:55.781 | INFO  | skyrl_train.inference_engines.inference_engine_client:__init__:61 - InferenceEngineClient initialized with 48 engines. (skyrl_entrypoint pid=487747) 2026-06-12 05:20:55.783 | INFO  | examples.terminal_bench.terminal_bench_generator:_configure_harbor_logging:242 - Harbor logging level set to INFO (skyrl_entrypoint pid=487747) 2026-06-12 05:20:55.784 | INFO  | examples.terminal_bench.terminal_bench_generator:__init__:142 - TerminalBenchGenerator initialized with HarborConfigBuilder. Exposed fields: ['name', 'max_episodes', 'enable_summarize', 'store_all_messages', 'trajectory_config', 'enable_episode_logging', 'record_terminal_session', 'enable_pane_logging', 'strict_json_parser', 'interleaved_thinking', 'extra_body', 'override_timeout_sec', 'override_cpus', 'override_memory_mb', 'override_storage_mb', 'auto_snapshot', 'verifier_override_timeout_sec', 'max_retries', 'min_wait_sec', 'max_wait_sec', 'wait_multiplier', 'exclude_exceptions', 'n_concurrent_trials', 'log_level', 'enable_reward_shaping', 'collect_rollout_details', 'enable_error_classification', 'mask_exceptions', 'default_error_treatment', 'passthrough_exceptions', 'zero_exceptions']. Retry config: max_retries=3, backoff=60.0-600.0s. Concurrent trials: 675. Reward shaping: enabled=False, shaper=pass_ratio. Error classification: enabled=True (skyrl_entrypoint pid=487747) 2026-06-12 05:20:55.785 | INFO  | examples.terminal_bench.terminal_bench_generator:__init__:158 - TerminalBenchGenerator initialized with custom chat template read from: chat_templates/qwen3_thinking_acc.jinja2 (skyrl_entrypoint pid=487747) 2026-06-12 05:20:55.785 | INFO  | skyrl_train.utils.trainer_utils:build_dataloader:656 - Total steps: 156 (skyrl_entrypoint pid=487747) 2026-06-12 05:20:55.786 | INFO  | skyrl_train.fully_async_trainer:_build_train_dataloader_and_compute_training_steps:357 - Length of train_dataloader: 5000 (skyrl_entrypoint pid=487747) 2026-06-12 05:20:55.786 | INFO  | skyrl_train.fully_async_trainer:_build_train_dataloader_and_compute_training_steps:358 - Number of steps per epoch: 78 (skyrl_entrypoint pid=487747) 2026-06-12 05:20:55.786 | INFO  | skyrl_train.fully_async_trainer:_build_train_dataloader_and_compute_training_steps:359 - Total training steps: 80 (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) 2026-06-12 05:20:57.227 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:152 - setup_envvars_for_vllm: distributed_executor_backend=uni, SKYRL_ENABLE_NUMA_AFFINITY=1, CUDA_VISIBLE_DEVICES=3, VLLM_ENABLE_V1_MULTIPROCESSING= [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) 2026-06-12 05:20:57.228 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:167 - setup_envvars_for_vllm: numa_enabled=True [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) 2026-06-12 05:20:57.228 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:170 - setup_envvars_for_vllm: set VLLM_ENABLE_V1_MULTIPROCESSING=0 for NUMA affinity [repeated 12x across cluster] (raylet, ip=10.128.17.43) 2026-06-12 05:20:56,871 WARNING node.py:1846 -- The object spilling config is specified from an unstable API - system config or environment variable. This is subject to change in the future. You can use the stable API - --object-spilling-directory in ray start or object_spilling_directory in ray.init() to specify the object spilling directory instead. If you need more advanced settings, please open a github issue with the Ray team. [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash: [repeated 21x across cluster] (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) No module named 'vllm._version' [repeated 21x across cluster] (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) from .version import __version__, __version_tuple__ # isort:skip [repeated 21x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) 2026-06-12 05:20:55.955 | INFO | skyrl_train.utils.numa:set_numa_affinity_for_gpu:340 - NUMA affinity: bound process to CPUs 216-287 (NUMA node 3) for GPU 3 [repeated 9x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) 2026-06-12 05:20:55.975 | DEBUG | skyrl_train.utils.numa:_set_membind_via_libnuma:375 - NUMA affinity: set memory preferred to NUMA node 3 [repeated 9x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) 2026-06-12 05:20:55.975 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:618 - BaseVLLMInferenceEngine: vllm_v1_disable_multiproc=True, vllm.__version__=dev, VLLM_ENABLE_V1_MULTIPROCESSING=0 [repeated 9x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) 2026-06-12 05:20:55.975 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:626 - BaseVLLMInferenceEngine: set VLLM_ENABLE_V1_MULTIPROCESSING=0 [repeated 9x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) 2026-06-12 05:20:55.975 | WARNING | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1192 - OpenAI API sampling params overridden: temperature=1.0, top_p=1.0, top_k=-1 [repeated 9x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) 2026-06-12 05:20:56.007 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1244 - Engine startup stagger: sleeping 1.97s (attempt 1/5) to avoid port collisions [repeated 9x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored. [repeated 24x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e [repeated 31x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) The tokenizer you are loading from '/e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue. [repeated 12x across cluster] (pid=332746, ip=10.128.17.43) Using blocking ray.get inside async actor. This blocks the event loop. Please use `await` on object ref with asyncio.gather if you want to yield execution to the event loop instead. (bundle_reservation_check_func pid=487823) [2026-06-12 05:21:01,156 E 487823 487863] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14 (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) 2026-06-12 05:21:01.831 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:152 - setup_envvars_for_vllm: distributed_executor_backend=uni, SKYRL_ENABLE_NUMA_AFFINITY=1, CUDA_VISIBLE_DEVICES=3, VLLM_ENABLE_V1_MULTIPROCESSING= [repeated 16x across cluster] (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) 2026-06-12 05:21:01.832 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:167 - setup_envvars_for_vllm: numa_enabled=True [repeated 16x across cluster] (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) 2026-06-12 05:21:01.832 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:170 - setup_envvars_for_vllm: set VLLM_ENABLE_V1_MULTIPROCESSING=0 for NUMA affinity [repeated 16x across cluster] (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash: [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) No module named 'vllm._version' [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) from .version import __version__, __version_tuple__ # isort:skip [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) 2026-06-12 05:21:02.982 | INFO | skyrl_train.utils.numa:set_numa_affinity_for_gpu:340 - NUMA affinity: bound process to CPUs 216-287 (NUMA node 3) for GPU 3 [repeated 19x across cluster] (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) 2026-06-12 05:21:03.003 | DEBUG | skyrl_train.utils.numa:_set_membind_via_libnuma:375 - NUMA affinity: set memory preferred to NUMA node 3 [repeated 19x across cluster] (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) 2026-06-12 05:21:03.003 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:618 - BaseVLLMInferenceEngine: vllm_v1_disable_multiproc=True, vllm.__version__=dev, VLLM_ENABLE_V1_MULTIPROCESSING=0 [repeated 19x across cluster] (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) 2026-06-12 05:21:03.003 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:626 - BaseVLLMInferenceEngine: set VLLM_ENABLE_V1_MULTIPROCESSING=0 [repeated 19x across cluster] (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) 2026-06-12 05:21:03.003 | WARNING | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1192 - OpenAI API sampling params overridden: temperature=1.0, top_p=1.0, top_k=-1 [repeated 19x across cluster] (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) 2026-06-12 05:21:00.877 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1244 - Engine startup stagger: sleeping 1.51s (attempt 1/5) to avoid port collisions [repeated 15x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) (EngineCore_DP0 pid=2073774) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/__init__.py:1617: UserWarning: Please use the new API settings to control TF32 behavior, such as torch.backends.cudnn.conv.fp32_precision = 'tf32' or torch.backends.cuda.matmul.fp32_precision = 'ieee'. Old settings, e.g, torch.backends.cuda.matmul.allow_tf32 = True, torch.backends.cudnn.allow_tf32 = True, allowTF32CuDNN() and allowTF32CuBLAS() will be deprecated after Pytorch 2.9. Please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:80.) (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) (EngineCore_DP0 pid=2073774) _C._set_float32_matmul_precision(precision) (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) (EngineCore_DP0 pid=309070) _C._set_float32_matmul_precision(precision) (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) (EngineCore_DP0 pid=309080) _C._set_float32_matmul_precision(precision) (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) [W612 05:21:03.547584079 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-007-23-interconnect-1.jupiter.internal]:54591 (errno: 97 - Address family not supported by protocol). (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) [W612 05:21:03.551177743 Utils.hpp:166] Warning: Environment variable NCCL_BLOCKING_WAIT is deprecated; use TORCH_NCCL_BLOCKING_WAIT instead (function operator()) (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) [rank0]:[W612 05:21:03.554852493 Utils.hpp:137] Warning: Environment variable TORCH_NCCL_TRACE_BUFFER_SIZE is deprecated; use TORCH_FR_BUFFER_SIZE instead (function operator()) (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) (EngineCore_DP0 pid=2073782) _C._set_float32_matmul_precision(precision) (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) (EngineCore_DP0 pid=2073795) _C._set_float32_matmul_precision(precision) (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) (EngineCore_DP0 pid=2073787) _C._set_float32_matmul_precision(precision) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored. [repeated 32x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) (EngineCore_DP0 pid=309096) _C._set_float32_matmul_precision(precision) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e [repeated 32x across cluster] (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) The tokenizer you are loading from '/e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue. [repeated 16x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) (EngineCore_DP0 pid=2062587) _C._set_float32_matmul_precision(precision) (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) [2026-06-12 05:21:05,430 E 2073418 2073520] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14 [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) (EngineCore_DP0 pid=309070) Loading safetensors checkpoint shards: 0% Completed | 0/4 [00:00 (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) 2026-06-12 05:21:08.141 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:167 - setup_envvars_for_vllm: numa_enabled=True (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) 2026-06-12 05:21:08.141 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:170 - setup_envvars_for_vllm: set VLLM_ENABLE_V1_MULTIPROCESSING=0 for NUMA affinity (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) (EngineCore_DP0 pid=2062603) _C._set_float32_matmul_precision(precision) (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) (EngineCore_DP0 pid=2062603) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/__init__.py:1617: UserWarning: Please use the new API settings to control TF32 behavior, such as torch.backends.cudnn.conv.fp32_precision = 'tf32' or torch.backends.cuda.matmul.fp32_precision = 'ieee'. Old settings, e.g, torch.backends.cuda.matmul.allow_tf32 = True, torch.backends.cudnn.allow_tf32 = True, allowTF32CuDNN() and allowTF32CuBLAS() will be deprecated after Pytorch 2.9. Please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:80.) [repeated 14x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) (EngineCore_DP0 pid=379917) _C._set_float32_matmul_precision(precision) (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) [W612 05:21:08.090693593 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-010-48.jupiter.internal]:40901 (errno: 97 - Address family not supported by protocol). [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) [W612 05:21:08.091121584 Utils.hpp:166] Warning: Environment variable NCCL_BLOCKING_WAIT is deprecated; use TORCH_NCCL_BLOCKING_WAIT instead (function operator()) [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) [rank0]:[W612 05:21:08.092953961 Utils.hpp:137] Warning: Environment variable TORCH_NCCL_TRACE_BUFFER_SIZE is deprecated; use TORCH_FR_BUFFER_SIZE instead (function operator()) [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) (EngineCore_DP0 pid=379901) _C._set_float32_matmul_precision(precision) (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) (EngineCore_DP0 pid=336335) _C._set_float32_matmul_precision(precision) (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) 2026-06-12 05:21:09.171 | INFO | skyrl_train.utils.numa:set_numa_affinity_for_gpu:340 - NUMA affinity: bound process to CPUs 0-71 (NUMA node 0) for GPU 0 (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) 2026-06-12 05:21:09.198 | DEBUG | skyrl_train.utils.numa:_set_membind_via_libnuma:375 - NUMA affinity: set memory preferred to NUMA node 0 (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) 2026-06-12 05:21:09.198 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:618 - BaseVLLMInferenceEngine: vllm_v1_disable_multiproc=True, vllm.__version__=dev, VLLM_ENABLE_V1_MULTIPROCESSING=0 (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) 2026-06-12 05:21:09.198 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:626 - BaseVLLMInferenceEngine: set VLLM_ENABLE_V1_MULTIPROCESSING=0 (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) 2026-06-12 05:21:09.198 | WARNING | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1192 - OpenAI API sampling params overridden: temperature=1.0, top_p=1.0, top_k=-1 (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) (EngineCore_DP0 pid=379913) _C._set_float32_matmul_precision(precision) (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) (EngineCore_DP0 pid=336355) _C._set_float32_matmul_precision(precision) (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored. [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) The tokenizer you are loading from '/e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue. [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) (EngineCore_DP0 pid=309070) Loading safetensors checkpoint shards: 25% Completed | 1/4 [00:04<00:13, 4.39s/it] (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) (EngineCore_DP0 pid=336331) _C._set_float32_matmul_precision(precision) (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) [2026-06-12 05:21:10,913 E 308849 309013] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14 [repeated 18x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) (EngineCore_DP0 pid=593291) _C._set_float32_matmul_precision(precision) (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) (EngineCore_DP0 pid=379897) Loading safetensors checkpoint shards: 0% Completed | 0/4 [00:00 [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) 2026-06-12 05:21:08.141 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:167 - setup_envvars_for_vllm: numa_enabled=True [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) 2026-06-12 05:21:08.141 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:170 - setup_envvars_for_vllm: set VLLM_ENABLE_V1_MULTIPROCESSING=0 for NUMA affinity [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) (EngineCore_DP0 pid=2080799) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/__init__.py:1617: UserWarning: Please use the new API settings to control TF32 behavior, such as torch.backends.cudnn.conv.fp32_precision = 'tf32' or torch.backends.cuda.matmul.fp32_precision = 'ieee'. Old settings, e.g, torch.backends.cuda.matmul.allow_tf32 = True, torch.backends.cudnn.allow_tf32 = True, allowTF32CuDNN() and allowTF32CuBLAS() will be deprecated after Pytorch 2.9. Please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:80.) [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) (EngineCore_DP0 pid=593303) _C._set_float32_matmul_precision(precision) (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) (EngineCore_DP0 pid=2080782) _C._set_float32_matmul_precision(precision) (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) [W612 05:21:13.158353191 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-051-33-interconnect-1.jupiter.internal]:54597 (errno: 97 - Address family not supported by protocol). [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) (EngineCore_DP0 pid=632066) _C._set_float32_matmul_precision(precision) (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) [W612 05:21:13.158871900 Utils.hpp:166] Warning: Environment variable NCCL_BLOCKING_WAIT is deprecated; use TORCH_NCCL_BLOCKING_WAIT instead (function operator()) [repeated 11x across cluster] (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) [rank0]:[W612 05:21:13.160962639 Utils.hpp:137] Warning: Environment variable TORCH_NCCL_TRACE_BUFFER_SIZE is deprecated; use TORCH_FR_BUFFER_SIZE instead (function operator()) [repeated 11x across cluster] (pid=332821, ip=10.128.17.43) Using blocking ray.get inside async actor. This blocks the event loop. Please use `await` on object ref with asyncio.gather if you want to yield execution to the event loop instead. (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) (EngineCore_DP0 pid=379652) _C._set_float32_matmul_precision(precision) (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) 2026-06-12 05:21:09.320 | INFO | skyrl_train.utils.numa:set_numa_affinity_for_gpu:340 - NUMA affinity: bound process to CPUs 144-215 (NUMA node 2) for GPU 2 [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) 2026-06-12 05:21:09.328 | DEBUG | skyrl_train.utils.numa:_set_membind_via_libnuma:375 - NUMA affinity: set memory preferred to NUMA node 3 [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) 2026-06-12 05:21:09.328 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:618 - BaseVLLMInferenceEngine: vllm_v1_disable_multiproc=True, vllm.__version__=dev, VLLM_ENABLE_V1_MULTIPROCESSING=0 [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) 2026-06-12 05:21:09.328 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:626 - BaseVLLMInferenceEngine: set VLLM_ENABLE_V1_MULTIPROCESSING=0 [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) 2026-06-12 05:21:09.328 | WARNING | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1192 - OpenAI API sampling params overridden: temperature=1.0, top_p=1.0, top_k=-1 [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) (EngineCore_DP0 pid=632077) _C._set_float32_matmul_precision(precision) (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) (EngineCore_DP0 pid=2080778) _C._set_float32_matmul_precision(precision) (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) (EngineCore_DP0 pid=388616) _C._set_float32_matmul_precision(precision) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) (EngineCore_DP0 pid=379664) _C._set_float32_matmul_precision(precision) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored. [repeated 8x across cluster] (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) (EngineCore_DP0 pid=632073) _C._set_float32_matmul_precision(precision) (raylet) [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e [repeated 50x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) The tokenizer you are loading from '/e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue. [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) (EngineCore_DP0 pid=388625) _C._set_float32_matmul_precision(precision) (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) (EngineCore_DP0 pid=593303) Loading safetensors checkpoint shards: 25% Completed | 1/4 [00:01<00:03, 1.26s/it] [repeated 28x across cluster] (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) [2026-06-12 05:21:14,846 E 379289 379391] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14 [repeated 10x across cluster] (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) (EngineCore_DP0 pid=388620) _C._set_float32_matmul_precision(precision) (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) (EngineCore_DP0 pid=632092) _C._set_float32_matmul_precision(precision) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) (EngineCore_DP0 pid=379664) Loading safetensors checkpoint shards: 0% Completed | 0/4 [00:00 10.128.17.29 (routable head IP) for coordinator litellm base_url connectivity (skyrl_entrypoint pid=487747) 2026-06-12 05:21:40.570 | INFO  | examples.terminal_bench.rollout_coordinator:__init__:405 - [RolloutDispatcher] configured num_coordinators=4, cpus_per_coordinator=8 (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) (EngineCore_DP0 pid=386887) Loading safetensors checkpoint shards: 0% Completed | 0/4 [00:00 168 (// 4) (RolloutCoordinator pid=336583, ip=10.128.17.53) The tokenizer you are loading from '/e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue. (RolloutCoordinator pid=336583, ip=10.128.17.53) 2026-06-12 05:21:48.155 | INFO | examples.terminal_bench.terminal_bench_generator:_configure_harbor_logging:242 - Harbor logging level set to INFO (RolloutCoordinator pid=336583, ip=10.128.17.53) 2026-06-12 05:21:48.155 | INFO | examples.terminal_bench.terminal_bench_generator:__init__:142 - TerminalBenchGenerator initialized with HarborConfigBuilder. Exposed fields: ['name', 'max_episodes', 'enable_summarize', 'store_all_messages', 'trajectory_config', 'enable_episode_logging', 'record_terminal_session', 'enable_pane_logging', 'strict_json_parser', 'interleaved_thinking', 'extra_body', 'override_timeout_sec', 'override_cpus', 'override_memory_mb', 'override_storage_mb', 'auto_snapshot', 'verifier_override_timeout_sec', 'max_retries', 'min_wait_sec', 'max_wait_sec', 'wait_multiplier', 'exclude_exceptions', 'n_concurrent_trials', 'log_level', 'enable_reward_shaping', 'collect_rollout_details', 'enable_error_classification', 'mask_exceptions', 'default_error_treatment', 'passthrough_exceptions', 'zero_exceptions']. Retry config: max_retries=3, backoff=60.0-600.0s. Concurrent trials: 168. Reward shaping: enabled=False, shaper=pass_ratio. Error classification: enabled=True (RolloutCoordinator pid=336583, ip=10.128.17.53) 2026-06-12 05:21:48.156 | INFO | examples.terminal_bench.terminal_bench_generator:__init__:158 - TerminalBenchGenerator initialized with custom chat template read from: chat_templates/qwen3_thinking_acc.jinja2 (RolloutCoordinator pid=336583, ip=10.128.17.53) 2026-06-12 05:21:48.156 | INFO | examples.terminal_bench.rollout_coordinator:__init__:246 - [RolloutCoordinator 0/4] constructed (http=10.128.17.29:8000) (RolloutCoordinator pid=336583, ip=10.128.17.53) 2026-06-12 05:21:48.159 | INFO | examples.terminal_bench.terminal_bench_generator:_create_orchestrator:298 - QueueOrchestrator created and started with n_concurrent_trials=168, rollback_hook registered for ContextLengthExceededError/AgentTimeoutError (RolloutCoordinator pid=336583, ip=10.128.17.53) 2026-06-12 05:21:48.159 | INFO | examples.terminal_bench.terminal_bench_generator:startup:255 - TerminalBenchGenerator startup complete. Shared orchestrator ready with n_concurrent_trials=168 (RolloutCoordinator pid=336583, ip=10.128.17.53) 2026-06-12 05:21:48.159 | INFO | examples.terminal_bench.rollout_coordinator:startup:254 - [RolloutCoordinator 0] startup complete (skyrl_entrypoint pid=487747) 2026-06-12 05:21:48.160 | INFO  | examples.terminal_bench.rollout_coordinator:startup:469 - [RolloutDispatcher] coordinator 1/4 started (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) (EngineCore_DP0 pid=386885) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) (EngineCore_DP0 pid=386894) (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) (EngineCore_DP0 pid=386886) (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) (EngineCore_DP0 pid=386887) (raylet, ip=10.128.17.55) 2026-06-12 05:21:51,155 WARNING node.py:1846 -- The object spilling config is specified from an unstable API - system config or environment variable. This is subject to change in the future. You can use the stable API - --object-spilling-directory in ray start or object_spilling_directory in ray.init() to specify the object spilling directory instead. If you need more advanced settings, please open a github issue with the Ray team. (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) 2026-06-12 05:21:53.078 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1301 - Initializing OpenAIServingChat with custom_chat_template read from: chat_templates/qwen3_thinking_acc.jinja2 (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) (EngineCore_DP0 pid=386887) Loading safetensors checkpoint shards: 100% Completed | 4/4 [00:14<00:00, 3.60s/it] [repeated 11x across cluster] (raylet, ip=10.128.17.55) [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) (EngineCore_DP0 pid=386887) The tokenizer you are loading from '/e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue. [repeated 4x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) 2026-06-12 05:21:57.222 | INFO | examples.terminal_bench.rollout_coordinator:_scale_terminal_bench_cfg:121 - [RolloutCoordinator] scaled n_concurrent_trials 675 -> 168 (// 4) (RolloutCoordinator pid=309396, ip=10.128.17.55) 2026-06-12 05:21:57.451 | INFO | examples.terminal_bench.terminal_bench_generator:_configure_harbor_logging:242 - Harbor logging level set to INFO (RolloutCoordinator pid=309396, ip=10.128.17.55) 2026-06-12 05:21:57.452 | INFO | examples.terminal_bench.terminal_bench_generator:__init__:142 - TerminalBenchGenerator initialized with HarborConfigBuilder. Exposed fields: ['name', 'max_episodes', 'enable_summarize', 'store_all_messages', 'trajectory_config', 'enable_episode_logging', 'record_terminal_session', 'enable_pane_logging', 'strict_json_parser', 'interleaved_thinking', 'extra_body', 'override_timeout_sec', 'override_cpus', 'override_memory_mb', 'override_storage_mb', 'auto_snapshot', 'verifier_override_timeout_sec', 'max_retries', 'min_wait_sec', 'max_wait_sec', 'wait_multiplier', 'exclude_exceptions', 'n_concurrent_trials', 'log_level', 'enable_reward_shaping', 'collect_rollout_details', 'enable_error_classification', 'mask_exceptions', 'default_error_treatment', 'passthrough_exceptions', 'zero_exceptions']. Retry config: max_retries=3, backoff=60.0-600.0s. Concurrent trials: 168. Reward shaping: enabled=False, shaper=pass_ratio. Error classification: enabled=True (RolloutCoordinator pid=309396, ip=10.128.17.55) 2026-06-12 05:21:57.452 | INFO | examples.terminal_bench.terminal_bench_generator:__init__:158 - TerminalBenchGenerator initialized with custom chat template read from: chat_templates/qwen3_thinking_acc.jinja2 (RolloutCoordinator pid=309396, ip=10.128.17.55) 2026-06-12 05:21:57.452 | INFO | examples.terminal_bench.rollout_coordinator:__init__:246 - [RolloutCoordinator 1/4] constructed (http=10.128.17.29:8000) (RolloutCoordinator pid=309396, ip=10.128.17.55) 2026-06-12 05:21:57.455 | INFO | examples.terminal_bench.terminal_bench_generator:_create_orchestrator:298 - QueueOrchestrator created and started with n_concurrent_trials=168, rollback_hook registered for ContextLengthExceededError/AgentTimeoutError (RolloutCoordinator pid=309396, ip=10.128.17.55) 2026-06-12 05:21:57.455 | INFO | examples.terminal_bench.terminal_bench_generator:startup:255 - TerminalBenchGenerator startup complete. Shared orchestrator ready with n_concurrent_trials=168 (RolloutCoordinator pid=309396, ip=10.128.17.55) 2026-06-12 05:21:57.455 | INFO | examples.terminal_bench.rollout_coordinator:startup:254 - [RolloutCoordinator 1] startup complete (skyrl_entrypoint pid=487747) 2026-06-12 05:21:57.457 | INFO  | examples.terminal_bench.rollout_coordinator:startup:469 - [RolloutDispatcher] coordinator 2/4 started (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) 2026-06-12 05:21:53.119 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1301 - Initializing OpenAIServingChat with custom_chat_template read from: chat_templates/qwen3_thinking_acc.jinja2 [repeated 3x across cluster] (raylet, ip=10.128.34.1) [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e (RolloutCoordinator pid=309396, ip=10.128.17.55) The tokenizer you are loading from '/e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue. (raylet, ip=10.128.34.1) [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e (raylet, ip=10.128.34.1) 2026-06-12 05:22:00,346 WARNING node.py:1846 -- The object spilling config is specified from an unstable API - system config or environment variable. This is subject to change in the future. You can use the stable API - --object-spilling-directory in ray start or object_spilling_directory in ray.init() to specify the object spilling directory instead. If you need more advanced settings, please open a github issue with the Ray team. (RolloutCoordinator pid=593552, ip=10.128.34.1) 2026-06-12 05:22:05.433 | INFO | examples.terminal_bench.rollout_coordinator:_scale_terminal_bench_cfg:121 - [RolloutCoordinator] scaled n_concurrent_trials 675 -> 168 (// 4) (raylet, ip=10.128.34.1) [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e [repeated 4x across cluster] (skyrl_entrypoint pid=487747) 2026-06-12 05:22:05.653 | INFO  | examples.terminal_bench.rollout_coordinator:startup:469 - [RolloutDispatcher] coordinator 3/4 started (RolloutCoordinator pid=593552, ip=10.128.34.1) The tokenizer you are loading from '/e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue. (RolloutCoordinator pid=593552, ip=10.128.34.1) 2026-06-12 05:22:05.649 | INFO | examples.terminal_bench.terminal_bench_generator:_configure_harbor_logging:242 - Harbor logging level set to INFO (RolloutCoordinator pid=593552, ip=10.128.34.1) 2026-06-12 05:22:05.649 | INFO | examples.terminal_bench.terminal_bench_generator:__init__:142 - TerminalBenchGenerator initialized with HarborConfigBuilder. Exposed fields: ['name', 'max_episodes', 'enable_summarize', 'store_all_messages', 'trajectory_config', 'enable_episode_logging', 'record_terminal_session', 'enable_pane_logging', 'strict_json_parser', 'interleaved_thinking', 'extra_body', 'override_timeout_sec', 'override_cpus', 'override_memory_mb', 'override_storage_mb', 'auto_snapshot', 'verifier_override_timeout_sec', 'max_retries', 'min_wait_sec', 'max_wait_sec', 'wait_multiplier', 'exclude_exceptions', 'n_concurrent_trials', 'log_level', 'enable_reward_shaping', 'collect_rollout_details', 'enable_error_classification', 'mask_exceptions', 'default_error_treatment', 'passthrough_exceptions', 'zero_exceptions']. Retry config: max_retries=3, backoff=60.0-600.0s. Concurrent trials: 168. Reward shaping: enabled=False, shaper=pass_ratio. Error classification: enabled=True (RolloutCoordinator pid=593552, ip=10.128.34.1) 2026-06-12 05:22:05.649 | INFO | examples.terminal_bench.terminal_bench_generator:__init__:158 - TerminalBenchGenerator initialized with custom chat template read from: chat_templates/qwen3_thinking_acc.jinja2 (RolloutCoordinator pid=593552, ip=10.128.34.1) 2026-06-12 05:22:05.650 | INFO | examples.terminal_bench.rollout_coordinator:__init__:246 - [RolloutCoordinator 2/4] constructed (http=10.128.17.29:8000) (RolloutCoordinator pid=593552, ip=10.128.34.1) 2026-06-12 05:22:05.653 | INFO | examples.terminal_bench.terminal_bench_generator:_create_orchestrator:298 - QueueOrchestrator created and started with n_concurrent_trials=168, rollback_hook registered for ContextLengthExceededError/AgentTimeoutError (RolloutCoordinator pid=593552, ip=10.128.34.1) 2026-06-12 05:22:05.653 | INFO | examples.terminal_bench.terminal_bench_generator:startup:255 - TerminalBenchGenerator startup complete. Shared orchestrator ready with n_concurrent_trials=168 (RolloutCoordinator pid=593552, ip=10.128.34.1) 2026-06-12 05:22:05.653 | INFO | examples.terminal_bench.rollout_coordinator:startup:254 - [RolloutCoordinator 2] startup complete (raylet, ip=10.128.17.224) 2026-06-12 05:22:08,553 WARNING node.py:1846 -- The object spilling config is specified from an unstable API - system config or environment variable. This is subject to change in the future. You can use the stable API - --object-spilling-directory in ray start or object_spilling_directory in ray.init() to specify the object spilling directory instead. If you need more advanced settings, please open a github issue with the Ray team. (RolloutCoordinator pid=336583, ip=10.128.17.53) [2026-06-12 05:22:11,525 E 336583 336623] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14 (raylet, ip=10.128.17.224) [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e [repeated 6x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) 2026-06-12 05:22:14.028 | INFO | examples.terminal_bench.rollout_coordinator:_scale_terminal_bench_cfg:121 - [RolloutCoordinator] scaled n_concurrent_trials 675 -> 168 (// 4) (RolloutCoordinator pid=2062840, ip=10.128.17.224) The tokenizer you are loading from '/e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6' with an incorrect regex pattern: https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503/discussions/84#69121093e8b480e709447d5e. This will lead to incorrect tokenization. You should set the `fix_mistral_regex=True` flag when loading this tokenizer to fix this issue. (RolloutCoordinator pid=2062840, ip=10.128.17.224) 2026-06-12 05:22:14.243 | INFO | examples.terminal_bench.terminal_bench_generator:_configure_harbor_logging:242 - Harbor logging level set to INFO (RolloutCoordinator pid=2062840, ip=10.128.17.224) 2026-06-12 05:22:14.244 | INFO | examples.terminal_bench.terminal_bench_generator:__init__:142 - TerminalBenchGenerator initialized with HarborConfigBuilder. Exposed fields: ['name', 'max_episodes', 'enable_summarize', 'store_all_messages', 'trajectory_config', 'enable_episode_logging', 'record_terminal_session', 'enable_pane_logging', 'strict_json_parser', 'interleaved_thinking', 'extra_body', 'override_timeout_sec', 'override_cpus', 'override_memory_mb', 'override_storage_mb', 'auto_snapshot', 'verifier_override_timeout_sec', 'max_retries', 'min_wait_sec', 'max_wait_sec', 'wait_multiplier', 'exclude_exceptions', 'n_concurrent_trials', 'log_level', 'enable_reward_shaping', 'collect_rollout_details', 'enable_error_classification', 'mask_exceptions', 'default_error_treatment', 'passthrough_exceptions', 'zero_exceptions']. Retry config: max_retries=3, backoff=60.0-600.0s. Concurrent trials: 168. Reward shaping: enabled=False, shaper=pass_ratio. Error classification: enabled=True (RolloutCoordinator pid=2062840, ip=10.128.17.224) 2026-06-12 05:22:14.244 | INFO | examples.terminal_bench.terminal_bench_generator:__init__:158 - TerminalBenchGenerator initialized with custom chat template read from: chat_templates/qwen3_thinking_acc.jinja2 (RolloutCoordinator pid=2062840, ip=10.128.17.224) 2026-06-12 05:22:14.244 | INFO | examples.terminal_bench.rollout_coordinator:__init__:246 - [RolloutCoordinator 3/4] constructed (http=10.128.17.29:8000) (RolloutCoordinator pid=2062840, ip=10.128.17.224) 2026-06-12 05:22:14.247 | INFO | examples.terminal_bench.terminal_bench_generator:_create_orchestrator:298 - QueueOrchestrator created and started with n_concurrent_trials=168, rollback_hook registered for ContextLengthExceededError/AgentTimeoutError (RolloutCoordinator pid=2062840, ip=10.128.17.224) 2026-06-12 05:22:14.247 | INFO | examples.terminal_bench.terminal_bench_generator:startup:255 - TerminalBenchGenerator startup complete. Shared orchestrator ready with n_concurrent_trials=168 (RolloutCoordinator pid=2062840, ip=10.128.17.224) 2026-06-12 05:22:14.247 | INFO | examples.terminal_bench.rollout_coordinator:startup:254 - [RolloutCoordinator 3] startup complete (skyrl_entrypoint pid=487747) 2026-06-12 05:22:14.247 | INFO  | examples.terminal_bench.rollout_coordinator:startup:469 - [RolloutDispatcher] coordinator 4/4 started (skyrl_entrypoint pid=487747) 2026-06-12 05:22:14.248 | INFO  | examples.terminal_bench.rollout_coordinator:startup:477 - [RolloutDispatcher] 4 coordinators started (skyrl_entrypoint pid=487747) 2026-06-12 05:22:14.248 | INFO  | skyrl_train.fully_async_trainer:train:485 - Generator startup complete (skyrl_entrypoint pid=487747) 2026-06-12 05:22:14.248 | INFO  | skyrl_train.fully_async_trainer:_train_loop:510 - Started: 'load_checkpoints' (skyrl_entrypoint pid=487747) 2026-06-12 05:22:14.383 | INFO  | skyrl_train.trainer:load_checkpoints:1735 - Loading checkpoint from: /e/data1/datasets/playground/ot-baf/explore-tis-minp/explore-tis-minp/checkpoints/global_step_81 (skyrl_entrypoint pid=487747) 2026-06-12 05:22:14.384 | INFO  | skyrl_train.trainer:load_checkpoints:1741 - Resuming from global_step: 81 (skyrl_entrypoint pid=487747) 2026-06-12 05:22:14.389 | INFO  | skyrl_train.trainer:load_checkpoints:1757 - Successfully loaded trainer state (skyrl_entrypoint pid=487747) 2026-06-12 05:22:14.391 | INFO  | skyrl_train.trainer:load_checkpoints:1767 - Successfully loaded dataloader state (skyrl_entrypoint pid=487747) 2026-06-12 05:22:14.391 | INFO  | skyrl_train.trainer:load_checkpoints:1776 - Loading policy checkpoint from /e/data1/datasets/playground/ot-baf/explore-tis-minp/explore-tis-minp/checkpoints/global_step_81/policy (RolloutCoordinator pid=309396, ip=10.128.17.55) [2026-06-12 05:22:21,175 E 309396 309436] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14 (RolloutCoordinator pid=593552, ip=10.128.34.1) [2026-06-12 05:22:30,364 E 593552 593592] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/distributed/distributed_c10d.py:4876: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning. (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) warnings.warn( # warn only once (skyrl_entrypoint pid=487747) 2026-06-12 05:22:37.072 | INFO  | skyrl_train.trainer:load_checkpoints:1786 - Successfully loaded policy checkpoint (skyrl_entrypoint pid=487747) 2026-06-12 05:22:37.072 | INFO  | skyrl_train.trainer:load_checkpoints:1802 - Successfully loaded complete checkpoint state from global_step_81 (skyrl_entrypoint pid=487747) 2026-06-12 05:22:37.072 | INFO  | skyrl_train.fully_async_trainer:_train_loop:512 - Resumed training from global_step 81 (skyrl_entrypoint pid=487747) 2026-06-12 05:22:37.073 | WARNING  | skyrl_train.fully_async_trainer:_train_loop:517 - No data consumption state found in checkpoint — resume may re-train on already-consumed data (skyrl_entrypoint pid=487747) 2026-06-12 05:22:37.074 | WARNING  | skyrl_train.fully_async_trainer:_train_loop:537 - Data consumption count mismatch on resume: expected 192, got 0. This can happen after epoch boundary transitions or error recovery. (skyrl_entrypoint pid=487747) 2026-06-12 05:22:37.074 | INFO  | skyrl_train.fully_async_trainer:_train_loop:510 - Finished: 'load_checkpoints', time cost: 22.83s (skyrl_entrypoint pid=487747) 2026-06-12 05:22:37.074 | INFO  | skyrl_train.fully_async_trainer:_train_loop:544 - Started: 'init_weight_sync_state' (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash: (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) No module named 'vllm._version' (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) from .version import __version__, __version_tuple__ # isort:skip (RolloutCoordinator pid=2062840, ip=10.128.17.224) [2026-06-12 05:22:38,571 E 2062840 2062882] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 2026-06-12 05:22:43.007 | INFO  | logging:info:2216 - [weight-sync] Using master_addr=10.128.17.43, master_port=60333 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 2026-06-12 05:22:43.007 | INFO  | logging:info:2216 - [weight-sync] get_node_ip_address()=10.128.17.43 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487913) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash: [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=487913) No module named 'vllm._version' [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=487913) from .version import __version__, __version_tuple__ # isort:skip [repeated 7x across cluster] (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) (EngineCore_DP0 pid=2062603) 2026-06-12 05:22:43.159 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:init_weight_update_communicator:223 - torch.distributed.get_rank(): 0, rank_offset: 5, rank: 5, world_size: 49, group_name: skyrl (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) [rank0]:[W612 05:22:43.882881151 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-007-11.jupiter.internal]:60333 (errno: 97 - Address family not supported by protocol). (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) (EngineCore_DP0 pid=2062603) 2026-06-12 05:22:43.162 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:init_weight_update_communicator:234 - init_weight_update_communicator: master_address=10.128.17.43, master_port=60333, (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/distributed/distributed_c10d.py:4876: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning. (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) warnings.warn( # warn only once (skyrl_entrypoint pid=487747) 2026-06-12 05:22:43.246 | INFO  | skyrl_train.trainer:init_weight_sync_state:874 - Initialized weight sync state for policy model and inference engines. (skyrl_entrypoint pid=487747) 2026-06-12 05:22:43.246 | INFO  | skyrl_train.fully_async_trainer:_train_loop:544 - Finished: 'init_weight_sync_state', time cost: 6.17s (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) [rank0]:[W612 05:22:47.048454997 Utils.hpp:137] Warning: Environment variable TORCH_NCCL_TRACE_BUFFER_SIZE is deprecated; use TORCH_FR_BUFFER_SIZE instead (function operator()) (RolloutCoordinator pid=336583, ip=10.128.17.53) 2026-06-12 05:22:51.765 | INFO | examples.terminal_bench.terminal_bench_generator:generate:560 - Starting batch generation for 8 trials (mode=training) (skyrl_entrypoint pid=487747) 2026-06-12 05:22:43.246 | INFO  | skyrl_train.fully_async_trainer:_train_loop:548 - Started: 'sync_weights_to_inference_engines' (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) (EngineCore_DP0 pid=593291) 2026-06-12 05:22:43.162 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:init_weight_update_communicator:223 - torch.distributed.get_rank(): 0, rank_offset: 40, rank: 40, world_size: 49, group_name: skyrl [repeated 47x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) [rank0]:[W612 05:22:43.623277380 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-007-11-interconnect-1.jupiter.internal]:60333 (errno: 97 - Address family not supported by protocol). [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) (EngineCore_DP0 pid=593291) 2026-06-12 05:22:43.166 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:init_weight_update_communicator:234 - init_weight_update_communicator: master_address=10.128.17.43, master_port=60333,  [repeated 47x across cluster] (skyrl_entrypoint pid=487747) 2026-06-12 05:22:51.759 | INFO  | skyrl_train.fully_async_trainer:_train_loop:548 - Finished: 'sync_weights_to_inference_engines', time cost: 8.51s (skyrl_entrypoint pid=487747) 2026-06-12 05:22:51.760 | INFO  | skyrl_train.callbacks.builtin:on_train_begin:270 - HFHubUploadCallback initialized: repo=laion/explore-tis-minp, upload_steps=5, export_path=/e/data1/datasets/playground/ot-baf/explore-tis-minp/explore-tis-minp/exports (skyrl_entrypoint pid=487747) Training Step Progress: 81it [00:00, ?it/s] (skyrl_entrypoint pid=487747) 2026-06-12 05:22:51.763 | INFO  | skyrl_train.fully_async_trainer:_train_loop:604 - Started: 'step' (skyrl_entrypoint pid=487747) Generation Buffer Progress: 0%| | 0/64 [00:00 cb=[_run_until_complete_cb() at /e/scratch/jureap59/feuer1/miniforge3/lib/python3.12/asyncio/base_events.py:181]> got Future > attached to a different loop (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.088 | INFO  | skyrl_train.trainer:_guarded_async:219 - Inference engine teardown complete (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.088 | INFO  | skyrl_train.trainer:_kill_ray_actors:252 - Killing policy_model actors... (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.096 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 0 died (ActorDiedError). 47/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.104 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 1 died (ActorDiedError). 46/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.111 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 2 died (ActorDiedError). 45/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.115 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 3 died (ActorDiedError). 44/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.123 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 4 died (ActorDiedError). 43/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.127 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 5 died (ActorDiedError). 42/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.148 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 6 died (ActorDiedError). 41/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.155 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 7 died (ActorDiedError). 40/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.158 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 8 died (ActorDiedError). 39/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.172 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 9 died (ActorDiedError). 38/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.183 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 10 died (ActorDiedError). 37/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.190 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 11 died (ActorDiedError). 36/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.193 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 12 died (ActorDiedError). 35/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.200 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 13 died (ActorDiedError). 34/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.208 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 14 died (ActorDiedError). 33/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.216 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 15 died (ActorDiedError). 32/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.218 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 16 died (ActorDiedError). 31/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.223 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 17 died (ActorDiedError). 30/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.226 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 18 died (ActorDiedError). 29/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.235 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 19 died (ActorDiedError). 28/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.245 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 20 died (ActorDiedError). 27/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.253 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 21 died (ActorDiedError). 26/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.258 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 22 died (ActorDiedError). 25/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.265 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 23 died (ActorDiedError). 24/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.271 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 24 died (ActorDiedError). 23/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.280 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 25 died (ActorDiedError). 22/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.291 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 26 died (ActorDiedError). 21/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.296 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 27 died (ActorDiedError). 20/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.300 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 28 died (ActorDiedError). 19/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.308 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 29 died (ActorDiedError). 18/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.311 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 30 died (ActorDiedError). 17/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.314 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 31 died (ActorDiedError). 16/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.318 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 32 died (ActorDiedError). 15/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.325 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 33 died (ActorDiedError). 14/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.335 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 34 died (ActorDiedError). 13/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.398 | INFO  | skyrl_train.trainer:_kill_ray_actors:271 - Killed 48 inference engine actor(s) (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.399 | INFO  | skyrl_train.trainer:_guarded_sync:230 - Ray actor cleanup complete (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.405 | INFO  | skyrl_train.trainer:_kill_ray_actors:252 - Killing policy_model actors... (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.408 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 35 died (ActorDiedError). 12/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.418 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 36 died (ActorDiedError). 11/48 engines remaining. (skyrl_entrypoint pid=487747) 2026-06-12 05:50:02.454 | WARNING  | skyrl_train.inference_engines.inference_engine_client:_mark_engine_dead:68 - Inference engine 37 died (ActorDiedError). 10/48 engines remaining. 2026-06-12 05:50:02.506 | INFO | __main__:main:137 - Shutting down Ray on head node... (pid=334175, ip=10.128.17.43) [2026-06-12 05:49:57,738 E 334175 334278] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14 (RolloutCoordinator pid=2062840, ip=10.128.17.224) 2026-06-12 05:50:01.527 | INFO | examples.terminal_bench.terminal_bench_generator:shutdown:369 - Shutting down shared QueueOrchestrator... [repeated 3x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) 2026-06-12 05:50:01.527 | INFO | examples.terminal_bench.terminal_bench_generator:shutdown:371 - QueueOrchestrator shutdown complete [repeated 3x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) 2026-06-12 05:50:01.527 | INFO | examples.terminal_bench.rollout_coordinator:shutdown:258 - [RolloutCoordinator 3] shutdown complete [repeated 3x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) Error closing AsyncDaytona client: Task cb=[_run_until_complete_cb() at /e/scratch/jureap59/feuer1/miniforge3/lib/python3.12/asyncio/base_events.py:181]> got Future > attached to a different loop [repeated 2x across cluster] (skyrl_entrypoint pid=487747) [fd-monitor] Started monitoring (every 120s) (skyrl_entrypoint pid=487747) [fd-monitor] [05:20:20] OK: 50 / 131,072 FDs open (0.0% of soft limit, hard limit: 131,072) (skyrl_entrypoint pid=487747) [fd-monitor] [05:20:20] OK: RSS 1.38 GiB | node mem 131.8/858.0 GiB used (15.4%), avail 726.1 GiB (skyrl_entrypoint pid=487747) ⚙️ Running in WANDB offline mode (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) WARNING 06-12 05:20:50 [arg_utils.py:1256] The global random seed is set to 44. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM. (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) INFO 06-12 05:20:50 [model.py:529] Resolved architecture: Qwen3ForCausalLM (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) INFO 06-12 05:20:50 [model.py:1549] Using max model len 32768 (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) INFO 06-12 05:20:50 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'NCCL_DEBUG': '***', 'NCCL_DEBUG_SUBSYS': '***', 'NCCL_SOCKET_FAMILY': '***', 'NCCL_SOCKET_IFNAME': '***', 'RAY_ADDRESS': '***', 'RAY_USE_UVLOOP': '***', 'SKYRL_WORKER_NCCL_TIMEOUT_IN_S': '***', 'TORCH_FR_BUFFER_SIZE': '***', 'TORCH_NCCL_ASYNC_ERROR_HANDLING': '***', 'TORCH_NCCL_DEBUG_INFO_TEMP_FILE': '***', 'TORCH_NCCL_DUMP_ON_TIMEOUT': '***', 'TORCH_NCCL_TRACE_BUFFER_SIZE': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}} (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) INFO 06-12 05:20:50 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=65536. (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) INFO 06-12 05:20:50 [vllm.py:690] Asynchronous scheduling is enabled. (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) WARNING 06-12 05:20:50 [vllm.py:728] Enforce eager set, overriding optimization level to -O0 (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) INFO 06-12 05:20:50 [vllm.py:846] Cudagraph is disabled under eager mode (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) WARNING 06-12 05:20:50 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1 (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) WARNING 06-12 05:20:50 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) WARNING 06-12 05:20:55 [arg_utils.py:1256] The global random seed is set to 51. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM. [repeated 14x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) INFO 06-12 05:20:55 [model.py:529] Resolved architecture: Qwen3ForCausalLM [repeated 14x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) INFO 06-12 05:20:55 [model.py:1549] Using max model len 32768 [repeated 14x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) INFO 06-12 05:20:55 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'NCCL_DEBUG': '***', 'NCCL_DEBUG_SUBSYS': '***', 'NCCL_SOCKET_FAMILY': '***', 'NCCL_SOCKET_IFNAME': '***', 'RAY_ADDRESS': '***', 'RAY_USE_UVLOOP': '***', 'SKYRL_WORKER_NCCL_TIMEOUT_IN_S': '***', 'TORCH_FR_BUFFER_SIZE': '***', 'TORCH_NCCL_ASYNC_ERROR_HANDLING': '***', 'TORCH_NCCL_DEBUG_INFO_TEMP_FILE': '***', 'TORCH_NCCL_DUMP_ON_TIMEOUT': '***', 'TORCH_NCCL_TRACE_BUFFER_SIZE': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}} [repeated 14x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) INFO 06-12 05:20:55 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=65536. [repeated 14x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) INFO 06-12 05:20:55 [vllm.py:690] Asynchronous scheduling is enabled. [repeated 14x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) WARNING 06-12 05:20:55 [vllm.py:728] Enforce eager set, overriding optimization level to -O0 [repeated 14x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) INFO 06-12 05:20:55 [vllm.py:846] Cudagraph is disabled under eager mode [repeated 14x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) WARNING 06-12 05:20:55 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1 [repeated 13x across cluster] (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) WARNING 06-12 05:20:55 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned [repeated 13x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) (EngineCore_DP0 pid=2073774) INFO 06-12 05:20:58 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', speculative_config=None, tokenizer='/e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=44, served_model_name=0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': , 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [65536], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': , 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': , 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []} (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) WARNING 06-12 05:21:00 [arg_utils.py:1256] The global random seed is set to 69. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM. [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) INFO 06-12 05:21:00 [model.py:529] Resolved architecture: Qwen3ForCausalLM [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) INFO 06-12 05:21:00 [model.py:1549] Using max model len 32768 [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) INFO 06-12 05:21:00 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'NCCL_DEBUG': '***', 'NCCL_DEBUG_SUBSYS': '***', 'NCCL_SOCKET_FAMILY': '***', 'NCCL_SOCKET_IFNAME': '***', 'RAY_ADDRESS': '***', 'RAY_USE_UVLOOP': '***', 'SKYRL_WORKER_NCCL_TIMEOUT_IN_S': '***', 'TORCH_FR_BUFFER_SIZE': '***', 'TORCH_NCCL_ASYNC_ERROR_HANDLING': '***', 'TORCH_NCCL_DEBUG_INFO_TEMP_FILE': '***', 'TORCH_NCCL_DUMP_ON_TIMEOUT': '***', 'TORCH_NCCL_TRACE_BUFFER_SIZE': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}} [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) INFO 06-12 05:21:00 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=65536. [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) INFO 06-12 05:21:00 [vllm.py:690] Asynchronous scheduling is enabled. [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) WARNING 06-12 05:21:00 [vllm.py:728] Enforce eager set, overriding optimization level to -O0 [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) INFO 06-12 05:21:00 [vllm.py:846] Cudagraph is disabled under eager mode [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) WARNING 06-12 05:21:00 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1 [repeated 13x across cluster] (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) WARNING 06-12 05:21:00 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned [repeated 13x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) (EngineCore_DP0 pid=2073774) INFO 06-12 05:21:03 [worker_base.py:289] Injected into for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'read_named_weights', 'set_numa_affinity', 'test_rpc'] (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) (EngineCore_DP0 pid=309070) INFO 06-12 05:21:03 [parallel_state.py:1234] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.128.17.55:34867 backend=nccl (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) (EngineCore_DP0 pid=309070) INFO 06-12 05:21:03 [parallel_state.py:1445] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) (EngineCore_DP0 pid=2062603) INFO 06-12 05:21:03 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', speculative_config=None, tokenizer='/e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=46, served_model_name=0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': , 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [65536], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': , 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': , 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []} [repeated 10x across cluster] (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) (EngineCore_DP0 pid=309070) INFO 06-12 05:21:05 [gpu_model_runner.py:4125] Starting to load model /e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6... (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) WARNING 06-12 05:21:05 [arg_utils.py:1256] The global random seed is set to 83. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM. [repeated 16x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) INFO 06-12 05:21:05 [model.py:529] Resolved architecture: Qwen3ForCausalLM [repeated 16x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) INFO 06-12 05:21:05 [model.py:1549] Using max model len 32768 [repeated 16x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) INFO 06-12 05:21:05 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'NCCL_DEBUG': '***', 'NCCL_DEBUG_SUBSYS': '***', 'NCCL_SOCKET_FAMILY': '***', 'NCCL_SOCKET_IFNAME': '***', 'RAY_ADDRESS': '***', 'RAY_USE_UVLOOP': '***', 'SKYRL_WORKER_NCCL_TIMEOUT_IN_S': '***', 'TORCH_FR_BUFFER_SIZE': '***', 'TORCH_NCCL_ASYNC_ERROR_HANDLING': '***', 'TORCH_NCCL_DEBUG_INFO_TEMP_FILE': '***', 'TORCH_NCCL_DUMP_ON_TIMEOUT': '***', 'TORCH_NCCL_TRACE_BUFFER_SIZE': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}} [repeated 16x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) INFO 06-12 05:21:05 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=65536. [repeated 16x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) INFO 06-12 05:21:05 [vllm.py:690] Asynchronous scheduling is enabled. [repeated 16x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) WARNING 06-12 05:21:05 [vllm.py:728] Enforce eager set, overriding optimization level to -O0 [repeated 16x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) INFO 06-12 05:21:05 [vllm.py:846] Cudagraph is disabled under eager mode [repeated 16x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) WARNING 06-12 05:21:05 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1 [repeated 16x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) WARNING 06-12 05:21:05 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned [repeated 16x across cluster] (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) (EngineCore_DP0 pid=309070) INFO 06-12 05:21:06 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION']. (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) (EngineCore_DP0 pid=2062603) INFO 06-12 05:21:08 [worker_base.py:289] Injected into for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'read_named_weights', 'set_numa_affinity', 'test_rpc'] [repeated 14x across cluster] (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) (EngineCore_DP0 pid=2062603) INFO 06-12 05:21:08 [parallel_state.py:1234] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.128.17.224:40901 backend=nccl [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) (EngineCore_DP0 pid=2062603) INFO 06-12 05:21:08 [parallel_state.py:1445] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) (EngineCore_DP0 pid=593291) INFO 06-12 05:21:08 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', speculative_config=None, tokenizer='/e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=81, served_model_name=0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': , 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [65536], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': , 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': , 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []} [repeated 13x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) (EngineCore_DP0 pid=379917) INFO 06-12 05:21:09 [gpu_model_runner.py:4125] Starting to load model /e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6... [repeated 15x across cluster] (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) WARNING 06-12 05:21:05 [arg_utils.py:1256] The global random seed is set to 85. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM. (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) INFO 06-12 05:21:05 [model.py:529] Resolved architecture: Qwen3ForCausalLM (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) INFO 06-12 05:21:05 [model.py:1549] Using max model len 32768 (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) INFO 06-12 05:21:05 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'NCCL_DEBUG': '***', 'NCCL_DEBUG_SUBSYS': '***', 'NCCL_SOCKET_FAMILY': '***', 'NCCL_SOCKET_IFNAME': '***', 'RAY_ADDRESS': '***', 'RAY_USE_UVLOOP': '***', 'SKYRL_WORKER_NCCL_TIMEOUT_IN_S': '***', 'TORCH_FR_BUFFER_SIZE': '***', 'TORCH_NCCL_ASYNC_ERROR_HANDLING': '***', 'TORCH_NCCL_DEBUG_INFO_TEMP_FILE': '***', 'TORCH_NCCL_DUMP_ON_TIMEOUT': '***', 'TORCH_NCCL_TRACE_BUFFER_SIZE': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}} (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) INFO 06-12 05:21:05 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=65536. (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) INFO 06-12 05:21:05 [vllm.py:690] Asynchronous scheduling is enabled. (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) WARNING 06-12 05:21:05 [vllm.py:728] Enforce eager set, overriding optimization level to -O0 (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) INFO 06-12 05:21:05 [vllm.py:846] Cudagraph is disabled under eager mode (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) WARNING 06-12 05:21:06 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1 (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) WARNING 06-12 05:21:06 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) WARNING 06-12 05:21:11 [arg_utils.py:1256] The global random seed is set to 86. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM. (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) INFO 06-12 05:21:11 [model.py:529] Resolved architecture: Qwen3ForCausalLM (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) INFO 06-12 05:21:11 [model.py:1549] Using max model len 32768 (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) INFO 06-12 05:21:11 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'NCCL_DEBUG': '***', 'NCCL_DEBUG_SUBSYS': '***', 'NCCL_SOCKET_FAMILY': '***', 'NCCL_SOCKET_IFNAME': '***', 'RAY_ADDRESS': '***', 'RAY_USE_UVLOOP': '***', 'SKYRL_WORKER_NCCL_TIMEOUT_IN_S': '***', 'TORCH_FR_BUFFER_SIZE': '***', 'TORCH_NCCL_ASYNC_ERROR_HANDLING': '***', 'TORCH_NCCL_DEBUG_INFO_TEMP_FILE': '***', 'TORCH_NCCL_DUMP_ON_TIMEOUT': '***', 'TORCH_NCCL_TRACE_BUFFER_SIZE': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}} (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) INFO 06-12 05:21:11 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=65536. (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) INFO 06-12 05:21:11 [vllm.py:690] Asynchronous scheduling is enabled. (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) WARNING 06-12 05:21:11 [vllm.py:728] Enforce eager set, overriding optimization level to -O0 (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) INFO 06-12 05:21:11 [vllm.py:846] Cudagraph is disabled under eager mode (pid=332746, ip=10.128.17.43) ⚙️ Running in WANDB offline mode (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) (EngineCore_DP0 pid=379913) INFO 06-12 05:21:11 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION']. [repeated 16x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) WARNING 06-12 05:21:11 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1 (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) WARNING 06-12 05:21:11 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) (EngineCore_DP0 pid=2080799) INFO 06-12 05:21:13 [worker_base.py:289] Injected into for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'read_named_weights', 'set_numa_affinity', 'test_rpc'] [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) (EngineCore_DP0 pid=593303) INFO 06-12 05:21:13 [parallel_state.py:1234] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.128.34.1:54597 backend=nccl [repeated 11x across cluster] (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) (EngineCore_DP0 pid=593303) INFO 06-12 05:21:13 [parallel_state.py:1445] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A [repeated 11x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) (EngineCore_DP0 pid=632092) INFO 06-12 05:21:12 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', speculative_config=None, tokenizer='/e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=76, served_model_name=0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': , 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [65536], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': , 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': , 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []} [repeated 16x across cluster] (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) (EngineCore_DP0 pid=379664) INFO 06-12 05:21:15 [gpu_model_runner.py:4125] Starting to load model /e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6... [repeated 16x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) WARNING 06-12 05:21:11 [arg_utils.py:1256] The global random seed is set to 87. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM. [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:21:11 [model.py:529] Resolved architecture: Qwen3ForCausalLM [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:21:11 [model.py:1549] Using max model len 32768 [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:21:11 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'NCCL_DEBUG': '***', 'NCCL_DEBUG_SUBSYS': '***', 'NCCL_SOCKET_FAMILY': '***', 'NCCL_SOCKET_IFNAME': '***', 'RAY_ADDRESS': '***', 'RAY_USE_UVLOOP': '***', 'SKYRL_WORKER_NCCL_TIMEOUT_IN_S': '***', 'TORCH_FR_BUFFER_SIZE': '***', 'TORCH_NCCL_ASYNC_ERROR_HANDLING': '***', 'TORCH_NCCL_DEBUG_INFO_TEMP_FILE': '***', 'TORCH_NCCL_DUMP_ON_TIMEOUT': '***', 'TORCH_NCCL_TRACE_BUFFER_SIZE': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}} [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:21:11 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=65536. [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:21:11 [vllm.py:690] Asynchronous scheduling is enabled. [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) WARNING 06-12 05:21:11 [vllm.py:728] Enforce eager set, overriding optimization level to -O0 [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:21:11 [vllm.py:846] Cudagraph is disabled under eager mode [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) (EngineCore_DP0 pid=379664) INFO 06-12 05:21:15 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION']. [repeated 15x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) WARNING 06-12 05:21:11 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1 [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) WARNING 06-12 05:21:11 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned [repeated 3x across cluster] (pid=332822, ip=10.128.17.43) ⚙️ Running in WANDB offline mode (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) (EngineCore_DP0 pid=389428) INFO 06-12 05:21:18 [worker_base.py:289] Injected into for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'read_named_weights', 'set_numa_affinity', 'test_rpc'] [repeated 15x across cluster] (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) (EngineCore_DP0 pid=388640) INFO 06-12 05:21:17 [parallel_state.py:1234] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.128.34.8:58309 backend=nccl [repeated 16x across cluster] (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) (EngineCore_DP0 pid=388640) INFO 06-12 05:21:17 [parallel_state.py:1445] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A [repeated 16x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) (EngineCore_DP0 pid=389432) INFO 06-12 05:21:15 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', speculative_config=None, tokenizer='/e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=83, served_model_name=0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': , 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [65536], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': , 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': , 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []} [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) (EngineCore_DP0 pid=389432) INFO 06-12 05:21:20 [gpu_model_runner.py:4125] Starting to load model /e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6... [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) (EngineCore_DP0 pid=309070) INFO 06-12 05:21:21 [default_loader.py:293] Loading weights took 14.45 seconds (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) (EngineCore_DP0 pid=309070) INFO 06-12 05:21:21 [gpu_model_runner.py:4222] Model loading took 15.27 GiB memory and 15.801878 seconds (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) (EngineCore_DP0 pid=389423) INFO 06-12 05:21:21 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION']. [repeated 12x across cluster] (pid=487915) ⚙️ Running in WANDB offline mode [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) (EngineCore_DP0 pid=389432) INFO 06-12 05:21:19 [worker_base.py:289] Injected into for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'read_named_weights', 'set_numa_affinity', 'test_rpc'] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) (EngineCore_DP0 pid=389432) INFO 06-12 05:21:20 [parallel_state.py:1234] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.128.17.63:35683 backend=nccl [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) (EngineCore_DP0 pid=389432) INFO 06-12 05:21:20 [parallel_state.py:1445] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) (EngineCore_DP0 pid=386894) INFO 06-12 05:21:21 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', speculative_config=None, tokenizer='/e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=87, served_model_name=0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': , 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [65536], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': , 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': , 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []} [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) (EngineCore_DP0 pid=309080) INFO 06-12 05:21:24 [gpu_worker.py:373] Available KV cache memory: 49.32 GiB (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) (EngineCore_DP0 pid=309080) INFO 06-12 05:21:24 [kv_cache_utils.py:1307] GPU KV cache size: 359,120 tokens (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) (EngineCore_DP0 pid=309080) INFO 06-12 05:21:24 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 10.96x (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) (EngineCore_DP0 pid=309070) INFO 06-12 05:21:24 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled. (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) (EngineCore_DP0 pid=593278) INFO 06-12 05:21:24 [core.py:278] init engine (profile, create kv cache, warmup model) took 3.36 seconds (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) (EngineCore_DP0 pid=593278) WARNING 06-12 05:21:25 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1 (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) (EngineCore_DP0 pid=593278) INFO 06-12 05:21:25 [vllm.py:690] Asynchronous scheduling is enabled. (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) (EngineCore_DP0 pid=593278) WARNING 06-12 05:21:25 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored. (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) (EngineCore_DP0 pid=593278) INFO 06-12 05:21:25 [vllm.py:846] Cudagraph is disabled under eager mode (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) WARNING 06-12 05:21:25 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`. (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) (EngineCore_DP0 pid=2062583) INFO 06-12 05:21:25 [default_loader.py:293] Loading weights took 15.64 seconds [repeated 23x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) (EngineCore_DP0 pid=336339) INFO 06-12 05:21:26 [gpu_model_runner.py:4222] Model loading took 15.27 GiB memory and 14.384757 seconds [repeated 23x across cluster] (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) (EngineCore_DP0 pid=2062603) INFO 06-12 05:21:28 [gpu_worker.py:373] Available KV cache memory: 49.32 GiB [repeated 23x across cluster] (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) (EngineCore_DP0 pid=2062603) INFO 06-12 05:21:28 [kv_cache_utils.py:1307] GPU KV cache size: 359,120 tokens [repeated 23x across cluster] (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) (EngineCore_DP0 pid=2062603) INFO 06-12 05:21:28 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 10.96x [repeated 23x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) (EngineCore_DP0 pid=2062587) INFO 06-12 05:21:28 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled. [repeated 23x across cluster] (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) (EngineCore_DP0 pid=2062603) INFO 06-12 05:21:28 [core.py:278] init engine (profile, create kv cache, warmup model) took 2.86 seconds [repeated 23x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) (EngineCore_DP0 pid=386885) INFO 06-12 05:21:29 [worker_base.py:289] Injected into for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'read_named_weights', 'set_numa_affinity', 'test_rpc'] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) (EngineCore_DP0 pid=386885) INFO 06-12 05:21:30 [parallel_state.py:1234] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.128.34.6:41369 backend=nccl (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) (EngineCore_DP0 pid=2062603) WARNING 06-12 05:21:29 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1 [repeated 23x across cluster] (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) (EngineCore_DP0 pid=2062603) INFO 06-12 05:21:29 [vllm.py:690] Asynchronous scheduling is enabled. [repeated 23x across cluster] (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) (EngineCore_DP0 pid=2062603) WARNING 06-12 05:21:29 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored. [repeated 23x across cluster] (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) (EngineCore_DP0 pid=2062603) INFO 06-12 05:21:29 [vllm.py:846] Cudagraph is disabled under eager mode [repeated 23x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) (EngineCore_DP0 pid=386885) INFO 06-12 05:21:30 [parallel_state.py:1445] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) WARNING 06-12 05:21:29 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`. [repeated 23x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) (EngineCore_DP0 pid=2080778) INFO 06-12 05:21:29 [default_loader.py:293] Loading weights took 13.42 seconds [repeated 8x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) (EngineCore_DP0 pid=2080778) INFO 06-12 05:21:29 [gpu_model_runner.py:4222] Model loading took 15.27 GiB memory and 14.216579 seconds [repeated 8x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) (EngineCore_DP0 pid=386885) INFO 06-12 05:21:32 [gpu_model_runner.py:4125] Starting to load model /e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6... (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) (EngineCore_DP0 pid=386885) INFO 06-12 05:21:34 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION']. (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) (EngineCore_DP0 pid=2080788) INFO 06-12 05:21:32 [gpu_worker.py:373] Available KV cache memory: 49.32 GiB [repeated 8x across cluster] (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) (EngineCore_DP0 pid=2080788) INFO 06-12 05:21:32 [kv_cache_utils.py:1307] GPU KV cache size: 359,120 tokens [repeated 8x across cluster] (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) (EngineCore_DP0 pid=2080788) INFO 06-12 05:21:32 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 10.96x [repeated 8x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) (EngineCore_DP0 pid=2080778) INFO 06-12 05:21:33 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled. [repeated 8x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) (EngineCore_DP0 pid=2080778) INFO 06-12 05:21:33 [core.py:278] init engine (profile, create kv cache, warmup model) took 3.34 seconds [repeated 8x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) (EngineCore_DP0 pid=386887) INFO 06-12 05:21:29 [worker_base.py:289] Injected into for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'read_named_weights', 'set_numa_affinity', 'test_rpc'] [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) (EngineCore_DP0 pid=386887) INFO 06-12 05:21:30 [parallel_state.py:1234] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.128.34.6:47819 backend=nccl [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) (EngineCore_DP0 pid=2080778) WARNING 06-12 05:21:33 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1 [repeated 8x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) (EngineCore_DP0 pid=2080778) INFO 06-12 05:21:33 [vllm.py:690] Asynchronous scheduling is enabled. [repeated 8x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) (EngineCore_DP0 pid=2080778) WARNING 06-12 05:21:33 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored. [repeated 8x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) (EngineCore_DP0 pid=2080778) INFO 06-12 05:21:33 [vllm.py:846] Cudagraph is disabled under eager mode [repeated 8x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) (EngineCore_DP0 pid=386887) INFO 06-12 05:21:30 [parallel_state.py:1445] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) WARNING 06-12 05:21:33 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`. [repeated 8x across cluster] (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) (EngineCore_DP0 pid=389444) INFO 06-12 05:21:35 [default_loader.py:293] Loading weights took 13.39 seconds [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) (EngineCore_DP0 pid=389444) INFO 06-12 05:21:35 [gpu_model_runner.py:4222] Model loading took 15.27 GiB memory and 14.582477 seconds [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) (EngineCore_DP0 pid=386887) INFO 06-12 05:21:32 [gpu_model_runner.py:4125] Starting to load model /e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6... [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO cudaDriverVersion 13020 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO NCCL version 2.27.7+cuda13.0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO NET/Plugin: Could not find: libnccl-net.so. (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO NET/IB : Using [0]mlx5_0:1/IB [1]mlx5_1:1/IB [2]mlx5_2:1/IB [3]mlx5_3:1/IB [RO]; OOB ib0:10.128.17.43<0> (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Initialized NET plugin IB (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Assigned NET plugin IB to comm (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Using network IB (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO DMA-BUF is available on GPU device 0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO ncclCommInitRankConfig comm 0x400e426b3c30 rank 0 nranks 8 cudaDev 0 nvmlDev 0 busId 901000 commId 0xb3c94e79582d448a - Init START (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO RAS client listening socket at 127.0.0.1<28028> (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO NCCL_NET_GDR_LEVEL set by environment to LOC (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Setting affinity for GPU 0 to 0-71 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO NVLS multicast support is not available on dev 0 (NVLS_NCHANNELS 0) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO comm 0x400e426b3c30 rank 0 nRanks 8 nNodes 2 localRanks 4 localRank 0 MNNVL 0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Channel 00/08 : 0 1 2 3 7 6 5 4 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Channel 01/08 : 0 4 5 6 7 3 2 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Channel 02/08 : 0 3 1 5 4 7 6 2 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Channel 03/08 : 0 3 2 6 4 7 5 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Channel 04/08 : 0 1 2 3 7 6 5 4 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Channel 05/08 : 0 4 5 6 7 3 2 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Channel 06/08 : 0 3 1 5 4 7 6 2 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Channel 07/08 : 0 3 2 6 4 7 5 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Trees [0] 1/4/-1->0->-1 [1] -1/-1/-1->0->3 [2] 1/-1/-1->0->2 [3] 2/-1/-1->0->3 [4] 1/-1/-1->0->4 [5] -1/-1/-1->0->3 [6] 1/-1/-1->0->2 [7] 2/-1/-1->0->3 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO P2P Chunksize set to 131072 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO PROFILER/Plugin: Could not find: libnccl-profiler.so. (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Check P2P Type isAllDirectP2p 1 directMode 0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO [Proxy Service] Device 0 CPU core 47 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333099 [0] NCCL INFO [Proxy Service UDS] Device 0 CPU core 48 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Channel 00/0 : 0[0] -> 1[1] via P2P/IPC (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x40118c000000 size 2097152 ipcDesc 0x400ed8008b70 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Channel 04/0 : 0[0] -> 1[1] via P2P/IPC (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x40118c200000 size 2097152 ipcDesc 0x400ed8009b30 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Channel 02/0 : 0[0] -> 3[3] via P2P/IPC (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x40118c400000 size 2097152 ipcDesc 0x400ed800a940 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Channel 03/0 : 0[0] -> 3[3] via P2P/IPC (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x40118c600000 size 2097152 ipcDesc 0x400ed800b730 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Channel 06/0 : 0[0] -> 3[3] via P2P/IPC (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x40118c800000 size 2097152 ipcDesc 0x400ed800c500 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Channel 07/0 : 0[0] -> 3[3] via P2P/IPC (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x40118ca00000 size 2097152 ipcDesc 0x400ed800d2d0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333103 [0] NCCL INFO [Proxy Progress] Device 0 CPU core 49 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Channel 00/0 : 4[0] -> 0[0] [receive] via NET/IB/0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Channel 04/0 : 4[0] -> 0[0] [receive] via NET/IB/0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Channel 01/0 : 0[0] -> 4[0] [send] via NET/IB/0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Channel 05/0 : 0[0] -> 4[0] [send] via NET/IB/0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x40118cc00000 size 10485760 ipcDesc 0x400ed8010180 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x40118d600000 size 10485760 ipcDesc 0x400ed8010f50 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x40118e000000 size 10485760 ipcDesc 0x400ed8010f50 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x40118ea00000 size 10485760 ipcDesc 0x400ed8010f50 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x40118f400000 size 10485760 ipcDesc 0x400ed8010f50 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x40118fe00000 size 10485760 ipcDesc 0x400ed8010f50 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 10485760 ptr 0x401190800000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 10485760 ptr 0x401191200000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 10485760 ptr 0x401191c00000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 10485760 ptr 0x401192600000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 10485760 ptr 0x401193000000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 10485760 ptr 0x401193a00000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 2097152 ptr 0x401194400000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 2097152 ptr 0x401194600000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 2097152 ptr 0x401194800000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 2097152 ptr 0x401194a00000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 2097152 ptr 0x401194c00000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 2097152 ptr 0x401194e00000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Connected all rings, use ring PXN 0 GDR 0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Channel 02/0 : 0[0] -> 1[1] via P2P/IPC (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x401195000000 size 2097152 ipcDesc 0x400ed8015c30 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Channel 06/0 : 0[0] -> 1[1] via P2P/IPC (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x401195200000 size 2097152 ipcDesc 0x400ed8015c30 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Channel 02/0 : 0[0] -> 2[2] via P2P/IPC (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x401195400000 size 2097152 ipcDesc 0x400ed8015c30 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Channel 03/0 : 0[0] -> 2[2] via P2P/IPC (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x401195600000 size 2097152 ipcDesc 0x400ed8015c30 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Channel 06/0 : 0[0] -> 2[2] via P2P/IPC (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x401195800000 size 2097152 ipcDesc 0x400ed8015c30 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Channel 07/0 : 0[0] -> 2[2] via P2P/IPC (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x401195a00000 size 2097152 ipcDesc 0x400ed8015c30 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Channel 01/0 : 0[0] -> 3[3] via P2P/IPC (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x401195c00000 size 2097152 ipcDesc 0x400ed8015c30 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Channel 05/0 : 0[0] -> 3[3] via P2P/IPC (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x401195e00000 size 2097152 ipcDesc 0x400ed8015c30 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Channel 00/0 : 0[0] -> 4[0] [send] via NET/IB/0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Channel 04/0 : 0[0] -> 4[0] [send] via NET/IB/0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x401196000000 size 10485760 ipcDesc 0x400ed8015c30 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x401196a00000 size 10485760 ipcDesc 0x400ed8015c30 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x401197400000 size 10485760 ipcDesc 0x400ed8015c30 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x401197e00000 size 10485760 ipcDesc 0x400ed8015c30 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x401198800000 size 10485760 ipcDesc 0x400ed8015c30 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x401199200000 size 10485760 ipcDesc 0x400ed8015c30 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x401199c00000 size 10485760 ipcDesc 0x400ed8015c30 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x40119a600000 size 10485760 ipcDesc 0x400ed8015c30 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x40119b000000 size 10485760 ipcDesc 0x400ed8015c30 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333097 [0] NCCL INFO Allocated shareable buffer 0x40119ba00000 size 10485760 ipcDesc 0x400ed8015c30 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 10485760 ptr 0x40119c400000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 10485760 ptr 0x40119ce00000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 10485760 ptr 0x40119d800000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 10485760 ptr 0x40119e200000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 10485760 ptr 0x40119ec00000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 10485760 ptr 0x40119f600000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 10485760 ptr 0x4011c0000000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 10485760 ptr 0x4011c0a00000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 2097152 ptr 0x4011c1400000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 2097152 ptr 0x4011c1600000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 2097152 ptr 0x4011c1800000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 2097152 ptr 0x4011c1a00000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 2097152 ptr 0x4011c1c00000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 2097152 ptr 0x4011c1e00000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 2097152 ptr 0x4011c2000000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 2097152 ptr 0x4011c2200000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 2097152 ptr 0x4011c2400000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Imported shareable buffer device 0 size 2097152 ptr 0x4011c2600000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Connected all trees (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO threadThresholds 8/8/64 | 64/8/64 | 512 | 512 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO 8 coll channels, 8 collnet channels, 0 nvls channels, 8 p2p channels, 2 p2p channels per peer (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO CC Off, workFifoBytes 1048576 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO TUNER/Plugin: Could not find: libnccl-tuner.so. Using internal tuner plugin. (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO ncclCommInitRankConfig comm 0x400e426b3c30 rank 0 nranks 8 cudaDev 0 nvmlDev 0 busId 901000 commId 0xb3c94e79582d448a - Init COMPLETE (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333065 [0] NCCL INFO Init timings - ncclCommInitRankConfig: rank 0 nranks 8 total 2.74 (kernels 0.32, alloc 1.76, bootstrap 0.02, allgathers 0.00, topo 0.02, graphs 0.01, connections 0.59, rest 0.00) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Broadcast: opCount 0 sendbuff 0x400ec0000000 recvbuff 0x400ec0000000 count 622329856 datatype 9 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-00 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:4 (FSDPPolicyWorkerBase pid=487913) jpbo (FSDPPolicyWorkerBase pid=487915) jpbo- (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 00/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 01/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 02/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 03/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 04/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 05/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 06/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 07/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 08/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 09/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 10/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 11/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 12/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 13/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 14/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 15/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 16/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 17/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 18/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 19/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 20/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 21/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 22/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 23/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 24/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 25/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 26/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 27/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 28/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 29/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 30/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 31/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 32/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 33/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 34/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 35/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 36/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 37/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 38/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Channel 39/40 : 0 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Check P2P Type isAllDirectP2p 0 directMode 0 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333112 [0] NCCL INFO Channel 01/0 : 0[2] -> 1[2] [sen (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333110 [0] NCCL INFO Channel 01/0 : 0[1] -> 1[1] [send] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333109 [0] NCCL INFO Channel 01/0 : 0[3] -> 1[3] [se (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) NET/IB/0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO Connected binomial trees (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333111 [0] NCCL INFO CC Off, workFifoBytes 1048576 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-00 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) nd] via NET/IB/3 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) d] via NET/IB/2 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333133 [0] NCCL INFO Allocated shareable (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) via NET/IB/1 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 00c00ed30 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333130 [0] NCCL INFO Alloc (FSDPPolicyWorkerBase pid=487914) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400f60000000 count 155582464 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400ec0000000 count 4194304 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x40108f000000 count 1048576 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x40108fa00000 count 1048576 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x40108f200000 count 4194304 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400eb9e00400 count 32 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400eb9e00200 count 32 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400ec6800000 count 12582912 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400ec0800000 count 12582912 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400ec2000000 count 12582912 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Broadcast: opCount a sendbuff 0x400eb9e00600 recvbuff 0x400eb9e00600 count 4096 datatype 9 op 0 root 0 comm 0x400e615aff00 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400eb9e02600 count 1024 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Broadcast: opCount b sendbuff 0x400eb9e02e00 recvbuff 0x400eb9e02e00 count 4096 datatype 9 op 0 root 0 comm 0x400e615aff00 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400eb9e00600 count 1024 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Broadcast: opCount c sendbuff 0x400ec3800000 recvbuff 0x400ec3800000 count 16777216 datatype 9 op 0 root 0 comm 0x400e615aff00 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400ec5800000 count 4194304 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Broadcast: opCount d sendbuff 0x400ec6000000 recvbuff 0x400ec6000000 count 4194304 datatype 9 op 0 root 0 comm 0x400e615aff00 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x40108fc00000 count 1048576 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Broadcast: opCount a sendbuff 0x40108fe00000 recvbuff 0x40108fe00000 count 4194304 datatype 9 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Broadcast: opCount e sendbuff 0x40108fe00000 recvbuff 0x40108fe00000 count 4194304 datatype 9 op 0 root 0 comm 0x400e615aff00 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400ec6000000 count 1048576 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Broadcast: opCount b sendbuff 0x400ec3800000 recvbuff 0x400ec3800000 count 16777216 datatype 9 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Broadcast: opCount f sendbuff 0x400ec3800000 recvbuff 0x400ec3800000 count 16777216 datatype 9 op 0 root 0 comm 0x400e615aff00 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x40108fe00000 count 4194304 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Broadcast: opCount c sendbuff 0x400eb9e00e00 recvbuff 0x400eb9e00e00 count 128 datatype 9 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400eb9e01000 count 32 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Broadcast: opCount c sendbuff 0x400eb9e01200 recvbuff 0x400eb9e01200 count 128 datatype 9 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400eb9e00e00 count 32 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Broadcast: opCount c sendbuff 0x400ec8000000 recvbuff 0x400ec8000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400ec3800000 count 12582912 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Broadcast: opCount d sendbuff 0x400ece000000 recvbuff 0x400ece000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400ec8000000 count 12582912 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Broadcast: opCount e sendbuff 0x400ed4000000 recvbuff 0x400ed4000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400ec9800000 count 12582912 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Broadcast: opCount f sendbuff 0x400eb9e02e00 recvbuff 0x400eb9e02e00 count 4096 datatype 9 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400eb9e01200 count 1024 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Broadcast: opCount f sendbuff 0x400eb9e04e00 recvbuff 0x400eb9e04e00 count 4096 datatype 9 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400eb9e01a00 count 1024 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Broadcast: opCount f sendbuff 0x400ecb000000 recvbuff 0x400ecb000000 count 16777216 datatype 9 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400ec5000000 count 4194304 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x401090e00000 count 1048576 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400ec6200000 count 1048576 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45 (FSDPPolicyWorkerBase pid=487912) c00000 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x4013528c0000 count 155582464 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401365180000 count 155582464 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401377a40000 count 155582464 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401100800000 count 4194304 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401101000000 count 4194304 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401101800000 count 4194304 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401116a00000 count 1048576 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401116c00000 count 1048576 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401116e00000 count 1048576 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401183a00000 count 1048576 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401183c00000 count 1048576 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401183e00000 count 1048576 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401100800000 count 4194304 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401101000000 count 4194304 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401101800000 count 4194304 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa02a40 count 32 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa02a80 count 32 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa02ac0 count 32 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa02e40 count 32 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa02e80 count 32 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa02ec0 count 32 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401109000000 count 12582912 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40110a800000 count 12582912 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40110c000000 count 12582912 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401135400000 count 12582912 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401136c00000 count 12582912 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401138400000 count 12582912 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401109000000 count 12582912 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40110a800000 count 12582912 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40110c000000 count 12582912 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa00800 count 1024 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa01000 count 1024 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa01800 count 1024 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa31c00 count 1024 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa32400 count 1024 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa32c00 count 1024 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401135c00000 count 4194304 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401136400000 count 4194304 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401136c00000 count 4194304 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401103200000 count 1048576 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401103400000 count 1048576 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401103600000 count 1048576 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401103a00000 count 1048576 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401103c00000 count 1048576 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401103e00000 count 1048576 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401135c00000 count 4194304 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401136400000 count 4194304 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401136c00000 count 4194304 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa03e40 count 32 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa03e80 count 32 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa03ec0 count 32 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa04240 count 32 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa04280 count 32 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa042c0 count 32 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401109000000 count 12582912 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40110a800000 count 12582912 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40110c000000 count 12582912 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40113cc00000 count 12582912 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40113e400000 count 12582912 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40113fc00000 count 12582912 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401109000000 count 12582912 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40110a800000 count 12582912 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40110c000000 count 12582912 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa00800 count 1024 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa01000 count 1024 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa01800 count 1024 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa32400 count 1024 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa32c00 count 1024 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa33400 count 1024 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401108000000 count 4194304 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401108800000 count 4194304 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401109000000 count 4194304 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401109a00000 count 1048576 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401109c00000 count 1048576 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401109e00000 count 1048576 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401107c00000 count 1048576 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401107e00000 count 1048576 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401108000000 count 1048576 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401108c00000 count 4194304 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401109400000 count 4194304 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401109c00000 count 4194304 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa04240 count 32 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa04280 count 32 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa042c0 count 32 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa00a40 count 32 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa00a80 count 32 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa00ac0 count 32 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40113cc00000 count 12582912 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40113e400000 count 12582912 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40113fc00000 count 12582912 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401341800000 count 12582912 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401343000000 count 12582912 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401344800000 count 12582912 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40113cc00000 count 12582912 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40113e400000 count 12582912 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40113fc00000 count 12582912 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa32400 count 1024 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa32c00 count 1024 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa33400 count 1024 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa34400 count 1024 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa34c00 count 1024 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa35400 count 1024 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40113bc00000 count 4194304 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40113c400000 count 4194304 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40113cc00000 count 4194304 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40113d600000 count 1048576 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40113d800000 count 1048576 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40113da00000 count 1048576 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40113b600000 count 1048576 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40113b800000 count 1048576 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40113ba00000 count 1048576 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40113c400000 count 4194304 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40113cc00000 count 4194304 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40113d400000 count 4194304 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa01a40 count 32 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa01a80 count 32 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa01ac0 count 32 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa01e40 count 32 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa01e80 count 32 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa01ec0 count 32 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401341800000 count 12582912 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401343000000 count 12582912 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401344800000 count 12582912 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401347800000 count 12582912 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401349000000 count 12582912 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134a800000 count 12582912 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401341800000 count 12582912 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401343000000 count 12582912 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401344800000 count 12582912 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa32400 count 1024 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa32c00 count 1024 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa33400 count 1024 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa34c00 count 1024 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa35400 count 1024 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa35c00 count 1024 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401340800000 count 4194304 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401341000000 count 4194304 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401341800000 count 4194304 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401140e00000 count 1048576 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401141000000 count 1048576 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401141200000 count 1048576 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401340400000 count 1048576 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401340600000 count 1048576 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401340800000 count 1048576 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401341200000 count 4194304 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401341a00000 count 4194304 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401342200000 count 4194304 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa01e40 count 32 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa01e80 count 32 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa01ec0 count 32 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa32640 count 32 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa32680 count 32 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa326c0 count 32 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401342200000 count 12582912 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401343a00000 count 12582912 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401345200000 count 12582912 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401349a00000 count 12582912 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134b200000 count 12582912 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134ca00000 count 12582912 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134fa00000 count 12582912 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401351200000 count 12582912 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401352a00000 count 12582912 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa34c00 count 1024 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa35400 count 1024 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa35c00 count 1024 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa36c00 count 1024 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa37400 count 1024 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa37c00 count 1024 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401344200000 count 4194304 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401344a00000 count 4194304 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401345200000 count 4194304 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401346400000 count 1048576 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401346600000 count 1048576 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401346800000 count 1048576 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401343c00000 count 1048576 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401343e00000 count 1048576 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401344000000 count 1048576 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401348a00000 count 4194304 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401349200000 count 4194304 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401349a00000 count 4194304 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa33640 count 32 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa33680 count 32 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa336c0 count 32 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa33a40 count 32 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa33a80 count 32 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa33ac0 count 32 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401349a00000 count 12582912 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134b200000 count 12582912 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134ca00000 count 12582912 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134fa00000 count 12582912 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401351200000 count 12582912 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401352a00000 count 12582912 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401355a00000 count 12582912 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401357200000 count 12582912 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401358a00000 count 12582912 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa34c00 count 1024 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa35400 count 1024 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa35c00 count 1024 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa37400 count 1024 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa37c00 count 1024 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa38400 count 1024 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134ba00000 count 4194304 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134c200000 count 4194304 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134ca00000 count 4194304 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO (FSDPPolicyWorkerBase pid=487913) 40100c00ed30 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 1f sendbuff 0x400ed9000000 recvbuff 0x400ed (FSDPPolicyWorkerBase pid=487915) 30 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO Recv: opCount 0 sendbuff (ni (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Broadcast: opCount 37 sendbuff 0x400eb9e07e00 recvbuff 0x400eb9e07e00 c (FSDPPolicyWorkerBase pid=487912) Broadcast: opCount 44 sendbuff 0x40134d200000 recvbuff 0x40134d200000 count 4194304 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134d400000 count 1048576 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134d600000 count 1048576 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134d800000 count 1048576 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134b400000 count 1048576 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134b600000 count 1048576 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134b800000 count 1048576 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134c400000 count 4194304 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134cc00000 count 4194304 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134d400000 count 4194304 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa33a40 count 32 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa33a80 count 32 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa33ac0 count 32 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa34e40 count 32 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa34e80 count 32 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa34ec0 count 32 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134d400000 count 12582912 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134ec00000 count 12582912 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401350400000 count 12582912 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401354c00000 count 12582912 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401356400000 count 12582912 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401357c00000 count 12582912 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40135ac00000 count 12582912 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40135c400000 count 12582912 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40135dc00000 count 12582912 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa37400 count 1024 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa37c00 count 1024 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa38400 count 1024 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa39400 count 1024 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa39c00 count 1024 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa3a400 count 1024 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134f400000 count 4194304 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134fc00000 count 4194304 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401350400000 count 4194304 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401351600000 count 1048576 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401351800000 count 1048576 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401351a00000 count 1048576 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134f000000 count 1048576 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134f200000 count 1048576 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134f400000 count 1048576 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401353c00000 count 4194304 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401354400000 count 4194304 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401354c00000 count 4194304 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa35e40 count 32 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa35e80 count 32 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa35ec0 count 32 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa36240 count 32 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa36280 count 32 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa362c0 count 32 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401354c00000 count 12582912 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401356400000 count 12582912 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401357c00000 count 12582912 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40135c400000 count 12582912 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40135dc00000 count 12582912 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40135f400000 count 12582912 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401362400000 count 12582912 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401363c00000 count 12582912 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401365400000 count 12582912 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa37400 count 1024 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa37c00 count 1024 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa38400 count 1024 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa39c00 count 1024 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa3a400 count 1024 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa3ac00 count 1024 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401356c00000 count 4194304 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401357400000 count 4194304 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401357c00000 count 4194304 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401358e00000 count 1048576 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401359000000 count 1048576 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401359200000 count 1048576 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134f800000 count 1048576 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134fa00000 count 1048576 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x40134fc00000 count 1048576 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401356c00000 count 4194304 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401357400000 count 4194304 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401357c00000 count 4194304 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa36240 count 32 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa36280 count 32 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa362c0 count 32 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa37640 count 32 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa37680 count 32 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa376c0 count 32 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-4 (FSDPPolicyWorkerBase pid=487913) 9000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpb (FSDPPolicyWorkerBase pid=487915) l) recvbuff 0x400ece800000 count 12582912 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO Broadcast: opCount 37 sendbuff 0x400eb9e07e00 recvbuff 0x400eb9e07e00 count 4 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount 8 sendbuff 0x400e9fa0 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) buffer 0x336400000 size 10485760 ipcDesc 0x400fe000ed30 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400e79e0b200 count 1024 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400eb9e0b200 count 1024 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nra (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) ated shareable buffer 0x32e600000 size 2097152 ipcDesc 0x40100c00ed30 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO Broadcast: opCo (FSDPPolicyWorkerBase pid=487914) ount 4096 datatype 9 op 0 root 0 comm 0x400e615aff00 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Broadcast: opCount 2e sendbuff 0x400ee1c00000 recvbuff 0x400ee1c00000 count 16777216 datatype 9 op 0 root 0 comm 0x400e (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) nks=4] stream (nil) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:3328 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x4013 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) unt 73 sendbuff 0x400eb9e0c200 recvbuff 0x400eb9e0c200 count 128 datatype 9 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400efbe00000 count 1258 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount (FSDPPolicyWorkerBase pid=487915) 096 datatype 9 op 0 root 0 comm 0x400e195a79c0 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff ( (FSDPPolicyWorkerBase pid=487913) 4b sendbuff 0x400eef400000 recvbuff 0x400eef400000 count 50331648 datatype 9 op 0 root 0 comm 0x400e195a95c0 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbu (FSDPPolicyWorkerBase pid=487914) recvbuff 0x400eb9e0a600 count 1024 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) nil) recvbuff 0x401379600000 count 12582912 datatype 9 op 0 root 2 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401373600000 count 4194304 datatype 9 op 0 root 2 comm 0x400dfaa (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:48 (FSDPPolicyWorkerBase pid=487915) ff 0x400eb9e0a600 count 1024 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Broadcast: opCount 53 sendbuff 0x401091000000 recvbuff 0x401091000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e612e3300 [nranks=8] st (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount aa sendbuff 0x400e9fa3fe00 recvbuff 0x400e9fa3fe00 count 128 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount ab sendbuff 0x400e9fa40200 recvbuff 0x400e9fa40200 count 128 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount ac sendbuff 0x401372600000 recvbuff 0x401372600000 count 50331648 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount ad sendbuff 0x40160f000000 recvbuff 0x40160f000000 count 50331648 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount ae sendbuff 0x401372600000 recvbuff 0x401372600000 count 50331648 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount af sendbuff 0x400e9fa40c00 recvbuff 0x400e9fa40c00 count 4096 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO Broadcast: opCount 53 sendbuff 0x401091000000 recvbuff 0x401091000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream ( (FSDPPolicyWorkerBase pid=487914) ream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Broadc (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount ba sendbuff 0x400e9fa43400 recvbuff 0x400e9fa43400 count 4096 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount bb sendbuff 0x400e9fa45400 recvbuff 0x400e9fa45400 count 4096 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount bc sendbuff 0x401372600000 recvbuff 0x401372600000 count 16777216 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount bd sendbuff 0x401374600000 recvbuff 0x401374600000 count 4194304 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount be sendbuff 0x401372800000 recvbuff 0x401372800000 count 4194304 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount bf sendbuff 0x401373200000 recvbuff 0x401373200000 count 16777216 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount ca sendbuff 0x4015c5800000 recvbuff 0x4015c5800000 count 16777216 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount cb sendbuff 0x400e9fa42a00 recvbuff 0x400e9fa42a00 count 128 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 95 sendbuff 0x400eb9e10200 recvbuff 0x400eb9e10200 count 128 datatype (FSDPPolicyWorkerBase pid=487915) nil) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbu (FSDPPolicyWorkerBase pid=487914) ast: opCount 5c sendbuff 0x401091000000 recvbuff 0x401091000000 count 16777216 datatype 9 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount cc sendbuff 0x400e9fa43e00 recvbuff 0x400e9fa43e00 count 128 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount cd sendbuff 0x401615000000 recvbuff 0x401615000000 count 50331648 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount ce sendbuff 0x4015cb000000 recvbuff 0x4015cb000000 count 50331648 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount cf sendbuff 0x401615000000 recvbuff 0x401615000000 count 50331648 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount da sendbuff 0x401615000000 recvbuff 0x401615000000 count 50331648 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount db sendbuff 0x400e9fa45c00 recvbuff 0x400e9fa45c00 count 4096 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount dc sendbuff 0x400e9fa48400 recvbuff 0x400e9fa48400 count 4096 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount dd sendbuff 0x4015d1000000 recvbuff 0x4015d1000000 count 16777216 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487913) 9 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x401074000000 count 12582912 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4 (FSDPPolicyWorkerBase pid=487915) ff 0x40109c000000 count 4194304 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO Broadcast: opCount ad sendbuff 0x401056c00000 recvbuff 0x401056c00000 count 50331648 datatype 9 op 0 ro (FSDPPolicyWorkerBase pid=487914) jp (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount de sendbuff 0x4015d3000000 recvbuff 0x4015d3000000 count 4194304 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount df sendbuff 0x4015d1000000 recvbuff 0x4015d1000000 count 4194304 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount ea sendbuff 0x4015d7000000 recvbuff 0x4015d7000000 count 4194304 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount eb sendbuff 0x4015d7800000 recvbuff 0x4015d7800000 count 16777216 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount ec sendbuff 0x400e9fa47600 recvbuff 0x400e9fa47600 count 128 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount ed sendbuff 0x400e9fa47a00 recvbuff 0x400e9fa47a00 count 128 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount ee sendbuff 0x4015d7000000 recvbuff 0x4015d7000000 count 50331648 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount ef sendbuff 0x4013a0000000 recvbuff 0x4013a0000000 count 50331648 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487913) ] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount c1 sendbuff 0x400eb9e15200 rec (FSDPPolicyWorkerBase pid=487915) ot 0 comm 0x400e195a79c0 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487915) jpbo-006 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400fa000000 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount fa sendbuff 0x4013a6000000 recvbuff 0x4013a6000000 count 50331648 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount fb sendbuff 0x4013a0000000 recvbuff 0x4013a0000000 count 50331648 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount fc sendbuff 0x400e9fa4ac00 recvbuff 0x400e9fa4ac00 count 4096 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount fd sendbuff 0x400e9fa4cc00 recvbuff 0x400e9fa4cc00 count 4096 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount fe sendbuff 0x4013a0000000 recvbuff 0x4013a0000000 count 16777216 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount ff sendbuff 0x4015db800000 recvbuff 0x4015db800000 count 4194304 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCount 10b sendbuff 0x4013ab000000 recvbuff 0x4013ab000000 count 4194304 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (n (FSDPPolicyWorkerBase pid=487913) vbuff 0x400eb9e15200 count 128 datatype 9 op 0 root 0 comm 0x400e195a95c0 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount d0 sendbuff 0x400eb9e16e00 recvbuff 0x400eb9e16e00 count 4096 datatype 9 o (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400fa0000000 coun (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Broadcast: opCount e8 sendbuff 0x400faf000000 recvbuff 0x400faf000000 count 16777216 datatype 9 op 0 root 0 co (FSDPPolicyWorkerBase pid=487912) il) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) rec (FSDPPolicyWorkerBase pid=487913) p 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487915) t 12582912 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487914) mm 0x400e615aff00 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x4013bf800000 count 12582912 datatype 9 op 0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Recv: opCo (FSDPPolicyWorkerBase pid=487915) jpbo-006-4 (FSDPPolicyWorkerBase pid=487913) unt 0 sendbuff (nil) recvbuff 0x400eb9e19e00 count 32 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount fc sendbuff 0x400eb9e1be00 recvbuf (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO Broadcast: opCount a8 se (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400 (FSDPPolicyWorkerBase pid=487913) f 0x400eb9e1be00 count 4096 datatype 9 op 0 root 0 comm 0x400e195a95c0 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 10b sendbuff 0x4010ab800000 recvbuff 0x4010ab800000 count 4194304 dat (FSDPPolicyWorkerBase pid=487915) ndbuff 0x4010c0000000 recvbuff 0x4010c0000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa4f240 count 32 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (ni (FSDPPolicyWorkerBase pid=487913) atype 9 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO Broadcast: opCount 123 sendbuff 0x400eb9e20200 recvbuff 0x400eb9e20200 count 128 da (FSDPPolicyWorkerBase pid=487915) tatype 9 op 0 root 0 comm 0x400e195a79c0 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:4 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCC (FSDPPolicyWorkerBase pid=487913) INFO Broadcast: opCount 146 sendbuff 0x4010b2000000 recvbuff 0x4010b2000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] N (FSDPPolicyWorkerBase pid=487914) L INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400eb9e25400 count 32 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487912) l) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45 (FSDPPolicyWorkerBase pid=487913) CCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400eb9e27e00 count 1024 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCou (FSDPPolicyWorkerBase pid=487915) Recv: opCount 0 sendbuff (nil) recvbuff 0x400eb9e25400 count 32 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487913) nt 163 sendbuff 0x4010b2000000 recvbuff 0x4010b2000000 count 4194304 datatype 9 op 0 root 0 comm 0x400e195a95c0 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO Broadcast: opCount 18a sendbuff 0x401100000000 recvbuff 0x401100000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e195a79c0 [nranks=2] stre (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL IN (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x4010efa00000 count 12582912 datatype 9 op 0 root 0 comm 0x400e196bcf10 [ (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: op (FSDPPolicyWorkerBase pid=487912) opCount 0 sendbuff (nil) recvbuff 0x400e9fa542c0 count 32 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) Count 0 sendbuff (nil) recvbuff 0x400e9fa02ac0 count 32 datatype 9 op 0 root 3 comm 0x400e42a9f600 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount 70 sendbuf (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) f 0x40137bc00000 recvbuff 0x40137bc00000 count 50331648 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount db sendbuff 0x400e9fa46400 recvbuff 0x400e9fa46400 coun (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) (EngineCore_DP0 pid=386887) INFO 06-12 05:21:34 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION']. [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) (EngineCore_DP0 pid=389432) INFO 06-12 05:21:37 [gpu_worker.py:373] Available KV cache memory: 49.32 GiB [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) (EngineCore_DP0 pid=389432) INFO 06-12 05:21:37 [kv_cache_utils.py:1307] GPU KV cache size: 359,120 tokens [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) (EngineCore_DP0 pid=389432) INFO 06-12 05:21:37 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 10.96x [repeated 12x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x4015d1c00000 count 1048576 datatype 9 op 0 ro (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401621c00000 count 4194304 datatype 9 op 0 root 3 comm 0x400e42a9f600 [nranks (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11: (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) (EngineCore_DP0 pid=389432) INFO 06-12 05:21:38 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled. [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) (EngineCore_DP0 pid=389432) INFO 06-12 05:21:38 [core.py:278] init engine (profile, create kv cache, warmup model) took 2.85 seconds [repeated 12x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] Started monitoring (every 120s) (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:21:47] OK: 46 / 131,072 FDs open (0.0% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:21:47] OK: RSS 0.78 GiB | node mem 369.8/858.0 GiB used (43.1%), avail 488.2 GiB (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) (EngineCore_DP0 pid=389432) WARNING 06-12 05:21:38 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1 [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) (EngineCore_DP0 pid=389432) INFO 06-12 05:21:38 [vllm.py:690] Asynchronous scheduling is enabled. [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) (EngineCore_DP0 pid=389432) WARNING 06-12 05:21:38 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored. [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) (EngineCore_DP0 pid=389432) INFO 06-12 05:21:38 [vllm.py:846] Cudagraph is disabled under eager mode [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) WARNING 06-12 05:21:38 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`. [repeated 12x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO cudaDriverVersion 13020 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO NCCL version 2.27.7+cuda13.0 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488213 [0] NCCL INFO NET/Plugin: Could not find: libnccl-net.so.  [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488213 [0] NCCL INFO NET/IB : Using [0]mlx5_0:1/IB [1]mlx5_1:1/IB [2]mlx5_2:1/IB [3]mlx5_3:1/IB [RO]; OOB ib0:10.128.17.29<0> [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488213 [0] NCCL INFO Initialized NET plugin IB [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333128 [0] NCCL INFO Assigned NET plugin IB to comm [repeated 23x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333128 [0] NCCL INFO Using network IB [repeated 23x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333128 [0] NCCL INFO DMA-BUF is available on GPU device 0 [repeated 23x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333128 [0] NCCL INFO ncclCommInitRankConfig comm 0x400e156cd8d0 rank 1 nranks 4 cudaDev 0 nvmlDev 1 busId 1901000 commId 0x7adbb8983205823a - Init START [repeated 23x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488213 [0] NCCL INFO RAS client listening socket at 127.0.0.1<28028> [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488213 [0] NCCL INFO NCCL_NET_GDR_LEVEL set by environment to LOC [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333128 [0] NCCL INFO Setting affinity for GPU 1 to 72-143 [repeated 23x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333128 [0] NCCL INFO NVLS multicast support is not available on dev 0 (NVLS_NCHANNELS 0) [repeated 15x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333128 [0] NCCL INFO comm 0x400e156cd8d0 rank 1 nRanks 4 nNodes 1 localRanks 4 localRank 1 MNNVL 0 [repeated 23x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488270 [0] NCCL INFO Channel 23/24 : 0 3 2 1 [repeated 168x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333128 [0] NCCL INFO Trees [0] 2/-1/-1->1->0 [1] 2/-1/-1->1->0 [2] 2/-1/-1->1->0 [3] 2/-1/-1->1->0 [4] 3/-1/-1->1->2 [5] 3/-1/-1->1->2 [6] 3/-1/-1->1->2 [7] 3/-1/-1->1->2 [8] 0/-1/-1->1->-1 [9] 0/-1/-1->1->-1 [10] 0/-1/-1->1->-1 [11] 0/-1/-1->1->-1 [12] 2/-1/-1->1->0 [13] 2/-1/-1->1->0 [14] 2/-1/-1->1->0 [15] 2/-1/-1->1->0 [16] 3/-1/-1->1->2 [17] 3/-1/-1->1->2 [18] 3/-1/-1->1->2 [19] 3/-1/-1->1->2 [20] 0/-1/-1->1->-1 [21] 0/-1/-1->1->-1 [22] 0/-1/-1->1->-1 [23] 0/-1/-1->1->-1 [repeated 23x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333128 [0] NCCL INFO P2P Chunksize set to 524288 [repeated 23x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488213 [0] NCCL INFO PROFILER/Plugin: Could not find: libnccl-profiler.so.  [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488270 [0] NCCL INFO Check P2P Type isAllDirectP2p 1 directMode 0 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333132 [0] NCCL INFO [Proxy Service] Device 0 CPU core 75 [repeated 23x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333135 [0] NCCL INFO [Proxy Service UDS] Device 0 CPU core 77 [repeated 23x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333127 [0] NCCL INFO Channel 22/0 : 3[3] -> 2[2] via P2P/IPC [repeated 586x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333130 [0] NCCL INFO Allocated shareable buffer 0x40103a200000 size 10485760 ipcDesc 0x40100c00ed30 [repeated 1168x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333138 [0] NCCL INFO [Proxy Progress] Device 0 CPU core 237 [repeated 23x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333109 [0] NCCL INFO Channel 39/0 : 1[3] -> 0[3] [receive] via NET/IB/3 [repeated 346x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333110 [0] NCCL INFO Channel 39/0 : 0[1] -> 1[1] [send] via NET/IB/1 [repeated 341x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333142 [0] NCCL INFO Imported shareable buffer device 0 size 2097152 ptr 0x40103ca00000 [repeated 1170x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488270 [0] NCCL INFO Connected all rings, use ring PXN 0 GDR 1 [repeated 23x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333127 [0] NCCL INFO Connected all trees [repeated 23x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333127 [0] NCCL INFO threadThresholds 8/8/64 | 32/8/64 | 512 | 512 [repeated 23x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333127 [0] NCCL INFO 24 coll channels, 24 collnet channels, 0 nvls channels, 32 p2p channels, 16 p2p channels per peer [repeated 23x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333126 [0] NCCL INFO CC Off, workFifoBytes 1048576 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488213 [0] NCCL INFO TUNER/Plugin: Could not find: libnccl-tuner.so. Using internal tuner plugin. [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333127 [0] NCCL INFO ncclCommInitRankConfig comm 0x400e196bcf10 rank 3 nranks 4 cudaDev 0 nvmlDev 3 busId 3901000 commId 0x7adbb8983205823a - Init COMPLETE [repeated 23x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333127 [0] NCCL INFO Init timings - ncclCommInitRankConfig: rank 3 nranks 4 total 0.17 (kernels 0.00, alloc 0.00, bootstrap 0.00, allgathers 0.00, topo 0.02, graphs 0.01, connections 0.14, rest 0.00) [repeated 23x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount 186 sendbuff 0x400e9fa02a00 recvbuff 0x400e9fa02a00 count 128 datatype 9 op 0 root 0 comm 0x400e42984290 [nranks=2] stream (nil) [repeated 5451x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333110 [0] NCCL INFO Connected binomial trees [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x4010ee200000 count 12582912 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream (nil) [repeated 2302x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount a sendbuff 0x401103000000 recvbuff 0x401103000000 count 4194304 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 14x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount b sendbuff 0x401135400000 recvbuff 0x401135400000 count 16777216 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 14x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount c sendbuff 0x401107800000 recvbuff 0x401107800000 count 50331648 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 24x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount d sendbuff 0x401107800000 recvbuff 0x401107800000 count 50331648 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 14x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount e sendbuff 0x401107800000 recvbuff 0x401107800000 count 50331648 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 14x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount f sendbuff 0x401107800000 recvbuff 0x401107800000 count 16777216 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 24x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa02ac0 count 32 datatype 9 op 0 root 3 comm 0x400e42a9f600 [nranks=4] stream (nil) [repeated 2057x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 172 sendbuff 0x4010e0000000 recvbuff [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487913) 0x4010e0000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Broadcast: opCount e9 sendbuff 0x4010c4000000 recvbuff 0x4010c4000000 count 4 [repeated 6x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Broadcast: opCount 15e sendbuff 0x [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Broadcast: opCo [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount f4 sendbuff 0x401630800000 recvbuff 0x401630800000 count 50331648 datatype 9 op 0 root 0 comm 0x400e426b3c30 [repeated 10x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) [nranks=8] stream (nil) [repeated 16x across cluster] (FSDPPolicyWorkerBase pid=487912) unt f0 sendbuff 0x4015d7000000 recvbuff 0x4015d7000000 count 50331648 datatype 9 op 0 root 0 comm 0x400dfa982e60 [nranks=2] stream (nil) [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Broadcast: opCount (FSDPPolicyWorkerBase pid=487914) 194304 datatype 9 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=487914) a8 sendbuff 0x4010c0000000 recvbuff 0x4010c0000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) 4010c0000000 recvbuff 0x4010c0000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e615aff00 [nranks=2] stream (nil) [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x4013ac400000 count 1048576 datatype 9 op 0 root 2 comm 0x400e4 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount aa sendbuff 0x4013a0000000 recvbuff 0x4013a0000000 count 4194304 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 15x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount ab sendbuff 0x4013a0000000 recvbuff 0x4013a0000000 count 4194304 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 15x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount ac sendbuff 0x4013a0000000 recvbuff 0x4013a0000000 count 16777216 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 15x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount ad sendbuff 0x4013a6000000 recvbuff 0x4013a6000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 26x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount ae sendbuff 0x4013a6000000 recvbuff 0x4013a6000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 15x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount af sendbuff 0x4013a6000000 recvbuff 0x4013a6000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 15x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount ba sendbuff 0x4013ac000000 recvbuff 0x4013ac000000 count 16777216 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 15x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount bb sendbuff 0x4013ac000000 recvbuff 0x4013ac000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 27x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount bc sendbuff 0x4013ac000000 recvbuff 0x4013ac000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 15x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount bd sendbuff 0x4013ac000000 recvbuff 0x4013ac000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 15x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount be sendbuff 0x4013ac000000 recvbuff 0x4013ac000000 count 16777216 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 27x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount bf sendbuff 0x4013ac000000 recvbuff 0x4013ac000000 count 4194304 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 15x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount ca sendbuff 0x4013b5000000 recvbuff 0x4013b5000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 15x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount cb sendbuff 0x4013b5000000 recvbuff 0x4013b5000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 15x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount cc sendbuff 0x4013b5000000 recvbuff 0x4013b5000000 count 16777216 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 27x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount cd sendbuff 0x4013b5000000 recvbuff 0x4013b5000000 count 4194304 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 15x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount ce sendbuff 0x4013b5000000 recvbuff 0x4013b5000000 count 4194304 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 15x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount cf sendbuff 0x4013b5c00000 recvbuff 0x4013b5c00000 count 16777216 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 15x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount da sendbuff 0x401620000000 recvbuff 0x401620000000 count 16777216 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 27x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount db sendbuff 0x4013c6800000 recvbuff 0x4013c6800000 count 4194304 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 14x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount dc sendbuff 0x4013c6800000 recvbuff 0x4013c6800000 count 4194304 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 15x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount dd sendbuff 0x401620400000 recvbuff 0x401620400000 count 16777216 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 15x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400eb9e2a400 count 32 datatype 9 op 0 root 0 comm 0x400e196c2 (FSDPPolicyWorkerBase pid=487913) jp (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount de sendbuff 0x401626000000 recvbuff 0x401626000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 27x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount df sendbuff 0x401626000000 recvbuff 0x401626000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 15x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount ea sendbuff 0x401626000000 recvbuff 0x401626000000 count 4194304 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 15x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount eb sendbuff 0x401628800000 recvbuff 0x401628800000 count 16777216 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 15x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount ec sendbuff 0x401630800000 recvbuff 0x401630800000 count 50331648 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 27x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount ed sendbuff 0x401630800000 recvbuff 0x401630800000 count 50331648 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 15x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount ee sendbuff 0x401630800000 recvbuff 0x401630800000 count 50331648 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 15x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount ef sendbuff 0x401630800000 recvbuff 0x401630800000 count 16777216 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 27x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) ot 1 comm 0x400e42a9f600 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400eb9e2c [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount fa sendbuff 0x400e9fa02a00 recvbuff 0x400e9fa02a00 count 128 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 20x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount fb sendbuff 0x4015d1a00000 recvbuff 0x4015d1a00000 count 50331648 datatype 9 op 0 root 0 comm 0x400e42984290 [nranks=2] stream (nil) [repeated 12x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount fc sendbuff 0x400e9fa4b400 recvbuff 0x400e9fa4b400 count 4096 datatype 9 op 0 root 0 comm 0x400e42984290 [nranks=2] stream (nil) [repeated 10x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount fd sendbuff 0x400e9fa4b400 recvbuff 0x400e9fa4b400 count 4096 datatype 9 op 0 root 0 comm 0x400e42984290 [nranks=2] stream (nil) [repeated 9x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount fe sendbuff 0x4015d1a00000 recvbuff 0x4015d1a00000 count 16777216 datatype 9 op 0 root 0 comm 0x400e42984290 [nranks=2] stream (nil) [repeated 9x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount ff sendbuff 0x4015d1a00000 recvbuff 0x4015d1a00000 count 4194304 datatype 9 op 0 root 0 comm 0x400e42984290 [nranks=2] stream (nil) [repeated 9x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 137 sendbuff 0x4010c7800000 recvbuff 0x4010c7800000 count 4194304 datatype 9 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (n (FSDPPolicyWorkerBase pid=487912) vbuff 0x400e9fa4dc00 count 1024 datatype 9 op 0 root 1 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487913) il) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) t 4096 datatype 9 op 0 root 0 comm 0x400e42984290 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] N (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) (EngineCore_DP0 pid=386885) INFO 06-12 05:21:49 [default_loader.py:293] Loading weights took 14.41 seconds (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) (EngineCore_DP0 pid=386885) INFO 06-12 05:21:49 [gpu_model_runner.py:4222] Model loading took 15.27 GiB memory and 16.847375 seconds (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) (EngineCore_DP0 pid=386886) INFO 06-12 05:21:51 [gpu_worker.py:373] Available KV cache memory: 49.32 GiB (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) (EngineCore_DP0 pid=386886) INFO 06-12 05:21:51 [kv_cache_utils.py:1307] GPU KV cache size: 359,120 tokens (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) (EngineCore_DP0 pid=386886) INFO 06-12 05:21:51 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 10.96x (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) (EngineCore_DP0 pid=386885) INFO 06-12 05:21:52 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled. (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) (EngineCore_DP0 pid=386894) INFO 06-12 05:21:52 [core.py:278] init engine (profile, create kv cache, warmup model) took 2.81 seconds (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) (EngineCore_DP0 pid=386885) WARNING 06-12 05:21:52 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1 (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) (EngineCore_DP0 pid=386885) INFO 06-12 05:21:52 [vllm.py:690] Asynchronous scheduling is enabled. (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) (EngineCore_DP0 pid=386885) WARNING 06-12 05:21:52 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored. (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) (EngineCore_DP0 pid=386885) INFO 06-12 05:21:52 [vllm.py:846] Cudagraph is disabled under eager mode (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) (EngineCore_DP0 pid=386894) WARNING 06-12 05:21:52 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1 (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) (EngineCore_DP0 pid=386894) INFO 06-12 05:21:52 [vllm.py:690] Asynchronous scheduling is enabled. (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) (EngineCore_DP0 pid=386894) WARNING 06-12 05:21:52 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored. (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) (EngineCore_DP0 pid=386894) INFO 06-12 05:21:52 [vllm.py:846] Cudagraph is disabled under eager mode (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) WARNING 06-12 05:21:53 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`. (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] Started monitoring (every 120s) (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:21:57] OK: 46 / 131,072 FDs open (0.0% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:21:57] OK: RSS 0.78 GiB | node mem 375.4/858.0 GiB used (43.8%), avail 482.6 GiB (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) (EngineCore_DP0 pid=386887) INFO 06-12 05:21:49 [default_loader.py:293] Loading weights took 14.41 seconds [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) (EngineCore_DP0 pid=386887) INFO 06-12 05:21:49 [gpu_model_runner.py:4222] Model loading took 15.27 GiB memory and 16.850045 seconds [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) (EngineCore_DP0 pid=386894) INFO 06-12 05:21:51 [gpu_worker.py:373] Available KV cache memory: 49.32 GiB [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) (EngineCore_DP0 pid=386894) INFO 06-12 05:21:51 [kv_cache_utils.py:1307] GPU KV cache size: 359,120 tokens [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) (EngineCore_DP0 pid=386894) INFO 06-12 05:21:51 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 10.96x [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) (EngineCore_DP0 pid=386887) INFO 06-12 05:21:52 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled. [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) (EngineCore_DP0 pid=386885) INFO 06-12 05:21:52 [core.py:278] init engine (profile, create kv cache, warmup model) took 2.84 seconds [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) (EngineCore_DP0 pid=386887) WARNING 06-12 05:21:52 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1 [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) (EngineCore_DP0 pid=386887) INFO 06-12 05:21:52 [vllm.py:690] Asynchronous scheduling is enabled. [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) (EngineCore_DP0 pid=386887) WARNING 06-12 05:21:52 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored. [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) (EngineCore_DP0 pid=386887) INFO 06-12 05:21:52 [vllm.py:846] Cudagraph is disabled under eager mode [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) WARNING 06-12 05:21:53 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`. [repeated 3x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] Started monitoring (every 120s) (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] [05:22:05] OK: 46 / 131,072 FDs open (0.0% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] [05:22:05] OK: RSS 0.78 GiB | node mem 387.7/858.0 GiB used (45.2%), avail 470.3 GiB (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] Started monitoring (every 120s) (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] [05:22:14] OK: 46 / 131,072 FDs open (0.0% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] [05:22:14] OK: RSS 0.78 GiB | node mem 379.2/858.0 GiB used (44.2%), avail 478.8 GiB (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) [rank-0]: Loading model from /e/data1/datasets/playground/ot-baf/explore-tis-minp/explore-tis-minp/checkpoints/global_step_81/policy/model_world_size_8_rank_0.pt (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) [rank-0]: Loading extra_state from /e/data1/datasets/playground/ot-baf/explore-tis-minp/explore-tis-minp/checkpoints/global_step_81/policy/extra_state_world_size_8_rank_0.pt (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) [rank-0]: Loading optim from /e/data1/datasets/playground/ot-baf/explore-tis-minp/explore-tis-minp/checkpoints/global_step_81/policy/optim_world_size_8_rank_0.pt (skyrl_entrypoint pid=487747) [fd-monitor] [05:22:20] OK: 102 / 131,072 FDs open (0.1% of soft limit, hard limit: 131,072) (skyrl_entrypoint pid=487747) [fd-monitor] [05:22:20] OK: RSS 1.64 GiB | node mem 229.9/858.0 GiB used (26.8%), avail 628.1 GiB (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) [rank-0]: Successfully loaded model state dict (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) [rank-0]: Successfully loaded optimizer state (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) [rank-0]: Successfully loaded scheduler state (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) [rank-0]: Checkpoint loaded successfully from /e/data1/datasets/playground/ot-baf/explore-tis-minp/explore-tis-minp/checkpoints/global_step_81/policy (FSDPPolicyWorkerBase pid=487913) e00 count 1024 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 181 sendbuff 0x400eb9e2d600 recvbuff 0x400eb9e2d600 count 4096 datatype 9 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 181 sendbuff 0x400eb9e2d600 recvbuff 0x400eb9e2d600 count 4096 datatype 9 op 0 root 0 comm 0x400e195a95c0 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400eb9e2ae00 count 1024 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 182 sendbuff 0x4010ec000000 recvbuff 0x4010ec000000 count 16777216 datatype 9 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 182 sendbuff 0x4010ec000000 recvbuff 0x4010ec000000 count 16777216 datatype 9 op 0 root 0 comm 0x400e195a95c0 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x4010e5000000 count 4194304 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 183 sendbuff 0x4010e5800000 recvbuff 0x4010e5800000 count 4194304 datatype 9 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 183 sendbuff 0x4010e5800000 recvbuff 0x4010e5800000 count 4194304 datatype 9 op 0 root 0 comm 0x400e195a95c0 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x4010ec000000 count 1048576 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 184 sendbuff 0x4010ec200000 recvbuff 0x4010ec200000 count 4194304 datatype 9 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 184 sendbuff 0x4010ec200000 recvbuff 0x4010ec200000 count 4194304 datatype 9 op 0 root 0 comm 0x400e195a95c0 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x4010e5800000 count 1048576 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 185 sendbuff 0x4010eca00000 recvbuff 0x4010eca00000 count 16777216 datatype 9 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 185 sendbuff 0x4010eca00000 recvbuff 0x4010eca00000 count 16777216 datatype 9 op 0 root 0 comm 0x400e195a95c0 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x4010ec200000 count 4194304 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 186 sendbuff 0x400eb9e2b600 recvbuff 0x400eb9e2b600 count 128 datatype 9 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 186 sendbuff 0x400eb9e2b600 recvbuff 0x400eb9e2b600 count 128 datatype 9 op 0 root 0 comm 0x400e195a95c0 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400eb9e2b800 count 32 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 187 sendbuff 0x400eb9e2ba00 recvbuff 0x400eb9e2ba00 count 128 datatype 9 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 187 sendbuff 0x400eb9e2ba00 recvbuff 0x400eb9e2ba00 count 128 datatype 9 op 0 root 0 comm 0x400e195a95c0 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400eb9e2b600 count 32 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 188 sendbuff 0x4010f2000000 recvbuff 0x4010f2000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 188 sendbuff 0x4010f2000000 recvbuff 0x4010f2000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e195a95c0 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x4010eca00000 count 12582912 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 189 sendbuff 0x4010f8000000 recvbuff 0x4010f8000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 189 sendbuff 0x4010f8000000 recvbuff 0x4010f8000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e195a95c0 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x4010ee200000 count 12582912 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 18a sendbuff 0x4010f2000000 recvbuff 0x4010f2000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 18a sendbuff 0x4010f2000000 recvbuff 0x4010f2000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e195a95c0 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x4010efa00000 count 12582912 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 18b sendbuff 0x400eb9e2d600 recvbuff 0x400eb9e2d600 count 4096 datatype 9 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 18b sendbuff 0x400eb9e2d600 recvbuff 0x400eb9e2d600 count 4096 datatype 9 op 0 root 0 comm 0x400e195a95c0 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400eb9e2ba00 count 1024 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 18c sendbuff 0x400eb9e2f600 recvbuff 0x400eb9e2f600 count 4096 datatype 9 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 18c sendbuff 0x400eb9e2f600 recvbuff 0x400eb9e2f600 count 4096 datatype 9 op 0 root 0 comm 0x400e195a95c0 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400eb9e2c200 count 1024 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 18d sendbuff 0x400eb9e2d600 recvbuff 0x400eb9e2d600 count 4096 datatype 9 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 18d sendbuff 0x400eb9e2d600 recvbuff 0x400eb9e2d600 count 4096 datatype 9 op 0 root 0 comm 0x400e195a95c0 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x400eb9e2f600 count 1024 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 18e sendbuff 0x401100000000 recvbuff 0x401100000000 count 622329856 datatype 9 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Broadcast: opCount 18e sendbuff 0x401100000000 recvbuff 0x401100000000 count 622329856 datatype 9 op 0 root 0 comm 0x400e195a95c0 [nranks=2] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x40114a400000 count 155582464 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa02a40 count 32 datatype 9 op 0 root 1 comm 0x400e42a9f600 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa02a80 count 32 datatype 9 op 0 root 2 comm 0x400e42a9f600 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa02ac0 count 32 datatype 9 op 0 root 3 comm 0x400e42a9f600 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount fa sendbuff 0x401640000000 recvbuff 0x401640000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401641800000 count 12582912 datatype 9 op 0 root 1 comm 0x400e42a9f600 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401643000000 count 12582912 datatype 9 op 0 root 2 comm 0x400e42a9f600 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401644800000 count 12582912 datatype 9 op 0 root 3 comm 0x400e42a9f600 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount fb sendbuff 0x401640000000 recvbuff 0x401640000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401641800000 count 12582912 datatype 9 op 0 root 1 comm 0x400e42a9f600 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401643000000 count 12582912 datatype 9 op 0 root 2 comm 0x400e42a9f600 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401644800000 count 12582912 datatype 9 op 0 root 3 comm 0x400e42a9f600 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount fc sendbuff 0x401640000000 recvbuff 0x401640000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401641800000 count 12582912 datatype 9 op 0 root 1 comm 0x400e42a9f600 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401643000000 count 12582912 datatype 9 op 0 root 2 comm 0x400e42a9f600 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401644800000 count 12582912 datatype 9 op 0 root 3 comm 0x400e42a9f600 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount fd sendbuff 0x400e9fa5a400 recvbuff 0x400e9fa5a400 count 4096 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa5ac00 count 1024 datatype 9 op 0 root 1 comm 0x400e42a9f600 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa5b400 count 1024 datatype 9 op 0 root 2 comm 0x400e42a9f600 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa5bc00 count 1024 datatype 9 op 0 root 3 comm 0x400e42a9f600 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount fd sendbuff 0x400e9fa5a400 recvbuff 0x400e9fa5a400 count 4096 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa5ac00 count 1024 datatype 9 op 0 root 1 comm 0x400e42a9f600 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa5b400 count 1024 datatype 9 op 0 root 2 comm 0x400e42a9f600 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa5bc00 count 1024 datatype 9 op 0 root 3 comm 0x400e42a9f600 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount fd sendbuff 0x400e9fa5a400 recvbuff 0x400e9fa5a400 count 4096 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa5ac00 count 1024 datatype 9 op 0 root 1 comm 0x400e42a9f600 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa5b400 count 1024 datatype 9 op 0 root 2 comm 0x400e42a9f600 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x400e9fa5bc00 count 1024 datatype 9 op 0 root 3 comm 0x400e42a9f600 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Broadcast: opCount fd sendbuff 0x401660000000 recvbuff 0x401660000000 count 622329856 datatype 9 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x4016728c0000 count 155582464 datatype 9 op 0 root 1 comm 0x400e42a9f600 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401685180000 count 155582464 datatype 9 op 0 root 2 comm 0x400e42a9f600 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401697a40000 count 155582464 datatype 9 op 0 root 3 comm 0x400e42a9f600 [nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2073787 [0] NCCL INFO cudaDriverVersion 13020 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2073787 [0] NCCL INFO NCCL version 2.27.7+cuda13.0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO NET/Plugin: Could not find: libnccl-net.so. (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO NET/IB : Using [0]mlx5_0:1/IB [1]mlx5_1:1/IB [2]mlx5_2:1/IB [3]mlx5_3:1/IB [RO]; OOB ib0:10.128.17.223<0> (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Initialized NET plugin IB (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Assigned NET plugin IB to comm (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Using network IB (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO DMA-BUF is available on GPU device 0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO ncclCommInitRankConfig comm 0xaaab37522a20 rank 2 nranks 49 cudaDev 0 nvmlDev 0 busId 901000 commId 0x9b0018c00ef5dafe - Init START (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO RAS client listening socket at 127.0.0.1<28028> (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO NCCL_NET_GDR_LEVEL set by environment to LOC (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Setting affinity for GPU 0 to 0-71 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO NVLS multicast support is not available on dev 0 (NVLS_NCHANNELS 0) (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO comm 0xaaab37522a20 rank 2 nRanks 49 nNodes 13 localRanks 4 localRank 0 MNNVL 0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Trees [0] 3/1/5->2->9 [1] -1/-1/-1->2->13 [2] 3/-1/-1->2->4 [3] 4/-1/-1->2->13 [4] 3/-1/-1->2->5 [5] -1/-1/-1->2->13 [6] 3/-1/-1->2->4 [7] 4/-1/-1->2->13 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO P2P Chunksize set to 131072 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO PROFILER/Plugin: Could not find: libnccl-profiler.so. (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074072 [0] NCCL INFO [Proxy Service] Device 0 CPU core 16 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074075 [0] NCCL INFO [Proxy Service UDS] Device 0 CPU core 17 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074077 [0] NCCL INFO [Proxy Progress] Device 0 CPU core 18 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Channel 00/0 : 1[0] -> 2[0] [receive] via NET/IB/0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Channel 04/0 : 1[0] -> 2[0] [receive] via NET/IB/0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Channel 00/0 : 2[0] -> 3[1] via P2P/IPC (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074072 [0] NCCL INFO Allocated shareable buffer 0x4017b0800000 size 2097152 ipcDesc 0x40097800a120 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Channel 04/0 : 2[0] -> 3[1] via P2P/IPC (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074072 [0] NCCL INFO Allocated shareable buffer 0x4017b0a00000 size 2097152 ipcDesc 0x40097800af80 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Channel 01/0 : 2[0] -> 5[0] [send] via NET/IB/0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Channel 05/0 : 2[0] -> 5[0] [send] via NET/IB/0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Channel 02/0 : 2[0] -> 13[3] via P2P/IPC (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074072 [0] NCCL INFO Allocated shareable buffer 0x4017b0c00000 size 2097152 ipcDesc 0x40097800ca30 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Channel 03/0 : 2[0] -> 13[3] via P2P/IPC (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074072 [0] NCCL INFO Allocated shareable buffer 0x4017b0e00000 size 2097152 ipcDesc 0x40097800d800 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Channel 06/0 : 2[0] -> 13[3] via P2P/IPC (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074072 [0] NCCL INFO Allocated shareable buffer 0x4017b1000000 size 2097152 ipcDesc 0x40097800e5d0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Channel 07/0 : 2[0] -> 13[3] via P2P/IPC (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074072 [0] NCCL INFO Allocated shareable buffer 0x4017b1200000 size 2097152 ipcDesc 0x40097800f3a0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074072 [0] NCCL INFO Allocated shareable buffer 0x4017b1400000 size 10485760 ipcDesc 0x400978010170 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074072 [0] NCCL INFO Allocated shareable buffer 0x4017b1e00000 size 10485760 ipcDesc 0x400978010f40 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074072 [0] NCCL INFO Allocated shareable buffer 0x4017b2800000 size 10485760 ipcDesc 0x400978010f40 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074072 [0] NCCL INFO Allocated shareable buffer 0x4017b3200000 size 10485760 ipcDesc 0x400978010f40 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074072 [0] NCCL INFO Allocated shareable buffer 0x4017b3c00000 size 10485760 ipcDesc 0x400978010f40 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074072 [0] NCCL INFO Allocated shareable buffer 0x4017b4600000 size 10485760 ipcDesc 0x400978010f40 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Imported shareable buffer device 0 size 10485760 ptr 0x4017b5000000 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Imported shareable buffer device 0 size 10485760 ptr 0x4017b5a00000 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Imported shareable buffer device 0 size 10485760 ptr 0x4017b6400000 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Imported shareable buffer device 0 size 10485760 ptr 0x4017b6e00000 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Imported shareable buffer device 0 size 10485760 ptr 0x4017b7800000 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Imported shareable buffer device 0 size 10485760 ptr 0x4017b8200000 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Imported shareable buffer device 0 size 2097152 ptr 0x4017b8c00000 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Imported shareable buffer device 0 size 2097152 ptr 0x4017b8e00000 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Imported shareable buffer device 0 size 2097152 ptr 0x4017b9000000 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Imported shareable buffer device 0 size 2097152 ptr 0x4017b9200000 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Imported shareable buffer device 0 size 2097152 ptr 0x4017b9400000 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Imported shareable buffer device 0 size 2097152 ptr 0x4017b9600000 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Connected all rings, use ring PXN 0 GDR 0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Channel 02/0 : 2[0] -> 3[1] via P2P/IPC (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074072 [0] NCCL INFO Allocated shareable buffer 0x4017b9800000 size 2097152 ipcDesc 0x4009780bedc0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Channel 06/0 : 2[0] -> 3[1] via P2P/IPC (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074072 [0] NCCL INFO Allocated shareable buffer 0x4017b9a00000 size 2097152 ipcDesc 0x4009780bedc0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Channel 02/0 : 2[0] -> 4[2] via P2P/IPC (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074072 [0] NCCL INFO Allocated shareable buffer 0x4017b9c00000 size 2097152 ipcDesc 0x4009780bedc0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Channel 03/0 : 2[0] -> 4[2] via P2P/IPC (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074072 [0] NCCL INFO Allocated shareable buffer 0x4017b9e00000 size 2097152 ipcDesc 0x4009780bedc0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Channel 06/0 : 2[0] -> 4[2] via P2P/IPC (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074072 [0] NCCL INFO Allocated shareable buffer 0x4017ba000000 size 2097152 ipcDesc 0x4009780bedc0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Channel 07/0 : 2[0] -> 4[2] via P2P/IPC (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074072 [0] NCCL INFO Allocated shareable buffer 0x4017ba200000 size 2097152 ipcDesc 0x4009780bedc0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Channel 00/0 : 2[0] -> 5[0] [send] via NET/IB/0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Channel 04/0 : 2[0] -> 5[0] [send] via NET/IB/0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Channel 00/0 : 2[0] -> 9[0] [send] via NET/IB/0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Channel 01/0 : 2[0] -> 13[3] via P2P/IPC (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074072 [0] NCCL INFO Allocated shareable buffer 0x4017ba400000 size 2097152 ipcDesc 0x4009780bf8c0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Channel 05/0 : 2[0] -> 13[3] via P2P/IPC (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074072 [0] NCCL INFO Allocated shareable buffer 0x4017ba600000 size 2097152 ipcDesc 0x4009780bf8c0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074072 [0] NCCL INFO Allocated shareable buffer 0x4017ba800000 size 10485760 ipcDesc 0x4009780bf8c0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074072 [0] NCCL INFO Allocated shareable buffer 0x4017bb200000 size 10485760 ipcDesc 0x4009780bf8c0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074072 [0] NCCL INFO Allocated shareable buffer 0x4017bbc00000 size 10485760 ipcDesc 0x4009780bf8c0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074072 [0] NCCL INFO Allocated shareable buffer 0x4017bc600000 size 10485760 ipcDesc 0x4009780bf8c0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Channel 00/0 : 9[0] -> 2[0] [receive] via NET/IB/0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2074044 [0] NCCL INFO Channel 00/0 : 5[0] -> 2[0] [receive] via NET/IB/0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47: (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) jpbo-010-47:20 (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) jpbo-010-47:2073795:2 (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) jpbo-010-47:2073782:2074042 [0] NCCL INFO Ch (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) jpbo-051 (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) jpbo- (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) jpbo-010-48:2062603:2062952 [0] NCCL INFO Allocated shareable buffer 0x40177ac00000 size 1048576 (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) jpbo-010-48:206 (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) jpbo-010-48:2062587:2062948 (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) jpb (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) jpbo-051 (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) jpbo-051 (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) jpbo-051-37:379664:379941 (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) jpbo-051-37 (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) jpbo-051-37:379657:379910 [0] NCCL INFO Channel 07/0 : 40[3] -> 37[3] [receive] via NET (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) jpbo-010 (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) jpbo-010-46:2080799:2081073 [0] NCCL INFO Allocated shareable buffer 0 (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) jpbo-010-46:2080778:2081044 [0] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) jpbo-007-02:632066:632361 [0] NCCL INFO Allocat (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) jpbo-00 (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) jpbo-051-34 (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) jpbo-051-34:379897:3801 (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) jpbo (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) jpbo-051-38:386885:387145 [0] NCCL INFO Connected all trees (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) jpbo-051-38:386885:387145 [0] NCCL INFO threadThresholds 8/8/64 | 392/8/64 | 512 | 512 (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) jpbo-051-38:386885:387145 [0] NCCL INFO 8 coll channels, 8 collnet channels, 0 nvls channels, 8 p2p channels, 1 p2p channels per peer (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) jpbo-051-38:386885:387145 [0] NCCL INFO TUNER/Plugin: Could not find: libnccl-tuner.so. Using internal tuner plugin. (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) jpbo-051-38:386885:387145 [0] NCCL INFO ncclCommInitRankConfig comm 0xaaab1e9d50e0 rank 45 nranks 49 cudaDev 0 nvmlDev 0 busId 901000 commId 0x9b0018c00ef5dafe - Init COMPLETE (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) jpbo-051-38:386885:387145 [0] NCCL INFO Init timings - ncclCommInitRankConfig: rank 45 nranks 49 total 3.28 (kernels 0.21, alloc 2.12, bootstrap 0.24, allgathers 0.01, topo 0.03, graphs 0.01, connections 0.66, rest 0.00) (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) annel 03/0 : 16[3] -> 13[3] [receive] via NET/IB/3 (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) c33f250 (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) 4833f6a0 (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) 0 ipcDesc 0x40095033f870 (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) 0 (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) 3fd10 (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) 0 (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) /IB/3 (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) x4017fcc00000 size 10485760 ipcDesc 0x40094c33fc90 (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ed shareable buffer 0x40179c600000 size 10485760 ipcDesc 0x40095433fc70 (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) 97433fdb0 (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) 6c33fd90 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount ff sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401820000000 recvbuff 0x401660000000 count 155582464 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7e80 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333171 [0] NCCL INFO Channel 00/08 : 0 16 15 14 1 2 3 4 13 8 7 6 5 9 10 11 12 20 19 18 17 21 22 23 24 40 39 38 25 26 27 28 37 32 31 30 ... (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333171 [0] NCCL INFO Channel 01/08 : 0 1 14 15 16 13 4 3 2 5 6 7 8 12 11 10 9 17 18 19 20 24 23 22 21 25 38 39 40 37 28 27 26 29 30 31 ... (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333171 [0] NCCL INFO Channel 02/08 : 0 14 1 16 15 4 2 13 3 6 5 8 7 11 9 12 10 18 17 20 19 23 21 24 22 38 25 40 39 28 26 37 27 30 29 32 ... (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333171 [0] NCCL INFO Channel 03/08 : 0 15 1 16 14 3 2 13 4 7 5 8 6 10 9 12 11 19 17 20 18 22 21 24 23 39 25 40 38 27 26 37 28 31 29 32 ... (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333171 [0] NCCL INFO Channel 04/08 : 0 16 15 14 1 2 3 4 13 8 7 6 5 9 10 11 12 20 19 18 17 21 22 23 24 40 39 38 25 26 27 28 37 32 31 30 ... (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333171 [0] NCCL INFO Channel 05/08 : 0 1 14 15 16 13 4 3 2 5 6 7 8 12 11 10 9 17 18 19 20 24 23 22 21 25 38 39 40 37 28 27 26 29 30 31 ... (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333171 [0] NCCL INFO Channel 06/08 : 0 14 1 16 15 4 2 13 3 6 5 8 7 11 9 12 10 18 17 20 19 23 21 24 22 38 25 40 39 28 26 37 27 30 29 32 ... (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333171 [0] NCCL INFO Channel 07/08 : 0 15 1 16 14 3 2 13 4 7 5 8 6 10 9 12 11 19 17 20 18 22 21 24 23 39 25 40 38 27 26 37 28 31 29 32 ... (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333171 [0] NCCL INFO Check P2P Type isAllDirectP2p 0 directMode 0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333171 [0] NCCL INFO CC Off, workFifoBytes 1048576 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401635c00000 recvbuff 0x401913c00000 count 4194304 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7e80 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 101 sendbuff 0x400e9fa8 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) FO Broadcast: opCount fa sendbuff 0x400e79e2b600 recvbuff 0x400e79e2b600 count 128 datatype 9 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) INFO Broadcast: opCount fa sendbuff 0x400eb9e2b600 recvbuff 0x400eb9e2b600 count 128 datatype 9 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4012f3e00000 recvbuff 0x401760000000 count 12582912 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] strea (FSDPPolicyWorkerBase pid=487912) CCL INFO Broadcast: opCount fc sendbuff 0x401632000000 recvbuff 0x401632000000 count 50331648 datatype 9 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCo (FSDPPolicyWorkerBase pid=487914) ] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: (FSDPPolicyWorkerBase pid=487915) am (nil) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCou (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 8c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount a sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 4096 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount b sendbuff 0x400e9fa8b000 recvbuff 0x400e9fa8b000 count 4096 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount c sendbuff 0x401913c00000 recvbuff 0x401913c00000 count 16777216 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount d sendbuff 0x401389580000 recvbuff 0x401389580000 count 4194304 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount e sendbuff 0x401913c00000 recvbuff 0x401913c00000 count 4194304 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount f sendbuff 0x401660000000 recvbuff 0x401660000000 count 16777216 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGa (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) nranks=4] stream (nil) (FSDPPolicyWorkerBase pid=487913) 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendb (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) ther: opCount 0 sendbuff 0x400e9fa5b600 recvbuff 0x400e9fa88e00 count 32 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7e80 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) x400df13e9a90 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e89000 recvbuff 0x400e79e89000 count 1 datatype 7 op 0 root 0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) m 0x400e153f9bc0 (FSDPPolicyWorkerBase pid=487914) opCount 0 sendbuff 0x4012f3a00000 recvbuff 0x4017ca400000 count 1048576 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6bd0 (FSDPPolicyWorkerBase pid=487912) unt 108 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487915) nt 0 sendbuff 0x4012f3a00000 recvbuff 0x4017ca400000 count 1048576 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee700 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) jpbo-007-02:632066:632066 [0] NCCL INFO Broadcast: opCount 2a sendbuff 0x401740000000 recvbuff 0x401740000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab2aa92db0 [nran (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) jpbo-007-02:632073:632073 [0] NCCL INFO Broadcast: opCount 1c sendbuff 0x4017c0000000 recvbuff 0x4017c0000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab28dd2810 [nranks=49] s (FSDPPolicyWorkerBase pid=487913) uff 0x400eb9e2ea00 recvbuff 0x400eb9e88e00 count 1024 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f11f0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2073787 [0] NCCL INFO Broadcast: opCount 2a sendbuff 0x401760000000 r (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) jpbo-010-47:2073774:2073774 [0] NCCL INFO Broadcast: opCount 1c sendbuff 0x401720000000 recvbuff (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 80 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 2c sendbuff 0x400e9fa8b000 recvbuff 0x400e9fa8b000 count 4096 datatype 9 op 0 ro (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400e79e32c00 recvbuff 0x400e79e88 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e330 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401820000000 recvbuff 0x401660000000 count 12582912 datatype 9 op (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nra (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) ot 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 13a sendbuff 0x400e9fa88c0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 3a sendbuff 0x401913c00000 recvbuff 0x401913c00000 count 4194304 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) [repeated 1727x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO Recv: opCount 0 sendbuff (nil) recvbuff 0x40114a400000 count 155582464 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream (nil) [repeated 32x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 139 sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 186x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4010f4a0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401309400000 recvbuff 0x4017ca400000 count 4194304 datat (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Send: opCount 0 sendbuff (nil) recvbuff 0x401697a40000 count 155582464 datatype 9 op 0 root 3 comm 0x400dfaa9d5a0 [nranks=4] stream (nil) [repeated 15x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO Broadcast: opCount fa sendbuff 0x401100000000 recvbuff 0x401100000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO Broadcast: opCount fb sendbuff 0x401106000000 recvbuff 0x401106000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO Broadcast: opCount fc sendbuff 0x401100000000 recvbuff 0x401100000000 count 50331648 datatype 9 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO Broadcast: opCount fd sendbuff 0x401120000000 recvbuff 0x401120000000 count 622329856 datatype 9 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 20x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) [repeated 240x across cluster] (FSDPPolicyWorkerBase pid=487912) 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40130ac00000 recvbuff 0x4017ca400000 count 4194304 datatype 9 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 0 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: o (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) c00 count 32 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df13e9a90 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount f (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCoun (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 0000 recvbuff 0x401718c00000 count 1048576 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e87d0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:33 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) ecvbuff 0x401760000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab37522a20 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2073787 [0] NCCL INFO Broadcast: opCount 54 sendbuff 0x401760000000 recvbuff 0x401760000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab37522a20 [ (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) 0x401720000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab46c65e00 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) jpbo-010-47:2073774:2073774 [0] NCCL INFO Broadcast: opCount 36 sendbuff 0x401720000000 recvbuff 0x401720000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab46c65e00 [nranks=4 (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) jpbo-051-33:593278:593278 [0] NCCL INFO Broadcast: opCount 55 sendbuff 0x401740000000 recvbuff 0x401740000000 count 50331648 datatype 9 op 0 root (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) tream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) pCount 0 sendbuff 0x40184e400000 recvbuff 0x401660000000 count 12582912 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7e80 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) e sendbuff 0x400e79e88c00 recvbuff 0x400e79e88c00 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:33303 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) t fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:33 (FSDPPolicyWorkerBase pid=487914) ype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6bd0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88e (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:4 (FSDPPolicyWorkerBase pid=487915) op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee700 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88e00 rec (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) x400e427c7e80 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 65 sendbuff 0x401389580000 recvbuff 0x401389580000 count 4194304 dataty (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:4879 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) pe 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4012e1800000 recvbuff 0x401740000000 count 4194304 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] s (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400faa000000 recvbuff 0x401760000000 count 4194304 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400e9fa67a00 recvbuff 0x400e9fa88e00 count 32 datatype 9 op 0 root 0 (FSDPPolicyWorkerBase pid=487915) vbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCo (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) jpbo-010-48:2062603:2062603 [0] NCCL INFO Broadcast: (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) jpbo-010-48:2062599:2062599 [0] NCCL INFO Broadcast: opCount (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400eb9e3ae00 recvbu (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) jpbo-051-37:379664:379664 [0] NCCL INFO Broadcast: opCount 80 sendbuff 0x401720000000 recvbuff 0x401720000000 count 503 (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) jpbo-051-37:379652:379652 [0] NCCL INFO Broadcast: opCount 52 sendbuff 0x401760000000 recvbuff 0x401760000000 count 50331648 da (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL IN (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) tream 0x400df13e9a90 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e88c00 recvbuff 0x400e79e88c00 count 1 datatype 7 op 0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) ] stream 0x400e153f9bc0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 o (FSDPPolicyWorkerBase pid=487914) : opCount 0 sendbuff 0x400eb9e3ae00 recvbuff 0x400eb9e8b000 count 1024 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6bd0 (FSDPPolicyWorkerBase pid=487912) comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487915) unt 0 sendbuff 0x400eb9e3ae00 recvbuff 0x400eb9e8b000 count 1024 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee700 (FSDPPolicyWorkerBase pid=487915) jpbo-006-4 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) FO AllGather: opCount 0 sendbuff 0x40186e800000 recvbuff 0x401660000000 count 12582912 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7e80 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) ff 0x400eb9e8b000 count 1024 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e87d0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400eb9e3e000 recvbuf (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e612 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 0x400e427c7e80 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) p 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) 31648 datatype 9 op 0 root 0 comm 0xaaab3461efd0 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) jpbo-007-21:336355:336355 [0] NCCL INFO Broadcast: opCount aa sendbuff 0x4003ddd82400 recvbuff 0x4003ddd82400 count 128 datatype 9 op 0 root 0 comm 0xaaab3461efd0 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) jpbo-007-21:336355:336355 [0] NCCL INFO Broadcast: opCount ab sendbuff 0x4003ddd82400 r (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) tatype 9 op 0 root 0 comm 0xaaab22fc1480 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) opCount 7f sendbuff 0x4003ddd82400 recvbuff 0x4003ddd82400 count 128 datatype 9 op 0 root 0 comm 0xaaab37522a20 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2073787 [0] NCCL INFO Broadcast: opCount a9 sendbuff 0x4016f9800000 recvbuff 0x4016f9800000 count 16777216 dat (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) 52 sendbuff 0x4003bdd82400 recvbuff 0x4003bdd82400 count 128 datatype 9 op 0 root 0 comm 0xaaab46c65e00 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) jpbo-010-47:2073774:2073774 [0] NCCL INFO Broadcast: opCount 6d sendbuff 0x4016b9800000 recvbuff 0x4016b9800000 count 16777216 datatype 9 (FSDPPolicyWorkerBase pid=487913) f 0x400eb9e88c00 count 32 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f11f0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount ab sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 128 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount ac sendbuff 0x401660000000 recvbuff 0x401660000000 count 50331648 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opC (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401345c00000 recvbuff 0x401780000000 count 12582912 d (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401348c00000 recvbuff 0x401780000000 count 12582912 datatyp (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount ad sendbuff 0x401666000000 recvbuff 0x401666000000 count 50331648 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount ae sendbuff 0x401660000000 recvbuff 0x401660000000 count 50331648 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount af sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 4096 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount ba sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 4096 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount bb sendbuff 0x400e9fa8b000 recvbuff 0x400e9fa8b000 count 4096 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount bc sendbuff 0x401913c00000 recvbuff 0x401913c00000 count 16777216 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount bd sendbuff 0x401389580000 recvbuff 0x401389580000 count 4194304 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount be sendbuff 0x401913c00000 recvbuff 0x401913c00000 count 4194304 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount bf sendbuff 0x401660000000 recvbuff 0x401660000000 count 16777216 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40188e200000 recvbuff 0x401913c00000 count 1048576 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nran (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) ount fe sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:3330 (FSDPPolicyWorkerBase pid=487912) 0dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 1c8 sendbuff 0x400e9fa88e00 re (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) atype 9 op 0 root 0 comm 0xaaab07142a30 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) jpbo-010-46:2080788:2080788 [0] NCCL INFO Broadcast: opCount ca sendbuff 0x401739800000 recvbuff 0x401739800000 count 16777216 datatype 9 op 0 root 0 comm 0xaaab07142a30 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) jpbo-010-46:2080788:2080788 [0] NCCL INFO Broadcast: opCount cb sendbuff 0x4003bdd82400 recvbuff 0x4003bdd82400 count 128 datatype 9 op 0 root 0 comm 0xaaab07142a30 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) jpbo-010-46:2080788:2080788 [0] NCCL INFO Broadcast: opCount cc sendbuff 0x4003bdd82400 recvbuff 0x4003bdd82400 count 128 datatype 9 op 0 root 0 comm 0xaaab07142a30 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) jpbo-010-46:2080788:2080788 [0] NCCL INFO Broadcast: opCount cd sendbuff 0x4017a0000000 recvbuff 0x4017a0000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab07142a30 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) jpbo-010-46:2080788:2080788 [0] NCCL INFO Broadcast: opCount ce sendbuff 0x4017a0000000 recvbuff 0x4017a0000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab07142a30 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) jpbo-010-46:2080788:2080788 [0] NCCL INFO Broadcast: opCount cf sendbuff 0x4017a0000000 recvbuff 0x4017a0000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab07142a30 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) jpbo-010-46:2080 (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) jpbo-010-46:2080799:2080 (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) jpbo-010-46:2080782:2080 (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) jpbo-010-46:2080 (FSDPPolicyWorkerBase pid=487914) atatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6bd0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) jpbo-007-02:632066:632066 [0] NCCL INFO Broadcast: opCo (FSDPPolicyWorkerBase pid=487915) e 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee700 (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) jpbo-007-02:632073:632073 [0] NCCL INFO Broadcast: opCount 8a s (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073 (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) jpbo-010-47:2073774:2073 (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) jpbo-010-47:2073795:2073 (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) jpbo-010-47:2073 (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) jpbo-010-48:2062 (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) jpbo-010-48:2062599:2062 (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) jpbo-010-48:2062587:2062 (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) jpbo-010-48:2062 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount d7 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 128 d (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401316400000 recvbuff 0x40178a400000 count 4194304 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] str (FSDPPolicyWorkerBase pid=487912) cvbuff 0x400e9fa88e00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL I (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) atatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount da sendbuff 0x401660000000 recvbuff 0x401660000000 count 50331648 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount db sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 4096 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount dc sendbuff 0x400e9fa8b000 recvbuff 0x400e9fa8b000 count 4096 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount dd sendbuff 0x401913c00000 recvbuff 0x401913c00000 count 16777216 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount de sendbuff 0x401389580000 recvbuff 0x401389580000 count 4194304 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount df sendbuff 0x401913c00000 recvbuff 0x401913c00000 count 4194304 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 1e5 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) FO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCou (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) eam 0x400df13e9a90 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e88e00 recvbuff 0x400e79e88e00 count 1 datatype 7 op 0 r (FSDPPolicyWorkerBase pid=487912) NFO AllGather: opCount 0 sendbuff 0x401871000000 recvbuff 0x401389180000 count 1048576 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401882000000 recvbuff 0x4018f2a00000 count 4194304 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nr (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount ea sendbuff 0x401913c00000 recvbuff 0x401913c00000 count 4194304 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount eb sendbuff 0x401660000000 recvbuff 0x401660000000 count 16777216 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount ec sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 128 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount ed sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 128 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount ee sendbuff 0x401660000000 recvbuff 0x401660000000 count 50331648 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount ef sendbuff 0x401666000000 recvbuff 0x401666000000 count 50331648 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:33281 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401339c00000 recvbuff 0x401760000000 count 1258 (FSDPPolicyWorkerBase pid=487914) opCount 0 sendbuff 0x40135b400000 recvbuff 0x401786000000 count 12582912 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6bd0 (FSDPPolicyWorkerBase pid=487913) INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nrank (FSDPPolicyWorkerBase pid=487915) nt 0 sendbuff 0x40135e400000 recvbuff 0x401786000000 count 12582912 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee700 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2073787 [0] NCCL INFO Broadcast: opCount fe sendbuff 0x4016f9800000 recvbuff (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) jpbo-051-33:593278:593278 [0] NCCL INFO Broadcast: opCount fe sendbuff 0x4016d9800000 recvbuff 0x4016d9800000 count 16777216 datatype 9 op 0 root 0 comm 0xaaab18407ce0 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) jpbo-051-33:593278:593278 [0] NCCL INFO Broadcast: opCount ff sendbuff 0x401739200000 recvbuff 0x401739200000 count 4194304 datatype 9 op 0 root 0 comm 0xaaab18407ce0 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) jpbo-051-33:593278:593278 [ (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) endbuff 0x40039dd82400 recvbuff 0x40039dd82400 count 128 datatype 9 op 0 root 0 comm 0xaaaafc6c8360 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 208 sendbuff 0x400e9fa88e00 recvbuff 0x400e9 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) oot 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 2912 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e87d0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: op (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e89000 recvbuff 0x400eb9e89000 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nran (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGath (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCC (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCou (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) jpbo-010-48:2062603:2062603 [0] NCCL INFO Broadcast: opCount 128 sendbuff 0x40075203c400 recvbuff 0x40075203c400 count 4096 datatyp (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) jpbo-010-48:2062599:2062599 [0] NCCL INFO Broadcast: opCount be sendbuff 0x40075203c400 recvbuff 0x40075203c400 count 4096 datatype 9 op 0 root 0 comm 0xaaab403c3100 [nranks=49] s (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400eb9e4e400 recvbuff 0x400eb9e88c00 count 32 dataty (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400eb9e4e400 recvbuff 0x400eb9e88c00 count 32 datatype 9 o (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) Count 18f sendbuff 0x400eb9e89000 recvbuff 0x400eb9e89000 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream ( (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) jpbo-007-02:632066:632066 [0] NCCL INFO Broadcast: opCount 12b sendbuff 0x401739200000 recvbuff 0x401739200000 count 4194304 datatype 9 op 0 root 0 c (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) NFO Broadcast: opCount a5 sendbuff 0x4016f9800000 recvbuff 0x4016f9800000 count 16777216 datatype 9 op 0 root 0 comm 0xaaab344c5740 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) jpbo-051 (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) jpbo-051 (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) jpbo-051 (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) jpbo-051 (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) jpbo-051 (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) jpbo-007 (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) jpbo-007 (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) jpbo-007 (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) jpbo-007 (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) jpbo-007 (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) jpbo-007 (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) jpbo-051 (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) jpbo-051 (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) jpbo-051 (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) jpbo-051 (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) jpbo-051 (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) jpbo-051 (FSDPPolicyWorkerBase pid=487912) er: opCount 0 sendbuff 0x400e9fa7b200 recvbuff 0x400e9fa8b000 count 1024 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) L INFO AllReduce: opCount 21e sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) nt fe sendbuff 0x400e79e88c00 recvbuff 0x400e79e88c00 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) Count fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822: (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) jpbo-007 (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) jpbo-007 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) nil) (FSDPPolicyWorkerBase pid=487914) pe 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6bd0 (FSDPPolicyWorkerBase pid=487913) : opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (n (FSDPPolicyWorkerBase pid=487915) p 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee700 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401391400000 recvbuff 0x401760000000 count 12582912 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 248 sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO All (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) omm 0xaaab05f7d2d0 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) jpbo-007-23:309070:309070 [0] NCCL INFO Broadcast: opCount 156 sendbuff 0x4016b (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) jpbo-007-23:309096:309096 [0] NCCL INFO Broadcast: opCount da sendbuff 0x4016d9800000 recvbuff 0x4016d9800000 count 16777216 datatype 9 op 0 root 0 comm 0xaaab38751670 [nran (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) e 9 op 0 root 0 comm 0xaaab37522a20 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) jpbo-010-47:2073774:2073774 [0] NCCL INFO Broadcast: opCount d9 se (FSDPPolicyWorkerBase pid=487913) il) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 (FSDPPolicyWorkerBase pid=487914) Gather: opCount 0 sendbuff 0x400eb9e53200 recvbuff 0x400eb9e88e00 count 32 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6bd0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 265 sendbuff 0x400e9fa88c00 recvbuf (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) t 18f sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) f 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) 0 (FSDPPolicyWorkerBase pid=487912) 0 sendbuff 0x4018c0000000 recvbuff 0x401660000000 count 12582912 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) jpbo-010-46:208 (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ndbuff 0x4017a0000000 recvbuff 0x4017a0000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab19362b60 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) jpbo-010-46:208 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:207 (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) jpbo-010-47:207 (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) jpbo-010-48:206 (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) jpbo-010-48:206 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) INFO AllGather: opCount 0 sendbuff 0x4018ee800000 recvbuff 0x401660000000 count 12582912 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7e80 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4018f0800000 recvbuff 0x401913c00000 count 4194304 datatype 9 op 0 root 0 comm 0x400e42a9f600 [ (FSDPPolicyWorkerBase pid=487912) 7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 288 sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 1 datatype (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opC (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) (EngineCore_DP0 pid=389428) INFO 06-12 05:22:51 [block_pool.py:452] Successfully reset prefix cache (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 data (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) jpbo-007-23:309090:309090 [0] NCCL INFO cudaDriverVersion 13020 [repeated 47x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) jpbo-007-23:309090:309090 [0] NCCL INFO NCCL version 2.27.7+cuda13.0 [repeated 47x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) jpbo-007-23:309090:309478 [0] NCCL INFO NET/Plugin: Could not find: libnccl-net.so.  [repeated 47x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) jpbo-007-23:309090:309478 [0] NCCL INFO NET/IB : Using [0]mlx5_0:1/IB [1]mlx5_1:1/IB [2]mlx5_2:1/IB [3]mlx5_3:1/IB [RO]; OOB ib0:10.128.17.55<0> [repeated 47x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) jpbo-007-23:309090:309478 [0] NCCL INFO Initialized NET plugin IB [repeated 47x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333171 [0] NCCL INFO Assigned NET plugin IB to comm [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333171 [0] NCCL INFO Using network IB [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333171 [0] NCCL INFO DMA-BUF is available on GPU device 0 [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333171 [0] NCCL INFO ncclCommInitRankConfig comm 0x400fa44ecc80 rank 0 nranks 49 cudaDev 0 nvmlDev 0 busId 901000 commId 0x9b0018c00ef5dafe - Init START [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) jpbo-007-23:309090:309478 [0] NCCL INFO RAS client listening socket at 127.0.0.1<28028> [repeated 47x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) jpbo-007-23:309090:309478 [0] NCCL INFO NCCL_NET_GDR_LEVEL set by environment to LOC [repeated 47x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333171 [0] NCCL INFO Setting affinity for GPU 0 to 0-71 [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) jpbo-007-23:309090:309478 [0] NCCL INFO NVLS multicast support is not available on dev 0 (NVLS_NCHANNELS 0) [repeated 47x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333171 [0] NCCL INFO comm 0x400fa44ecc80 rank 0 nRanks 49 nNodes 13 localRanks 1 localRank 0 MNNVL 0 [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333171 [0] NCCL INFO Trees [0] 26/-1/-1->0->-1 [1] 27/-1/-1->0->-1 [2] 28/-1/-1->0->-1 [3] 37/-1/-1->0->-1 [4] 41/-1/-1->0->29 [5] 42/-1/-1->0->30 [6] 43/-1/-1->0->31 [7] 44/-1/-1->0->32 [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333171 [0] NCCL INFO P2P Chunksize set to 131072 [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) jpbo-007-23:309090:309478 [0] NCCL INFO PROFILER/Plugin: Could not find: libnccl-profiler.so.  [repeated 47x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333172 [0] NCCL INFO [Proxy Service] Device 0 CPU core 1 [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333173 [0] NCCL INFO [Proxy Service UDS] Device 0 CPU core 2 [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333174 [0] NCCL INFO [Proxy Progress] Device 0 CPU core 3 [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333171 [0] NCCL INFO Channel 00/0 : 26[0] -> 0[0] [receive] via NET/IB/0 [repeated 278x across cluster] (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) jpbo-051-34:379897:380166 [0] NCCL INFO Channel 07/0 : 11[2] -> 10[1] via P2P/IPC [repeated 658x across cluster] (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) jpbo-051-34:379897:380194 [0] NCCL INFO Allocated shareable buffer 0x40171fe00000 size 2097152 ipcDesc 0x40094433fcb0 [repeated 1315x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333171 [0] NCCL INFO Channel 07/0 : 0[0] -> 44[3] [send] via NET/IB/0 [repeated 279x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) jpbo-051-34:379917:380164 [0] NCCL INFO Imported shareable buffer device 0 size 10485760 ptr 0x4016bd800000 [repeated 1332x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333171 [0] NCCL INFO Connected all rings, use ring PXN 0 GDR 0 [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) jpbo- [repeated 19x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) jpbo-051-34:379901:380196 [0] NCCL INFO Allocated shareable buffer 0x40179c600000 size 104 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487912) jpb (FSDPPolicyWorkerBase pid=487913) jpbo [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333171 [0] NCCL INFO Connected all trees [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333171 [0] NCCL INFO threadThresholds 8/8/64 | 392/8/64 | 512 | 512 [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333171 [0] NCCL INFO 8 coll channels, 8 collnet channels, 0 nvls channels, 8 p2p channels, 1 p2p channels per peer [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) jpbo-051-34:379917:380164 [0] NCCL INFO TUNER/Plugin: Could not find: libnccl-tuner.so. Using internal tuner plugin. [repeated 47x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333171 [0] NCCL INFO ncclCommInitRankConfig comm 0x400fa44ecc80 rank 0 nranks 49 cudaDev 0 nvmlDev 0 busId 901000 commId 0x9b0018c00ef5dafe - Init COMPLETE [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333171 [0] NCCL INFO Init timings - ncclCommInitRankConfig: rank 0 nranks 49 total 3.30 (kernels 0.00, alloc 0.00, bootstrap 2.59, allgathers 0.04, topo 0.01, graphs 0.00, connections 0.66, rest 0.00) [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) 85760 ipcDesc 0x40095c33fd20 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount ff sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4013c0000000 recvbuff 0x401780000000 count 155582464 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee700 [repeated 3088x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff  [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) ] stream (nil) (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) 9800000 recvbuff 0x4016d9800000 count 16777216 datatype 9 op 0 root 0 comm 0xaaab1e9d50e0 [nranks=49] stream (nil) [repeated 26x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) jpbo-010-46:2080778:2080778 [0] NCCL INFO Broadcast: opCount a sendbuff 0x40079203c400 recvbuff 0x40079203c400 count 4096 datatype 9 op 0 root 0 comm 0xaaab2ebd6ee0 [nranks=49] stream (nil) [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) jpbo-010-46:2080778:2080778 [0] NCCL INFO Broadcast: opCount b sendbuff 0x40079203c400 recvbuff 0x40079203c400 count 4096 datatype 9 op 0 root 0 comm 0xaaab2ebd6ee0 [nranks=49] stream (nil) [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) jpbo-010-46:2080778:2080778 [0] NCCL INFO Broadcast: opCount c sendbuff 0x4016f9800000 recvbuff 0x4016f9800000 count 16777216 datatype 9 op 0 root 0 comm 0xaaab2ebd6ee0 [nranks=49] stream (nil) [repeated 98x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) jpbo-010-46:2080778:2080778 [0] NCCL INFO Broadcast: opCount d sendbuff 0x401759200000 recvbuff 0x401759200000 count 4194304 datatype 9 op 0 root 0 comm 0xaaab2ebd6ee0 [nranks=49] stream (nil) [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) jpbo-010-46:2080778:2080778 [0] NCCL INFO Broadcast: opCount e sendbuff 0x401759200000 recvbuff 0x401759200000 count 4194304 datatype 9 op 0 root 0 comm 0xaaab2ebd6ee0 [nranks=49] stream (nil) [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) jpbo-010-46:2080778:2080778 [0] NCCL INFO Broadcast: opCount f sendbuff 0x4016f9800000 recvbuff 0x4016f9800000 count 16777216 datatype 9 op 0 root 0 comm 0xaaab2ebd6ee0 [nranks=49] stream (nil) [repeated 98x across cluster] (FSDPPolicyWorkerBase pid=487915) 20 [nranks=8] stream (nil) [repeated 88x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e89000 recvbuff 0x400eb9e89000 count 1 datatype 7 op 0 root (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) unt d6 sendbuff 0x4003bdd82400 recvbuff 0x4003bdd82400 count 128 datatype 9 op 0 root 0 comm 0xaaaae79c06a0 [nranks=49] stream (nil) [repeated 17x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) jpbo-051-37:379657:379657 [0] NCCL INFO Broadcast: opCount 2a sendbuff 0x401740000000 recvbuff 0x401740000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab1d13b270 [nran [repeated 16x across cluster] (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) jpbo-051-37:379644:379644 [0] NCCL INFO Broadcast: opCount 1c sendbuff 0x401740000000 recvbuff 0x401740000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab25666d20 [nranks=49] s [repeated 18x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) jpbo-010-46:2080778:2080778 [0] NCCL INFO Broadcast: opCount 2a sendbuff 0x401760000000 r [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 110 sendbuff 0x401666000000 recvbuff [repeated 31x across cluster] (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) 0 comm 0xaaab1c571b40 [nranks=49] stream (nil) [repeated 18x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4013a7400000 recvbuff 0x401 [repeated 10x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401379800000 recvbuff 0x4017aa400000 count 1048576 datatype 9 op (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) jpbo-051-38:386887:386887 [0] NCCL INFO Broadcast: opCount f6 sendbuff 0x40075203c400 recvbuff 0x40075203c400 count 4096 datatype 9 op 0 root 0 comm 0xaaab215f1670 [nranks=49] stream (nil) [repeated 14436x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) [repeated 1346x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4012e74000 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400eb9e50a00 recvbuff 0x400eb9e88e00 count 32 datat (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) jpbo-007-23:309090:309090 [0] NCCL INFO Broadcast: opCount fa sendbuff 0x401740000000 recvbuff 0x401740000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaaae79c06a0 [nranks=49] stream (nil) [repeated 24x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) jpbo-007-23:309090:309090 [0] NCCL INFO Broadcast: opCount fb sendbuff 0x401740000000 recvbuff 0x401740000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaaae79c06a0 [nranks=49] stream (nil) [repeated 24x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) jpbo-007-23:309090:309090 [0] NCCL INFO Broadcast: opCount fc sendbuff 0x40077203c400 recvbuff 0x40077203c400 count 4096 datatype 9 op 0 root 0 comm 0xaaaae79c06a0 [nranks=49] stream (nil) [repeated 24x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) jpbo-007-23:309090:309090 [0] NCCL INFO Broadcast: opCount fd sendbuff 0x40077203c400 recvbuff 0x40077203c400 count 4096 datatype 9 op 0 root 0 comm 0xaaaae79c06a0 [nranks=49] stream (nil) [repeated 24x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 1323x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e87d0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 766000000 count 12582912 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9bc0 [repeated 10x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCoun [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487913) 00 recvbuff 0x401766000000 count 12582912 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f11f0 (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ecvbuff 0x4003ddd82400 count 128 datatype 9 op 0 root 0 comm 0xaaab1d13b270 [nranks=49] stream (nil) [repeated 22x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) jpbo-010-46:2080778:2080778 [0] NCCL INFO Broadcast: opCount 54 sendbuff 0x401760000000 recvbuff 0x401760000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab2ebd6ee0 [ [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=487912) count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) [repeated 39x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 157 sendbuff 0x401389580000 recvbuff 0x401389580000 count 4194304 datatype 9 op 0 root 0 comm 0x40 [repeated 31x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) jpbo-051-38:386887:386887 [0] NCCL INFO Broadcast: opCount f6 sendbuff 0x40075203c400 recvbuff 0x40075203c400 count 4096 datatype 9 op 0 root  [repeated 41x across cluster] (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) tream (nil) [repeated 24x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) t fe sendbuff 0x400e79e88e00 recvbuff 0x400e79e88e00 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) ype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e87d0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 22c sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e4 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401354400000 recvbuff 0x401740000000 count 12582912 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) jpbo-010-47:2073782:2073782 [0] NCCL INFO Broadcast:  [repeated 5x across cluster] (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) jpbo-010-47:2073795:2073795 [0] NCCL INFO Broadcast: opCount  [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 9e sendbuff 0x401660000000 recvbuff 0x401660000000 count 167 [repeated 17x across cluster] (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) jpbo-051-40:388620:388620 [0] NCCL INFO Broadcast: opCount 52 sendbuff 0x401740000000 recvbuff 0x401740000000 count 50331648 da [repeated 18x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL IN (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e88c00 recvbuff 0x400e79e88c00 count 1 datatype 7 op  [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487915) : opCount 0 sendbuff 0x400eb9e53200 recvbuff 0x400eb9e88e00 count 32 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee700 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) 31648 datatype 9 op 0 root 0 comm 0xaaab1f476500 [nranks=49] stream (nil) [repeated 17x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) jpbo-007-02:632092:632092 [0] NCCL INFO Broadcast: opCount aa sendbuff 0x401759200000 recvbuff 0x401759200000 count 4194304 datatype 9 op 0 root 0 comm 0xaaab0bda6fa0 [nranks=49] stream (nil) [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) jpbo-007-31:389444:389444 [0] NCCL INFO Broadcast: opCount ab sendbuff 0x4003ddd82400 r [repeated 16x across cluster] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) tatype 9 op 0 root 0 comm 0xaaab32ee8f60 [nranks=49] stream (nil) [repeated 18x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) opCount 7f sendbuff 0x4003ddd82400 recvbuff 0x4003ddd82400 count 128 datatype 9 op 0 root 0 comm 0xaaab2ebd6ee0 [nranks=49] stream (nil) [repeated 5x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) jpbo-010-46:2080778:2080778 [0] NCCL INFO Broadcast: opCount a9 sendbuff 0x4016f9800000 recvbuff 0x4016f9800000 count 16777216 dat [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=487913) 8f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) [repeated 7x across cluster] (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) jpbo-010-46:2080782:2080782 [0] NCCL INFO Broadcast: opCount 6d sendbuff 0x4016d9800000 recvbuff 0x4016d9800000 count 16777216 datatype 9  [repeated 5x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) jpbo-007-02:632092:632092 [0] NCCL INFO Broadcast: opCount ab sendbuff 0x401759200000 recvbuff 0x401759200000 count 4194304 datatype 9 op 0 root 0 comm 0xaaab0bda6fa0 [nranks=49] stream (nil) [repeated 31x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) jpbo-007-02:632092:632092 [0] NCCL INFO Broadcast: opCount ac sendbuff 0x4016f9800000 recvbuff 0x4016f9800000 count 16777216 datatype 9 op 0 root 0 comm 0xaaab0bda6fa0 [nranks=49] stream (nil) [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) jpbo-007-02:632092:632092 [0] NCCL INFO Broadcast: opCount ad sendbuff 0x401760000000 recvbuff 0x401760000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab0bda6fa0 [nranks=49] stream (nil) [repeated 98x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) jpbo-007-02:632092:632092 [0] NCCL INFO Broadcast: opCount ae sendbuff 0x401760000000 recvbuff 0x401760000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab0bda6fa0 [nranks=49] stream (nil) [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) jpbo-007-02:632092:632092 [0] NCCL INFO Broadcast: opCount af sendbuff 0x401760000000 recvbuff 0x401760000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab0bda6fa0 [nranks=49] stream (nil) [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) jpbo-007-02:632092:632092 [0] NCCL INFO Broadcast: opCount ba sendbuff 0x4016f9800000 recvbuff 0x4016f9800000 count 16777216 datatype 9 op 0 root 0 comm 0xaaab0bda6fa0 [nranks=49] stream (nil) [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) jpbo-007-02:632092:632092 [0] NCCL INFO Broadcast: opCount bb sendbuff 0x401760000000 recvbuff 0x401760000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab0bda6fa0 [nranks=49] stream (nil) [repeated 98x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) jpbo-007-02:632092:632092 [0] NCCL INFO Broadcast: opCount bc sendbuff 0x401760000000 recvbuff 0x401760000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab0bda6fa0 [nranks=49] stream (nil) [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) jpbo-007-02:632092:632092 [0] NCCL INFO Broadcast: opCount bd sendbuff 0x401760000000 recvbuff 0x401760000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab0bda6fa0 [nranks=49] stream (nil) [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) jpbo-010-46:2080782:2080782 [0] NCCL INFO Broadcast: opCount be sendbuff 0x4016d9800000 recvbuff 0x4016d9800000 count 16777216 datatype 9 op 0 root 0 comm 0xaaaae5ee1730 [nranks=49] stream (nil) [repeated 92x across cluster] (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) jpbo-010-46:2080782:2080782 [0] NCCL INFO Broadcast: opCount bf sendbuff 0x401739200000 recvbuff 0x401739200000 count 4194304 datatype 9 op 0 root 0 comm 0xaaaae5ee1730 [nranks=49] stream (nil) [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) =4] stream 0x400e153f9bc0 [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) atype 9 op 0 root 0 comm 0xaaab07846660 [nranks=49] stream (nil) [repeated 5x across cluster] (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) jpbo-007-21:336331:336331 [0] NCCL INFO Broadcast: opCount ca sendbuff 0x401740000000 recvbuff 0x401740000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab51e80c30 [nranks=49] stream (nil) [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) jpbo-007-21:336331:336331 [0] NCCL INFO Broadcast: opCount cb sendbuff 0x401740000000 recvbuff 0x401740000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab51e80c30 [nranks=49] stream (nil) [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) jpbo-007-21:336331:336331 [0] NCCL INFO Broadcast: opCount cc sendbuff 0x4016d9800000 recvbuff 0x4016d9800000 count 16777216 datatype 9 op 0 root 0 comm 0xaaab51e80c30 [nranks=49] stream (nil) [repeated 98x across cluster] (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) jpbo-007-21:336331:336331 [0] NCCL INFO Broadcast: opCount cd sendbuff 0x401739200000 recvbuff 0x401739200000 count 4194304 datatype 9 op 0 root 0 comm 0xaaab51e80c30 [nranks=49] stream (nil) [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) jpbo-007-21:336331:336331 [0] NCCL INFO Broadcast: opCount ce sendbuff 0x401739200000 recvbuff 0x401739200000 count 4194304 datatype 9 op 0 root 0 comm 0xaaab51e80c30 [nranks=49] stream (nil) [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) jpbo-007-21:336331:336331 [0] NCCL INFO Broadcast: opCount cf sendbuff 0x4016d9800000 recvbuff 0x4016d9800000 count 16777216 datatype 9 op 0 root 0 comm 0xaaab51e80c30 [nranks=49] stream (nil) [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) jpbo-051-37:379657:379657 [0] NCCL INFO Broadcast: opCo [repeated 16x across cluster] (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) jpbo-051-37:379644:379644 [0] NCCL INFO Broadcast: opCount 8a s [repeated 18x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400e9fa7d200 recvbuff 0x400e9fa88e00 count 1024 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] str (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) jpbo-007-23:309080:309080 [0] NCCL I [repeated 19x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) jpbo-010-48:2062587:2062587 [0] NCCL INFO Broadcast: opCount da sendbuff 0x4016f9800000 recvbuff 0x4016f9800000 count 16777216 datatype 9 op 0 root 0 comm 0xaaab0645a9d0 [nranks=49] stream (nil) [repeated 79x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) jpbo-051-38:386887:386887 [0] NCCL INFO Broadcast: opCount db sendbuff 0x401739200000 recvbuff 0x401739200000 count 4194304 datatype 9 op 0 root 0 comm 0xaaab215f1670 [nranks=49] stream (nil) [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) jpbo-051-38:386887:386887 [0] NCCL INFO Broadcast: opCount dc sendbuff 0x401739200000 recvbuff 0x401739200000 count 4194304 datatype 9 op 0 root 0 comm 0xaaab215f1670 [nranks=49] stream (nil) [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) jpbo-051-38:386887:386887 [0] NCCL INFO Broadcast: opCount dd sendbuff 0x4016d9800000 recvbuff 0x4016d9800000 count 16777216 datatype 9 op 0 root 0 comm 0xaaab215f1670 [nranks=49] stream (nil) [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) jpbo-051-38:386887:386887 [0] NCCL INFO Broadcast: opCount de sendbuff 0x401740000000 recvbuff 0x401740000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab215f1670 [nranks=49] stream (nil) [repeated 98x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) jpbo-051-38:386887:386887 [0] NCCL INFO Broadcast: opCount df sendbuff 0x401740000000 recvbuff 0x401740000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab215f1670 [nranks=49] stream (nil) [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL (FSDPPolicyWorkerBase pid=487912) eam 0x400dfa7c67e0 (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) jpbo-051-38:386887:386887 [0] NCCL INFO Broadcast: opCount ea sendbuff 0x401739200000 recvbuff 0x401739200000 count 4194304 datatype 9 op 0 root 0 comm 0xaaab215f1670 [nranks=49] stream (nil) [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) jpbo-051-38:386887:386887 [0] NCCL INFO Broadcast: opCount eb sendbuff 0x4016d9800000 recvbuff 0x4016d9800000 count 16777216 datatype 9 op 0 root 0 comm 0xaaab215f1670 [nranks=49] stream (nil) [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) jpbo-051-38:386887:386887 [0] NCCL INFO Broadcast: opCount ec sendbuff 0x401740000000 recvbuff 0x401740000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab215f1670 [nranks=49] stream (nil) [repeated 98x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) jpbo-051-38:386887:386887 [0] NCCL INFO Broadcast: opCount ed sendbuff 0x401740000000 recvbuff 0x401740000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab215f1670 [nranks=49] stream (nil) [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) jpbo-051-38:386887:386887 [0] NCCL INFO Broadcast: opCount ee sendbuff 0x401740000000 recvbuff 0x401740000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab215f1670 [nranks=49] stream (nil) [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) jpbo-051-38:386887:386887 [0] NCCL INFO Broadcast: opCount ef sendbuff 0x4016d9800000 recvbuff 0x4016d9800000 count 16777216 datatype 9 op 0 root 0 comm 0xaaab215f1670 [nranks=49] stream (nil) [repeated 98x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401342a00000 recvbuff 0x40171a400000 count (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) jpbo-010-46:2080778:2080778 [0] NCCL INFO Broadcast: opCount fe sendbuff 0x4016f9800000 recvbuff  [repeated 5x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) jpbo-007-23:309090:309090 [0] NCCL INFO Broadcast: opCount fe sendbuff 0x4016d9800000 recvbuff 0x4016d9800000 count 16777216 datatype 9 op 0 root 0 comm 0xaaaae79c06a0 [nranks=49] stream (nil) [repeated 17x across cluster] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) jpbo-010-47:2073782:2073782 [0] NCCL INFO Broadcast: opCount ff sendbuff 0x401719200000 recvbuff 0x401719200000 count 4194304 datatype 9 op 0 root 0 comm 0xaaab0bce14a0 [nranks=49] stream (nil) [repeated 23x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [ [repeated 17x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) endbuff 0x4003bdd82400 recvbuff 0x4003bdd82400 count 128 datatype 9 op 0 root 0 comm 0xaaab2452b110 [nranks=49] stream (nil) [repeated 18x across cluster] (FSDPPolicyWorkerBase pid=487913) 1048576 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f11f0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: op (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) jpbo-010-47:2073782:2073782 [0] NCCL INFO Broadcast: opCount 128 sendbuff 0x40075203c400 recvbuff 0x40075203c400 count 4096 datatyp [repeated 5x across cluster] (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) jpbo-010-47:2073795:2073795 [0] NCCL INFO Broadcast: opCount be sendbuff 0x40077203c400 recvbuff 0x40077203c400 count 4096 datatype 9 op 0 root 0 comm 0xaaab26735d60 [nranks=49] s [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400eb9e53e00 recvbuff 0x400eb9e8b000 count 1024 dataty (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) jpbo-051-37:379657:379657 [0] NCCL INFO Broadcast: opCount 12b sendbuff 0x401739200000 recvbuff 0x401739200000 count 4194304 datatype 9 op 0 root 0 c [repeated 16x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) NFO Broadcast: opCount a5 sendbuff 0x4016f9800000 recvbuff 0x4016f9800000 count 16777216 datatype 9 op 0 root 0 comm 0xaaab0bda6fa0 [nranks=49] stream (nil) [repeated 18x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487913) pe 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f11f0 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) omm 0xaaaaf5247510 [nranks=49] stream (nil) [repeated 16x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) jpbo-007-21:336339:336339 [0] NCCL INFO Broadcast: opCount 156 sendbuff 0x4016d [repeated 16x across cluster] (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) jpbo-007-21:336331:336331 [0] NCCL INFO Broadcast: opCount da sendbuff 0x4016d9800000 recvbuff 0x4016d9800000 count 16777216 datatype 9 op 0 root 0 comm 0xaaab51e80c30 [nran [repeated 18x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) e 9 op 0 root 0 comm 0xaaab2ebd6ee0 [nranks=49] stream (nil) [repeated 5x across cluster] (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) jpbo-010-46:2080782:2080782 [0] NCCL INFO Broadcast: opCount d9 se [repeated 5x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ndbuff 0x401760000000 recvbuff 0x401760000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab0645a9d0 [nranks=49] stream (nil) [repeated 5x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) (EngineCore_DP0 pid=386887) INFO 06-12 05:22:51 [block_pool.py:452] Successfully reset prefix cache [repeated 47x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 1008x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 225x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 782x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) INFO 06-12 05:23:10 [hf.py:318] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this. (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) INFO 06-12 05:23:15 [loggers.py:259] Engine 000: Avg prompt throughput: 9.4 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.3%, Prefix cache hit rate: 0.0% (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:23:14 [hf.py:318] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this. [repeated 17x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) INFO 06-12 05:23:20 [loggers.py:259] Engine 000: Avg prompt throughput: 73.4 tokens/s, Avg generation throughput: 59.1 tokens/s, Running: 2 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.5%, Prefix cache hit rate: 31.6% [repeated 19x across cluster] (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) INFO 06-12 05:23:20 [hf.py:318] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this. [repeated 5x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) INFO 06-12 05:23:24 [loggers.py:259] Engine 000: Avg prompt throughput: 8.9 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.3%, Prefix cache hit rate: 0.0% [repeated 22x across cluster] (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) INFO 06-12 05:23:25 [hf.py:318] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this. [repeated 5x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) INFO 06-12 05:23:30 [loggers.py:259] Engine 000: Avg prompt throughput: 50.8 tokens/s, Avg generation throughput: 121.0 tokens/s, Running: 2 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.7%, Prefix cache hit rate: 44.2% [repeated 34x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) INFO 06-12 05:23:29 [hf.py:318] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this. [repeated 8x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) INFO 06-12 05:23:35 [loggers.py:259] Engine 000: Avg prompt throughput: 135.3 tokens/s, Avg generation throughput: 108.4 tokens/s, Running: 4 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.2%, Prefix cache hit rate: 56.9% [repeated 36x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) INFO 06-12 05:23:40 [loggers.py:259] Engine 000: Avg prompt throughput: 109.2 tokens/s, Avg generation throughput: 143.1 tokens/s, Running: 4 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.5%, Prefix cache hit rate: 62.9% [repeated 35x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) INFO 06-12 05:23:45 [loggers.py:259] Engine 000: Avg prompt throughput: 33.9 tokens/s, Avg generation throughput: 132.0 tokens/s, Running: 4 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.7%, Prefix cache hit rate: 67.4% [repeated 35x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:23:47] OK: 243 / 131,072 FDs open (0.2% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:23:47] OK: RSS 1.13 GiB | node mem 378.8/858.0 GiB used (44.2%), avail 479.2 GiB (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) INFO 06-12 05:23:50 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 159.1 tokens/s, Running: 4 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.0%, Prefix cache hit rate: 67.4% [repeated 34x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) INFO 06-12 05:23:55 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 154.9 tokens/s, Running: 4 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.2%, Prefix cache hit rate: 67.4% [repeated 35x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:23:57] OK: 245 / 131,072 FDs open (0.2% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:23:57] OK: RSS 1.12 GiB | node mem 384.5/858.0 GiB used (44.8%), avail 473.4 GiB (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) INFO 06-12 05:24:00 [loggers.py:259] Engine 000: Avg prompt throughput: 153.0 tokens/s, Avg generation throughput: 144.8 tokens/s, Running: 4 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.6%, Prefix cache hit rate: 71.2% [repeated 34x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 3x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] [05:24:05] OK: 405 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] [05:24:05] OK: RSS 1.14 GiB | node mem 397.0/858.0 GiB used (46.3%), avail 461.0 GiB (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) INFO 06-12 05:24:05 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 158.6 tokens/s, Running: 4 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.8%, Prefix cache hit rate: 71.2% [repeated 31x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 7x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) INFO 06-12 05:24:10 [loggers.py:259] Engine 000: Avg prompt throughput: 27.6 tokens/s, Avg generation throughput: 33.3 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.8%, Prefix cache hit rate: 74.7% [repeated 35x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 11x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) INFO 06-12 05:24:13 [hf.py:318] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this. (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] [05:24:14] OK: 418 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] [05:24:14] OK: RSS 1.09 GiB | node mem 388.2/858.0 GiB used (45.3%), avail 469.7 GiB (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) INFO 06-12 05:24:15 [loggers.py:259] Engine 000: Avg prompt throughput: 137.5 tokens/s, Avg generation throughput: 89.7 tokens/s, Running: 3 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.7%, Prefix cache hit rate: 72.7% [repeated 32x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) INFO 06-12 05:24:17 [hf.py:318] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this. [repeated 4x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) INFO 06-12 05:24:20 [loggers.py:259] Engine 000: Avg prompt throughput: 117.1 tokens/s, Avg generation throughput: 85.2 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.4%, Prefix cache hit rate: 69.3% [repeated 40x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (skyrl_entrypoint pid=487747) [fd-monitor] [05:24:20] OK: 241 / 131,072 FDs open (0.2% of soft limit, hard limit: 131,072) (skyrl_entrypoint pid=487747) [fd-monitor] [05:24:20] OK: RSS 1.75 GiB | node mem 241.2/858.0 GiB used (28.1%), avail 616.8 GiB (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 7x across cluster] (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) INFO 06-12 05:24:20 [hf.py:318] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) INFO 06-12 05:24:25 [loggers.py:259] Engine 000: Avg prompt throughput: 257.6 tokens/s, Avg generation throughput: 133.7 tokens/s, Running: 4 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.3%, Prefix cache hit rate: 75.0% [repeated 42x across cluster] (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) INFO 06-12 05:24:27 [hf.py:318] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this. (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 5x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) INFO 06-12 05:24:30 [loggers.py:259] Engine 000: Avg prompt throughput: 78.7 tokens/s, Avg generation throughput: 199.7 tokens/s, Running: 4 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.5%, Prefix cache hit rate: 74.4% [repeated 45x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) INFO 06-12 05:24:28 [hf.py:318] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this. [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) INFO 06-12 05:24:35 [loggers.py:259] Engine 000: Avg prompt throughput: 34.6 tokens/s, Avg generation throughput: 218.0 tokens/s, Running: 4 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.8%, Prefix cache hit rate: 75.2% [repeated 47x across cluster] (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) INFO 06-12 05:24:34 [hf.py:318] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this. [repeated 2x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) INFO 06-12 05:24:40 [loggers.py:259] Engine 000: Avg prompt throughput: 45.5 tokens/s, Avg generation throughput: 241.3 tokens/s, Running: 4 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.6%, Prefix cache hit rate: 78.0% [repeated 47x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 5x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) INFO 06-12 05:24:45 [loggers.py:259] Engine 000: Avg prompt throughput: 194.5 tokens/s, Avg generation throughput: 235.7 tokens/s, Running: 4 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.6%, Prefix cache hit rate: 80.2% [repeated 47x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 20x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) INFO 06-12 05:24:50 [loggers.py:259] Engine 000: Avg prompt throughput: 110.1 tokens/s, Avg generation throughput: 193.1 tokens/s, Running: 4 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 83.7% [repeated 47x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 15x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) INFO 06-12 05:24:55 [loggers.py:259] Engine 000: Avg prompt throughput: 99.9 tokens/s, Avg generation throughput: 252.6 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.9%, Prefix cache hit rate: 83.6% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 20x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) INFO 06-12 05:25:00 [loggers.py:259] Engine 000: Avg prompt throughput: 218.3 tokens/s, Avg generation throughput: 215.7 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 5.1%, Prefix cache hit rate: 84.6% [repeated 46x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 28x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) INFO 06-12 05:25:05 [loggers.py:259] Engine 000: Avg prompt throughput: 261.1 tokens/s, Avg generation throughput: 225.1 tokens/s, Running: 4 reqs, Waiting: 0 reqs, GPU KV cache usage: 5.3%, Prefix cache hit rate: 83.1% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) INFO 06-12 05:25:10 [loggers.py:259] Engine 000: Avg prompt throughput: 433.8 tokens/s, Avg generation throughput: 169.1 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 6.9%, Prefix cache hit rate: 84.6% [repeated 46x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 20x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) INFO 06-12 05:25:16 [loggers.py:259] Engine 000: Avg prompt throughput: 179.4 tokens/s, Avg generation throughput: 134.5 tokens/s, Running: 4 reqs, Waiting: 0 reqs, GPU KV cache usage: 5.6%, Prefix cache hit rate: 84.4% [repeated 49x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 25x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) INFO 06-12 05:25:21 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 157.5 tokens/s, Running: 2 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.1%, Prefix cache hit rate: 84.4% [repeated 46x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 30x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) INFO 06-12 05:25:26 [loggers.py:259] Engine 000: Avg prompt throughput: 69.7 tokens/s, Avg generation throughput: 166.8 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 5.4%, Prefix cache hit rate: 86.8% [repeated 47x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 18x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) INFO 06-12 05:25:31 [loggers.py:259] Engine 000: Avg prompt throughput: 84.4 tokens/s, Avg generation throughput: 221.4 tokens/s, Running: 4 reqs, Waiting: 0 reqs, GPU KV cache usage: 5.3%, Prefix cache hit rate: 86.8% [repeated 47x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 30x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) INFO 06-12 05:25:36 [loggers.py:259] Engine 000: Avg prompt throughput: 26.5 tokens/s, Avg generation throughput: 179.6 tokens/s, Running: 4 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.9%, Prefix cache hit rate: 87.0% [repeated 46x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 13x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) INFO 06-12 05:25:41 [loggers.py:259] Engine 000: Avg prompt throughput: 106.9 tokens/s, Avg generation throughput: 216.4 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 6.4%, Prefix cache hit rate: 87.1% [repeated 47x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) INFO 06-12 05:25:43 [hf.py:318] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this. (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 28x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) INFO 06-12 05:25:46 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 219.7 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 6.7%, Prefix cache hit rate: 87.1% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:25:47] OK: 370 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:25:47] OK: RSS 1.22 GiB | node mem 379.2/858.0 GiB used (44.2%), avail 478.8 GiB (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 13x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) INFO 06-12 05:25:51 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 220.1 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 7.1%, Prefix cache hit rate: 87.1% [repeated 49x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 12x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) INFO 06-12 05:25:56 [loggers.py:259] Engine 000: Avg prompt throughput: 334.5 tokens/s, Avg generation throughput: 123.0 tokens/s, Running: 3 reqs, Waiting: 0 reqs, GPU KV cache usage: 5.2%, Prefix cache hit rate: 82.8% [repeated 47x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:25:57] OK: 276 / 131,072 FDs open (0.2% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:25:57] OK: RSS 1.19 GiB | node mem 384.7/858.0 GiB used (44.8%), avail 473.3 GiB (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 15x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) INFO 06-12 05:26:01 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 219.9 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 7.7%, Prefix cache hit rate: 87.5% [repeated 49x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 5x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] [05:26:05] OK: 368 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] [05:26:05] OK: RSS 1.19 GiB | node mem 397.3/858.0 GiB used (46.3%), avail 460.6 GiB (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) INFO 06-12 05:26:06 [loggers.py:259] Engine 000: Avg prompt throughput: 145.8 tokens/s, Avg generation throughput: 186.4 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 8.2%, Prefix cache hit rate: 87.6% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) INFO 06-12 05:26:11 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 219.8 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 8.5%, Prefix cache hit rate: 87.6% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 5x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] [05:26:14] OK: 419 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] [05:26:14] OK: RSS 1.17 GiB | node mem 388.5/858.0 GiB used (45.3%), avail 469.4 GiB (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) INFO 06-12 05:26:16 [loggers.py:259] Engine 000: Avg prompt throughput: 83.3 tokens/s, Avg generation throughput: 197.7 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 7.0%, Prefix cache hit rate: 87.3% [repeated 47x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 28x across cluster] (skyrl_entrypoint pid=487747) [fd-monitor] [05:26:20] OK: 378 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (skyrl_entrypoint pid=487747) [fd-monitor] [05:26:20] OK: RSS 1.78 GiB | node mem 241.2/858.0 GiB used (28.1%), avail 616.7 GiB (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) INFO 06-12 05:26:21 [loggers.py:259] Engine 000: Avg prompt throughput: 78.5 tokens/s, Avg generation throughput: 287.7 tokens/s, Running: 7 reqs, Waiting: 0 reqs, GPU KV cache usage: 9.4%, Prefix cache hit rate: 87.9% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:26:26 [loggers.py:259] Engine 000: Avg prompt throughput: 321.1 tokens/s, Avg generation throughput: 347.5 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 14.5%, Prefix cache hit rate: 87.8% [repeated 49x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 5x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:26:31 [loggers.py:259] Engine 000: Avg prompt throughput: 127.5 tokens/s, Avg generation throughput: 453.1 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 18.1%, Prefix cache hit rate: 88.5% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 31x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:26:36 [loggers.py:259] Engine 000: Avg prompt throughput: 99.4 tokens/s, Avg generation throughput: 530.3 tokens/s, Running: 12 reqs, Waiting: 0 reqs, GPU KV cache usage: 19.0%, Prefix cache hit rate: 88.4% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 14x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:26:41 [loggers.py:259] Engine 000: Avg prompt throughput: 304.8 tokens/s, Avg generation throughput: 432.7 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 18.9%, Prefix cache hit rate: 89.0% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 23x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:26:46 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 458.0 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 16.4%, Prefix cache hit rate: 89.0% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 23x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:26:51 [loggers.py:259] Engine 000: Avg prompt throughput: 237.0 tokens/s, Avg generation throughput: 378.8 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 17.0%, Prefix cache hit rate: 89.2% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 38x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:26:56 [loggers.py:259] Engine 000: Avg prompt throughput: 49.8 tokens/s, Avg generation throughput: 441.3 tokens/s, Running: 10 reqs, Waiting: 0 reqs, GPU KV cache usage: 19.9%, Prefix cache hit rate: 89.5% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 22x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:27:01 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 411.7 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 16.1%, Prefix cache hit rate: 89.5% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 5x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:27:06 [loggers.py:259] Engine 000: Avg prompt throughput: 20.5 tokens/s, Avg generation throughput: 399.8 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 19.1%, Prefix cache hit rate: 89.9% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 23x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:27:11 [loggers.py:259] Engine 000: Avg prompt throughput: 90.8 tokens/s, Avg generation throughput: 407.7 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 17.2%, Prefix cache hit rate: 89.8% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 15x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:27:16 [loggers.py:259] Engine 000: Avg prompt throughput: 212.3 tokens/s, Avg generation throughput: 386.5 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 18.3%, Prefix cache hit rate: 90.3% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 18x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:27:21 [loggers.py:259] Engine 000: Avg prompt throughput: 295.6 tokens/s, Avg generation throughput: 419.1 tokens/s, Running: 10 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.6%, Prefix cache hit rate: 90.5% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 33x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:27:26 [loggers.py:259] Engine 000: Avg prompt throughput: 62.7 tokens/s, Avg generation throughput: 428.6 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 19.3%, Prefix cache hit rate: 90.6% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 15x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:27:31 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 402.3 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 19.9%, Prefix cache hit rate: 90.6% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 25x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:27:36 [loggers.py:259] Engine 000: Avg prompt throughput: 129.6 tokens/s, Avg generation throughput: 384.1 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 20.6%, Prefix cache hit rate: 90.8% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 28x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) INFO 06-12 05:27:41 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 319.1 tokens/s, Running: 7 reqs, Waiting: 0 reqs, GPU KV cache usage: 15.3%, Prefix cache hit rate: 89.5% [repeated 47x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 18x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:27:46 [loggers.py:259] Engine 000: Avg prompt throughput: 120.3 tokens/s, Avg generation throughput: 445.7 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 16.1%, Prefix cache hit rate: 91.0% [repeated 49x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:27:47] OK: 377 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:27:47] OK: RSS 1.39 GiB | node mem 379.5/858.0 GiB used (44.2%), avail 478.5 GiB (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 27x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:27:51 [loggers.py:259] Engine 000: Avg prompt throughput: 419.0 tokens/s, Avg generation throughput: 418.2 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 23.0%, Prefix cache hit rate: 91.6% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 18x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:27:56 [loggers.py:259] Engine 000: Avg prompt throughput: 209.6 tokens/s, Avg generation throughput: 490.4 tokens/s, Running: 12 reqs, Waiting: 0 reqs, GPU KV cache usage: 24.2%, Prefix cache hit rate: 91.7% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:27:57] OK: 311 / 131,072 FDs open (0.2% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:27:57] OK: RSS 1.33 GiB | node mem 384.8/858.0 GiB used (44.9%), avail 473.1 GiB (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 21x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:28:01 [loggers.py:259] Engine 000: Avg prompt throughput: 83.6 tokens/s, Avg generation throughput: 463.9 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 18.1%, Prefix cache hit rate: 91.6% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 17x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] [05:28:05] OK: 322 / 131,072 FDs open (0.2% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] [05:28:05] OK: RSS 1.35 GiB | node mem 397.4/858.0 GiB used (46.3%), avail 460.5 GiB (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:28:06 [loggers.py:259] Engine 000: Avg prompt throughput: 115.1 tokens/s, Avg generation throughput: 390.6 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 17.1%, Prefix cache hit rate: 91.8% [repeated 47x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 38x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:28:11 [loggers.py:259] Engine 000: Avg prompt throughput: 265.6 tokens/s, Avg generation throughput: 433.7 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 19.0%, Prefix cache hit rate: 91.9% [repeated 47x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] [05:28:14] OK: 318 / 131,072 FDs open (0.2% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] [05:28:14] OK: RSS 1.34 GiB | node mem 388.6/858.0 GiB used (45.3%), avail 469.3 GiB (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 15x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:28:16 [loggers.py:259] Engine 000: Avg prompt throughput: 135.0 tokens/s, Avg generation throughput: 378.8 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 20.3%, Prefix cache hit rate: 92.0% [repeated 47x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 21x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (skyrl_entrypoint pid=487747) [fd-monitor] [05:28:20] OK: 417 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (skyrl_entrypoint pid=487747) [fd-monitor] [05:28:20] OK: RSS 1.82 GiB | node mem 241.2/858.0 GiB used (28.1%), avail 616.7 GiB (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:28:21 [loggers.py:259] Engine 000: Avg prompt throughput: 166.5 tokens/s, Avg generation throughput: 444.3 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 24.7%, Prefix cache hit rate: 92.1% [repeated 46x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 20x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:28:26 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 445.7 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.2%, Prefix cache hit rate: 92.1% [repeated 46x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 20x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:28:31 [loggers.py:259] Engine 000: Avg prompt throughput: 164.7 tokens/s, Avg generation throughput: 371.0 tokens/s, Running: 10 reqs, Waiting: 0 reqs, GPU KV cache usage: 25.3%, Prefix cache hit rate: 92.2% [repeated 46x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 33x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:28:36 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 438.5 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 24.9%, Prefix cache hit rate: 92.2% [repeated 47x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 35x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) INFO 06-12 05:28:42 [loggers.py:259] Engine 000: Avg prompt throughput: 47.3 tokens/s, Avg generation throughput: 430.1 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 25.3%, Prefix cache hit rate: 92.2% [repeated 47x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 26x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) INFO 06-12 05:28:47 [loggers.py:259] Engine 000: Avg prompt throughput: 103.4 tokens/s, Avg generation throughput: 216.6 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 13.0%, Prefix cache hit rate: 91.8% [repeated 49x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 35x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) INFO 06-12 05:28:52 [loggers.py:259] Engine 000: Avg prompt throughput: 86.9 tokens/s, Avg generation throughput: 225.8 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 13.4%, Prefix cache hit rate: 91.7% [repeated 49x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 23x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) INFO 06-12 05:28:57 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 247.7 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 13.8%, Prefix cache hit rate: 91.7% [repeated 47x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 15x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) INFO 06-12 05:29:02 [loggers.py:259] Engine 000: Avg prompt throughput: 72.2 tokens/s, Avg generation throughput: 199.5 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 13.8%, Prefix cache hit rate: 85.9% [repeated 46x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 16x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) INFO 06-12 05:29:07 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 216.5 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 14.1%, Prefix cache hit rate: 85.9% [repeated 47x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 17x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) INFO 06-12 05:29:12 [loggers.py:259] Engine 000: Avg prompt throughput: 173.9 tokens/s, Avg generation throughput: 167.9 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 14.6%, Prefix cache hit rate: 87.4% [repeated 47x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 13x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) INFO 06-12 05:29:17 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 149.6 tokens/s, Running: 3 reqs, Waiting: 0 reqs, GPU KV cache usage: 9.2%, Prefix cache hit rate: 91.8% [repeated 47x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) INFO 06-12 05:29:22 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 148.4 tokens/s, Running: 3 reqs, Waiting: 0 reqs, GPU KV cache usage: 9.4%, Prefix cache hit rate: 91.8% [repeated 47x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 25x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) INFO 06-12 05:29:27 [loggers.py:259] Engine 000: Avg prompt throughput: 59.4 tokens/s, Avg generation throughput: 154.8 tokens/s, Running: 4 reqs, Waiting: 0 reqs, GPU KV cache usage: 9.7%, Prefix cache hit rate: 91.7% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 13x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) INFO 06-12 05:29:32 [loggers.py:259] Engine 000: Avg prompt throughput: 68.4 tokens/s, Avg generation throughput: 294.3 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 10.3%, Prefix cache hit rate: 87.4% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 20x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) INFO 06-12 05:29:37 [loggers.py:259] Engine 000: Avg prompt throughput: 75.5 tokens/s, Avg generation throughput: 228.1 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 14.5%, Prefix cache hit rate: 88.2% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 7x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) INFO 06-12 05:29:42 [loggers.py:259] Engine 000: Avg prompt throughput: 978.5 tokens/s, Avg generation throughput: 340.1 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 13.8%, Prefix cache hit rate: 86.4% [repeated 47x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 18x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) INFO 06-12 05:29:47 [loggers.py:259] Engine 000: Avg prompt throughput: 107.4 tokens/s, Avg generation throughput: 326.3 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.2%, Prefix cache hit rate: 88.9% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:29:47] OK: 404 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:29:47] OK: RSS 1.56 GiB | node mem 379.7/858.0 GiB used (44.3%), avail 478.3 GiB (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 5x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) INFO 06-12 05:29:52 [loggers.py:259] Engine 000: Avg prompt throughput: 22.5 tokens/s, Avg generation throughput: 314.1 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 23.1%, Prefix cache hit rate: 89.1% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 23x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:29:57] OK: 348 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:29:57] OK: RSS 1.44 GiB | node mem 384.9/858.0 GiB used (44.9%), avail 473.0 GiB (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) INFO 06-12 05:29:57 [loggers.py:259] Engine 000: Avg prompt throughput: 153.4 tokens/s, Avg generation throughput: 287.7 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 20.5%, Prefix cache hit rate: 89.3% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) INFO 06-12 05:30:02 [loggers.py:259] Engine 000: Avg prompt throughput: 121.2 tokens/s, Avg generation throughput: 233.9 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 22.6%, Prefix cache hit rate: 89.4% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 16x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] [05:30:05] OK: 403 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] [05:30:05] OK: RSS 1.53 GiB | node mem 397.7/858.0 GiB used (46.4%), avail 460.3 GiB (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) INFO 06-12 05:30:07 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 300.8 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 23.0%, Prefix cache hit rate: 89.4% [repeated 49x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 14x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) INFO 06-12 05:30:12 [loggers.py:259] Engine 000: Avg prompt throughput: 23.2 tokens/s, Avg generation throughput: 284.3 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 22.7%, Prefix cache hit rate: 89.6% [repeated 47x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] [05:30:14] OK: 316 / 131,072 FDs open (0.2% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] [05:30:14] OK: RSS 1.49 GiB | node mem 388.8/858.0 GiB used (45.3%), avail 469.2 GiB (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 36x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) INFO 06-12 05:30:17 [loggers.py:259] Engine 000: Avg prompt throughput: 280.8 tokens/s, Avg generation throughput: 237.3 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 24.2%, Prefix cache hit rate: 89.6% [repeated 47x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 28x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (skyrl_entrypoint pid=487747) [fd-monitor] [05:30:20] OK: 497 / 131,072 FDs open (0.4% of soft limit, hard limit: 131,072) (skyrl_entrypoint pid=487747) [fd-monitor] [05:30:20] OK: RSS 1.85 GiB | node mem 241.3/858.0 GiB used (28.1%), avail 616.7 GiB (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) INFO 06-12 05:30:22 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 461.5 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 25.7%, Prefix cache hit rate: 94.8% [repeated 52x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 23x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) INFO 06-12 05:30:27 [loggers.py:259] Engine 000: Avg prompt throughput: 34.8 tokens/s, Avg generation throughput: 334.2 tokens/s, Running: 7 reqs, Waiting: 0 reqs, GPU KV cache usage: 13.9%, Prefix cache hit rate: 88.1% [repeated 49x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 30x across cluster] (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) INFO 06-12 05:30:32 [loggers.py:259] Engine 000: Avg prompt throughput: 104.6 tokens/s, Avg generation throughput: 284.5 tokens/s, Running: 7 reqs, Waiting: 0 reqs, GPU KV cache usage: 14.3%, Prefix cache hit rate: 89.5% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 13x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) INFO 06-12 05:30:37 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 327.3 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 13.9%, Prefix cache hit rate: 89.5% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 17x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) INFO 06-12 05:30:42 [loggers.py:259] Engine 000: Avg prompt throughput: 245.0 tokens/s, Avg generation throughput: 227.5 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 14.1%, Prefix cache hit rate: 89.4% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 20x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) INFO 06-12 05:30:47 [loggers.py:259] Engine 000: Avg prompt throughput: 75.0 tokens/s, Avg generation throughput: 228.3 tokens/s, Running: 4 reqs, Waiting: 0 reqs, GPU KV cache usage: 11.8%, Prefix cache hit rate: 89.2% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) INFO 06-12 05:30:52 [loggers.py:259] Engine 000: Avg prompt throughput: 193.7 tokens/s, Avg generation throughput: 223.3 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 14.3%, Prefix cache hit rate: 89.2% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 37x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) INFO 06-12 05:30:57 [loggers.py:259] Engine 000: Avg prompt throughput: 136.7 tokens/s, Avg generation throughput: 316.9 tokens/s, Running: 7 reqs, Waiting: 0 reqs, GPU KV cache usage: 16.4%, Prefix cache hit rate: 89.5% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) INFO 06-12 05:31:02 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 335.9 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 14.8%, Prefix cache hit rate: 89.5% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 41x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) INFO 06-12 05:31:07 [loggers.py:259] Engine 000: Avg prompt throughput: 191.9 tokens/s, Avg generation throughput: 356.0 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 16.8%, Prefix cache hit rate: 89.7% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 40x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) INFO 06-12 05:31:12 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 374.9 tokens/s, Running: 7 reqs, Waiting: 0 reqs, GPU KV cache usage: 17.1%, Prefix cache hit rate: 89.7% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 23x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) INFO 06-12 05:31:17 [loggers.py:259] Engine 000: Avg prompt throughput: 252.6 tokens/s, Avg generation throughput: 357.7 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 12.9%, Prefix cache hit rate: 89.5% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 22x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) INFO 06-12 05:31:22 [loggers.py:259] Engine 000: Avg prompt throughput: 61.2 tokens/s, Avg generation throughput: 315.7 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 13.8%, Prefix cache hit rate: 90.3% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) INFO 06-12 05:31:27 [loggers.py:259] Engine 000: Avg prompt throughput: 137.4 tokens/s, Avg generation throughput: 294.9 tokens/s, Running: 7 reqs, Waiting: 0 reqs, GPU KV cache usage: 17.7%, Prefix cache hit rate: 90.6% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 5x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) INFO 06-12 05:31:32 [loggers.py:259] Engine 000: Avg prompt throughput: 68.2 tokens/s, Avg generation throughput: 378.6 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 20.4%, Prefix cache hit rate: 91.0% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 35x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) INFO 06-12 05:31:37 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 389.2 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 20.9%, Prefix cache hit rate: 91.0% [repeated 49x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) INFO 06-12 05:31:42 [loggers.py:259] Engine 000: Avg prompt throughput: 41.7 tokens/s, Avg generation throughput: 364.8 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 17.3%, Prefix cache hit rate: 91.1% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 35x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) INFO 06-12 05:31:47 [loggers.py:259] Engine 000: Avg prompt throughput: 93.9 tokens/s, Avg generation throughput: 303.2 tokens/s, Running: 7 reqs, Waiting: 0 reqs, GPU KV cache usage: 20.2%, Prefix cache hit rate: 92.0% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:31:47] OK: 415 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:31:47] OK: RSS 1.76 GiB | node mem 380.0/858.0 GiB used (44.3%), avail 478.0 GiB (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 31x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) INFO 06-12 05:31:52 [loggers.py:259] Engine 000: Avg prompt throughput: 16.0 tokens/s, Avg generation throughput: 331.4 tokens/s, Running: 7 reqs, Waiting: 0 reqs, GPU KV cache usage: 20.7%, Prefix cache hit rate: 92.2% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 13x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:31:57] OK: 333 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:31:57] OK: RSS 1.59 GiB | node mem 385.1/858.0 GiB used (44.9%), avail 472.8 GiB (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) INFO 06-12 05:31:57 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 292.2 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 16.9%, Prefix cache hit rate: 92.2% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 7x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) INFO 06-12 05:32:02 [loggers.py:259] Engine 000: Avg prompt throughput: 189.0 tokens/s, Avg generation throughput: 299.3 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 20.3%, Prefix cache hit rate: 92.6% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] [05:32:05] OK: 365 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] [05:32:05] OK: RSS 1.75 GiB | node mem 398.1/858.0 GiB used (46.4%), avail 459.9 GiB (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 23x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) INFO 06-12 05:32:07 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 292.6 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 20.7%, Prefix cache hit rate: 92.6% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 31x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) INFO 06-12 05:32:12 [loggers.py:259] Engine 000: Avg prompt throughput: 319.2 tokens/s, Avg generation throughput: 264.8 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.3%, Prefix cache hit rate: 92.9% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] [05:32:14] OK: 314 / 131,072 FDs open (0.2% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] [05:32:14] OK: RSS 1.71 GiB | node mem 389.0/858.0 GiB used (45.3%), avail 469.0 GiB (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 27x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) INFO 06-12 05:32:18 [loggers.py:259] Engine 000: Avg prompt throughput: 161.5 tokens/s, Avg generation throughput: 264.5 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.3%, Prefix cache hit rate: 93.4% [repeated 49x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (skyrl_entrypoint pid=487747) [fd-monitor] [05:32:20] OK: 505 / 131,072 FDs open (0.4% of soft limit, hard limit: 131,072) (skyrl_entrypoint pid=487747) [fd-monitor] [05:32:20] OK: RSS 1.88 GiB | node mem 241.3/858.0 GiB used (28.1%), avail 616.6 GiB (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 31x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) INFO 06-12 05:32:23 [loggers.py:259] Engine 000: Avg prompt throughput: 339.5 tokens/s, Avg generation throughput: 372.5 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 24.0%, Prefix cache hit rate: 93.4% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 24x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) INFO 06-12 05:32:28 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 397.1 tokens/s, Running: 7 reqs, Waiting: 0 reqs, GPU KV cache usage: 23.7%, Prefix cache hit rate: 93.4% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 25x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) INFO 06-12 05:32:33 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 234.9 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 13.7%, Prefix cache hit rate: 94.4% [repeated 49x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 15x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) INFO 06-12 05:32:38 [loggers.py:259] Engine 000: Avg prompt throughput: 68.7 tokens/s, Avg generation throughput: 208.6 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 16.8%, Prefix cache hit rate: 94.5% [repeated 49x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:32:41 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:32:41 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:32:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:32:41 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:32:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:32:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:32:41 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:32:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:32:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:32:41 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:32:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:32:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:32:41 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:32:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:32:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:32:41 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:32:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:32:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:32:41 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:32:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:32:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:32:41 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:32:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:32:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:32:41 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:32:41 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) INFO 06-12 05:32:43 [loggers.py:259] Engine 000: Avg prompt throughput: 165.4 tokens/s, Avg generation throughput: 213.1 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 17.3%, Prefix cache hit rate: 94.4% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:32:41 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:32:41 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:32:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:32:41 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:32:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:32:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:32:41 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:32:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:32:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:32:41 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:32:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:32:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:32:41 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:32:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:32:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:32:41 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:32:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:32:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:32:41 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:32:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:32:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:32:41 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:32:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:32:41 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:32:41 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 8x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:32:48 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 317.3 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 27.9%, Prefix cache hit rate: 91.7% [repeated 49x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 32x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:32:53 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 317.3 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 28.4%, Prefix cache hit rate: 91.7% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 41x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:32:58 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 309.4 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 28.8%, Prefix cache hit rate: 91.7% [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:32:59 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:32:59 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:32:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:32:59 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:32:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:32:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:32:59 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:32:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:32:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:32:59 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:32:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:32:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:32:59 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:32:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:32:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:32:59 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:32:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:32:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:32:59 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:32:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:32:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:32:59 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:32:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:32:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:32:59 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:32:59 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 18x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:33:03 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 309.7 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 29.2%, Prefix cache hit rate: 91.7% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 22x across cluster] (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:33:08 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 308.5 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 29.7%, Prefix cache hit rate: 91.7% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 5x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) INFO 06-12 05:33:13 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 181.5 tokens/s, Running: 4 reqs, Waiting: 0 reqs, GPU KV cache usage: 19.9%, Prefix cache hit rate: 92.5% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 23x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) INFO 06-12 05:33:18 [loggers.py:259] Engine 000: Avg prompt throughput: 70.7 tokens/s, Avg generation throughput: 181.6 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 22.4%, Prefix cache hit rate: 92.5% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 20x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) INFO 06-12 05:33:23 [loggers.py:259] Engine 000: Avg prompt throughput: 69.6 tokens/s, Avg generation throughput: 202.7 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 22.8%, Prefix cache hit rate: 92.5% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 27x across cluster] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) INFO 06-12 05:33:28 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 229.5 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 23.1%, Prefix cache hit rate: 92.5% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 13x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) INFO 06-12 05:33:33 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 232.4 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 23.4%, Prefix cache hit rate: 92.5% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 7x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) INFO 06-12 05:33:38 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 214.5 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 23.5%, Prefix cache hit rate: 92.5% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 20x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:33:42 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:33:42 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:33:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:33:42 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:33:42 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:33:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:33:42 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:33:42 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:33:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:33:42 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:33:42 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:33:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:33:42 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:33:42 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:33:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:33:42 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:33:42 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:33:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:33:42 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:33:42 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:33:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:33:42 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:33:42 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:33:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:33:42 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:33:42 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) INFO 06-12 05:33:43 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 194.0 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 23.8%, Prefix cache hit rate: 92.5% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 16x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:33:42 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:33:42 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:33:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:33:42 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:33:42 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:33:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:33:42 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:33:42 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:33:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:33:42 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:33:42 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:33:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:33:42 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:33:42 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:33:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:33:42 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:33:42 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:33:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:33:42 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:33:42 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:33:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:33:42 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:33:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:33:42 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:33:42 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:33:47] OK: 398 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:33:47] OK: RSS 1.94 GiB | node mem 380.2/858.0 GiB used (44.3%), avail 477.8 GiB (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) INFO 06-12 05:33:48 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 191.5 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 24.0%, Prefix cache hit rate: 92.5% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) INFO 06-12 05:33:53 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 189.1 tokens/s, Running: 4 reqs, Waiting: 0 reqs, GPU KV cache usage: 15.0%, Prefix cache hit rate: 91.8% [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:33:54 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:33:54 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:33:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:33:54 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:33:54 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:33:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:33:54 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:33:54 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:33:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:33:54 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:33:54 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:33:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:33:54 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:33:54 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:33:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:33:54 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:33:54 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:33:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:33:54 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:33:54 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:33:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:33:54 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:33:54 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:33:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:33:54 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:33:54 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 5x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:33:57] OK: 379 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:33:57] OK: RSS 1.81 GiB | node mem 385.4/858.0 GiB used (44.9%), avail 472.6 GiB (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) INFO 06-12 05:33:58 [loggers.py:259] Engine 000: Avg prompt throughput: 140.1 tokens/s, Avg generation throughput: 208.1 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.7%, Prefix cache hit rate: 92.0% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:33:56 [serving.py:315] Error in preprocessing prompt inputs [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:33:56 [serving.py:315] Traceback (most recent call last): [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:33:56 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:33:56 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:33:56 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:33:56 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:33:56 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:33:56 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:33:56 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:33:56 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:33:56 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:33:56 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:33:56 [serving.py:315] return await asyncio.gather( [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:33:56 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:33:56 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:33:56 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:33:56 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:33:56 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:33:56 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:33:56 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:33:56 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:33:56 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:33:56 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:33:56 [serving.py:315] raise VLLMValidationError( [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:33:56 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 3x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 3x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 3x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 28x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) INFO 06-12 05:34:03 [loggers.py:259] Engine 000: Avg prompt throughput: 20.2 tokens/s, Avg generation throughput: 281.2 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 22.4%, Prefix cache hit rate: 92.0% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:33:59 [serving.py:315] Error in preprocessing prompt inputs [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:33:59 [serving.py:315] Traceback (most recent call last): [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:33:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:33:59 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:33:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:33:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:33:59 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:33:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:33:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:33:59 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:33:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:33:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:33:59 [serving.py:315] return await asyncio.gather( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:33:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:33:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:33:59 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:33:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:33:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:33:59 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:33:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:33:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:33:59 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:33:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:33:59 [serving.py:315] raise VLLMValidationError( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:33:59 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 2x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] [05:34:05] OK: 386 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] [05:34:05] OK: RSS 2.14 GiB | node mem 398.5/858.0 GiB used (46.4%), avail 459.5 GiB (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 23x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) INFO 06-12 05:34:08 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 260.7 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 22.4%, Prefix cache hit rate: 92.0% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:04 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:04 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:04 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:04 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:04 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:04 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:04 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:04 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:04 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:04 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:04 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:04 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:04 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:04 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:04 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:04 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:04 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:04 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:04 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:04 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:04 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:04 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:04 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:04 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:04 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 22x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) INFO 06-12 05:34:13 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 232.0 tokens/s, Running: 4 reqs, Waiting: 0 reqs, GPU KV cache usage: 18.6%, Prefix cache hit rate: 92.0% [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:34:13 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:34:13 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:34:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:34:13 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:34:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:34:13 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:34:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:34:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:34:13 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:34:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:34:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:34:13 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:34:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:34:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:34:13 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:34:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:34:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:34:13 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:34:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:34:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:34:13 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:34:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:34:13 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:34:13 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] [05:34:14] OK: 416 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] [05:34:14] OK: RSS 2.05 GiB | node mem 389.5/858.0 GiB used (45.4%), avail 468.5 GiB (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:34:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 15x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) INFO 06-12 05:34:18 [loggers.py:259] Engine 000: Avg prompt throughput: 124.7 tokens/s, Avg generation throughput: 205.0 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 23.1%, Prefix cache hit rate: 92.1% [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:18 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:18 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:18 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:18 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:18 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:18 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:18 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:18 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:18 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:18 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:18 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:18 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:18 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:18 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:18 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:18 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:18 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:18 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:18 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:18 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:18 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:18 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:18 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:18 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:18 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (skyrl_entrypoint pid=487747) [fd-monitor] [05:34:20] OK: 555 / 131,072 FDs open (0.4% of soft limit, hard limit: 131,072) (skyrl_entrypoint pid=487747) [fd-monitor] [05:34:20] OK: RSS 1.91 GiB | node mem 241.5/858.0 GiB used (28.1%), avail 616.5 GiB (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:21 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:21 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:21 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:21 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:21 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:21 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:21 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:21 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:21 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:21 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:34:21 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 33x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) INFO 06-12 05:34:23 [loggers.py:259] Engine 000: Avg prompt throughput: 165.2 tokens/s, Avg generation throughput: 256.4 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 15.2%, Prefix cache hit rate: 92.0% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:34:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 6x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 3x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:34:25 [serving.py:315] Error in preprocessing prompt inputs [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:34:25 [serving.py:315] Traceback (most recent call last): [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:34:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:34:25 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:34:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:34:25 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:34:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:34:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:34:25 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:34:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:34:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:34:25 [serving.py:315] return await asyncio.gather( [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:34:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:34:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:34:25 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:34:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:34:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:34:25 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:34:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:34:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:34:25 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:34:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:34:25 [serving.py:315] raise VLLMValidationError( [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:34:25 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 3x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 7x across cluster] (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) INFO 06-12 05:34:28 [loggers.py:259] Engine 000: Avg prompt throughput: 23.2 tokens/s, Avg generation throughput: 259.1 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 19.3%, Prefix cache hit rate: 92.2% [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:34:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:34:31 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:34:31 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:34:31 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:34:31 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:34:31 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:34:31 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:34:31 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:34:31 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:34:31 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:34:31 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:34:31 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:34:31 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:34:31 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:34:31 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:34:31 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:34:31 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:34:31 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:34:31 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:34:31 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:34:31 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:34:31 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:34:31 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:34:31 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:34:31 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 5x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) INFO 06-12 05:34:33 [loggers.py:259] Engine 000: Avg prompt throughput: 59.6 tokens/s, Avg generation throughput: 236.1 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 15.3%, Prefix cache hit rate: 92.2% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:34:31 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 74x across cluster] (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:37 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:37 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:37 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:37 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:37 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:37 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:37 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:37 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:37 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:37 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:37 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:37 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:37 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:37 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:37 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:37 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:34:39 [loggers.py:259] Engine 000: Avg prompt throughput: 264.3 tokens/s, Avg generation throughput: 384.4 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 32.6%, Prefix cache hit rate: 91.7% [repeated 50x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:34:37 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 18x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:34:41 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:34:41 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:34:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:34:41 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:34:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:34:41 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:34:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:34:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:34:41 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:34:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:34:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:34:41 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:34:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:34:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:34:41 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:34:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:34:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:34:41 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:34:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:34:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:34:41 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:34:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:34:41 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:34:41 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:34:44 [loggers.py:259] Engine 000: Avg prompt throughput: 183.2 tokens/s, Avg generation throughput: 391.2 tokens/s, Running: 10 reqs, Waiting: 0 reqs, GPU KV cache usage: 26.9%, Prefix cache hit rate: 91.7% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:34:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:34:45 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:34:45 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:34:45 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:34:45 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:34:45 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:34:45 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:34:45 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:34:45 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:34:45 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:34:45 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:34:45 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:34:45 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:34:45 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:34:45 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:34:45 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:34:45 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:34:45 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:34:45 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:34:45 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:34:45 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:34:45 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:34:45 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:34:45 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:34:45 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 22x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:34:49 [loggers.py:259] Engine 000: Avg prompt throughput: 96.6 tokens/s, Avg generation throughput: 330.2 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 23.4%, Prefix cache hit rate: 91.7% [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:34:49 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 6x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:34:49 [serving.py:315] Error in preprocessing prompt inputs [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:34:49 [serving.py:315] Traceback (most recent call last): [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:34:49 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:34:49 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:34:49 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:34:49 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:34:49 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:34:49 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:34:49 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:34:49 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:34:49 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:34:49 [serving.py:315] return await asyncio.gather( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:34:49 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:34:49 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:34:49 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:34:49 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:34:49 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:34:49 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:34:49 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:34:49 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:34:49 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:34:49 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:34:49 [serving.py:315] raise VLLMValidationError( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:34:49 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 2x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 8x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:34:54 [loggers.py:259] Engine 000: Avg prompt throughput: 38.8 tokens/s, Avg generation throughput: 290.2 tokens/s, Running: 7 reqs, Waiting: 0 reqs, GPU KV cache usage: 27.6%, Prefix cache hit rate: 91.9% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:34:54 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:34:54 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:34:54 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:34:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:34:54 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:34:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:34:54 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:34:54 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:34:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:34:54 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:34:54 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:34:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:34:54 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:34:54 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:34:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:34:54 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:34:54 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:34:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:34:54 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:34:54 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:34:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:34:54 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:34:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:34:54 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:34:54 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 27x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:34:59 [loggers.py:259] Engine 000: Avg prompt throughput: 507.4 tokens/s, Avg generation throughput: 301.0 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 39.4%, Prefix cache hit rate: 92.0% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 44x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:35:02 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:35:02 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:35:02 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:35:02 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:35:02 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:35:02 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:35:02 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:35:02 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:35:02 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:35:02 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:35:02 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:35:02 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:35:02 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:35:02 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:35:02 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:35:02 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:35:02 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:35:02 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:35:02 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:35:02 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:35:02 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:35:02 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:35:02 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:35:02 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:35:02 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:35:02 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:35:04 [loggers.py:259] Engine 000: Avg prompt throughput: 103.5 tokens/s, Avg generation throughput: 427.3 tokens/s, Running: 12 reqs, Waiting: 0 reqs, GPU KV cache usage: 40.6%, Prefix cache hit rate: 92.0% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 24x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:35:07 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:35:07 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:35:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:35:07 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:35:07 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:35:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:35:07 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:35:07 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:35:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:35:07 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:35:07 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:35:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:35:07 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:35:07 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:35:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:35:07 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:35:07 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:35:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:35:07 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:35:07 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:35:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:35:07 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:35:07 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:35:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:35:07 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:35:07 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:35:09 [loggers.py:259] Engine 000: Avg prompt throughput: 93.3 tokens/s, Avg generation throughput: 438.8 tokens/s, Running: 12 reqs, Waiting: 0 reqs, GPU KV cache usage: 38.2%, Prefix cache hit rate: 92.0% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 23x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:35:11 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:35:11 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:35:11 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:35:11 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:35:11 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:35:11 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:35:11 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:35:11 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:35:11 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:35:11 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:35:11 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:35:11 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:35:11 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:35:11 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:35:11 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:35:11 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:35:11 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:35:11 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:35:11 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:35:11 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:35:11 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:35:11 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:35:11 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:35:11 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:35:11 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:35:14 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 386.6 tokens/s, Running: 7 reqs, Waiting: 0 reqs, GPU KV cache usage: 29.0%, Prefix cache hit rate: 92.0% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:14 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:14 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:14 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:14 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:14 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:14 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:14 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:14 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:14 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:14 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:14 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:14 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:14 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:14 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:14 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:14 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:14 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:14 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:14 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:14 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:14 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:14 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:14 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:14 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 36x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:35:17 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 6x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 4x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 4x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:35:19 [loggers.py:259] Engine 000: Avg prompt throughput: 190.3 tokens/s, Avg generation throughput: 304.3 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 32.6%, Prefix cache hit rate: 92.0% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:35:17 [serving.py:315] Error in preprocessing prompt inputs [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:35:17 [serving.py:315] Traceback (most recent call last): [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:35:17 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:35:17 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:35:17 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:35:17 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:35:17 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:35:17 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:35:17 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:35:17 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:35:17 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:35:17 [serving.py:315] return await asyncio.gather( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:35:17 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:35:17 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:35:17 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:35:17 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:35:17 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:35:17 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:35:17 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:35:17 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:35:17 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:35:17 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:35:17 [serving.py:315] raise VLLMValidationError( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:35:17 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 2x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 22x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:23 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 3x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 3x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:35:24 [loggers.py:259] Engine 000: Avg prompt throughput: 504.5 tokens/s, Avg generation throughput: 344.6 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 34.5%, Prefix cache hit rate: 92.1% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:23 [serving.py:315] Error in preprocessing prompt inputs [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:23 [serving.py:315] Traceback (most recent call last): [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:23 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:23 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:23 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:23 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:23 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:23 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:23 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:23 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:23 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:23 [serving.py:315] return await asyncio.gather( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:23 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:23 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:23 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:23 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:23 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:23 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:23 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:23 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:23 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:23 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:23 [serving.py:315] raise VLLMValidationError( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:23 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 2x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 38x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:27 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 3x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 3x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:35:29 [loggers.py:259] Engine 000: Avg prompt throughput: 88.5 tokens/s, Avg generation throughput: 340.2 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 34.4%, Prefix cache hit rate: 92.2% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:27 [serving.py:315] Error in preprocessing prompt inputs [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:27 [serving.py:315] Traceback (most recent call last): [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:27 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:27 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:27 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:27 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:27 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:27 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:27 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:27 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:27 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:27 [serving.py:315] return await asyncio.gather( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:27 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:27 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:27 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:27 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:27 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:27 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:27 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:27 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:27 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:27 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:27 [serving.py:315] raise VLLMValidationError( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:35:27 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 2x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 56x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:35:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:35:34 [loggers.py:259] Engine 000: Avg prompt throughput: 122.7 tokens/s, Avg generation throughput: 340.5 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 33.5%, Prefix cache hit rate: 92.2% [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:35:33 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:35:33 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:35:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:35:33 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:35:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:35:33 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:35:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:35:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:35:33 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:35:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:35:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:35:33 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:35:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:35:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:35:33 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:35:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:35:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:35:33 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:35:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:35:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:35:33 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:35:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:35:33 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:35:33 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 23x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:35:39 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:35:39 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:35:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:35:39 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:35:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:35:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:35:39 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:35:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:35:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:35:39 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:35:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:35:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:35:39 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:35:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:35:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:35:39 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:35:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:35:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:35:39 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:35:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:35:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:35:39 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:35:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:35:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:35:39 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:35:39 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:35:39 [loggers.py:259] Engine 000: Avg prompt throughput: 52.3 tokens/s, Avg generation throughput: 325.3 tokens/s, Running: 7 reqs, Waiting: 0 reqs, GPU KV cache usage: 27.2%, Prefix cache hit rate: 92.3% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 35x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:35:42 [serving.py:315] Error in preprocessing prompt inputs [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:35:42 [serving.py:315] Traceback (most recent call last): [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:35:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:35:42 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:35:42 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:35:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:35:42 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:35:42 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:35:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:35:42 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:35:42 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:35:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:35:42 [serving.py:315] return await asyncio.gather( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:35:42 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:35:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:35:42 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:35:42 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:35:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:35:42 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:35:42 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:35:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:35:42 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:35:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:35:42 [serving.py:315] raise VLLMValidationError( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:35:42 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 2x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:35:44 [loggers.py:259] Engine 000: Avg prompt throughput: 51.5 tokens/s, Avg generation throughput: 275.9 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 31.4%, Prefix cache hit rate: 92.4% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 23x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:35:47] OK: 411 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:35:47] OK: RSS 2.32 GiB | node mem 380.7/858.0 GiB used (44.4%), avail 477.3 GiB (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:35:44 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:35:44 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:35:44 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:35:44 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:35:44 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:35:44 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:35:44 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:35:44 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:35:44 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:35:44 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:35:44 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:35:44 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:35:44 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:35:44 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:35:44 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:35:44 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:35:44 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:35:44 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:35:44 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:35:44 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:35:44 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:35:44 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:35:44 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:35:44 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:35:44 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:35:49 [loggers.py:259] Engine 000: Avg prompt throughput: 108.2 tokens/s, Avg generation throughput: 321.7 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 33.0%, Prefix cache hit rate: 92.4% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 25x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) INFO 06-12 05:35:53 [loggers.py:259] Engine 000: Avg prompt throughput: 51.6 tokens/s, Avg generation throughput: 253.8 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 92.1% [repeated 47x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:35:57] OK: 370 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:35:57] OK: RSS 2.01 GiB | node mem 385.6/858.0 GiB used (44.9%), avail 472.3 GiB (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 45x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:35:57 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:35:57 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:35:57 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:35:57 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:35:57 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:35:57 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:35:57 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:35:57 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:35:57 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:35:57 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:35:57 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:35:57 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:35:57 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:35:57 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:35:57 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:35:57 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:35:57 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:35:57 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:35:57 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:35:57 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:35:57 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:35:57 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:35:57 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:35:57 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:35:57 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:35:57 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:35:59 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 309.8 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 33.0%, Prefix cache hit rate: 92.4% [repeated 49x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 15x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:36:03 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:36:03 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:36:03 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:36:03 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:36:03 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:36:03 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:36:03 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:36:03 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:36:03 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:36:03 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:36:03 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:36:03 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:36:03 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:36:03 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:36:03 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:36:03 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:36:03 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:36:03 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:36:03 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:36:03 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:36:03 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:36:03 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:36:03 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:36:03 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:36:03 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:36:03 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:36:04 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 303.4 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 33.5%, Prefix cache hit rate: 92.4% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] [05:36:05] OK: 381 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] [05:36:05] OK: RSS 2.28 GiB | node mem 398.6/858.0 GiB used (46.5%), avail 459.3 GiB (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 33x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:05 [serving.py:315] Error in preprocessing prompt inputs [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:05 [serving.py:315] Traceback (most recent call last): [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:05 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:05 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:05 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:05 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:05 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:05 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:05 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:05 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:05 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:05 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:05 [serving.py:315] return await asyncio.gather( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:05 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:05 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:05 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:05 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:05 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:05 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:05 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:05 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:05 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:05 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:05 [serving.py:315] raise VLLMValidationError( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:05 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 2x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:36:09 [loggers.py:259] Engine 000: Avg prompt throughput: 137.3 tokens/s, Avg generation throughput: 277.6 tokens/s, Running: 7 reqs, Waiting: 0 reqs, GPU KV cache usage: 23.2%, Prefix cache hit rate: 92.3% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 13x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:36:13 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:36:13 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:36:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:36:13 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:36:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:36:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:36:13 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:36:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:36:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:36:13 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:36:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:36:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:36:13 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:36:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:36:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:36:13 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:36:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:36:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:36:13 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:36:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:36:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:36:13 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:36:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:36:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:36:13 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:36:13 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) INFO 06-12 05:36:14 [loggers.py:259] Engine 000: Avg prompt throughput: 72.5 tokens/s, Avg generation throughput: 369.5 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 30.5%, Prefix cache hit rate: 92.1% [repeated 47x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] [05:36:14] OK: 390 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] [05:36:14] OK: RSS 2.19 GiB | node mem 389.6/858.0 GiB used (45.4%), avail 468.4 GiB (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 25x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:36:19 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:36:19 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:36:19 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:36:19 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:36:19 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:36:19 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:36:19 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:36:19 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:36:19 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:36:19 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:36:19 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:36:19 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:36:19 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:36:19 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:36:19 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:36:19 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:36:19 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:36:19 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:36:19 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:36:19 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:36:19 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:36:19 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:36:19 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:36:19 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:36:19 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:36:19 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:36:19 [loggers.py:259] Engine 000: Avg prompt throughput: 150.3 tokens/s, Avg generation throughput: 232.6 tokens/s, Running: 7 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.5%, Prefix cache hit rate: 92.3% [repeated 49x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (skyrl_entrypoint pid=487747) [fd-monitor] [05:36:20] OK: 564 / 131,072 FDs open (0.4% of soft limit, hard limit: 131,072) (skyrl_entrypoint pid=487747) [fd-monitor] [05:36:20] OK: RSS 1.94 GiB | node mem 241.6/858.0 GiB used (28.2%), avail 616.4 GiB (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 46x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:36:23 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:36:23 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:36:23 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:36:23 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:36:23 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:36:23 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:36:23 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:36:23 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:36:23 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:36:23 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:36:23 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:36:23 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:36:23 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:36:23 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:36:23 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:36:23 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:36:23 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:36:23 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:36:23 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:36:23 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:36:23 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:36:23 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:36:23 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:36:23 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) ERROR 06-12 05:36:23 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:36:24 [loggers.py:259] Engine 000: Avg prompt throughput: 112.4 tokens/s, Avg generation throughput: 314.4 tokens/s, Running: 7 reqs, Waiting: 0 reqs, GPU KV cache usage: 26.6%, Prefix cache hit rate: 92.5% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:36:25 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:36:25 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:36:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:36:25 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:36:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:36:25 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:36:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:36:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:36:25 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:36:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:36:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:36:25 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:36:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:36:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:36:25 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:36:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:36:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:36:25 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:36:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:36:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:36:25 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:36:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:36:25 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:36:25 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 37x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:36:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:36:29 [loggers.py:259] Engine 000: Avg prompt throughput: 120.9 tokens/s, Avg generation throughput: 278.5 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 28.6%, Prefix cache hit rate: 92.5% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:36:30 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:36:30 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:36:30 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:36:30 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:36:30 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:36:30 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:36:30 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:36:30 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:36:30 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:36:30 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:36:30 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:36:30 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:36:30 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:36:30 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:36:30 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:36:30 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:36:30 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:36:30 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:36:30 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:36:30 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:36:30 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:36:30 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:36:30 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:36:30 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 40x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:36:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:36:34 [loggers.py:259] Engine 000: Avg prompt throughput: 20.9 tokens/s, Avg generation throughput: 300.9 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 25.9%, Prefix cache hit rate: 92.5% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:36:34 [serving.py:315] Error in preprocessing prompt inputs [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:36:34 [serving.py:315] Traceback (most recent call last): [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:36:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:36:34 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:36:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:36:34 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:36:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:36:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:36:34 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:36:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:36:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:36:34 [serving.py:315] return await asyncio.gather( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:36:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:36:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:36:34 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:36:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:36:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:36:34 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:36:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:36:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:36:34 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:36:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:36:34 [serving.py:315] raise VLLMValidationError( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:36:34 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 2x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 20x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:36:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:36:39 [loggers.py:259] Engine 000: Avg prompt throughput: 97.5 tokens/s, Avg generation throughput: 339.1 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 28.1%, Prefix cache hit rate: 92.5% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:39 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:39 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:39 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:39 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:39 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:39 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:39 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:39 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:39 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:39 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:36:39 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 35x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:36:43 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:36:43 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:36:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:36:43 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:36:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:36:43 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:36:43 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:36:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:36:43 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:36:43 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:36:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:36:43 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:36:43 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:36:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:36:43 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:36:43 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:36:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:36:43 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:36:43 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:36:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:36:43 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:36:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:36:43 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:36:43 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:36:43 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:36:44 [loggers.py:259] Engine 000: Avg prompt throughput: 134.8 tokens/s, Avg generation throughput: 379.0 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 33.2%, Prefix cache hit rate: 92.6% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 28x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:36:46 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:36:46 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:36:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:36:46 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:36:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:36:46 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:36:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:36:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:36:46 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:36:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:36:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:36:46 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:36:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:36:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:36:46 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:36:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:36:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:36:46 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:36:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:36:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:36:46 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:36:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:36:46 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:36:46 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:36:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:36:49 [loggers.py:259] Engine 000: Avg prompt throughput: 170.9 tokens/s, Avg generation throughput: 389.5 tokens/s, Running: 10 reqs, Waiting: 0 reqs, GPU KV cache usage: 36.4%, Prefix cache hit rate: 92.7% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:36:50 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:36:50 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:36:50 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:36:50 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:36:50 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:36:50 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:36:50 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:36:50 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:36:50 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:36:50 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:36:50 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:36:50 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:36:50 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:36:50 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:36:50 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:36:50 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:36:50 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:36:50 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:36:50 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:36:50 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:36:50 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:36:50 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:36:50 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:36:50 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 23x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:36:51 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:36:54 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 351.3 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 35.1%, Prefix cache hit rate: 92.7% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:36:54 [serving.py:315] Error in preprocessing prompt inputs [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:36:54 [serving.py:315] Traceback (most recent call last): [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:36:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:36:54 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:36:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:36:54 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:36:54 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:36:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:36:54 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:36:54 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:36:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:36:54 [serving.py:315] return await asyncio.gather( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:36:54 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:36:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:36:54 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:36:54 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:36:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:36:54 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:36:54 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:36:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:36:54 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:36:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:36:54 [serving.py:315] raise VLLMValidationError( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:36:54 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 2x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 7x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:36:54 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:36:59 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 300.3 tokens/s, Running: 7 reqs, Waiting: 0 reqs, GPU KV cache usage: 27.2%, Prefix cache hit rate: 92.7% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:37:01 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:37:01 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:37:01 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:37:01 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:37:01 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:37:01 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:37:01 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:37:01 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:37:01 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:37:01 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:37:01 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:37:01 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:37:01 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:37:01 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:37:01 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:37:01 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:37:01 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:37:01 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:37:01 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:37:01 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:37:01 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:37:01 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:37:01 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:37:01 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 43x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:37:03 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:37:04 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 274.5 tokens/s, Running: 7 reqs, Waiting: 0 reqs, GPU KV cache usage: 27.6%, Prefix cache hit rate: 92.7% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:37:03 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:37:03 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:37:03 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:37:03 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:37:03 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:37:03 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:37:03 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:37:03 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:37:03 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:37:03 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:37:03 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:37:03 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:37:03 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:37:03 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:37:03 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:37:03 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:37:03 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:37:03 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:37:03 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:37:03 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:37:03 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:37:03 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:37:03 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:37:03 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 39x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:37:08 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:37:08 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:37:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:37:08 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:37:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:37:08 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:37:08 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:37:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:37:08 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:37:08 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:37:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:37:08 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:37:08 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:37:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:37:08 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:37:08 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:37:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:37:08 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:37:08 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:37:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:37:08 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:37:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:37:08 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:37:08 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) ERROR 06-12 05:37:08 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) INFO 06-12 05:37:09 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 224.9 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 17.2%, Prefix cache hit rate: 92.7% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 32x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:37:10 [serving.py:315] Error in preprocessing prompt inputs [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:37:10 [serving.py:315] Traceback (most recent call last): [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:37:10 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:37:10 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:37:10 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:37:10 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:37:10 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:37:10 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:37:10 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:37:10 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:37:10 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:37:10 [serving.py:315] return await asyncio.gather( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:37:10 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:37:10 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:37:10 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:37:10 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:37:10 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:37:10 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:37:10 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:37:10 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:37:10 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:37:10 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:37:10 [serving.py:315] raise VLLMValidationError( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:37:10 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 2x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:37:10 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) INFO 06-12 05:37:14 [loggers.py:259] Engine 000: Avg prompt throughput: 171.2 tokens/s, Avg generation throughput: 240.0 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 10.8%, Prefix cache hit rate: 93.0% [repeated 49x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 5x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) INFO 06-12 05:37:19 [loggers.py:259] Engine 000: Avg prompt throughput: 241.8 tokens/s, Avg generation throughput: 294.9 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 11.2%, Prefix cache hit rate: 92.9% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 20x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:37:24 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:37:24 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:37:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:37:24 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:37:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:37:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:37:24 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:37:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:37:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:37:24 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:37:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:37:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:37:24 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:37:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:37:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:37:24 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:37:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:37:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:37:24 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:37:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:37:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:37:24 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:37:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:37:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:37:24 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:37:24 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) INFO 06-12 05:37:24 [loggers.py:259] Engine 000: Avg prompt throughput: 151.5 tokens/s, Avg generation throughput: 358.6 tokens/s, Running: 7 reqs, Waiting: 0 reqs, GPU KV cache usage: 24.2%, Prefix cache hit rate: 93.2% [repeated 49x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 33x across cluster] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) INFO 06-12 05:37:29 [loggers.py:259] Engine 000: Avg prompt throughput: 203.0 tokens/s, Avg generation throughput: 349.7 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 29.2%, Prefix cache hit rate: 93.3% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 13x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) INFO 06-12 05:37:34 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 373.3 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 29.7%, Prefix cache hit rate: 93.3% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 35x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) INFO 06-12 05:37:39 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 330.0 tokens/s, Running: 7 reqs, Waiting: 0 reqs, GPU KV cache usage: 25.6%, Prefix cache hit rate: 93.3% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 28x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:37:45 [loggers.py:259] Engine 000: Avg prompt throughput: 262.3 tokens/s, Avg generation throughput: 500.5 tokens/s, Running: 13 reqs, Waiting: 0 reqs, GPU KV cache usage: 47.8%, Prefix cache hit rate: 94.3% [repeated 49x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 24x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:37:47] OK: 401 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:37:47] OK: RSS 2.46 GiB | node mem 380.9/858.0 GiB used (44.4%), avail 477.0 GiB (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:37:50 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 486.7 tokens/s, Running: 12 reqs, Waiting: 0 reqs, GPU KV cache usage: 45.3%, Prefix cache hit rate: 94.3% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 13x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:37:55 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 427.5 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 39.1%, Prefix cache hit rate: 94.3% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:37:57] OK: 363 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:37:57] OK: RSS 2.11 GiB | node mem 385.8/858.0 GiB used (45.0%), avail 472.2 GiB (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 12x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) INFO 06-12 05:37:59 [loggers.py:259] Engine 000: Avg prompt throughput: 40.3 tokens/s, Avg generation throughput: 343.6 tokens/s, Running: 7 reqs, Waiting: 0 reqs, GPU KV cache usage: 20.7%, Prefix cache hit rate: 93.0% [repeated 47x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 30x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:38:05 [loggers.py:259] Engine 000: Avg prompt throughput: 80.0 tokens/s, Avg generation throughput: 398.4 tokens/s, Running: 10 reqs, Waiting: 0 reqs, GPU KV cache usage: 35.6%, Prefix cache hit rate: 94.2% [repeated 49x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] [05:38:05] OK: 363 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] [05:38:05] OK: RSS 2.48 GiB | node mem 398.9/858.0 GiB used (46.5%), avail 459.1 GiB (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 30x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:38:10 [loggers.py:259] Engine 000: Avg prompt throughput: 158.5 tokens/s, Avg generation throughput: 371.3 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 32.5%, Prefix cache hit rate: 94.2% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 35x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) ERROR 06-12 05:38:13 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) ERROR 06-12 05:38:13 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) ERROR 06-12 05:38:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) ERROR 06-12 05:38:13 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) ERROR 06-12 05:38:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) ERROR 06-12 05:38:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) ERROR 06-12 05:38:13 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) ERROR 06-12 05:38:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) ERROR 06-12 05:38:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) ERROR 06-12 05:38:13 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) ERROR 06-12 05:38:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) ERROR 06-12 05:38:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) ERROR 06-12 05:38:13 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) ERROR 06-12 05:38:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) ERROR 06-12 05:38:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) ERROR 06-12 05:38:13 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) ERROR 06-12 05:38:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) ERROR 06-12 05:38:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) ERROR 06-12 05:38:13 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) ERROR 06-12 05:38:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) ERROR 06-12 05:38:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) ERROR 06-12 05:38:13 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) ERROR 06-12 05:38:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) ERROR 06-12 05:38:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) ERROR 06-12 05:38:13 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) ERROR 06-12 05:38:13 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] [05:38:14] OK: 380 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] [05:38:14] OK: RSS 2.34 GiB | node mem 389.8/858.0 GiB used (45.4%), avail 468.2 GiB (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:38:15 [loggers.py:259] Engine 000: Avg prompt throughput: 41.5 tokens/s, Avg generation throughput: 339.6 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 39.5%, Prefix cache hit rate: 94.3% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:38:20 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 349.0 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 40.0%, Prefix cache hit rate: 94.3% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (skyrl_entrypoint pid=487747) [fd-monitor] [05:38:20] OK: 565 / 131,072 FDs open (0.4% of soft limit, hard limit: 131,072) (skyrl_entrypoint pid=487747) [fd-monitor] [05:38:20] OK: RSS 1.98 GiB | node mem 241.8/858.0 GiB used (28.2%), avail 616.2 GiB (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 20x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) INFO 06-12 05:38:25 [loggers.py:259] Engine 000: Avg prompt throughput: 755.7 tokens/s, Avg generation throughput: 433.6 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 42.8%, Prefix cache hit rate: 94.1% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:38:27 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:38:27 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:38:27 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:38:27 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:38:27 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:38:27 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:38:27 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:38:27 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:38:27 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:38:27 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:38:27 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:38:27 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:38:27 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:38:27 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:38:27 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:38:27 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:38:27 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:38:27 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:38:27 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:38:27 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:38:27 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:38:27 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:38:27 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:38:27 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:38:27 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:38:27 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 28x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:38:30 [loggers.py:259] Engine 000: Avg prompt throughput: 287.9 tokens/s, Avg generation throughput: 340.6 tokens/s, Running: 12 reqs, Waiting: 0 reqs, GPU KV cache usage: 47.1%, Prefix cache hit rate: 94.4% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 22x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:38:34 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:38:34 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:38:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:38:34 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:38:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:38:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:38:34 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:38:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:38:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:38:34 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:38:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:38:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:38:34 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:38:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:38:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:38:34 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:38:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:38:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:38:34 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:38:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:38:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:38:34 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:38:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:38:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:38:34 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:38:34 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:38:35 [loggers.py:259] Engine 000: Avg prompt throughput: 102.5 tokens/s, Avg generation throughput: 461.3 tokens/s, Running: 13 reqs, Waiting: 0 reqs, GPU KV cache usage: 51.1%, Prefix cache hit rate: 94.4% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:38:40 [loggers.py:259] Engine 000: Avg prompt throughput: 64.4 tokens/s, Avg generation throughput: 476.3 tokens/s, Running: 14 reqs, Waiting: 0 reqs, GPU KV cache usage: 48.6%, Prefix cache hit rate: 94.4% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:38:41 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:38:41 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:38:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:38:41 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:38:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:38:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:38:41 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:38:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:38:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:38:41 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:38:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:38:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:38:41 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:38:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:38:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:38:41 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:38:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:38:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:38:41 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:38:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:38:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:38:41 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:38:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:38:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:38:41 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:38:41 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 15x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:38:45 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 519.3 tokens/s, Running: 12 reqs, Waiting: 0 reqs, GPU KV cache usage: 40.9%, Prefix cache hit rate: 94.4% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:38:50 [loggers.py:259] Engine 000: Avg prompt throughput: 137.1 tokens/s, Avg generation throughput: 434.5 tokens/s, Running: 12 reqs, Waiting: 0 reqs, GPU KV cache usage: 44.1%, Prefix cache hit rate: 94.4% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 23x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) ERROR 06-12 05:38:54 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) ERROR 06-12 05:38:54 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) ERROR 06-12 05:38:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) ERROR 06-12 05:38:54 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) ERROR 06-12 05:38:54 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) ERROR 06-12 05:38:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) ERROR 06-12 05:38:54 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) ERROR 06-12 05:38:54 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) ERROR 06-12 05:38:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) ERROR 06-12 05:38:54 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) ERROR 06-12 05:38:54 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) ERROR 06-12 05:38:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) ERROR 06-12 05:38:54 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) ERROR 06-12 05:38:54 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) ERROR 06-12 05:38:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) ERROR 06-12 05:38:54 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) ERROR 06-12 05:38:54 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) ERROR 06-12 05:38:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) ERROR 06-12 05:38:54 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) ERROR 06-12 05:38:54 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) ERROR 06-12 05:38:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) ERROR 06-12 05:38:54 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) ERROR 06-12 05:38:54 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) ERROR 06-12 05:38:54 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) ERROR 06-12 05:38:54 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) ERROR 06-12 05:38:54 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:38:55 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 428.8 tokens/s, Running: 10 reqs, Waiting: 0 reqs, GPU KV cache usage: 36.8%, Prefix cache hit rate: 94.4% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 25x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:38:59 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:38:59 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:38:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:38:59 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:38:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:38:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:38:59 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:38:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:38:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:38:59 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:38:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:38:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:38:59 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:38:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:38:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:38:59 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:38:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:38:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:38:59 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:38:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:38:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:38:59 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:38:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:38:59 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:38:59 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:39:00 [loggers.py:259] Engine 000: Avg prompt throughput: 115.3 tokens/s, Avg generation throughput: 383.0 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 32.1%, Prefix cache hit rate: 94.4% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 33x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:39:05 [loggers.py:259] Engine 000: Avg prompt throughput: 36.0 tokens/s, Avg generation throughput: 357.8 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 30.9%, Prefix cache hit rate: 94.4% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 5x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:39:10 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 337.2 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 31.4%, Prefix cache hit rate: 94.4% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:39:15 [loggers.py:259] Engine 000: Avg prompt throughput: 87.6 tokens/s, Avg generation throughput: 350.6 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 33.3%, Prefix cache hit rate: 94.5% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 15x across cluster] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:39:18 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:39:18 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:39:18 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:39:18 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:39:18 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:39:18 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:39:18 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:39:18 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:39:18 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:39:18 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:39:18 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:39:18 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:39:18 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:39:18 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:39:18 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:39:18 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:39:18 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:39:18 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:39:18 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:39:18 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:39:18 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:39:18 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:39:18 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:39:18 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:39:18 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:39:18 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:39:20 [loggers.py:259] Engine 000: Avg prompt throughput: 135.1 tokens/s, Avg generation throughput: 396.4 tokens/s, Running: 10 reqs, Waiting: 0 reqs, GPU KV cache usage: 32.7%, Prefix cache hit rate: 94.5% [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:39:24 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:39:24 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:39:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:39:24 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:39:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:39:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:39:24 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:39:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:39:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:39:24 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:39:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:39:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:39:24 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:39:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:39:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:39:24 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:39:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:39:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:39:24 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:39:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:39:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:39:24 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:39:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:39:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:39:24 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) ERROR 06-12 05:39:24 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:39:25 [loggers.py:259] Engine 000: Avg prompt throughput: 192.7 tokens/s, Avg generation throughput: 380.4 tokens/s, Running: 10 reqs, Waiting: 0 reqs, GPU KV cache usage: 36.1%, Prefix cache hit rate: 94.5% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:39:25 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:39:25 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:39:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:39:25 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:39:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:39:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:39:25 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:39:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:39:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:39:25 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:39:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:39:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:39:25 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:39:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:39:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:39:25 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:39:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:39:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:39:25 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:39:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:39:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:39:25 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:39:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:39:25 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:39:25 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 25x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:39:30 [loggers.py:259] Engine 000: Avg prompt throughput: 168.4 tokens/s, Avg generation throughput: 389.3 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 29.3%, Prefix cache hit rate: 94.5% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:39:33 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:39:33 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:39:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:39:33 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:39:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:39:33 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:39:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:39:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:39:33 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:39:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:39:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:39:33 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:39:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:39:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:39:33 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:39:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:39:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:39:33 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:39:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:39:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:39:33 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:39:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:39:33 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:39:33 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:39:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 15x across cluster] (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:39:35 [loggers.py:259] Engine 000: Avg prompt throughput: 111.3 tokens/s, Avg generation throughput: 374.0 tokens/s, Running: 10 reqs, Waiting: 0 reqs, GPU KV cache usage: 33.6%, Prefix cache hit rate: 94.6% [repeated 49x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:39:37 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:39:37 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:39:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:39:37 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:39:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:39:37 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:39:37 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:39:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:39:37 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:39:37 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:39:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:39:37 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:39:37 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:39:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:39:37 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:39:37 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:39:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:39:37 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:39:37 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:39:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:39:37 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:39:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:39:37 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:39:37 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:39:37 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 20x across cluster] (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:39:40 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:39:40 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:39:40 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:39:40 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:39:40 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:39:40 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:39:40 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:39:40 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:39:40 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:39:40 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:39:40 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:39:40 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:39:40 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:39:40 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:39:40 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:39:40 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:39:40 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:39:40 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:39:40 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:39:40 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:39:40 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:39:40 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:39:40 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:39:40 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:39:40 [loggers.py:259] Engine 000: Avg prompt throughput: 150.2 tokens/s, Avg generation throughput: 392.6 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 38.3%, Prefix cache hit rate: 94.6% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) ERROR 06-12 05:39:43 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 20x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:39:44 [serving.py:315] Error in preprocessing prompt inputs [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:39:44 [serving.py:315] Traceback (most recent call last): [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:39:44 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:39:44 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:39:44 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:39:44 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:39:44 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:39:44 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:39:44 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:39:44 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:39:44 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:39:44 [serving.py:315] return await asyncio.gather( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:39:44 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:39:44 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:39:44 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:39:44 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:39:44 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:39:44 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:39:44 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:39:44 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:39:44 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:39:44 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:39:44 [serving.py:315] raise VLLMValidationError( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:39:44 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:39:45 [loggers.py:259] Engine 000: Avg prompt throughput: 112.0 tokens/s, Avg generation throughput: 431.4 tokens/s, Running: 12 reqs, Waiting: 0 reqs, GPU KV cache usage: 37.5%, Prefix cache hit rate: 94.6% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:39:47] OK: 386 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:39:47] OK: RSS 2.73 GiB | node mem 381.3/858.0 GiB used (44.4%), avail 476.7 GiB (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:39:44 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 18x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:39:50 [loggers.py:259] Engine 000: Avg prompt throughput: 207.0 tokens/s, Avg generation throughput: 478.9 tokens/s, Running: 12 reqs, Waiting: 0 reqs, GPU KV cache usage: 38.2%, Prefix cache hit rate: 94.5% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 22x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:39:55 [loggers.py:259] Engine 000: Avg prompt throughput: 480.2 tokens/s, Avg generation throughput: 458.3 tokens/s, Running: 13 reqs, Waiting: 0 reqs, GPU KV cache usage: 45.8%, Prefix cache hit rate: 94.5% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:39:57] OK: 360 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:39:57] OK: RSS 2.32 GiB | node mem 386.0/858.0 GiB used (45.0%), avail 472.0 GiB (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 38x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) INFO 06-12 05:40:00 [loggers.py:259] Engine 000: Avg prompt throughput: 23.2 tokens/s, Avg generation throughput: 453.2 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 42.3%, Prefix cache hit rate: 94.3% [repeated 47x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:00 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:00 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:00 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:00 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:00 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:00 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:00 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:00 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:00 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:00 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:00 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:00 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:00 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:00 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:00 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:00 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:00 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:00 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:00 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:00 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:00 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:00 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:00 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:00 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:00 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:00 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 28x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:40:05 [loggers.py:259] Engine 000: Avg prompt throughput: 25.7 tokens/s, Avg generation throughput: 489.5 tokens/s, Running: 13 reqs, Waiting: 0 reqs, GPU KV cache usage: 43.1%, Prefix cache hit rate: 94.5% [repeated 49x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] [05:40:05] OK: 361 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] [05:40:05] OK: RSS 2.70 GiB | node mem 399.1/858.0 GiB used (46.5%), avail 458.9 GiB (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:40:08 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:40:08 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:40:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:40:08 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:40:08 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:40:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:40:08 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:40:08 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:40:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:40:08 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:40:08 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:40:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:40:08 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:40:08 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:40:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:40:08 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:40:08 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:40:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:40:08 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:40:08 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:40:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:40:08 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:40:08 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:40:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:40:08 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) ERROR 06-12 05:40:08 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 22x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:40:10 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 488.4 tokens/s, Running: 12 reqs, Waiting: 0 reqs, GPU KV cache usage: 41.3%, Prefix cache hit rate: 94.5% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] [05:40:14] OK: 431 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] [05:40:14] OK: RSS 2.49 GiB | node mem 389.9/858.0 GiB used (45.4%), avail 468.1 GiB (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 21x across cluster] (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:40:15 [loggers.py:259] Engine 000: Avg prompt throughput: 62.0 tokens/s, Avg generation throughput: 496.4 tokens/s, Running: 14 reqs, Waiting: 0 reqs, GPU KV cache usage: 49.6%, Prefix cache hit rate: 94.6% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:40:17 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:40:17 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:40:17 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:40:17 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:40:17 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:40:17 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:40:17 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:40:17 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:40:17 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:40:17 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:40:17 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:40:17 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:40:17 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:40:17 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:40:17 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:40:17 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:40:17 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:40:17 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:40:17 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:40:17 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:40:17 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:40:17 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:40:17 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:40:17 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:40:17 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:40:17 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:40:20 [loggers.py:259] Engine 000: Avg prompt throughput: 147.8 tokens/s, Avg generation throughput: 545.5 tokens/s, Running: 13 reqs, Waiting: 0 reqs, GPU KV cache usage: 44.5%, Prefix cache hit rate: 94.7% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (skyrl_entrypoint pid=487747) [fd-monitor] [05:40:20] OK: 563 / 131,072 FDs open (0.4% of soft limit, hard limit: 131,072) (skyrl_entrypoint pid=487747) [fd-monitor] [05:40:20] OK: RSS 2.03 GiB | node mem 241.9/858.0 GiB used (28.2%), avail 616.1 GiB (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:40:20 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:40:20 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:40:20 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:40:20 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:40:20 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:40:20 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:40:20 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:40:20 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:40:20 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:40:20 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:40:20 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:40:20 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:40:20 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:40:20 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:40:20 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:40:20 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:40:20 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:40:20 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:40:20 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:40:20 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:40:20 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:40:20 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:40:20 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:40:20 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:40:20 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:40:24 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:40:24 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:40:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:40:24 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:40:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:40:24 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:40:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:40:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:40:24 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:40:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:40:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:40:24 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:40:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:40:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:40:24 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:40:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:40:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:40:24 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:40:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:40:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:40:24 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:40:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:40:24 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:40:24 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 29x across cluster] (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:40:25 [loggers.py:259] Engine 000: Avg prompt throughput: 10.2 tokens/s, Avg generation throughput: 471.4 tokens/s, Running: 10 reqs, Waiting: 0 reqs, GPU KV cache usage: 32.4%, Prefix cache hit rate: 94.7% [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:40:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:28 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:28 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:28 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:28 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:28 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:28 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:28 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:28 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:28 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:28 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:28 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:28 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:28 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:28 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:28 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:28 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:28 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:28 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:28 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:28 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:28 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:28 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:28 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:28 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 13x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:40:30 [loggers.py:259] Engine 000: Avg prompt throughput: 164.0 tokens/s, Avg generation throughput: 375.6 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 37.6%, Prefix cache hit rate: 94.7% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:28 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:40:32 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:40:32 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:40:32 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:40:32 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:40:32 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:40:32 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:40:32 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:40:32 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:40:32 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:40:32 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:40:32 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:40:32 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:40:32 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:40:32 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:40:32 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:40:32 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:40:32 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:40:32 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:40:32 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:40:32 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:40:32 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:40:32 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:40:32 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:40:32 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 26x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:40:35 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 347.3 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 38.0%, Prefix cache hit rate: 94.7% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:40:32 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 38x across cluster] (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:40:40 [loggers.py:259] Engine 000: Avg prompt throughput: 165.3 tokens/s, Avg generation throughput: 365.2 tokens/s, Running: 10 reqs, Waiting: 0 reqs, GPU KV cache usage: 40.7%, Prefix cache hit rate: 94.7% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:43 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:43 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:43 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:43 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:43 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:43 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:43 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:43 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:43 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:43 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:43 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:43 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:43 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:43 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:43 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:43 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:43 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:43 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 17x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) INFO 06-12 05:40:45 [loggers.py:259] Engine 000: Avg prompt throughput: 9.8 tokens/s, Avg generation throughput: 578.2 tokens/s, Running: 12 reqs, Waiting: 0 reqs, GPU KV cache usage: 47.0%, Prefix cache hit rate: 94.2% [repeated 47x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:46 [serving.py:315] Error in preprocessing prompt inputs [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:46 [serving.py:315] Traceback (most recent call last): [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:46 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:46 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:46 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:46 [serving.py:315] return await asyncio.gather( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:46 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:46 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:46 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:46 [serving.py:315] raise VLLMValidationError( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:40:46 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 20x across cluster] (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:40:50 [loggers.py:259] Engine 000: Avg prompt throughput: 204.8 tokens/s, Avg generation throughput: 504.2 tokens/s, Running: 14 reqs, Waiting: 0 reqs, GPU KV cache usage: 51.2%, Prefix cache hit rate: 94.8% [repeated 49x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 3x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:48 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:48 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:48 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:48 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:48 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:48 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:48 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:48 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:48 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:48 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:48 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:48 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:48 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:48 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:48 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:48 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:48 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:48 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:48 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:48 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:48 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:48 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:48 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:48 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) ERROR 06-12 05:40:48 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:40:55 [loggers.py:259] Engine 000: Avg prompt throughput: 11.4 tokens/s, Avg generation throughput: 557.9 tokens/s, Running: 15 reqs, Waiting: 0 reqs, GPU KV cache usage: 53.5%, Prefix cache hit rate: 94.8% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:41:00 [loggers.py:259] Engine 000: Avg prompt throughput: 98.4 tokens/s, Avg generation throughput: 601.0 tokens/s, Running: 15 reqs, Waiting: 0 reqs, GPU KV cache usage: 59.9%, Prefix cache hit rate: 94.9% [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:41:00 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:41:00 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:41:00 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:41:00 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:41:00 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:41:00 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:41:00 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:41:00 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:41:00 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:41:00 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:41:00 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:41:00 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:41:00 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:41:00 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:41:00 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:41:00 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:41:00 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:41:00 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:41:00 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:41:00 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:41:00 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:41:00 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:41:00 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:41:00 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:41:00 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:41:00 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:41:05 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 573.6 tokens/s, Running: 15 reqs, Waiting: 0 reqs, GPU KV cache usage: 60.7%, Prefix cache hit rate: 94.9% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:07 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:07 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:07 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:07 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:07 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:07 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:07 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:07 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:07 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:07 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:07 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:07 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:07 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:07 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:07 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:07 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:07 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:07 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:41:10 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 521.4 tokens/s, Running: 12 reqs, Waiting: 0 reqs, GPU KV cache usage: 51.6%, Prefix cache hit rate: 94.9% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:41:13 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:41:13 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:41:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:41:13 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:41:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:41:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:41:13 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:41:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:41:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:41:13 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:41:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:41:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:41:13 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:41:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:41:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:41:13 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:41:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:41:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:41:13 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:41:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:41:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:41:13 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:41:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:41:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:41:13 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:41:13 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 19x across cluster] (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:41:15 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 447.1 tokens/s, Running: 10 reqs, Waiting: 0 reqs, GPU KV cache usage: 42.3%, Prefix cache hit rate: 94.9% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:41:16 [serving.py:315] Error in preprocessing prompt inputs [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:41:16 [serving.py:315] Traceback (most recent call last): [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:41:16 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:41:16 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:41:16 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:41:16 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:41:16 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:41:16 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:41:16 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:41:16 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:41:16 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:41:16 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:41:16 [serving.py:315] return await asyncio.gather( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:41:16 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:41:16 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:41:16 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:41:16 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:41:16 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:41:16 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:41:16 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:41:16 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:41:16 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:41:16 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:41:16 [serving.py:315] raise VLLMValidationError( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:41:16 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 20x across cluster] (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:41:20 [loggers.py:259] Engine 000: Avg prompt throughput: 93.6 tokens/s, Avg generation throughput: 399.0 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 47.3%, Prefix cache hit rate: 94.9% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:41:23 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:41:23 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:41:23 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:41:23 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:41:23 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:41:23 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:41:23 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:41:23 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:41:23 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:41:23 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:41:23 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:41:23 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:41:23 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:41:23 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:41:23 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:41:23 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:41:23 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:41:23 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:41:23 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:41:23 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:41:23 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:41:23 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:41:23 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:41:23 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:41:23 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:41:23 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 23x across cluster] (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:41:25 [loggers.py:259] Engine 000: Avg prompt throughput: 89.2 tokens/s, Avg generation throughput: 421.6 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 42.9%, Prefix cache hit rate: 94.9% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 4x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 4x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:41:26 [serving.py:315] Error in preprocessing prompt inputs [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:41:26 [serving.py:315] Traceback (most recent call last): [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:41:26 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:41:26 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:41:26 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:41:26 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:41:26 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:41:26 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:41:26 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:41:26 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:41:26 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:41:26 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:41:26 [serving.py:315] return await asyncio.gather( [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:41:26 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:41:26 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:41:26 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:41:26 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:41:26 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:41:26 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:41:26 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:41:26 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:41:26 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:41:26 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:41:26 [serving.py:315] raise VLLMValidationError( [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:41:26 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 3x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 12x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:41:30 [loggers.py:259] Engine 000: Avg prompt throughput: 258.7 tokens/s, Avg generation throughput: 465.3 tokens/s, Running: 13 reqs, Waiting: 0 reqs, GPU KV cache usage: 49.4%, Prefix cache hit rate: 94.9% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 26x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:41:35 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:41:35 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:41:35 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:41:35 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:41:35 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:41:35 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:41:35 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:41:35 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:41:35 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:41:35 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:41:35 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:41:35 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:41:35 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:41:35 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:41:35 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:41:35 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:41:35 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:41:35 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:41:35 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:41:35 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:41:35 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:41:35 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:41:35 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:41:35 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:41:35 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:41:35 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:41:35 [loggers.py:259] Engine 000: Avg prompt throughput: 1.8 tokens/s, Avg generation throughput: 502.5 tokens/s, Running: 13 reqs, Waiting: 0 reqs, GPU KV cache usage: 49.0%, Prefix cache hit rate: 94.9% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 14x across cluster] (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:41:39 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:41:39 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:41:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:41:39 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:41:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:41:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:41:39 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:41:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:41:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:41:39 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:41:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:41:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:41:39 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:41:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:41:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:41:39 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:41:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:41:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:41:39 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:41:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:41:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:41:39 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:41:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:41:39 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:41:39 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:41:40 [loggers.py:259] Engine 000: Avg prompt throughput: 212.2 tokens/s, Avg generation throughput: 529.7 tokens/s, Running: 14 reqs, Waiting: 0 reqs, GPU KV cache usage: 46.8%, Prefix cache hit rate: 94.9% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 24x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:41:45 [loggers.py:259] Engine 000: Avg prompt throughput: 117.6 tokens/s, Avg generation throughput: 563.6 tokens/s, Running: 15 reqs, Waiting: 0 reqs, GPU KV cache usage: 53.8%, Prefix cache hit rate: 94.9% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:41:47] OK: 388 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:41:47] OK: RSS 3.01 GiB | node mem 381.8/858.0 GiB used (44.5%), avail 476.2 GiB (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:41:49 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:41:49 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:41:49 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:41:49 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:41:49 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:41:49 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:41:49 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:41:49 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:41:49 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:41:49 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:41:49 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:41:49 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:41:49 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:41:49 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:41:49 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:41:49 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:41:49 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:41:49 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:41:49 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:41:49 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:41:49 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:41:49 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:41:49 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:41:49 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:41:49 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:41:49 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 23x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) INFO 06-12 05:41:50 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 624.4 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 48.6%, Prefix cache hit rate: 94.1% [repeated 47x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 3x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:51 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:51 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:51 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:51 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:51 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:51 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:51 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:51 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:51 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:51 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:51 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:51 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:51 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:51 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:51 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:51 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:41:51 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 25x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:41:55 [loggers.py:259] Engine 000: Avg prompt throughput: 113.6 tokens/s, Avg generation throughput: 505.9 tokens/s, Running: 13 reqs, Waiting: 0 reqs, GPU KV cache usage: 44.5%, Prefix cache hit rate: 94.9% [repeated 49x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:41:57] OK: 363 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:41:57] OK: RSS 2.56 GiB | node mem 386.3/858.0 GiB used (45.0%), avail 471.6 GiB (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) ERROR 06-12 05:41:59 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) ERROR 06-12 05:41:59 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) ERROR 06-12 05:41:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) ERROR 06-12 05:41:59 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) ERROR 06-12 05:41:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) ERROR 06-12 05:41:59 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) ERROR 06-12 05:41:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) ERROR 06-12 05:41:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) ERROR 06-12 05:41:59 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) ERROR 06-12 05:41:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) ERROR 06-12 05:41:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) ERROR 06-12 05:41:59 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) ERROR 06-12 05:41:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) ERROR 06-12 05:41:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) ERROR 06-12 05:41:59 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) ERROR 06-12 05:41:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) ERROR 06-12 05:41:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) ERROR 06-12 05:41:59 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) ERROR 06-12 05:41:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) ERROR 06-12 05:41:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) ERROR 06-12 05:41:59 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) ERROR 06-12 05:41:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) ERROR 06-12 05:41:59 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) ERROR 06-12 05:41:59 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) ERROR 06-12 05:41:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 22x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:42:00 [loggers.py:259] Engine 000: Avg prompt throughput: 34.4 tokens/s, Avg generation throughput: 519.4 tokens/s, Running: 13 reqs, Waiting: 0 reqs, GPU KV cache usage: 47.2%, Prefix cache hit rate: 94.9% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:02 [serving.py:315] Error in preprocessing prompt inputs [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:02 [serving.py:315] Traceback (most recent call last): [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:02 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:02 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:02 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:02 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:02 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:02 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:02 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:02 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:02 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:02 [serving.py:315] return await asyncio.gather( [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:02 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:02 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:02 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:02 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:02 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:02 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:02 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:02 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:02 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:02 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:02 [serving.py:315] raise VLLMValidationError( [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:02 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:02 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 8x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 3x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 3x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 18x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] [05:42:05] OK: 355 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] [05:42:05] OK: RSS 2.88 GiB | node mem 399.3/858.0 GiB used (46.5%), avail 458.7 GiB (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:42:05 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 508.7 tokens/s, Running: 13 reqs, Waiting: 0 reqs, GPU KV cache usage: 48.0%, Prefix cache hit rate: 94.9% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:42:09 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:42:09 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:42:09 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:42:09 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:42:09 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:42:09 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:42:09 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:42:09 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:42:09 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:42:09 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:42:09 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:42:09 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:42:09 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:42:09 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:42:09 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:42:09 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:42:09 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:42:09 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:42:09 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:42:09 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:42:09 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:42:09 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:42:09 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:42:09 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:42:09 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) INFO 06-12 05:42:10 [loggers.py:259] Engine 000: Avg prompt throughput: 91.4 tokens/s, Avg generation throughput: 473.9 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 36.4%, Prefix cache hit rate: 94.1% [repeated 47x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:42:12 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:42:12 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:42:12 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:42:12 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:42:12 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:42:12 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:42:12 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:42:12 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:42:12 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:42:12 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:42:12 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:42:12 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:42:12 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:42:12 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:42:12 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:42:12 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:42:12 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:42:12 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:42:12 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:42:12 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:42:12 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:42:12 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:42:12 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) ERROR 06-12 05:42:12 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] [05:42:14] OK: 358 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] [05:42:14] OK: RSS 2.55 GiB | node mem 390.0/858.0 GiB used (45.5%), avail 468.0 GiB (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:42:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 6x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 20x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:42:16 [loggers.py:259] Engine 000: Avg prompt throughput: 100.4 tokens/s, Avg generation throughput: 317.7 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 42.5%, Prefix cache hit rate: 93.5% [repeated 49x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:42:13 [serving.py:315] Error in preprocessing prompt inputs [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:42:13 [serving.py:315] Traceback (most recent call last): [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:42:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:42:13 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:42:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:42:13 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:42:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:42:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:42:13 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:42:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:42:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:42:13 [serving.py:315] return await asyncio.gather( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:42:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:42:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:42:13 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:42:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:42:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:42:13 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:42:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:42:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:42:13 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:42:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:42:13 [serving.py:315] raise VLLMValidationError( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:42:13 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 28x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:42:21 [loggers.py:259] Engine 000: Avg prompt throughput: 355.9 tokens/s, Avg generation throughput: 391.7 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 42.9%, Prefix cache hit rate: 93.4% [repeated 48x across cluster] (skyrl_entrypoint pid=487747) [fd-monitor] [05:42:20] OK: 564 / 131,072 FDs open (0.4% of soft limit, hard limit: 131,072) (skyrl_entrypoint pid=487747) [fd-monitor] [05:42:20] OK: RSS 2.06 GiB | node mem 241.9/858.0 GiB used (28.2%), avail 616.0 GiB (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:22 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:22 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:22 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:22 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:22 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:22 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:22 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:22 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:22 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:22 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:22 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:22 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:22 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:22 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:22 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:22 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:22 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:22 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:22 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:22 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:22 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:22 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:22 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:22 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:22 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:42:22 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 18x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:42:26 [loggers.py:259] Engine 000: Avg prompt throughput: 228.5 tokens/s, Avg generation throughput: 452.4 tokens/s, Running: 12 reqs, Waiting: 0 reqs, GPU KV cache usage: 45.2%, Prefix cache hit rate: 93.4% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:42:24 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:42:24 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:42:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:42:24 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:42:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:42:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:42:24 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:42:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:42:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:42:24 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:42:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:42:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:42:24 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:42:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:42:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:42:24 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:42:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:42:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:42:24 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:42:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:42:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:42:24 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:42:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:42:24 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:42:24 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 51x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:42:31 [loggers.py:259] Engine 000: Avg prompt throughput: 73.4 tokens/s, Avg generation throughput: 490.6 tokens/s, Running: 13 reqs, Waiting: 0 reqs, GPU KV cache usage: 51.8%, Prefix cache hit rate: 93.5% [repeated 47x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) ERROR 06-12 05:42:32 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) ERROR 06-12 05:42:32 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) ERROR 06-12 05:42:32 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) ERROR 06-12 05:42:32 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) ERROR 06-12 05:42:32 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) ERROR 06-12 05:42:32 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) ERROR 06-12 05:42:32 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) ERROR 06-12 05:42:32 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) ERROR 06-12 05:42:32 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) ERROR 06-12 05:42:32 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) ERROR 06-12 05:42:32 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) ERROR 06-12 05:42:32 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) ERROR 06-12 05:42:32 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) ERROR 06-12 05:42:32 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) ERROR 06-12 05:42:32 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) ERROR 06-12 05:42:32 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) ERROR 06-12 05:42:32 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) ERROR 06-12 05:42:32 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) ERROR 06-12 05:42:32 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) ERROR 06-12 05:42:32 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) ERROR 06-12 05:42:32 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) ERROR 06-12 05:42:32 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) ERROR 06-12 05:42:32 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) ERROR 06-12 05:42:32 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) ERROR 06-12 05:42:32 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 33x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:42:36 [loggers.py:259] Engine 000: Avg prompt throughput: 397.7 tokens/s, Avg generation throughput: 536.9 tokens/s, Running: 14 reqs, Waiting: 0 reqs, GPU KV cache usage: 53.9%, Prefix cache hit rate: 93.4% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:35 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:35 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:35 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:35 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:35 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:35 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:35 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:35 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:35 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:35 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:35 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:35 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:35 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:35 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:35 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:35 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:35 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:35 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:35 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:35 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:35 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:35 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:35 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:35 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:35 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 25x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:42:41 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 528.1 tokens/s, Running: 13 reqs, Waiting: 0 reqs, GPU KV cache usage: 48.0%, Prefix cache hit rate: 93.4% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:45 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:45 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:45 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:45 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:45 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:45 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:45 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:45 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:45 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:45 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:45 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:45 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:45 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:45 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:45 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:45 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:45 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:45 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:45 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:45 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:45 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:45 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:45 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:45 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:45 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:42:45 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:42:46 [loggers.py:259] Engine 000: Avg prompt throughput: 9.2 tokens/s, Avg generation throughput: 502.9 tokens/s, Running: 12 reqs, Waiting: 0 reqs, GPU KV cache usage: 38.5%, Prefix cache hit rate: 93.4% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 5x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:42:51 [loggers.py:259] Engine 000: Avg prompt throughput: 53.6 tokens/s, Avg generation throughput: 457.3 tokens/s, Running: 12 reqs, Waiting: 0 reqs, GPU KV cache usage: 42.4%, Prefix cache hit rate: 93.4% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:42:52 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:42:52 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:42:52 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:42:52 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:42:52 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:42:52 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:42:52 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:42:52 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:42:52 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:42:52 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:42:52 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:42:52 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:42:52 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:42:52 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:42:52 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:42:52 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:42:52 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:42:52 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:42:52 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:42:52 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:42:52 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:42:52 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:42:52 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:42:52 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:42:52 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) ERROR 06-12 05:42:52 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 35x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:42:56 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 455.6 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 41.1%, Prefix cache hit rate: 93.4% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:42:57 [serving.py:315] Error in preprocessing prompt inputs [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:42:57 [serving.py:315] Traceback (most recent call last): [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:42:57 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:42:57 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:42:57 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:42:57 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:42:57 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:42:57 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:42:57 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:42:57 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:42:57 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:42:57 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:42:57 [serving.py:315] return await asyncio.gather( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:42:57 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:42:57 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:42:57 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:42:57 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:42:57 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:42:57 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:42:57 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:42:57 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:42:57 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:42:57 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:42:57 [serving.py:315] raise VLLMValidationError( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:42:57 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:43:01 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 441.2 tokens/s, Running: 10 reqs, Waiting: 0 reqs, GPU KV cache usage: 39.3%, Prefix cache hit rate: 93.4% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 35x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:43:06 [loggers.py:259] Engine 000: Avg prompt throughput: 309.9 tokens/s, Avg generation throughput: 406.9 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 37.3%, Prefix cache hit rate: 93.4% [repeated 47x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:43:07 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:43:07 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:43:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:43:07 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:43:07 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:43:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:43:07 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:43:07 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:43:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:43:07 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:43:07 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:43:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:43:07 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:43:07 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:43:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:43:07 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:43:07 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:43:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:43:07 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:43:07 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:43:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:43:07 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:43:07 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:43:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:43:07 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:43:07 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 25x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:43:11 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 370.0 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 37.2%, Prefix cache hit rate: 93.4% [repeated 47x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:43:13 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:43:13 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:43:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:43:13 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:43:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:43:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:43:13 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:43:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:43:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:43:13 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:43:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:43:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:43:13 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:43:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:43:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:43:13 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:43:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:43:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:43:13 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:43:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:43:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:43:13 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:43:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:43:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:43:13 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:43:13 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 13x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:43:16 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 327.1 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 37.7%, Prefix cache hit rate: 93.4% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 5x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 5x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:43:13 [serving.py:315] Error in preprocessing prompt inputs [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:43:13 [serving.py:315] Traceback (most recent call last): [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:43:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:43:13 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:43:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:43:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:43:13 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:43:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:43:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:43:13 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:43:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:43:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:43:13 [serving.py:315] return await asyncio.gather( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:43:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:43:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:43:13 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:43:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:43:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:43:13 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:43:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:43:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:43:13 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:43:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:43:13 [serving.py:315] raise VLLMValidationError( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:43:13 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 27x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:43:21 [loggers.py:259] Engine 000: Avg prompt throughput: 286.0 tokens/s, Avg generation throughput: 344.0 tokens/s, Running: 10 reqs, Waiting: 0 reqs, GPU KV cache usage: 40.5%, Prefix cache hit rate: 93.4% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:43:25 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:43:25 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:43:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:43:25 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:43:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:43:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:43:25 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:43:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:43:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:43:25 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:43:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:43:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:43:25 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:43:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:43:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:43:25 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:43:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:43:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:43:25 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:43:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:43:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:43:25 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:43:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:43:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:43:25 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) ERROR 06-12 05:43:25 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 8x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:43:26 [loggers.py:259] Engine 000: Avg prompt throughput: 448.6 tokens/s, Avg generation throughput: 462.9 tokens/s, Running: 13 reqs, Waiting: 0 reqs, GPU KV cache usage: 46.0%, Prefix cache hit rate: 93.3% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 27x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:43:31 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 510.0 tokens/s, Running: 12 reqs, Waiting: 0 reqs, GPU KV cache usage: 46.4%, Prefix cache hit rate: 93.3% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 34x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:43:36 [loggers.py:259] Engine 000: Avg prompt throughput: 156.5 tokens/s, Avg generation throughput: 466.4 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 45.6%, Prefix cache hit rate: 93.3% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:39 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:39 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:39 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:39 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:39 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:39 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:39 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:39 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:39 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:39 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:39 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 17x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:43:41 [loggers.py:259] Engine 000: Avg prompt throughput: 65.2 tokens/s, Avg generation throughput: 451.4 tokens/s, Running: 12 reqs, Waiting: 0 reqs, GPU KV cache usage: 46.3%, Prefix cache hit rate: 93.3% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:43:41 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:43:41 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:43:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:43:41 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:43:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:43:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:43:41 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:43:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:43:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:43:41 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:43:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:43:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:43:41 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:43:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:43:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:43:41 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:43:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:43:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:43:41 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:43:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:43:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:43:41 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:43:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:43:41 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:43:41 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:43:46 [loggers.py:259] Engine 000: Avg prompt throughput: 45.3 tokens/s, Avg generation throughput: 471.8 tokens/s, Running: 13 reqs, Waiting: 0 reqs, GPU KV cache usage: 47.5%, Prefix cache hit rate: 93.3% [repeated 49x across cluster] (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:43:47 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:43:47 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:43:47 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:43:47 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:43:47 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:43:47 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:43:47 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:43:47 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:43:47 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:43:47 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:43:47 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:43:47 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:43:47 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:43:47 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:43:47 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:43:47 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:43:47 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:43:47 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:43:47 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:43:47 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:43:47 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:43:47 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:43:47 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:43:47 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:43:47] OK: 394 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:43:47] OK: RSS 3.16 GiB | node mem 382.0/858.0 GiB used (44.5%), avail 476.0 GiB (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) ERROR 06-12 05:43:47 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 15x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:43:51 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 491.4 tokens/s, Running: 12 reqs, Waiting: 0 reqs, GPU KV cache usage: 46.3%, Prefix cache hit rate: 93.3% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:55 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:55 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:55 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:55 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:55 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:55 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:55 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:55 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:55 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:55 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:55 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:55 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:55 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:55 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:55 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:55 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:55 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:55 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:55 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:55 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:55 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:55 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:55 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:55 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:55 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:43:55 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 25x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:43:56 [loggers.py:259] Engine 000: Avg prompt throughput: 46.2 tokens/s, Avg generation throughput: 497.0 tokens/s, Running: 13 reqs, Waiting: 0 reqs, GPU KV cache usage: 47.8%, Prefix cache hit rate: 93.3% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:43:57] OK: 356 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:43:57] OK: RSS 2.79 GiB | node mem 386.6/858.0 GiB used (45.1%), avail 471.3 GiB (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) ERROR 06-12 05:43:58 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) ERROR 06-12 05:43:58 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) ERROR 06-12 05:43:58 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) ERROR 06-12 05:43:58 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) ERROR 06-12 05:43:58 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) ERROR 06-12 05:43:58 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) ERROR 06-12 05:43:58 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) ERROR 06-12 05:43:58 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) ERROR 06-12 05:43:58 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) ERROR 06-12 05:43:58 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) ERROR 06-12 05:43:58 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) ERROR 06-12 05:43:58 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) ERROR 06-12 05:43:58 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) ERROR 06-12 05:43:58 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) ERROR 06-12 05:43:58 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) ERROR 06-12 05:43:58 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) ERROR 06-12 05:43:58 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) ERROR 06-12 05:43:58 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) ERROR 06-12 05:43:58 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) ERROR 06-12 05:43:58 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) ERROR 06-12 05:43:58 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) ERROR 06-12 05:43:58 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) ERROR 06-12 05:43:58 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) ERROR 06-12 05:43:58 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) ERROR 06-12 05:43:58 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 34x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:44:01 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 511.0 tokens/s, Running: 13 reqs, Waiting: 0 reqs, GPU KV cache usage: 48.5%, Prefix cache hit rate: 93.3% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] [05:44:05] OK: 445 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] [05:44:05] OK: RSS 3.06 GiB | node mem 399.4/858.0 GiB used (46.6%), avail 458.6 GiB (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:44:06 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 483.6 tokens/s, Running: 12 reqs, Waiting: 0 reqs, GPU KV cache usage: 48.3%, Prefix cache hit rate: 93.3% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:07 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:07 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:07 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:07 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:07 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:07 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:07 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:07 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:07 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:07 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:07 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:07 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:07 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:07 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:07 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:07 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:07 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:07 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:07 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 23x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:44:11 [loggers.py:259] Engine 000: Avg prompt throughput: 121.5 tokens/s, Avg generation throughput: 462.5 tokens/s, Running: 12 reqs, Waiting: 0 reqs, GPU KV cache usage: 49.3%, Prefix cache hit rate: 93.3% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:44:12 [serving.py:315] Error in preprocessing prompt inputs [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:44:12 [serving.py:315] Traceback (most recent call last): [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:44:12 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:44:12 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:44:12 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:44:12 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:44:12 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:44:12 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:44:12 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:44:12 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:44:12 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:44:12 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:44:12 [serving.py:315] return await asyncio.gather( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:44:12 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:44:12 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:44:12 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:44:12 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:44:12 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:44:12 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:44:12 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:44:12 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:44:12 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:44:12 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:44:12 [serving.py:315] raise VLLMValidationError( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:44:12 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 2x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] [05:44:14] OK: 362 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] [05:44:14] OK: RSS 2.64 GiB | node mem 390.1/858.0 GiB used (45.5%), avail 467.9 GiB (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 35x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:44:16 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 480.4 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 49.2%, Prefix cache hit rate: 93.3% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (skyrl_entrypoint pid=487747) [fd-monitor] [05:44:20] OK: 564 / 131,072 FDs open (0.4% of soft limit, hard limit: 131,072) (skyrl_entrypoint pid=487747) [fd-monitor] [05:44:20] OK: RSS 2.08 GiB | node mem 242.0/858.0 GiB used (28.2%), avail 616.0 GiB (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 15x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:44:21 [loggers.py:259] Engine 000: Avg prompt throughput: 20.9 tokens/s, Avg generation throughput: 432.5 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 50.2%, Prefix cache hit rate: 93.3% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:21 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:21 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:21 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:21 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:21 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:21 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:21 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:21 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:21 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:21 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:44:21 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 23x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:44:26 [loggers.py:259] Engine 000: Avg prompt throughput: 79.4 tokens/s, Avg generation throughput: 435.4 tokens/s, Running: 12 reqs, Waiting: 0 reqs, GPU KV cache usage: 52.8%, Prefix cache hit rate: 93.3% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) ERROR 06-12 05:44:22 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) ERROR 06-12 05:44:22 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) ERROR 06-12 05:44:22 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) ERROR 06-12 05:44:22 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) ERROR 06-12 05:44:22 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) ERROR 06-12 05:44:22 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) ERROR 06-12 05:44:22 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) ERROR 06-12 05:44:22 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) ERROR 06-12 05:44:22 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) ERROR 06-12 05:44:22 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) ERROR 06-12 05:44:22 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) ERROR 06-12 05:44:22 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) ERROR 06-12 05:44:22 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) ERROR 06-12 05:44:22 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) ERROR 06-12 05:44:22 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) ERROR 06-12 05:44:22 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) ERROR 06-12 05:44:22 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) ERROR 06-12 05:44:22 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) ERROR 06-12 05:44:22 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) ERROR 06-12 05:44:22 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) ERROR 06-12 05:44:22 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) ERROR 06-12 05:44:22 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) ERROR 06-12 05:44:22 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) ERROR 06-12 05:44:22 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) ERROR 06-12 05:44:22 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 30x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:44:31 [loggers.py:259] Engine 000: Avg prompt throughput: 258.5 tokens/s, Avg generation throughput: 491.7 tokens/s, Running: 14 reqs, Waiting: 0 reqs, GPU KV cache usage: 54.1%, Prefix cache hit rate: 93.3% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:44:33 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:44:33 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:44:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:44:33 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:44:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:44:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:44:33 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:44:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:44:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:44:33 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:44:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:44:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:44:33 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:44:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:44:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:44:33 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:44:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:44:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:44:33 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:44:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:44:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:44:33 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:44:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:44:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:44:33 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ERROR 06-12 05:44:33 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:44:36 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 556.5 tokens/s, Running: 14 reqs, Waiting: 0 reqs, GPU KV cache usage: 54.9%, Prefix cache hit rate: 93.3% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:44:39 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:44:39 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:44:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:44:39 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:44:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:44:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:44:39 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:44:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:44:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:44:39 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:44:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:44:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:44:39 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:44:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:44:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:44:39 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:44:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:44:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:44:39 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:44:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:44:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:44:39 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:44:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:44:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:44:39 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:44:39 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 15x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:44:41 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 494.4 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 52.5%, Prefix cache hit rate: 93.3% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) nranks=4] stream 0x400e427c7e80 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 182 sendbuff 0x401913c00000 recvbuff 0x401913c00000 count 16777216 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 282 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4018f1000000 recvbuff 0x401389580000 count 1048576 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7e80 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 183 sendbuff 0x401389580000 recvbuff 0x401389580000 count 4194304 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 283 sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4018f1200000 recvbuff 0x401913c00000 count 1048576 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7e80 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 184 sendbuff 0x401913c00000 recvbuff 0x401913c00000 count 4194304 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 284 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4018f5c00000 recvbuff 0x401660000000 count 4194304 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7e80 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 185 sendbuff 0x401660000000 recvbuff 0x401660000000 count 16777216 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 285 sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400e9fa86e00 recvbuff 0x400e9fa88e00 count 32 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7e80 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 186 sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 128 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 286 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400e9fa87000 recvbuff 0x400e9fa88c00 count 32 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7e80 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 187 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 128 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 287 sendbuff 0x400e9fa89000 recvbuff 0x400e9fa89000 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4018f7000000 recvbuff 0x401660000000 count 12582912 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7e80 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 188 sendbuff 0x401660000000 recvbuff 0x401660000000 count 50331648 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 288 sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4018f8800000 recvbuff 0x401666000000 count 12582912 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7e80 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 189 sendbuff 0x401666000000 recvbuff 0x401666000000 count 50331648 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 289 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4018fa000000 recvbuff 0x401660000000 count 12582912 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7e80 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 18a sendbuff 0x401660000000 recvbuff 0x401660000000 count 50331648 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 28a sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400e9fa87200 recvbuff 0x400e9fa88e00 count 1024 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7e80 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 18b sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 4096 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 28b sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400e9fa87a00 recvbuff 0x400e9fa8b000 count 1024 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7e80 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 18c sendbuff 0x400e9fa8b000 recvbuff 0x400e9fa8b000 count 4096 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 28c sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400e9fa88200 recvbuff 0x400e9fa88e00 count 1024 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7e80 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 18d sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 4096 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 28d sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401900000000 recvbuff 0x401660000000 count 155582464 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7e80 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 18e sendbuff 0x401660000000 recvbuff 0x401660000000 count 622329856 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 28e sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 28f sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401420000000 recvbuff 0x401420000000 count 155583488 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401480000000 recvbuff 0x401480000000 count 155582464 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401420000000 recvbuff 0x401420000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401437004 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: op (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487912) 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) t fe sendbuff 0x400e79e88e00 recvbuff 0x400e79e88e00 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e88c00 recvbuff 0x400e79e88c00 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e88e00 recvbuff 0x400e79e88e00 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e88c00 recvbuff 0x400e79e88c00 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e88e00 recvbuff 0x400e79e88e00 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e88c00 recvbuff 0x400e79e88c00 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e88e00 recvbuff 0x400e79e88e00 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e88c00 recvbuff 0x400e79e88c00 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017c2806300 recvbuff 0x4017b7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) ount fe sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017dcc05280 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks= (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401857a40000 recv (FSDPPolicyWorkerBase pid=487914) type 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e5c01080 recvbuff 0x4017e0000000 count 48236608 d (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017eb802100 recvbuff 0x4017e0000000 count 48236608 datatyp (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 200 recvbuff 0x401437004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0 (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 33x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:44:46 [loggers.py:259] Engine 000: Avg prompt throughput: 219.9 tokens/s, Avg generation throughput: 422.1 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 52.6%, Prefix cache hit rate: 93.2% [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=487913) Count 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO All (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) [repeated 24x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 [repeated 325x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017ab802100 recvbuff 0x4017a0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [ (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487912) AllGather: opCount 0 sendbuff 0x401420000000 recvbuff 0x401420000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) buff 0x401820000000 count 155582464 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487915) op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) x401420000000 recvbuff 0x401420000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017c5c01080 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d (FSDPPolicyWorkerBase pid=487914) atatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017fcc05280 recvbuff 0x4017f7004200 count (FSDPPolicyWorkerBase pid=487915) e 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017c2806300 recvbuff 0x4017b7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400d (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuf (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) sendbuff 0x401437004200 recvbuff 0x401437004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487913) Gather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCC (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017ab802100 recvbuff 0x4017a0000000 count 48236608 datatype 9 op 0 root 0 c (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017c5c01080 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) f 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) opCount 0 sendbuff 0x401420000000 recvbuff 0x401420000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO A (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:4 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4018328c0000 recvbuff 0x401820000000 count 155582464 datatype (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) omm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401825180000 recvbuff 0x401800000000 count 155582464 datatype 9 o (FSDPPolicyWorkerBase pid=487914) 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e5c01080 recvbuff 0x4017e0000 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017eb802100 recvbuff 0x4017e0000000 co (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:44:51 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:44:51 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:44:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:44:51 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:44:51 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:44:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:44:51 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:44:51 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:44:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:44:51 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:44:51 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:44:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:44:51 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:44:51 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:44:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:44:51 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:44:51 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:44:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:44:51 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:44:51 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:44:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:44:51 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:44:51 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:44:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:44:51 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) ERROR 06-12 05:44:51 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 25x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:44:51 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 425.5 tokens/s, Running: 10 reqs, Waiting: 0 reqs, GPU KV cache usage: 52.1%, Prefix cache hit rate: 93.2% [repeated 49x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGathe (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 [repeated 881x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) llGather: opCount 0 sendbuff 0x401437004200 recvbuff 0x401437004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] N (FSDPPolicyWorkerBase pid=487913) L INFO AllGather: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:4881 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 00e156cd8d0 [nranks=4] stream 0x400e152ce4c0 [repeated 5x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) p 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017ab802100 recvbuff 0x4017a0000000 count 48236608 da (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017c5c01080 recvbuff 0x4017c0000000 count 48236608 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017dcc05280 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x4 (FSDPPolicyWorkerBase pid=487912) jpbo-006-4 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) r: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INF (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) CCL INFO AllGather: opCount 0 sendbuff 0x401420000000 recvbuff 0x401420000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:33 (FSDPPolicyWorkerBase pid=487914) 000 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017fcc05280 recvbuff 0 (FSDPPolicyWorkerBase pid=487915) unt 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:4 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017dcc05280 recvbuff 0x4017d7004200 coun (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) tatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=487914) x4017f7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e5c01080 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017eb802100 recvbu (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) O AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11 (FSDPPolicyWorkerBase pid=487913) jpb (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401802806300 recvbuff 0x4017f (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:44:56 [loggers.py:259] Engine 000: Avg prompt throughput: 82.2 tokens/s, Avg generation throughput: 414.0 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 53.5%, Prefix cache hit rate: 93.2% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487912) 0 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:44:56 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:44:56 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:44:56 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:44:56 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:44:56 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:44:56 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:44:56 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:44:56 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:44:56 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:44:56 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:44:56 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:44:56 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:44:56 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:44:56 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:44:56 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:44:56 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:44:56 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:44:56 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:44:56 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:44:56 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:44:56 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:44:56 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:44:56 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:44:56 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:44:56 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) ERROR 06-12 05:44:56 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017ab802100 recvbuff 0x4017a00000 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017c2806300 recvbuff 0x4017b7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 [repeated 806x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) t 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487914) recvbuff 0x4017e0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401 (FSDPPolicyWorkerBase pid=487915) ff 0x4017e0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) ] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) j (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017c2806300 recvbuff 0x4017b7004200 count 4 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 00000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487912) 00df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401437004200 recvbuff 0x401437004200 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] (FSDPPolicyWorkerBase pid=487914) 7fcc05280 recvbuff 0x4017f7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 send (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401802806 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017ab802100 r (FSDPPolicyWorkerBase pid=487913) 324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] strea (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 8236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 15x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:45:01 [loggers.py:259] Engine 000: Avg prompt throughput: 43.5 tokens/s, Avg generation throughput: 417.6 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 46.8%, Prefix cache hit rate: 93.2% [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=487912) stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401420000000 recvbuff 0x401420000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [ (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401420000000 recvbuff 0x401420000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] str (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:01 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:01 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:01 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:01 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:01 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:01 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:01 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:01 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:01 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:01 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:01 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:01 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:01 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:01 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:01 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:01 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:01 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:01 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:01 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:01 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:01 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:01 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:01 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:01 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:01 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017dcc05280 recvbuff [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401437004200 recvbuff 0x401437004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 [repeated 765x across cluster] (FSDPPolicyWorkerBase pid=487914) buff 0x4017e5c01080 recvbuff 0x4017e0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCo (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487913) m 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nrank (FSDPPolicyWorkerBase pid=487915) x4017eb802100 recvbuff 0x4017e0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) ecvbuff 0x4017a0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=487912) nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401437004200 recvbuff 0x401437004200 count 48236608 datatype 9 op 0 root 0 comm 0x400d (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) eam 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401437004200 recvbuff 0x401437004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nra (FSDPPolicyWorkerBase pid=487914) unt 0 sendbuff 0x4017fcc05280 recvbuff 0x4017f7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGa (FSDPPolicyWorkerBase pid=487915) sendbuff 0x401802806300 recvbuff 0x4017f7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendb (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 se (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 00e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401420000000 recvbuff 0x401420000000 count 48236608 datatype 9 op 0 root 0 c (FSDPPolicyWorkerBase pid=487914) ther: opCount 0 sendbuff 0x4017e5c01080 recvbuff 0x4017e0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487915) opCount 0 sendbuff 0x4017eb802100 recvbuff 0x4017e0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO A (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:45:06 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:45:06 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:45:06 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:45:06 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:45:06 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:45:06 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:45:06 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:45:06 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:45:06 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:45:06 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:45:06 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:45:06 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:45:06 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:45:06 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:45:06 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:45:06 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:45:06 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:45:06 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:45:06 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:45:06 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:45:06 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:45:06 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:45:06 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:45:06 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) uff 0x4017ab802100 recvbuff 0x4017a0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCou (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) ndbuff 0x4017c5c01080 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: op (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 18x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:45:06 [loggers.py:259] Engine 000: Avg prompt throughput: 137.3 tokens/s, Avg generation throughput: 401.3 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 44.6%, Prefix cache hit rate: 93.2% [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=487913) c90 [nranks=4] stream 0xaaaafb324d40 [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [ (FSDPPolicyWorkerBase pid=487912) omm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401437004200 recvbuff 0x401437004200 count 48236608 datatype 9 op (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401437004200 recvbuff 0x401437004200 count 48236608 datatype 9 op 0 root 0 co (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) ERROR 06-12 05:45:06 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401420000000 recvbuff 0x401420000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 [repeated 961x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487914) INFO AllGather: opCount 0 sendbuff 0x4017fcc05280 recvbuff 0x4017f7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 (FSDPPolicyWorkerBase pid=487915) llGather: opCount 0 sendbuff 0x401802806300 recvbuff 0x4017f7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] N (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 017dcc05280 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 [repeated 3x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 r (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) nt 0 sendbuff 0x4017c2806300 recvbuff 0x4017b7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGat (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) Count 0 sendbuff 0x4017dcc05280 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO All (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) mm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=487912) 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401420000000 recvbuff 0x401420000000 count 48236608 data (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401857a40000 recvbuff 0x401820000000 count 155582464 datatype 9 op 0 root 0 comm 0x40 [repeated 3x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) her: opCount 0 sendbuff 0x4017ab802100 recvbuff 0x4017a0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL I (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) Gather: opCount 0 sendbuff 0x4017c5c01080 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCC (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:08 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:08 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:08 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:08 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:08 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:08 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:08 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:08 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:08 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:08 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:08 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:08 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:08 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:08 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:08 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:08 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 8x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:45:11 [loggers.py:259] Engine 000: Avg prompt throughput: 380.6 tokens/s, Avg generation throughput: 465.6 tokens/s, Running: 13 reqs, Waiting: 0 reqs, GPU KV cache usage: 47.1%, Prefix cache hit rate: 93.2% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487 (FSDPPolicyWorkerBase pid=487915) CCL INFO AllGather: opCount 0 sendbuff 0x4017eb802100 recvbuff 0x4017e0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:48 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) nranks=4] stream 0x400e192bd380 [repeated 4x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401420000000 recvbuff 0x401420000000 count 48236608 datatype 9 op 0 (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:08 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017eb802100 recvbuff 0x4017e0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 [repeated 481x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487913) oot 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatyp (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:45:13 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:45:13 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:45:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:45:13 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:45:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:45:13 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:45:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:45:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:45:13 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:45:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:45:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:45:13 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:45:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:45:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:45:13 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:45:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:45:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:45:13 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:45:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:45:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:45:13 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:45:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:45:13 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:45:13 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487912) type 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401437004200 recvbuff 0x401437004200 count 482 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487914) jpbo- (FSDPPolicyWorkerBase pid=487915) jpbo-006-45 (FSDPPolicyWorkerBase pid=487913) e 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 27x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:45:16 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 513.3 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 44.1%, Prefix cache hit rate: 93.2% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:45:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 [repeated 202x across cluster] (FSDPPolicyWorkerBase pid=487912) 36608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401420000000 recvbuff 0x401420000000 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:45:17 [serving.py:315] Error in preprocessing prompt inputs [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:45:17 [serving.py:315] Traceback (most recent call last): [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:45:17 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:45:17 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:45:17 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:45:17 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:45:17 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:45:17 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:45:17 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:45:17 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:45:17 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:45:17 [serving.py:315] return await asyncio.gather( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:45:17 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:45:17 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:45:17 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:45:17 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:45:17 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:45:17 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:45:17 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:45:17 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:45:17 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:45:17 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:45:17 [serving.py:315] raise VLLMValidationError( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ERROR 06-12 05:45:17 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 4x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 482366 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 co (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487912) count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401437004200 recvbuff 0x40 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 26x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) INFO 06-12 05:45:21 [loggers.py:259] Engine 000: Avg prompt throughput: 43.1 tokens/s, Avg generation throughput: 430.6 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 48.5%, Prefix cache hit rate: 93.3% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:20 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 6x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401420000000 recvbuff 0x401420000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 [repeated 202x across cluster] (FSDPPolicyWorkerBase pid=487913) 08 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:20 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:20 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:20 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:20 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:20 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:20 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:20 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:20 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:20 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:20 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:20 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:20 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:20 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:20 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:20 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:20 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:20 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:20 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:20 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:20 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:20 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:20 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:20 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:20 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:45:24 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:45:24 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:45:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:45:24 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:45:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:45:24 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:45:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:45:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:45:24 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:45:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:45:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:45:24 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:45:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:45:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:45:24 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:45:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:45:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:45:24 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:45:24 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:45:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:45:24 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:45:24 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:45:24 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) ERROR 06-12 05:45:24 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 22x across cluster] (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) INFO 06-12 05:45:26 [loggers.py:259] Engine 000: Avg prompt throughput: 167.3 tokens/s, Avg generation throughput: 339.6 tokens/s, Running: 7 reqs, Waiting: 0 reqs, GPU KV cache usage: 25.4%, Prefix cache hit rate: 93.1% [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 0e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401857a41800 recvbuff 0x401820000000 count 155583488 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401997a40000 recvbuff 0x401960000000 count 155582464 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401848407380 recvbuff 0x401837004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e19 (FSDPPolicyWorkerBase pid=487914) 0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487915) (FSDPPolicyWorkerBase pid=487915) jpbo (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 roo (FSDPPolicyWorkerBase pid=487913) unt 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count (FSDPPolicyWorkerBase pid=487912) 1437004200 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401677004200 recvbuff 0x401677 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) NFO AllGather: opCount 0 sendbuff 0x4017c2806300 recvbuff 0x4017b7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO A (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) L INFO AllGather: opCount 0 sendbuff 0x4017dcc05280 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INF (FSDPPolicyWorkerBase pid=487914) 0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401a80000000 recvbuff 0x401b4e600000 count 48236608 datatype 7 op 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x401b4e600000 recvbuff 0x401b4e600000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401a80000000 recvbuff 0x401ae9200000 count 48236608 datatype 7 op 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount 190 sendbuff 0x401ae9200000 recvbuff 0x401ae9200000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 190 sendbuff 0x401ae3400000 recvbuff 0x401ae3400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] strea (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) t 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 6bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401b00000000 recvbuff 0x401bce000000 count 48236608 datatype 7 op 4 r (FSDPPolicyWorkerBase pid=487913) 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 194 sendbuff (FSDPPolicyWorkerBase pid=487912) 004200 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [ (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401a80000000 re (FSDPPolicyWorkerBase pid=487915) m 0x400e193f14b0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401a80000000 recvbuff (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) llGather: opCount 0 sendbuff 0x401822806300 recvbuff 0x401817004200 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) O AllGather: opCount 0 sendbuff 0x40177cc05280 recvbuff 0x401777004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=487913) 0x401a11e00000 recvbuff 0x401a11e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 1a7 sendbuff 0x401bb1800000 recvbuff 0x401bb1800000 count 48236608 datatype 7 (FSDPPolicyWorkerBase pid=487914) cvbuff 0x401b60000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (FSDPPolicyWorkerBase pid=487915) jpbo- (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) oot 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: op (FSDPPolicyWorkerBase pid=487913) 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 1af sendbuff 0x401ca5e00000 recvbuf (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) Count 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4019cb19c000 recvbuff 0x401db1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bc (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:25 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:25 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:25 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:25 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:25 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:25 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:25 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:25 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:25 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:25 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:25 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:25 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:45:25 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 (FSDPPolicyWorkerBase pid=487914) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 send (FSDPPolicyWorkerBase pid=487912) op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) j (FSDPPolicyWorkerBase pid=487913) f 0x401ca5e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nrank (FSDPPolicyWorkerBase pid=487912) recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) buff 0x40179cc05280 recvbuff 0x401797004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllRed (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) ] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e88c00 recvbuff 0x400e79e88c00 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401ae0000000 recvbuff 0x401845200000 count 48236608 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 1b8 sendbuff 0x4014a5200000 recvbuff 0x4014a5200000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [ (FSDPPolicyWorkerBase pid=487915) uce: opCount 1b7 sendbuff 0x401680000000 recvbuff 0x401680000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 (FSDPPolicyWorkerBase pid=487914) AllReduce: opCount 1b7 sendbuff 0x401685e00000 recvbuff 0x401685e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) INFO 06-12 05:45:31 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 274.0 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 19.9%, Prefix cache hit rate: 93.1% [repeated 50x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL I (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:33403 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 1d2 sendb (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount 1d2 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 80 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 1c4 sendbuff 0x401cc0000000 recvbuff 0x401cc0000 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401ae0000000 recvbuff 0x401cc0000000 count 48236608 datatype 7 op 4 root 0 (FSDPPolicyWorkerBase pid=487914) 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 [repeated 8x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 [repeated 1210x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 299 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) [repeated 38x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401785c01080 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm  [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) [repeated 38x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401420 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 4 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401e9545ac00 recvbuff 0x401d65e00000 count 48236608 datatype 7 op (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 200 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40177cc05280 recvbuff 0x401777004 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487913) NFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401d7306b800 recvbuff 0x401be0000000 count 48236608 datatype 7 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:33 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:33 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:33 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:33 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:33 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:33 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:33 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:33 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:33 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:33 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:33 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823: (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401db1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 [repeated 525x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 1d1 sendbuff 0x401db1800000 recvbuff 0x401db1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 [repeated 518x across cluster] (FSDPPolicyWorkerBase pid=487913) datatype 7 op 4 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 294 sendbuff 0x400e9fa88c00 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [ (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 4 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df13e9b50 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e88c00 recvbuff 0x400e79e88c00 count 1 datatyp (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 000000 recvbuff 0x401b31800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 data (FSDPPolicyWorkerBase pid=487912) 0dfa7c68a0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 29f sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) e427d8400 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount (FSDPPolicyWorkerBase pid=487914) sendbuff 0x401c71800000 recvbuff 0x401c71800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40179cc05280 recvbuff 0x401797004200 count 48236608 datatype (FSDPPolicyWorkerBase pid=487915) uff 0x401c60000000 recvbuff 0x401c60000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 (FSDPPolicyWorkerBase pid=487913) comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] N (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) e 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401822806300 recvbuff 0x401817004200 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks= (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) type 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40177cc05280 recvbuff 0x401777004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nran (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334 (FSDPPolicyWorkerBase pid=487912) stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017a2806300 recvbuff 0x401797004200 count 48236608 datatype 9 op (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401cf9188000 recvbuff 0x40168ba0000 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount 1e7 sendbuff 0x401c71800000 rec (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487913) CCL INFO AllGather: opCount 0 sendbuff 0x401788407380 recvbuff 0x401777004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] st (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) : opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401cd1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e1 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 1e7 sendbuff 0x401ace600000 recvbuff 0x401ace600000 count 48236608 datatyp (FSDPPolicyWorkerBase pid=487914) 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4018 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) vbuff 0x401c71800000 count 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jp (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:4 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915: (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllR (FSDPPolicyWorkerBase pid=487913) ream 0xaaaafb324d40 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487912) e 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO Reduc (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4018a2000000 recvbuff 0x401c71800000 count 48236608 datat (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401960000000 recvbuff 0x401c80000000 count 48236608 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL IN (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL (FSDPPolicyWorkerBase pid=487912) eScatter: opCount 0 sendbuff 0x4014e2000000 recvbuff 0x401bd1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=487915) ype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 (FSDPPolicyWorkerBase pid=487915) jpbo- (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) INFO 06-12 05:45:36 [loggers.py:259] Engine 000: Avg prompt throughput: 370.3 tokens/s, Avg generation throughput: 299.9 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 27.1%, Prefix cache hit rate: 93.1% [repeated 47x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) educe: opCount 1fa sendbuff 0x401c40000000 recvbuff 0x401c40000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401677004200 recvbuff 0x401677004200 count 48 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401788407380 recv (FSDPPolicyWorkerBase pid=487912) 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) FO AllReduce: opCount fe sendbuff 0x400e79e88c00 recvbuff 0x400e79e88c00 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO A (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INF (FSDPPolicyWorkerBase pid=487913) buff 0x401777004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NC (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 204 sendbuff 0x4017e5200000 recvbuff 0x4017e (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x40 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 96bcf10 [nranks=4] stream 0x400e193e8890 [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 [repeated 1064x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) [repeated 70x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401785c01080 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 68x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) llGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) O AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4017c0000000 recvbuff 0x401aee600000 count 48236608 datatype 7 op 4 root 0 c (FSDPPolicyWorkerBase pid=487913) CL INFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401958408000 recvbuff 0x401bd1800000 count 48236608 datatype 7 op 4 ro (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 5200000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=487912) 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x4 (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:45:35 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:45:35 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:45:35 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:45:35 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:45:35 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:45:35 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:45:35 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:45:35 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:45:35 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:45:35 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:45:35 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:45:35 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:45:35 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:45:35 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:45:35 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:45:35 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:45:35 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:45:35 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:45:35 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:45:35 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:45:35 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:45:35 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:45:35 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:45:35 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:45:35 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (FSDPPolicyWorkerBase pid=487913) ot 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opC (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather: opCount 0 se (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) omm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9c80 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401bf1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9c80 [repeated 580x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount 21b sendbuff 0x401bf1800000 recvbuff 0x401bf1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 [repeated 586x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 21f sendbuff 0x401de0000000 recvbuff 0x401de0000000 count 48236608 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 212 sendbuff 0x401b7180 [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) x400e427d8400 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) ndbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=487915) 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) 0000 recvbuff 0x401b71800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401900000000 recvbuff 0x401900000000 count 155582464 datatype 9 op 0 ro (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGather: opCount 0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401677004200 recvbuff 0x401677004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks= (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nran (FSDPPolicyWorkerBase pid=487914) aaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendb (FSDPPolicyWorkerBase pid=487915) 967e0 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount 227 sendbuff 0x401490a00000 recvbuff 0x401490 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 dataty (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 92bd380 (FSDPPolicyWorkerBase pid=487913) ount 0 sendbuff 0x401797a40000 recvbuff 0x401760000000 count 155582464 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401788407380 recvbuff 0x401777004200 count 48236608 datatype 9 op 0 root 0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401433b18000 recvbuff 0x [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 1ef sendbuff 0x401bc0000000 rec (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4018925d2000 recvbuff 0x4018e0000000 count 48236608 datatype 7 op 4 root 0 comm 0x4 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:3328 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:3 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) pe 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO (FSDPPolicyWorkerBase pid=487913) vbuff 0x401bc0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (FSDPPolicyWorkerBase pid=487913) jp (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 sendbu (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 227 sendbuff 0x401060000000 recvbuff 0x401060000000 count 48236608 dat (FSDPPolicyWorkerBase pid=487914) uff 0x401785c01080 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [ (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCC (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) AllReduce: opCount 23a sendbuff 0x401b60000000 recvbuff 0x401b60000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) ff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=487913) comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) INFO 06-12 05:45:41 [loggers.py:259] Engine 000: Avg prompt throughput: 121.6 tokens/s, Avg generation throughput: 313.8 tokens/s, Running: 7 reqs, Waiting: 0 reqs, GPU KV cache usage: 30.1%, Prefix cache hit rate: 94.4% [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4018f28c0000 recvbuff 0x4018e0000000 count  [repeated 2x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487912) atype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO R (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 13x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 00e196bcf10 [nranks=4] stream 0x400e193e8890 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 [repeated 718x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) [repeated 36x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 36x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401b00000000 recvbuff 0x401ba0000000 count 48236608 datatype 7 op (FSDPPolicyWorkerBase pid=487915) L INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4018a2000000 recvbuff 0x401b85e00000 count 48236608 datatype 7 op 4 roo (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:45:42 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:45:42 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:45:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:45:42 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:45:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:45:42 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:45:42 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:45:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:45:42 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:45:42 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:45:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:45:42 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:45:42 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:45:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:45:42 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:45:42 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:45:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:45:42 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:45:42 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:45:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:45:42 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:45:42 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:45:42 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:45:42 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 4a0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) ERROR 06-12 05:45:42 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487912) ot 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401ae0000000 recvbuff 0x401bd1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 [repeated 280x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 23c sendbuff 0x401bd1800000 recvbuff 0x401bd1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 [repeated 274x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount 242 sendbuff 0x401c00000000 recvbuff 0x401c00000000 count 482366 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 155582464 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 22f sendbuff 0x4018ca400000 (FSDPPolicyWorkerBase pid=487912) educeScatter: opCount 0 sendbuff 0x4019c0000000 recvbuff 0x401af1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 242 sendbuff 0x401be5e00000 recvbuff 0x401be5e00000 count 48236608 datatype 7 op 4 root 0 co (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487914) 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllGather (FSDPPolicyWorkerBase pid=487913) recvbuff 0x4018ca400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487915) t 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCo (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487913) 0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 24a sendbuff 0x401ce0000000 recvbuff 0x401ce0000 [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40177cc05280 recvbuff 0x401777004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] str (FSDPPolicyWorkerBase pid=487912) mm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487914) : opCount 0 sendbuff 0x401785c01080 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401785c01080 recvbuff 0x401780000000 count 48236608 datatype (FSDPPolicyWorkerBase pid=487915) unt 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: (FSDPPolicyWorkerBase pid=487913) jpbo- (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401be0000000 recvbuff 0x401df1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount 252 sendbuff 0x4017d0200000 rec (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401a8bf70000 recvbuff 0x40106ba00000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] str (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) eam 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4017c0000000 recvbuff 0x4014b1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e156c (FSDPPolicyWorkerBase pid=487914) 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 0x400df12be390 (FSDPPolicyWorkerBase pid=487914) vbuff 0x4017d0200000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487914) jp (FSDPPolicyWorkerBase pid=487915) jpbo-006 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 253 sendbuff 0x401565e00000 recvbuff 0x401565e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stre (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) opCount 0 sendbuff 0x401848407380 recvbuff 0x401837004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487912) jpb (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount 26d sendbuff 0x401ca5e00000 recvbuff 0x401ca5e00000 c (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 271 sendbuff 0x401e00000000 recvbuff 0x401e00000000 count 155583488 datatype 7 op 4 root 0 comm 0x4 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCC (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) am 0x400e427d8400 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401420000000 recvbuf (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) f 0x401b71800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCC (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) INFO 06-12 05:45:46 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 287.6 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 26.6%, Prefix cache hit rate: 94.4% [repeated 49x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 400df13e9b50 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0x400e153f9c80 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGathe (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) L INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401e00000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 (FSDPPolicyWorkerBase pid=487914) ount 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487915) jpbo-0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbu (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 se (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 30x across cluster] (FSDPPolicyWorkerBase pid=487912) 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL IN (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) r: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [ (FSDPPolicyWorkerBase pid=487913) 00e195a95c0 [nranks=2] stream 0x400e193f4220 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 [repeated 925x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) [repeated 72x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 72x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) ff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL I (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401ae91be000 recvbuff 0x401a11e00000 count 48236608 datatype 7 op (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) L INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:45:46 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:45:46 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:45:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:45:46 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:45:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:45:46 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:45:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:45:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:45:46 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:45:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:45:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:45:46 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:45:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:45:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:45:46 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:45:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:45:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:45:46 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:45:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:45:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:45:46 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:45:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:45:46 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:45:46 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (FSDPPolicyWorkerBase pid=487913) ndbuff 0x401771403180 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (FSDPPolicyWorkerBase pid=487912) FO ReduceScatter: opCount 0 sendbuff 0x401a80000000 recvbuff 0x401300000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 282 sendbuff 0x401b00000000 recvbuff 0x401b00000000 count 48236608 datatype 7 op 4 root (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount 282 sendbuff 0x401ca5e00000 recvbuff 0x401ca5e00000 count 48236608 datatype 7 op (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount 282 sendbuff 0x401b20000000 recvbuff 0x401b20000000 count 48236608 datatype 7 (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:45:47] OK: 391 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:45:47] OK: RSS 3.29 GiB | node mem 382.2/858.0 GiB used (44.5%), avail 475.8 GiB (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401d05e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nrank (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) NFO AllReduce: opCount 27a sendbuff 0x401913c00000 recvbuff 0x401913c00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 c (FSDPPolicyWorkerBase pid=487913) 24d40 (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:45:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:45:48 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:45:48 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:45:48 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:45:48 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:45:48 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:45:48 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:45:48 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:45:48 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:45:48 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:45:48 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:45:48 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:45:48 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:45:48 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:45:48 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:45:48 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:45:48 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:45:48 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:45:48 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:45:48 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:45:48 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:45:48 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:45:48 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:45:48 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:45:48 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (FSDPPolicyWorkerBase pid=487912) 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCo (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatte (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceSca (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401d7bdd7000 recvbuff 0x401c11800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9c80 [repeated 594x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount 28f sendbuff 0x401c11800000 recvbuff 0x401c11800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 [repeated 593x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4018a2000000 recvbuff 0x401ca5e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] s (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) omm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 25f sendbuff 0x401cf1800000 recvbuff 0x401cf1800000 count 4823 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487915) 8236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo- (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 482 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) tter: opCount 0 sendbuff 0x401d7bdd7000 recvbuff 0x401c25e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9c80 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401848407380 recvbuff 0x401837004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 co (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbu (FSDPPolicyWorkerBase pid=487915) tream 0x400e193ee7c0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 28a sendbuff 0x401c00000000 recvbuff 0x401c00 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 rec (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487912) unt 0 sendbuff 0x4014e2000000 recvbuff 0x401c11800000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recv (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather:  [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) mm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401420000000 rec (FSDPPolicyWorkerBase pid=487914) ff 0x40179cc05280 recvbuff 0x401797004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:4879 (FSDPPolicyWorkerBase pid=487912) eam 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401ce0000000 recvbuff 0x4017ca400000 count 48236608 datatype 7 op 4 root 0 comm 0x400 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401848407380 recvbuff 0x401837004200 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGat (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) vbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) j (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4 [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] (FSDPPolicyWorkerBase pid=487912) buff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 2ad sendbuff 0x401bb1800000 recvbuff 0x401bb1800000 count 48236608 dataty (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) INFO 06-12 05:45:51 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 249.5 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 26.6%, Prefix cache hit rate: 94.4% [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=487913) her: opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 coun (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 c (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487912) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 co (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbu (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4014f033c000 recvbuff 0x401c91800000 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllGather: opCount 0 se [repeated 2x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 295 sendbuff 0x401c80000000 recvbuff 0x401c80000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) r: opCount 0 sendbuff 0x401866726000 recvbuff 0x401db1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df13e9b50 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) [nranks=4] stream 0x400e192bd380 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 [repeated 935x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) [repeated 36x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 36x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INF (FSDPPolicyWorkerBase pid=487915) pe 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO Redu (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) ff 0x401800000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=487914) ndbuff 0x401785c01080 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) t 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) ount 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stre (FSDPPolicyWorkerBase pid=487912) unt 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recv (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:33304 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 2ba sendbuff 0x401e11800000 recvbuff 0x401e11800000 count 48236608 datatype 7 op (FSDPPolicyWorkerBase pid=487914) O ReduceScatter: opCount 0 sendbuff 0x401960000000 recvbuff 0x401cd1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40179cc05280 recvbuff 0x401797004200 count 48236608 datatype 9 (FSDPPolicyWorkerBase pid=487915) ceScatter: opCount 0 sendbuff 0x4018a2000000 recvbuff 0x401cc0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017a2806300 recvbuff 0x401797004200 count 48236608 datatype 9 op 0 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401b00000000 recvbuff 0x401cc0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e616ca9c0 [nrank (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487913) buff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] s (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:45:51 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:45:51 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:45:51 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:45:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:45:51 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:45:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:45:51 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:45:51 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:45:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:45:51 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:45:51 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:45:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:45:51 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:45:51 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:45:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:45:51 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:45:51 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:45:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:45:51 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:45:51 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:45:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:45:51 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:45:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:45:51 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:45:51 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount 2c2 sendbuff 0x4017b0a00000 recvbuff 0x4017b0a00000 count 48236608 datatype 7 op 4 root 0 comm (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) am 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount 2c2 sendbuff 0x401732a00000 recvbuff 0x401732a00000 count 48236608 datatype 7 op 4 root 0 co (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 2c2 sendbuff 0x401311800000 recvbuff 0x401311800000 count 48236608 datatype 7 op 4 r (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401760000000 recvbuff 0x401bce000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [n (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487914) op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4019f0b3a000 recvbuff 0x40196a400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 [repeated 450x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 2c7 sendbuff 0x40196a400000 recvbuff 0x40196a400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 [repeated 453x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount 2ad sendbuff 0x401bc0000000 recvbuff 0x401bc0000000 count 48236608  [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487914) datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 [repeated 6x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487912) oot 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 s (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatter: opCount (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opC (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count 482366 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45: (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) endbuff 0x401740000000 recvbuff 0x401cd1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df13e9b50 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount 2dd sendbuff 0x401dc5e00000 recvbuff 0x401dc5e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0 sendbuff 0x401882000000 recvbuff 0x401b40000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9c80 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount 2dd sendbuff 0x401c40000000 recvbuff 0x401c40000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nra (FSDPPolicyWorkerBase pid=487915) root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=487914) 400e613f6c90 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017928c0000 recvbuff 0x401 (FSDPPolicyWorkerBase pid=487915) 3ee7c0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) ount 2d5 sendbuff 0x401b91800000 recvbuff 0x401b91800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 (FSDPPolicyWorkerBase pid=487913) tream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stre (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 send (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: o (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 193e8890 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NC (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) pCount 0 sendbuff 0x401822806300 recvbuff 0x401817004200 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) : opCount 0 sendbuff 0x40177cc05280 recvbuff 0x401777004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11: (FSDPPolicyWorkerBase pid=487915) jpbo-00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) buff 0x401677004200 recvbuff 0x401677004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 2ca sendbuff 0x401acc000000 recvbuff 0x40 (FSDPPolicyWorkerBase pid=487913) am (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL IN (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 dat (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 400 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401420000000 recvbuff 0x401ba5e000 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 2fa sendbuff 0x401d31800000 recvbuff 0x401d31800000 count 48236608 datatype (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487912) CL INFO AllGather: opCount 0 sendbuff 0x401677004200 recvbuff 0x401677004200 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) atype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-1 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) mm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO Reduce (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount 2ed sendbuff 0x401870c00000 recvbuff 0x401870c00000 count 48236608 datatype 7 op (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 2ed sendbuff 0x40186b000000 recvbuff 0x40186b000000 count 48236608 datatype 7 op 4 root (FSDPPolicyWorkerBase pid=487913) FO AllGather: opCount 0 sendbuff 0x401788407380 recvbuff 0x401777004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401d00000000 recvbuff 0x401b2e800000 count 48236608 datatype 7 op 4 root 0 comm  [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487912) opCount 0 sendbuff 0x401a8ab38000 recvbuff 0x401b25e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487912) 68a0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatte (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) vbuff 0x40196a400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 (FSDPPolicyWorkerBase pid=487914) j (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opC (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sen (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4018314031 [repeated 4x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:56 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:56 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:56 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:56 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:56 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:56 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:56 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:56 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:56 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:56 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:56 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:56 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:56 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:56 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:56 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:56 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:56 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:56 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:56 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:56 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:56 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:56 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:56 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:45:56 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:48791 (FSDPPolicyWorkerBase pid=487914) r: opCount 0 sendbuff 0x401b00000000 recvbuff 0x401be5e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (FSDPPolicyWorkerBase pid=487913) dbuff 0x401771403180 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) FO AllReduce: opCount 304 sendbuff 0x401de0000000 recvbuff 0x401de0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) INFO AllReduce: opCount 304 sendbuff 0x401c51800000 recvbuff 0x401c51800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) Scatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401e40000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) INFO 06-12 05:45:56 [loggers.py:259] Engine 000: Avg prompt throughput: 128.6 tokens/s, Avg generation throughput: 423.9 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 28.1%, Prefix cache hit rate: 93.2% [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401677004200 recvb (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4018 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) opCount 315 sendbuff 0x401a20c00000 recvbuff 0x401a20c00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuf (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 23x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:45:57] OK: 366 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:45:57] OK: RSS 2.98 GiB | node mem 387.6/858.0 GiB used (45.2%), avail 470.4 GiB (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 00df13eca20 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401862000000 recvbuff 0x40 (FSDPPolicyWorkerBase pid=487915) 0 [nranks=2] stream 0x400e193f14b0 [repeated 11x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 [repeated 1578x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 2ea sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) [repeated 104x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 108x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0x400e153fcbf0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 400e193e8890 (FSDPPolicyWorkerBase pid=487913) eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 (FSDPPolicyWorkerBase pid=487912) 0 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 [repeated 5x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:45:58 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=487912) uff 0x401677004200 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatt (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 308 sendbuff 0x401cd1800000 recvbuff 0x401cd1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c [repeated 3x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401ca0000000 recvbuff 0x40146a400000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 [repeated 681x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 30f sendbuff 0x40146a400000 recvbuff 0x40146a400000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 [repeated 680x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 [repeated 4x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487914) f 0x401785c01080 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) j (FSDPPolicyWorkerBase pid=487915) jpbo-00 (FSDPPolicyWorkerBase pid=487915) 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401677004200 recvbuff 0x401677004200 count 48236608 [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487915) 178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuff  [repeated 3x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487913) jpbo-006 (FSDPPolicyWorkerBase pid=487915) ount 0 sendbuff 0x401b00000000 recvbuff 0x401bd1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487912) er: opCount 0 sendbuff 0x401a98564000 recvbuff 0x4014a5200000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL IN (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401785c01080 recvbuff 0x401780000000 coun (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 325 sendbuff 0x401c31800000 recvbuff 0x401c31800000 count 4 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487914) t 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:48926 (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:59 [serving.py:315] Error in preprocessing prompt inputs [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:59 [serving.py:315] Traceback (most recent call last): [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:59 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:59 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:59 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:59 [serving.py:315] return await asyncio.gather( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:59 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:59 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:59 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:59 [serving.py:315] raise VLLMValidationError( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:59 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40177cc05280 recvbuff 0x401777004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] strea (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401677004200 recvbuff 0x401677004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] str (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype 9 op 0 r (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 2ec sendbuff 0x400e9 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) m 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401677004200 recvbuff 0x401677004200 count 48236608 datatype 9 op 0 root 0 c (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) INFO 06-12 05:46:01 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 455.7 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 40.0%, Prefix cache hit rate: 93.8% [repeated 47x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) eam 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 33a sendbuff 0x401c45e00000 recvbuff 0x401c45e00000 count 48236608 datat (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) oot 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduc (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) ype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO Red (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:33 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) e: opCount 344 sendbuff 0x401b65e00000 recvbuff 0x401b65e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) uceScatter: opCount 0 sendbuff 0x401760000000 recvbuff 0x401d51800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401a20000000 recvbuff 0x401f00000000 count 155582464 datatype 7 op 4 roo (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401785c01080 recvbuff 0x401780000000 count 48236608 datatype 9 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) opCount 344 sendbuff 0x401cf1800000 recvbuff 0x401cf1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e15 [repeated 3x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401a9b97c000 recvbuff 0x401480000000 count 48236608 datatype 7 op 4 root 0 c (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401420000000 recvbuff 0x401a2c [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487913) 3c90 [nranks=4] stream 0xaaaafb324d40 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 [repeated 671x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 2f3 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) [repeated 36x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 36x across cluster] (FSDPPolicyWorkerBase pid=487912) omm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 600000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 [repeated 5x across cluster] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:45:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487914) op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 s (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 31d sendbuff 0x401b40000000 recvbuff 0x401b40000000 count 48236608 datatype 7 op 4 root 0 comm 0x400df (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuf (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401ce0000000 recvbuff 0x40194a400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 [repeated 322x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 33a sendbuff 0x40194a400000 recvbuff 0x40194a400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 [repeated 324x across cluster] (FSDPPolicyWorkerBase pid=487913) 8236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 4823 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:4 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=487912) 0x4014e2000000 recvbuff 0x401c45e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x4016600 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487914) endbuff 0x4017e0000000 recvbuff 0x401950200000 count 48236608 datatype 7 op 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount 348 sendbuff 0x401c00000000 recvbuff 0x401c00000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks (FSDPPolicyWorkerBase pid=487915) f 0x401ce0000000 recvbuff 0x4019b2800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 348 sendbuff 0x401be5e00000 recvbuff 0x401be5e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] st (FSDPPolicyWorkerBase pid=487913) sendbuff 0x401771403180 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x40194eeea000 recvbuff 0x401c45e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c3c90 [nr (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 d (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 2f7 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nr (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913: (FSDPPolicyWorkerBase pid=487915) ream 0x400e193f14b0 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: op (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) anks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) t 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCou (FSDPPolicyWorkerBase pid=487912) atatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-0 (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] [05:46:05] OK: 382 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] [05:46:05] OK: RSS 3.21 GiB | node mem 399.6/858.0 GiB used (46.6%), avail 458.4 GiB (FSDPPolicyWorkerBase pid=487914) b9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jp (FSDPPolicyWorkerBase pid=487915) jpbo-006 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) Count 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) nt 0 sendbuff 0x401848407380 recvbuff 0x401837004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401760000000 recvbuff 0x401c60000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] strea (FSDPPolicyWorkerBase pid=487913) jpbo (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 35d sendbuff 0x40152a400000 recvbuff 0x40152a400000 count 48236608 datatype 7 op 4 root 0 comm 0x4 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401e9fec1000 recvbuff 0x401d05e00000 cou (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401785c01080 recvbuff 0x401780000000 count 48236608 dat (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 370 sendb (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 365 sendbuff 0x401b51800000 recvbuff 0x401b51800000 cou (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) nt 48236608 datatype 7 op 4 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df13e9b50 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:33405 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 r (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) m 0x400e193e8890 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:33 (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) INFO 06-12 05:46:06 [loggers.py:259] Engine 000: Avg prompt throughput: 76.1 tokens/s, Avg generation throughput: 412.7 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 27.0%, Prefix cache hit rate: 93.2% [repeated 49x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 send (FSDPPolicyWorkerBase pid=487914) atype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL IN (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO All (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487915) oot 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype (FSDPPolicyWorkerBase pid=487912) buff 0x401ca0000000 recvbuff 0x401b60000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 378 sendbuff 0x401c60000000 recvbuff 0x401c60000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) uff 0x401bd1800000 recvbuff 0x401bd1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 303 se (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11: (FSDPPolicyWorkerBase pid=487914) FO AllReduce: opCount 36f sendbuff 0x401c11800000 recvbuff 0x401c11800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] st (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178 (FSDPPolicyWorkerBase pid=487915) Reduce: opCount 36f sendbuff 0x401c00000000 recvbuff 0x401c00000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream ( (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 33x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401b80000000 (FSDPPolicyWorkerBase pid=487912) 00dfa982e60 [nranks=2] stream 0x400dfa7c98e0 [repeated 6x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401822806300 recvbuff 0x401817004200 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 [repeated 818x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) [repeated 51x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e88c00 recvbuff 0x400e79e88c00 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) [repeated 47x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) ndbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sen (FSDPPolicyWorkerBase pid=487913) omm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount 384 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-1 (FSDPPolicyWorkerBase pid=487915) 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487914) ream (nil) (FSDPPolicyWorkerBase pid=487915) nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401420000000 recvbuff 0x401a8a400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 [repeated 485x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 389 sendbuff 0x401a8a400000 recvbuff 0x401a8a400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 [repeated 484x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) ndbuff 0x401acd800000 recvbuff 0x401acd800000 count 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks (FSDPPolicyWorkerBase pid=487913) 8407380 recvbuff 0x401777004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) ReduceScatter: opCount 0 sendbuff 0x401760000000 recvbuff 0x401c71800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 395 sendbuff 0x401d65e00000 recvbuff 0x401d65e00000 count 48236608 datatype 7 op 4 root 0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401420000000 recvbuff 0x401be0000000 count 48236 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount 39f sendbuff 0x401e11800000 (FSDPPolicyWorkerBase pid=487912) dbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006- (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) sendbuff 0x401845200000 recvbuff 0x401845200000 count 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401ae0000000 recvbuff 0x401b60000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c3c90 (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:46:10 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:46:10 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:46:10 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:46:10 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:46:10 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:46:10 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:46:10 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:46:10 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:46:10 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:46:10 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:46:10 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:46:10 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:46:10 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:46:10 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:46:10 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:46:10 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:46:10 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:46:10 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:46:10 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:46:10 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:46:10 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:46:10 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:46:10 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:46:10 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:46:10 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:46:10 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487914) 00e613f9c00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4017e0000000 recvbuff 0x40 (FSDPPolicyWorkerBase pid=487915) f14b0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCou (FSDPPolicyWorkerBase pid=487915) 00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jp (FSDPPolicyWorkerBase pid=487913) nt 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x4 (FSDPPolicyWorkerBase pid=487913) nt 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] N (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) INFO 06-12 05:46:11 [loggers.py:259] Engine 000: Avg prompt throughput: 330.2 tokens/s, Avg generation throughput: 306.1 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.8%, Prefix cache hit rate: 93.2% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 30e sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x40 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401788407380 recvbuff 0x401777004200 count 48236608 datatype 9 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbu (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 0e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401848407380 recvbuff 0x401837004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nran (FSDPPolicyWorkerBase pid=487912) CCL INFO AllReduce: opCount 39f sendbuff 0x401c71800000 recvbuff 0x401c71800000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 3b0 s (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401740000000 recvbuff 0x401b05200000 count 48236608 da (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbu (FSDPPolicyWorkerBase pid=487914) recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40179cc05280 recvbuff 0x401797004200 count 48236608 datatype 9 o (FSDPPolicyWorkerBase pid=487915) ff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017a2806300 recvbuff 0x401797004200 count 48236608 datatype 9 op 0 ro (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount 384 se (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771 [repeated 3x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 21x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) endbuff 0x401aee600000 recvbuff 0x401aee600000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) tatype 7 op 4 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df13e9b50 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INF (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) ff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) j (FSDPPolicyWorkerBase pid=487913) op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 4823660 [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) ks=4] stream 0x400e192bd380 [repeated 6x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 [repeated 966x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 30e sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) [repeated 53x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 55x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) O AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401740000000 recvbuff 0x401e25e00000 count 48236608 datatype 7 op 4 root 0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 (FSDPPolicyWorkerBase pid=487914) p 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: (FSDPPolicyWorkerBase pid=487915) ot 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCoun (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487913) count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 [repeated 4x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487912) sendbuff 0x401a20000000 recvbuff 0x40194a400000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 3b8 sendbuff 0x401b71800000 recvbuff 0x401b71800000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nrank (FSDPPolicyWorkerBase pid=487914) opCount 3af sendbuff 0x401a65200000 recvbuff 0x401a65200000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487915) t 3af sendbuff 0x401a65200000 recvbuff 0x401a65200000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [n (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4018a2000000 recvbuff 0x401c11800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 [repeated 469x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 3bc sendbuff 0x401c11800000 recvbuff 0x401c11800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 [repeated 467x across cluster] (FSDPPolicyWorkerBase pid=487912) 8 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 [repeated 4x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 3c0 sendbuff 0x401c71800000 recvbuff 0x401c71800000 count 48236608 dataty (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401785c01080 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 00 recvbuff 0x401c85e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4014e20000 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x40192a400000 count 48236608 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) ] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:33 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil (FSDPPolicyWorkerBase pid=487913) pe 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (FSDPPolicyWorkerBase pid=487915) jpbo (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount 39f sendbuff 0x401c85e000 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opC (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) ) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] st (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) ount 0 sendbuff 0x401822806300 recvbuff 0x401817004200 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401420000000 recvbuff 0x401b00000000 count 4 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 3d5 sendbuff 0x401c80000000 recvbuff 0x401c80000000 count 48236608 datatype 7 op 4 roo (FSDPPolicyWorkerBase pid=487914) jp (FSDPPolicyWorkerBase pid=487915) jpbo-006 (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] [05:46:14] OK: 353 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] [05:46:14] OK: RSS 2.83 GiB | node mem 390.3/858.0 GiB used (45.5%), avail 467.6 GiB (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) f12be390 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 00e152ce4c0 (FSDPPolicyWorkerBase pid=487913) 324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401ae0000000 recvbuff 0x40166ba00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) t 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: op (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-00 (FSDPPolicyWorkerBase pid=487912) ream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4017e0000000 recvbuff 0x401ae9200000 cou (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [ (FSDPPolicyWorkerBase pid=487914) nt 48236608 datatype 7 op 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:48926 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401d00000000 recvbuff 0x401da0000000 count 155583488 datatype (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4018a2000000 recvbuff 0x401c1180 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) ] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) Count 0 sendbuff 0x401760000000 recvbuff 0x401d91800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c (FSDPPolicyWorkerBase pid=487913) 0x400e193f12b0 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 3 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 re (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount 3ef se (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather: o (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGather (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) cvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 3f8 sendbuff 0x4019aa600000 recvbuff 0x4019aa600000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [n (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) INFO 06-12 05:46:16 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 252.3 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 18.5%, Prefix cache hit rate: 93.2% [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) pCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401862000000 recvbuff 0x401d40000000 count 48236608 datatype 7 op 4 root 0 comm 0x400df16b (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) : opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9c80 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount 3df sendbuff 0x401ba0000000 recvbuff 0x [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487914) ndbuff 0x4014a0000000 recvbuff 0x4014a0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] strea (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 90 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 400 sendbuff 0x401b85e00000 recvbuff 0x401b85e00000 count 48236608 da (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount 414 sen (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nran (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4014e20 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487915) m 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 40a sendbuff 0x401c31800000 recvbu (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401925180000  [repeated 4x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 25x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INF (FSDPPolicyWorkerBase pid=487913) tatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO (FSDPPolicyWorkerBase pid=487912) 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401d00000000 recvbuff 0x401da0000000 count 155582464 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] strea (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40177cc05280 recvbuff 0x401777004200 count 482 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487914) stream 0xaaaaef0966d0 [repeated 9x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 155582464 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 [repeated 1363x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 320 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) [repeated 70x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) [repeated 71x across cluster] (FSDPPolicyWorkerBase pid=487912) O AllGather: opCount 0 sendbuff 0x401677004200 recvbuff 0x401677004200 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401ca0000000 count 48236608 datatype 7 op 4 root (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40179cc05280 recvbuff 0x401797004200 count 48236608 datatype 9 op 0 root (FSDPPolicyWorkerBase pid=487913) ReduceScatter: opCount 0 sendbuff 0x401940000000 recvbuff 0x401c91800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 (FSDPPolicyWorkerBase pid=487915) ff 0x401c31800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) dbuff 0x401ee0000000 recvbuff 0x401ee0000000 count 155582464 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 415 sendbuff 0x401490c00000 recvbuff 0x401490c00000 count 48236608 datatype 7 op 4 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) sendbuff 0x401dc0000000 recvbuff 0x401dc0000000 count 155582464 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401d00000000 recvbuff 0x4014b1800000 count 48236608 da (FSDPPolicyWorkerBase pid=487913) opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 [repeated 2x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 00df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount 41f sendbuff 0x401c4e600000 recvbuff 0x401c4e600000 coun (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount 41f sendbuff 0x4019f3000000 recvbuff 0x4019f3000000 c (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:4892 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:33403 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401420000000 recvbuff 0x401b11800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 [repeated 741x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 424 sendbuff 0x401b11800000 recvbuff 0x401b11800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 [repeated 739x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 36608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INF (FSDPPolicyWorkerBase pid=487915) 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllG (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017a2806300 recvbuff 0x401797004200 count 48236608 datatype 9 op 0 root 0 comm [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:46:18 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:46:18 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:46:18 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:46:18 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:46:18 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:46:18 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:46:18 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:46:18 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:46:18 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:46:18 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:46:18 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:46:18 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:46:18 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:46:18 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:46:18 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:46:18 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:46:18 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:46:18 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:46:18 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:46:18 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:46:18 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:46:18 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:46:18 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:46:18 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:46:18 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:46:18 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401ae0000000 recvbuff 0x401ba0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c3c90 [nranks=4] st (FSDPPolicyWorkerBase pid=487913) 0x401771403180 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 [repeated 10x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) t 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) ount 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:33403 (FSDPPolicyWorkerBase pid=487914) O AllGather: opCount 0 sendbuff 0x401785c01080 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401d00000000 recvbuff 0x401c60000000 count 48236608 datatype 7 op 4 root 0 c (FSDPPolicyWorkerBase pid=487915) ather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 430 sendbuff 0x401da0000000 recvbuff 0x401da0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuf (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487914) ) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: o (FSDPPolicyWorkerBase pid=487914) omm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45: (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e88c00 recvbuff 0x400e79e88c00 count 1 datatype 7 op 0 root 0 comm (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 c (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11: (FSDPPolicyWorkerBase pid=487913) f 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401cf81b6000 recvbuff 0x401a70c00000 count 482 (FSDPPolicyWorkerBase pid=487912) pCount 32b sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 440 sendbuff 0x40180a400000 recvbuff 0x40180a400000 count 4823660 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCo (FSDPPolicyWorkerBase pid=487913) ream 0x400e193f12b0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) omm 0x400e152e6370 [nranks=8] stream (nil) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401 (FSDPPolicyWorkerBase pid=487914) ef0966d0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401420000000 recvbuff 0x40146a400000 cou (FSDPPolicyWorkerBase pid=487915) e0 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [ (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL I (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [ (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:4 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915: (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) nt 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 (FSDPPolicyWorkerBase pid=487913) NFO ReduceScatter: opCount 0 sendbuff 0x401ae0000000 recvbuff 0x401bb1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 45b sendbuff 0x401ca5e00000 recvbuff 0x401ca5e00000 count 48236608 datatype 7 op 4 root (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) unt 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401a2b000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e156cd8d0 [nrank (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 3e8890 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4014f033c000 recvbuff 0x401c20000000 count 48236608 datatype 7 [repeated 2x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 44b sendbuff 0x401b [repeated 4x across cluster] (skyrl_entrypoint pid=487747) [fd-monitor] [05:46:20] OK: 564 / 131,072 FDs open (0.4% of soft limit, hard limit: 131,072) (skyrl_entrypoint pid=487747) [fd-monitor] [05:46:20] OK: RSS 14.62 GiB | node mem 441.6/858.0 GiB used (51.5%), avail 416.3 GiB (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) ] stream 0x400df13e9b50 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 rec (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount 414 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: o (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) vbuff 0x401800000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCoun (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opC (FSDPPolicyWorkerBase pid=487912) 400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 336 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 roo (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: op (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) INFO 06-12 05:46:21 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 228.2 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 18.8%, Prefix cache hit rate: 93.2% [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) pCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401862000000 recvbuff 0x401c60000000 count 48236608 datatype 7 op 4 root 0 comm 0x400df16bdc50 [nranks=4 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) : opCount 0 sendbuff 0x401882000000 recvbuff 0x401ca5e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 (FSDPPolicyWorkerBase pid=487912) t 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45: (FSDPPolicyWorkerBase pid=487913) 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 44a sendbuff 0x401b4e600000 recvbuff 0x401b4e600 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllGather: o (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) t fe sendbuff 0x400e79e88c00 recvbuff 0x400e79e88c00 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) ount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatyp (FSDPPolicyWorkerBase pid=487912) m 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff  [repeated 5x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 37x across cluster] (FSDPPolicyWorkerBase pid=487914) pCount 0 sendbuff 0x401785c01080 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401b00000000 recvbuff 0x401b60000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [ (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) Count 0 sendbuff 0x401677004200 recvbuff 0x401677004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:3 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceSca (FSDPPolicyWorkerBase pid=487914) tatype 7 op 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487912) e 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO A (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 4823 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) s=4] stream 0x400e153f9c80 [repeated 8x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 [repeated 1624x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 33b sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) [repeated 106x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) [repeated 104x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO R (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) tter: opCount 0 sendbuff 0x4019d464c000 recvbuff 0x401752a00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 470 sendbuff 0x401cc0000000 recvbuff 0x401cc0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401a40000000 recvbuff 0x401b31800000 count 48236608 dataty (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 (FSDPPolicyWorkerBase pid=487915) 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401a40000000 recvbuff 0x401b20000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 [repeated 713x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 471 sendbuff 0x401b20000000 recvbuff 0x401b20000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 [repeated 720x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 6608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INF (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 4x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401677004200 recvbuff 0x401677004200 count 48236608 datatype 9 op 0 root 0 comm 0x [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:21 [serving.py:315] Error in preprocessing prompt inputs [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:21 [serving.py:315] Traceback (most recent call last): [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:21 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:21 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:21 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:21 [serving.py:315] return await asyncio.gather( [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:21 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:21 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:21 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:21 [serving.py:315] raise VLLMValidationError( [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:21 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) educeScatter: opCount 0 sendbuff 0x401740000000 recvbuff 0x401c71800000 count 48236608 datatype 7 op 4 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df13e9b50 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount 47a sendbuff 0x401d65e00000 recvbuff 0x401d65e00000 count 48236608 datatype 7 op 4 root 0 co (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) O ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401a93000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9c80 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 [repeated 6x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) pe 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO Al (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbu (FSDPPolicyWorkerBase pid=487912) llReduce: opCount 47a sendbuff 0x401bc0000000 recvbuff 0x401bc0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INF (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) mm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e88c00 recvbuff 0x400e79e88c00 count 1 datatype 7 op 0 root 0 comm 0x400df1 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) ff 0x401882000000 recvbuff 0x401dc5e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 data (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401785c01080 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) lGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401677004200 recvbuff 0x401677004200 co (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487913) 0 (FSDPPolicyWorkerBase pid=487912) O AllReduce: opCount 342 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 2d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) unt 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) type 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount 48a sendbuff 0x40172e800000 recvbuff 0x40172e800000 coun (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 48a sendbuff 0x40168ba00000 recvbuff 0x40168ba00000 count 4823 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 00df12be390 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NC (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 493 sendbuff 0x401ace600000 recvbuff 0x401ace600000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] strea (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount 47a sendbuff 0x401be0000000 recvbuff 0x401be0000000 count 48236608 datatype 7 op 4 root 0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11 (FSDPPolicyWorkerBase pid=487914) t 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [ (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCC (FSDPPolicyWorkerBase pid=487913) CL INFO ReduceScatter: opCount 0 sendbuff 0x401946000000 recvbuff 0x401850c00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 49b sendbuff 0x401bc0000000 recvbuff 0x401bc0000000 count 48236608 datatype 7 op 4 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11: (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007- (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 48b sendbuff 0x (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) x400e40fa4c00 (FSDPPolicyWorkerBase pid=487912) m 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4014e2000000 recvbuff (FSDPPolicyWorkerBase pid=487913) ] stream 0x400e193f12b0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 rec (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount 4a5 sendbuff 0x401c80000000 recvbuff 0x401c80000000 count 48236608 datatype 7 (FSDPPolicyWorkerBase pid=487915) L INFO ReduceScatter: opCount 0 sendbuff 0x4018a2000000 recvbuff 0x401b71800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 4a5 sendbuff 0x401c71800000 recvbuff 0x401c71800000 count 48236608 datatype 7 op 4 r (FSDPPolicyWorkerBase pid=487913) root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 34d sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datat (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendb (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 se (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:4 (FSDPPolicyWorkerBase pid=487913) vbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487915) oot 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915: (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401677004200 recvbuff 0x401677004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [n (FSDPPolicyWorkerBase pid=487913) opCount 0 sendbuff 0x401940000000 recvbuff 0x401cd1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401788407380 recvbuff 0x401777004200 count 48236608 datatype 9 op 0 r (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) INFO 06-12 05:46:26 [loggers.py:259] Engine 000: Avg prompt throughput: 630.9 tokens/s, Avg generation throughput: 330.1 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 34.2%, Prefix cache hit rate: 93.2% [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401b00000000 recvbuff 0x401b71800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e616c [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487912) ype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 dat (FSDPPolicyWorkerBase pid=487914) comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 4a6 sendbuff 0x401c40000000 recvbuff (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401ce974c000 recvbuff 0x401752a00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stre (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487912) atype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL IN (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487914) ] stream 0x400e613f6c90 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401785c01080 rec (FSDPPolicyWorkerBase pid=487915) am 0x400e193ee7c0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 [repeated 5x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 13x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) uff 0x401825181000 recvbuff 0x401800000000 count 155583488 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) ndbuff 0x4017728c0800 recvbuff 0x401760000000 count 155583488 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatte (FSDPPolicyWorkerBase pid=487912) FO AllReduce: opCount 4ba sendbuff 0x401ae0000000 recvbuff 0x401ae0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 co (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 4b0 sendbuff 0x4017e5200000 recvbuff 0x4017e5200000 count 48236608 datatype 7 op 4 root 0 comm 0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 (FSDPPolicyWorkerBase pid=487915) 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 [repeated 1182x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) [repeated 70x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 70x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401aa0000000 recvbuff 0x401551800000 count 48236608 da (FSDPPolicyWorkerBase pid=487914) vbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487913) f12b0 (FSDPPolicyWorkerBase pid=487915) jpbo (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) r: opCount 0 sendbuff 0x4017c0000000 recvbuff 0x4014a0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9c80 (FSDPPolicyWorkerBase pid=487915) 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 [repeated 6x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) tatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 se (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401cd1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 [repeated 608x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 4bd sendbuff 0x401cd1800000 recvbuff 0x401cd1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 [repeated 606x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487914) op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm (FSDPPolicyWorkerBase pid=487913) recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 [repeated 6x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) O AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401420000000 recvbuff 0x401c51800000 count 48236608 datatype 7 op 4 root 0 c (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 4cb sendbuff 0x401de0000000 recvbuff 0x401de0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks= (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INF (FSDPPolicyWorkerBase pid=487912) mm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatyp (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 data (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) omm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllG (FSDPPolicyWorkerBase pid=487914) 0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017a2806300 r (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) e 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 4 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) e152e6370 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) type 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 coun (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stre (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) ather: opCount 0 sendbuff 0x401848407380 recvbuff 0x401837004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-00 (FSDPPolicyWorkerBase pid=487912) a4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 4d3 sendbuff 0x401060000000 recvbuff 0x401060000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] s (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:3340 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO R (FSDPPolicyWorkerBase pid=487915) ecvbuff 0x401797004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceS (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) t 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [ (FSDPPolicyWorkerBase pid=487913) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 4db sendbuff 0x4018ca400000 recvbuff 0x4018ca400000 count 48236608 datatype 7 o (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 4e6 sendbuff 0x401b60000000 recvb (FSDPPolicyWorkerBase pid=487912) tream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401ca0000000 recv (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype 9 op (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] (FSDPPolicyWorkerBase pid=487914) educeScatter: opCount 0 sendbuff 0x4017e0000000 recvbuff 0x4017d0200000 count 48236608 datatype 7 op 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount 4e5 sendbuff 0x401ba0000000 recvbuff 0x401ba0000000 count 48236608 datatype 7 op 4 root 0 co (FSDPPolicyWorkerBase pid=487915) catter: opCount 0 sendbuff 0x401a80000000 recvbuff 0x401772a00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) uff 0x401b60000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 (FSDPPolicyWorkerBase pid=487913) p 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScat (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: (FSDPPolicyWorkerBase pid=487913) oot 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) opCount 0 sendbuff 0x401740000000 recvbuff 0x401490a00000 count 48236608 datatype 7 op 4 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df13e9b50 (FSDPPolicyWorkerBase pid=487912) buff 0x401af1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=487912) j (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) INFO 06-12 05:46:31 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 462.4 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 31.2%, Prefix cache hit rate: 93.2% [repeated 49x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401b00000000 recvbuff 0x401772a00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e616ca9c0 [nranks=4 (FSDPPolicyWorkerBase pid=487914) mm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401797a40000 recvbuff 0x4017600000 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 send (FSDPPolicyWorkerBase pid=487913) ter: opCount 0 sendbuff 0x401ae0000000 recvbuff 0x401be5e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: o (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=487914) 0 sendbuff 0x4017e0000000 recvbuff 0x401ca5e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40179cc05280 recvbuff 0x401797004200 count 48236608 datatype 9 op 0 root 0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) am 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000  [repeated 2x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 15x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) ndbuff 0x401882000000 recvbuff 0x401ce0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 (FSDPPolicyWorkerBase pid=487915) buff 0x401960000000 recvbuff 0x401c91800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017a2806300 recvbuff 0x401797004200 count 48236608 datatype 9 op 0 root 0 comm 0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCoun (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 0 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount 4fa sendbuff 0x4016f0000000 recvbuff 0x4016f0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks=2] [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) e4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount 4fa sendbuff 0x4014b1800000 recvbuff 0x4014b1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks (FSDPPolicyWorkerBase pid=487913) e195a95c0 [nranks=2] stream 0x400e193f4220 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 [repeated 998x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) [repeated 72x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 72x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487912) pCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCC (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401ea19fea00 re (FSDPPolicyWorkerBase pid=487914) comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x40142bb82000 recvbuff 0x401b71800000 count 48236608 datatype 7 op 4 root (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 50b sendbuff 0x401cf1800000 recvbuff 0x401cf1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nra (FSDPPolicyWorkerBase pid=487913) t 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) cvbuff 0x401ca5e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df13e9b50 (FSDPPolicyWorkerBase pid=487912) L INFO AllReduce: opCount 4fa sendbuff 0x40106ba00000 recvbuff 0x40106ba00000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCou (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401df1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 [repeated 565x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 518 sendbuff 0x401df1800000 recvbuff 0x401df1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 [repeated 569x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 6608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45: (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) recvbuff 0x401b20000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9c80 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=487913) e193f12b0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000 [repeated 3x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:46:34 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:46:34 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:46:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:46:34 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:46:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:46:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:46:34 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:46:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:46:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:46:34 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:46:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:46:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:46:34 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:46:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:46:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:46:34 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:46:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:46:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:46:34 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:46:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:46:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:46:34 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:46:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:46:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:46:34 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) ERROR 06-12 05:46:34 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCo (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 dataty (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487913) jpbo-00 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487912) unt 515 sendbuff 0x401c00000000 recvbuff 0x401c00000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 c (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 send (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) bf0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 s (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:48 (FSDPPolicyWorkerBase pid=487915) pe 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:4 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) nt 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 roo (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count  [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487912) ount 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 525 sendbuff 0x4017ca400000 recvbuff 0x4017ca400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [n (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) endbuff 0x40177cc05280 recvbuff 0x401777004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL IN (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO Reduce (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) t 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 526 sendbuff 0x401913c00000 r (FSDPPolicyWorkerBase pid=487913) Scatter: opCount 0 sendbuff 0x401940000000 recvbuff 0x401a45200000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) FO AllReduce: opCount 521 sendbuff 0x40170e800000 recvbuff 0x40170e800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 co (FSDPPolicyWorkerBase pid=487912) e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4014e2000000 recvbuff 0x401c1180000 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) AllReduce: opCount 521 sendbuff 0x401732a00000 recvbuff 0x401732a00000 count 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) ecvbuff 0x401913c00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) INFO 06-12 05:46:36 [loggers.py:259] Engine 000: Avg prompt throughput: 201.2 tokens/s, Avg generation throughput: 466.0 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 34.0%, Prefix cache hit rate: 93.2% [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount 53c (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) mm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuff 0x401760 [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 541 sendbuff 0x401c80000000 recvbuff 0x401c80000000 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCoun (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 0 sendbuff 0x401882000000 recvbuff 0x4018f1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 26x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) buff 0x401822806300 recvbuff 0x401817004200 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype 9 op 0 root 0 comm  [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) sendbuff 0x401db1800000 recvbuff 0x401db1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401822806300 recvbuff 0x401817004200 count 48236608 datatype (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40177cc05280 recvbuff 0x401777004200 count 48236608 datat (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [ (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) j (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 536 sendbuff 0x401c00000000 recvbuff 0x401c00000000 count 48236608 datatype 7 op 4 root 0 comm 0x [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487913) 400e195a95c0 [nranks=2] stream 0x400e193f4220 [repeated 9x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 [repeated 1074x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 371 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 63x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) [repeated 54x across cluster] (FSDPPolicyWorkerBase pid=487914) 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401785c01 (FSDPPolicyWorkerBase pid=487915) 193f14b0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 re (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487913) t 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nrank (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) ype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=487912) 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 r (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401a52606000 recvbuff 0x40196a400000 count 48236608 datatype 7 op 4 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 (FSDPPolicyWorkerBase pid=487915) cvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NC (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401b00000000 recvbuff 0x40182a400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 [repeated 481x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 54b sendbuff 0x4017ca400000 recvbuff 0x4017ca400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 [repeated 475x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487914) 080 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-1 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:3 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x40188 [repeated 4x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: o (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:46:34 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:46:34 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:46:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:46:34 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:46:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:46:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:46:34 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:46:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:46:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:46:34 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:46:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:46:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:46:34 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:46:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:46:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:46:34 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:46:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:46:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:46:34 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:46:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:46:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:46:34 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:46:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:46:34 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:46:34 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (FSDPPolicyWorkerBase pid=487912) oot 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: (FSDPPolicyWorkerBase pid=487915) CL INFO AllReduce: opCount 54c sendbuff 0x40182a400000 recvbuff 0x40182a400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401877036000 recvbuff 0x401dc5e00000 count 48236608 (FSDPPolicyWorkerBase pid=487913) jpbo-006 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) pCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4019e0000000 recvbuff 0x401c91800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a (FSDPPolicyWorkerBase pid=487912) opCount 555 sendbuff 0x401b11800000 recvbuff 0x401b11800000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduc (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opC (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 c (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 0e193eb800 (FSDPPolicyWorkerBase pid=487914) e: opCount 567 sendbuff 0x401cd1800000 recvbuff 0x401cd1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401785c01080 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 c (FSDPPolicyWorkerBase pid=487915) ount 567 sendbuff 0x401cc0000000 recvbuff 0x401cc0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 (FSDPPolicyWorkerBase pid=487913) omm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] N (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 4c0 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401862000000 recvbuff 0x4017b [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) (FSDPPolicyWorkerBase pid=487914) omm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401b00000000 re (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:33 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487913) CCL INFO AllGather: opCount 0 sendbuff 0x401788407380 recvbuff 0x401777004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 576 sendbuff 0x401acc000000 recvbuff 0x401acc000000 count 48236608 datatype 7 op 4 root 0 com (FSDPPolicyWorkerBase pid=487912) 7c98e0 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount 57c sendbuff 0x40 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) INFO 06-12 05:46:41 [loggers.py:259] Engine 000: Avg prompt throughput: 199.2 tokens/s, Avg generation throughput: 474.6 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 35.5%, Prefix cache hit rate: 93.3% [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 581 sendbuff 0x401b91800000 recvbuff 0x401b9180 (FSDPPolicyWorkerBase pid=487915) jpbo- (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount (FSDPPolicyWorkerBase pid=487913) m 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuff 0x40 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401785c01080 recvbuff 0x401780000000 count 48236608 datatype 9 op (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11: (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO All (FSDPPolicyWorkerBase pid=487914) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 25x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=487913) sendbuff 0x401951b40000 recvbuff 0x401c20000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 00df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendb (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 se (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401677004200 recvbuff 0x401677004200 count 48 (FSDPPolicyWorkerBase pid=487914) 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 54b sendbuff 0x401aee000000 recvbuff 0x401aee000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 8d0 [nranks=4] stream 0x400e152ce4c0 [repeated 8x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 [repeated 1180x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 383 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 57x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) [repeated 72x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4018 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) Reduce: opCount 58d sendbuff 0x401e25e00000 recvbuff 0x401e25e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401677004200 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487915) 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) ndbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:33403 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 0000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 [repeated 6x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) uff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [ (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGathe (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:48 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487914) cvbuff 0x401870c00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401ec0000000 recvbuff 0x401de0000000 count 155582464 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 [repeated 583x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 590 sendbuff 0x401de0000000 recvbuff 0x401de0000000 count 155582464 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 [repeated 589x across cluster] (FSDPPolicyWorkerBase pid=487915) 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) x401b40000000 recvbuff 0x401b40000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4018669f2000 recvbuff 0x401ce0000000 count 48236608 datatype 7 op (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401aa0000000 recvbuff 0x401b51800000 count 48236608 datatype 7 (FSDPPolicyWorkerBase pid=487913) Reduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401788407380 recv (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) r: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4019e0000000 recvbuff 0x401ba5e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGat (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllRedu (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401960000000 recvbuff 0x401cd1800000 count [repeated 2x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 5b6 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 0x400e193eb800 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487913) buff 0x401777004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401ebba5cc00 recvbuff  [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=487912) ce: opCount 595 sendbuff 0x40146a400000 recvbuff 0x40146a400000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount 57c sendbuff 0 (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) INFO 06-12 05:46:46 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 390.9 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 35.2%, Prefix cache hit rate: 93.3% [repeated 47x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 5a7 sendb (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO All (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 27x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df [repeated 3x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count 4 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 4 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df13e9b50 (FSDPPolicyWorkerBase pid=487915) 196c2490 [nranks=4] stream 0xaaaabca967e0 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 [repeated 486x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) [repeated 21x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 18x across cluster] (FSDPPolicyWorkerBase pid=487914) count 48236608 datatype 7 op 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:46:48 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:46:48 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:46:48 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:46:48 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:46:48 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:46:48 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:46:48 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:46:48 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:46:48 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:46:48 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:46:48 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:46:48 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:46:48 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:46:48 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:46:48 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:46:48 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:46:48 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:46:48 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:46:48 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:46:48 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:46:48 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:46:48 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:46:48 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:46:48 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:46:48 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:46:48 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 5b0 sendbu (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) op 4 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9c80 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4014e2000000 recvbuff 0x401c31800000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 [repeated 209x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 5af sendbuff 0x401c20000000 recvbuff 0x401c20000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 [repeated 208x across cluster] (FSDPPolicyWorkerBase pid=487912) recvbuff 0x401677004200 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487914) sendbuff 0x401be5e00000 recvbuff 0x401be5e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401d60000000 recvbuff 0x401e00000000 count 155583488 datatype 7 op 4 root 0 comm 0x400e616ca9c0 [ (FSDPPolicyWorkerBase pid=487915) uff 0x401bd1800000 recvbuff 0x401bd1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401d40000000 recvbuff 0x401de0000000 count 155583488 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487913) 8236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) : opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) her: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 dataty (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401677004200 recvbuff 0x401677004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks= (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 0x401e40000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 (FSDPPolicyWorkerBase pid=487912) ff 0x401c31800000 recvbuff 0x401c31800000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401677004200 recvbuff 0x401677004200 count 48236608 datatype 9 op 0 ro (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount 5a7 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40179cc052 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017a2806300 rec (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) pe 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount 5bc sendbuff 0x4017aa400000 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount 5bc sendbuff 0x4017d1800000 rec (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 92bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datat (FSDPPolicyWorkerBase pid=487913) op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 5c1 sendbuff 0x401a20c00000 recvbuff 0x401a (FSDPPolicyWorkerBase pid=487912) ot 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) vbuff 0x4017d1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jp (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) ype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:33 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915: (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:4 (FSDPPolicyWorkerBase pid=487913) comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCou (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount 5d7 sendbuff 0x401df1800000 recvbuff 0x401df1800000 c (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) AllReduce: opCount 5cd sendbuff 0x401d40000000 recvbuff 0x401d40000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401b00000000 recvbuff 0x401f00000000 count 155583488 datatype 7 op 4 root (FSDPPolicyWorkerBase pid=487912) jpbo-006-45 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4018a2000000 recvbuff 0x401be5e00000 count 48236608 datat (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x40142f638000 recvbuff 0x401 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401b00000000 recvbuff 0x401c00000000 count 48236608 (FSDPPolicyWorkerBase pid=487913) nt 0 sendbuff 0x401ae0000000 recvbuff 0x401b40000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 5d1 sendbuff 0x401c31800000 recvbuff 0x401c31800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [ (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) INFO 06-12 05:46:51 [loggers.py:259] Engine 000: Avg prompt throughput: 25.5 tokens/s, Avg generation throughput: 348.9 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 32.2%, Prefix cache hit rate: 93.3% [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCC (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [ (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 38e sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 o (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:48 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) ount 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40179cc05280 recvbuff 0x401797004200 count 48236608 datatype 9 op 0 root 0 co (FSDPPolicyWorkerBase pid=487915) L INFO AllReduce: opCount 5dc sendbuff 0x401de0000000 recvbuff 0x401de0000000 count 155582464 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) p 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 15x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017a2806300 recvbuff 0x401797004200 count 48236608 datatype 9 op 0 root 0 comm 0x4 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401aa0000000 recvbuff 0x40180a400000 count 48236608 datatype 7 op 4 root 0 c (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x4 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401862000000 recvbuff 0x4018c5e00000 count 48236608 datatype 7 op 4 root 0 comm (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) opCount 0 sendbuff 0x401848407380 recvbuff 0x401837004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 5e6 sendbuff 0x401c45e00000 recvbuff 0x401c45e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] str (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) L INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=487913) nranks=2] stream 0x400e193f4220 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 [repeated 1001x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 395 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 70x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 72x across cluster] (FSDPPolicyWorkerBase pid=487914) mm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000 cou (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) omm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9c80 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGather: opCount 0 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather: opCount 0 se (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 5f0 se (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendb (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) eam 0x400e193eb800 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbu (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4018a2000000 recvbuff 0x401be5e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 [repeated 525x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 5f4 sendbuff 0x401be5e00000 recvbuff 0x401be5e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 [repeated 525x across cluster] (FSDPPolicyWorkerBase pid=487915) 0000 recvbuff 0x4019b2800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=487913) nt 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401d7d2fec00 recvbuff 0x401c71800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e156cd8d0 [nr (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487912) ndbuff 0x401b51800000 recvbuff 0x401b51800000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401a9bac9a00 recvbuff 0x401e00000000 count 48236608 datatype 7 op 4 root 0 comm 0x400df16bdc50 [nrank (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) ff 0x401d51800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 (FSDPPolicyWorkerBase pid=487912) 08 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487914) aaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88 (FSDPPolicyWorkerBase pid=487915) 967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 co (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) uff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] N (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0 count 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbu (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sen (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 8890 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401848407380 recvbuff 0x401837004200 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) CCL INFO AllReduce: opCount 39a sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbu (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 5e7 sendbuff 0x4019b280 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487914) c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:4879 (FSDPPolicyWorkerBase pid=487915) unt 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) dbuff 0x40177cc05280 recvbuff 0x401777004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:3 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount 5d7 sendbuff 0x401c60000000 recvbuff 0x401c6000000 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) ff 0x401822806300 recvbuff 0x401817004200 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:3328 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: o (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-1 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL (FSDPPolicyWorkerBase pid=487912) jpbo-00 (FSDPPolicyWorkerBase pid=487915) vbuff 0x401797004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) ype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount 617 sendbuff 0x401d05e00000 recvbuff 0x401d05e00000 count 48236608 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4017e0000000 recvbuff 0x401a2f000000 count 48236608 datatype 7 op (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401b06000000 recvbuff 0x401a23400000 count 48236608 datatype 7 op 4 roo (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) INFO AllReduce: opCount 60d sendbuff 0x401c60000000 recvbuff 0x401c60000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 (FSDPPolicyWorkerBase pid=487913) pCount 0 sendbuff 0x4021c923c000 recvbuff 0x401491800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 611 sendbuff 0x401b51800000 recvbuff 0x401b51800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INF (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4014e2000000 recvbuff 0x401b60000000 count 48 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487914) 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllGather (FSDPPolicyWorkerBase pid=487915) t 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCo (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) INFO 06-12 05:46:56 [loggers.py:259] Engine 000: Avg prompt throughput: 135.3 tokens/s, Avg generation throughput: 390.3 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 34.5%, Prefix cache hit rate: 93.3% [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x40 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCC (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487914) : opCount 0 sendbuff 0x401785c01080 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914 (FSDPPolicyWorkerBase pid=487915) unt 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:48817 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f s (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 20x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401788407380 recvbuff 0x401777004200 count 48236608 datat (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) O ReduceScatter: opCount 0 sendbuff 0x401c0c000000 recvbuff 0x401e11800000 count 48236608 datatype 7 op 4 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df13e9b50 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401822806300 recvbuff 0x401817004200 count 48236608 datatype 9 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401c85e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9c80 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40177cc05280 recvbuff 0x401777004200 count 48236608 datatyp (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 4 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) endbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401848407380 recvbuff 0x401837004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401420000000 recvbuff 0x401a2c600000 count 48236608 datatype 7 op 4 root 0 comm 0x (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 155583488 datatype 9 op 0 root 0 c (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401785c0108 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recv (FSDPPolicyWorkerBase pid=487913) c0 [nranks=2] stream 0x400e193f4220 [repeated 8x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017a2806300 recvbuff 0x401797004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 [repeated 1070x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 3a7 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 69x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 70x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) e 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=487913) ype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) omm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 s (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 re (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 63 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401420000000 recvbuff 0x40156b000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 [repeated 478x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 62f sendbuff 0x40152a400000 recvbuff 0x40152a400000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 [repeated 482x across cluster] (FSDPPolicyWorkerBase pid=487914) jp (FSDPPolicyWorkerBase pid=487915) buff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006 (FSDPPolicyWorkerBase pid=487914) 0 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913 (FSDPPolicyWorkerBase pid=487912) 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) cvbuff 0x401c71800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) ndbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) endbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4019e0000000 recvbuff 0x401be0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nran (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:59 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:59 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:59 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:59 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:59 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:59 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:59 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:59 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:59 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:59 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:59 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:59 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:46:59 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) ff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=487912) 8236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 [repeated 5x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487912) 0 sendbuff 0x40156b000000 recvbuff 0x40156b000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nrank (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0x400e153f9c80 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487913) count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 400df13e9b50 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount 642 sendbuff 0x401c20000000 recvbuff 0x401c20000000 c (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401825180000 recvbuff 0x401 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count 48236608 dataty (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 64b sendbuff 0x401c71800 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:33274 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) (FSDPPolicyWorkerBase pid=487914) ount 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL (FSDPPolicyWorkerBase pid=487913) pe 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 dat (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 642 sendbuff 0x401c11800000 recvbuff 0x401c11800000 count 4 [repeated 4x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) j (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017928c0800 recvbuff 0x401780000000 count 155583488 datatype 9 op 0 root 0 co (FSDPPolicyWorkerBase pid=487915) atype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017a5181000 recvbuff 0x401780000000 count 155583488 datatype 9 op 0 root 0 comm 0x4 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487913) aaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 651 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e19 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 00e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:33304 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401ca0000000 recvbuff 0x40194a400000 coun (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-0 (FSDPPolicyWorkerBase pid=487914) mm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllGather: opCount 0 se (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) INFO 06-12 05:47:01 [loggers.py:259] Engine 000: Avg prompt throughput: 172.5 tokens/s, Avg generation throughput: 485.5 tokens/s, Running: 10 reqs, Waiting: 0 reqs, GPU KV cache usage: 35.7%, Prefix cache hit rate: 93.3% [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=487912) t 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount 657 sendbuff 0x401b05200000 recvbuff 0x401b05200000 count 48236608 datatype 7 op (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount 657 sendbuff 0x40192a400000 recvbuff 0x40192a400000 count 48236608 datatype 7 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 roo (FSDPPolicyWorkerBase pid=487914) ndbuff 0x401785c01080 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4019731ac000 recvbuff 0x401c31800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e616ca9c0 [nrank (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401b00000000 recvbuff 0x401c20000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] s (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 66c sendbuff 0x401c71800000 recvbuff 0x401c71800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stre (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 35x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4014e2000000 recvbuff 0x401c80000000 count 48236608 datatype 7 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatte (FSDPPolicyWorkerBase pid=487914) 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount 0 sendbuff [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceSca (FSDPPolicyWorkerBase pid=487913) 5a95c0 [nranks=2] stream 0x400e193f4220 [repeated 6x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 [repeated 883x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) [repeated 35x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 34x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) t 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: op (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4014364b8000 recvbuff 0x401be0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 [repeated 464x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 669 sendbuff 0x401be0000000 recvbuff 0x401be0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 [repeated 459x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487913) 0x401940000000 recvbuff 0x401b71800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 [repeated 5x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) r: opCount 0 sendbuff 0x401740000000 recvbuff 0x401d31800000 count 48236608 datatype 7 op 4 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df13e9b50 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) tter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401ba0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9c80 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) ] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4014f633c000 recvbuff 0x401bf1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) Count 668 sendbuff 0x401d80000000 recvbuff 0x401d80000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [ (FSDPPolicyWorkerBase pid=487915) tream 0x400e193ee7c0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCC (FSDPPolicyWorkerBase pid=487913) am 0x400e193f4220 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGa (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 401760000000 count 155582464 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 [repeated 2x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuf (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 send (FSDPPolicyWorkerBase pid=487912) op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NC (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:4879 (FSDPPolicyWorkerBase pid=487915) L INFO AllGather: opCount 0 sendbuff 0x4017a5181000 recvbuff 0x401780000000 count 155583488 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489 (FSDPPolicyWorkerBase pid=487913) ther: opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:48 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) f 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) j (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) buff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=487912) CL INFO AllGather: opCount 0 sendbuff 0x401677004200 recvbuff 0x401677004200 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [n (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4019e0000000 recvbuff 0x401b00000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [ (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 682 sendbuff 0x401a70c00000 recvbuff 0x401a70c00000 count 48236608 dataty (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count 48236608 da (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 coun (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 c (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 68b sendbuff 0x401b8 (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:47:05 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:47:05 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:47:05 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:47:05 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:47:05 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:47:05 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:47:05 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:47:05 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:47:05 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:47:05 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:47:05 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:47:05 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:47:05 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:47:05 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:47:05 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:47:05 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:47:05 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:47:05 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:47:05 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:47:05 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:47:05 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:47:05 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:47:05 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:47:05 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:47:05 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:47:05 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount 682 sendbuff 0x401ae9200000 recvbuff 0x401ae9200000 count 48236608  [repeated 2x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) b800 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401abe484a00 recvbuff 0x401d91800 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INF (FSDPPolicyWorkerBase pid=487915) pe 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO Redu (FSDPPolicyWorkerBase pid=487913) tatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL I (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount 672 sendbuff 0x401ca0000000 recvbuff 0x401ca0000000 count 48236608 datatype 7 op 4 root 0 comm 0x40 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) ount 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatyp (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 r (FSDPPolicyWorkerBase pid=487914) O ReduceScatter: opCount 0 sendbuff 0x401b00000000 recvbuff 0x401c45e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401785c01080 recvbuf (FSDPPolicyWorkerBase pid=487915) ceScatter: opCount 0 sendbuff 0x4018a2000000 recvbuff 0x401c31800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x40 (FSDPPolicyWorkerBase pid=487913) NFO AllReduce: opCount 693 sendbuff 0x401c80000000 recvbuff 0x401c80000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype 9 (FSDPPolicyWorkerBase pid=487912) 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401677004200 recvb (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) oot 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) e 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 69c sendbuff 0x401311800000 recvbuff 0x401311800000 c (FSDPPolicyWorkerBase pid=487914) f 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) INFO 06-12 05:47:06 [loggers.py:259] Engine 000: Avg prompt throughput: 69.0 tokens/s, Avg generation throughput: 493.3 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 40.1%, Prefix cache hit rate: 93.3% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendb (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) t 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 s (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatter: opCount (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:33403 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce (FSDPPolicyWorkerBase pid=487912) uff 0x401677004200 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:48 (FSDPPolicyWorkerBase pid=487913) uff 0x401ae6000000 recvbuff 0x4017aa400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) endbuff 0x401740000000 recvbuff 0x401c4e600000 count 48236608 datatype 7 op 4 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df13e9b50 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount 6b2 sendbuff 0x401d40000000 recvbuff 0x401d40000000 count 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount 6b2 sendbuff 0x401bb1800000 recvbuff 0x401bb1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nra (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 34x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401420000000 recvbuff 0x401ba0000000 count 48236608 datatyp (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 6b7 sendbuff 0x401c05e00000 recvbuff 0x401c05e00000 count 48236608 datatype (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) : opCount 6a8 sendbuff 0x401c91800000 recvbuff 0x401c91800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487914) 400e613f6c90 (FSDPPolicyWorkerBase pid=487915) 3ee7c0 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401ae0000000 rec (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b80210 [repeated 5x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487913) stream 0x400e193f4220 [repeated 6x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 155582464 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 [repeated 1228x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 3c2 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) [repeated 70x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) [repeated 72x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-0 (FSDPPolicyWorkerBase pid=487913) vbuff 0x401c91800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 79e88c00 recvbuff 0x400e79e88c00 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jp (FSDPPolicyWorkerBase pid=487912) INFO AllReduce: opCount 6c0 sendbuff 0x401da0000000 recvbuff 0x401da0000000 count 155582464 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401677004200 recvbuff 0x401677004200 count 48236608 datatype 9 op 0 root 0 c (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount 6c2 sendbuff 0x4014b1800000 recvbuff 0x4014b1800000 count 48236608 datatype 7 op (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 0e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401a40000000 recvbuff 0x40146a400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f6 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count 4823660 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 6c2 sendbuff 0x4014b1800000 recvbuff 0x4014b1800000 count 48236608 datatype 7 op 4 root (FSDPPolicyWorkerBase pid=487912) omm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 6cb sendbuff 0x4 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401420000000 recvbuff 0x401a12000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 [repeated 706x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 6ca sendbuff 0x4019aa600000 recvbuff 0x4019aa600000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 [repeated 705x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487915) 0 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 [repeated 6x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 dat (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 193eb800 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401ca (FSDPPolicyWorkerBase pid=487914) 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatte (FSDPPolicyWorkerBase pid=487915) 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opC (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) atype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL IN (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 [repeated 3x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:3 (FSDPPolicyWorkerBase pid=487914) r: opCount 0 sendbuff 0x401b050e0000 recvbuff 0x401b60000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (FSDPPolicyWorkerBase pid=487913) CL INFO AllReduce: opCount 6d3 sendbuff 0x401ba0000000 recvbuff 0x401ba0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root (FSDPPolicyWorkerBase pid=487915) ount 0 sendbuff 0x401ce0000000 recvbuff 0x401b4e600000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 5e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 [repeated 6x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) FO AllReduce: opCount 6d9 sendbuff 0x401d51800000 recvbuff 0x401d51800000 count 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e88c00 recvbuff 0x400e79e88c00 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] st (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NC (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] (FSDPPolicyWorkerBase pid=487912) 400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 3cd sendbuff 0x400e9fa88c00 recvbuff 0x400e9f (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) x400e427c7f40 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGath (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sen (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913: (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NC (FSDPPolicyWorkerBase pid=487912) a88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-0 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487914) j (FSDPPolicyWorkerBase pid=487915) jpbo-00 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) ream (nil) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) stream (nil) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendb (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487915) CL INFO AllGather: opCount 0 sendbuff 0x4017a2806300 recvbuff 0x401797004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401a80000000 recvbuff 0x401a72000000 count 48236608 dat (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) er: opCount 0 sendbuff 0x401677004200 recvbuff 0x401677004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:3 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) dbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllRe (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0 sendbuff 0x401882000000 recvbuff 0x4019f3000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9c80 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401ae0000000 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO (FSDPPolicyWorkerBase pid=487914) t 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:4892 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401785c01080 recvbuff 0x401780000000 coun (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) duce: opCount 6e8 sendbuff 0x401670c00000 recvbuff 0x401670c00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 6f7 sendbuff 0x401b20000000 recvbuff 0x401b20000000 count 48236608 data (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 00df13eca20 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401740000000 recvbuff 0x40 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0x400e153fcbf0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401d746bc000 recvbuff 0 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 703 send (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 6dd sendbuff 0x401c45e00000 recvbuff 0x401c45e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) ount 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) type 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO Re (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487912) AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) duceScatter: opCount 0 sendbuff 0x4014fb1f0000 recvbuff 0x401c31800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) buff 0x401db1800000 recvbuff 0x401db1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype (FSDPPolicyWorkerBase pid=487913) 220 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401822806300 recvbuff 0x401817004200 count 48236608 datatype 9 o (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40177cc05280 recvbuff 0x401777004200 count 48236608 datatype (FSDPPolicyWorkerBase pid=487912) e 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) INFO 06-12 05:47:11 [loggers.py:259] Engine 000: Avg prompt throughput: 133.7 tokens/s, Avg generation throughput: 530.5 tokens/s, Running: 14 reqs, Waiting: 0 reqs, GPU KV cache usage: 51.6%, Prefix cache hit rate: 93.4% [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) p 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduc (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007- (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 s (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuf (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4019ed470000 recvbuff 0x401b31800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stre (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) opCount 719 sendbuff 0x401c71800000 recvbuff 0x401c71800000 count 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) e: opCount 719 sendbuff 0x401a93000000 recvbuff 0x401a93000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401ebe009200 recvbuff 0x401dc5e00000 count 4823 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) uff 0x40177cc05280 recvbuff 0x401777004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 6ec sendbuff 0x40180a400000 recvbuff 0x40180a400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) am 0x400e427c7f40 (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 40x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] strea (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487915) f 0x4017e0000000 recvbuff 0x40168ba00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 71d sendbuff 0x401b60000000 recvbuff 0x401b60000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] st (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e15 [repeated 6x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487914) endbuff 0x401b06000000 recvbuff 0x40172e800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount 71d sendbuff 0x401b71800000 recvbuff 0x401b71800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks (FSDPPolicyWorkerBase pid=487913) ] NCCL INFO AllReduce: opCount 713 sendbuff 0x401845200000 recvbuff 0x401845200000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 (FSDPPolicyWorkerBase pid=487912) m 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 726 sendbuff 0x401bc0000000 recvbu (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) e196bcf10 [nranks=4] stream 0x400e192bd380 [repeated 11x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 [repeated 1758x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 3dd sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) [repeated 107x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) [repeated 102x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487915) ream 0x400e193f14b0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4018a2000000 recvb (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) INFO AllReduce: opCount 6d9 sendbuff 0x401bc5e00000 recvbuff 0x401bc5e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e88c00 recvbuff 0x400e79e88c00 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream ( (FSDPPolicyWorkerBase pid=487912) ff 0x401bc0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 3e4 sendbuff 0x400e9fa88c00 rec (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 4823660 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] st (FSDPPolicyWorkerBase pid=487915) uff 0x401c71800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401d40000000 recvbuff 0x401de0000000 count 155582464 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 [repeated 862x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 732 sendbuff 0x401de0000000 recvbuff 0x401de0000000 count 155582464 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 [repeated 862x across cluster] (FSDPPolicyWorkerBase pid=487914) recvbuff 0x401c80000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 [repeated 7x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487913) root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487912) atype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 6608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) ) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jp (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sen (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) nil) (FSDPPolicyWorkerBase pid=487912) vbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401785c01080 recvbuff 0x401780000000 count 48236608 dat (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 737 sendbuff 0x401540000000 recvbuff 0x401540000000 count 48236608 (FSDPPolicyWorkerBase pid=487912) dbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487913) op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401aebd2 [repeated 2x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL (FSDPPolicyWorkerBase pid=487914) atype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL IN (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO All (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401740000000 recvbuff 0x401c80000000 cou (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 743 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401aee600000  [repeated 2x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INF (FSDPPolicyWorkerBase pid=487912) INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4014e2000000 recvbuff 0x401bd1800000 count 48236608 datatype 7 op 4 root (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487914) FO AllReduce: opCount 744 sendbuff 0x401b85e00000 recvbuff 0x401b85e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401785c01080 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 co (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff  [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=487915) Reduce: opCount 744 sendbuff 0x401b71800000 recvbuff 0x401b71800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) nt 48236608 datatype 7 op 4 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df13e9b50 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:33405 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype  [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:33 (FSDPPolicyWorkerBase pid=487913) 93f4220 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) sendbuff 0x401cd1800000 recvbuff 0x401cd1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks= (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) INFO 06-12 05:47:16 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 601.3 tokens/s, Running: 12 reqs, Waiting: 0 reqs, GPU KV cache usage: 44.8%, Prefix cache hit rate: 93.4% [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) O ReduceScatter: opCount 0 sendbuff 0x40198f430000 recvbuff 0x401b45e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 752 sendbuff 0x401c40000000 recvbuff 0x401c40000000 count 48236608 datatype 7 op 4 root 0 (FSDPPolicyWorkerBase pid=487913) opCount 72e sendbuff 0x401cc0000000 recvbuff 0x401cc0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 30x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487912) 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 3ef sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] str (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:33 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x4 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=487914) mm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [ (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCC (FSDPPolicyWorkerBase pid=487912) eam (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] s (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount 759 se (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401435b40000 recvbuff 0x401551800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) O AllGather: opCount 0 sendbuff 0x401848407380 recvbuff 0x401837004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo- (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4018 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487915) L INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) 00e196c2490 [nranks=4] stream 0xaaaabca967e0 [repeated 9x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 [repeated 1008x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) [repeated 70x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 70x across cluster] (FSDPPolicyWorkerBase pid=487913) 324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count 48236608 datatype 9 o (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) ndbuff 0x401490a00000 recvbuff 0x401490a00000 count 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) ream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvb (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401ce0000000 count (FSDPPolicyWorkerBase pid=487912) tream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 766 sendbuff 0x401ae0000000 re (FSDPPolicyWorkerBase pid=487914) 00e613f9c00 (FSDPPolicyWorkerBase pid=487915) f14b0 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount 774 sendbuff 0x401d91800000 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 1555834 (FSDPPolicyWorkerBase pid=487913) p 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllRed (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) uff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401434220000 recvbuff 0x401c51800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 [repeated 486x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 778 sendbuff 0x401c40000000 recvbuff 0x401c40000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 [repeated 481x across cluster] (FSDPPolicyWorkerBase pid=487912) cvbuff 0x401ae0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) j (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) f28c0000 recvbuff 0x4018e0000000 count 155582464 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo-006-4 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:4879 (FSDPPolicyWorkerBase pid=487913) uce: opCount 76e sendbuff 0x401bd1800000 recvbuff 0x401bd1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401822806300 recvbuff 0x401817004200 count 48236608 datatype 9 op 0 root (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40177cc05280 recvbuff 0x401777004200 count 48236608 datatype 9 op 0 ro (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGat (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: (FSDPPolicyWorkerBase pid=487915) 0000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=487912) her: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] N (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCoun (FSDPPolicyWorkerBase pid=487913) opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401862000000 recvbuff 0x4016f0000000 count 48236608 da (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) ot 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=487912) CCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401e33fc7800 recvbuff 0x401af1800000 count 48236608 datatype 7 op 4 r (FSDPPolicyWorkerBase pid=487915) t 784 sendbuff 0x401772a00000 recvbuff 0x401772a00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [n (FSDPPolicyWorkerBase pid=487913) 00e193f4220 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) tatype 7 op 4 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df13e9b50 (FSDPPolicyWorkerBase pid=487912) oot 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: op (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nra (FSDPPolicyWorkerBase pid=487913) jpbo-006- (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) INFO ReduceScatter: opCount 0 sendbuff 0x401aa0000000 recvbuff 0x401565e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 792 sendbuff 0x401b60000000 recvbuff 0x401b60000000 count 48236608 datatype 7 op 4 ro (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487915) 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401862000000 recvbuff 0x401da0000000 count 48236608 datatype 7 op 4 root 0 c (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL  [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: o (FSDPPolicyWorkerBase pid=487912) Count 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) omm 0x400df16bdc50 [nranks=4] stream 0x400df13e9b50 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11: (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007- (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) pCount 0 sendbuff 0x4019e0000000 recvbuff 0x401c65e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount  [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4017c0000000 recvbuff 0x40 [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INF [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401b00000000 recvbuff 0x401830200000 cou (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401c11800000 count 48236608 datatype 7 op 4 root (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401ec22fa800 recvbuff 0x4018f1800000 co (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO Al (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 r (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000  [repeated 5x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:47:21 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:47:21 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:47:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:47:21 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:47:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:47:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:47:21 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:47:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:47:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:47:21 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:47:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:47:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:47:21 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:47:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:47:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:47:21 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:47:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:47:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:47:21 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:47:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:47:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:47:21 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:47:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:47:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:47:21 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:47:21 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:48926 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) f12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount 7b4 sendbuff 0x401ca5e00000 recvbuff 0x401 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487913) lReduce: opCount 7ae sendbuff 0x401a11e00000 recvbuff 0x401a11e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (FSDPPolicyWorkerBase pid=487914) nt 48236608 datatype 7 op 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 00e152ce4c0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) unt 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:3340 (FSDPPolicyWorkerBase pid=487912) 50 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) ecvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 783 sendbuff 0x4018e0000000 recvbuff 0x4018e0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) INFO 06-12 05:47:22 [loggers.py:259] Engine 000: Avg prompt throughput: 42.7 tokens/s, Avg generation throughput: 316.4 tokens/s, Running: 7 reqs, Waiting: 0 reqs, GPU KV cache usage: 29.0%, Prefix cache hit rate: 94.4% [repeated 49x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401b00000000 recvbuff 0x401cc0000000 count 48236608 datatype 7 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCo (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4018a2000000 recvbuff 0x401ca5e00000 count 48236608 datatype 7 op 4 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:3 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-1 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) f40 (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 21x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487914) op 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9c80 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9 [repeated 8x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [ (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11: (FSDPPolicyWorkerBase pid=487913) unt 0 sendbuff 0x401d60000000 recvbuff 0x401e00000000 count 155582464 datatype 7 op 4 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) O AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487913) 400e196c3c90 [nranks=4] stream 0xaaaafb324d40 [repeated 11x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa [repeated 1603x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 40a sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 108x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 108x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401ce0000000 recvbuff 0x4017ca400000 count 482 [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487915) INFO AllGather: opCount 0 sendbuff 0x4017a2806300 recvbuff 0x401797004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] strea (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 7a6 sendbuff 0x40106ba0000 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401b00000000 recvbuff 0x4017ca400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 [repeated 725x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 7d1 sendbuff 0x4017ca400000 recvbuff 0x4017ca400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 [repeated 732x across cluster] (FSDPPolicyWorkerBase pid=487912) 0 recvbuff 0x40106ba00000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 [repeated 7x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487914) ] NCCL INFO AllGather: opCount 0 sendbuff 0x40179cc05280 recvbuff 0x401797004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather: o (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGather (FSDPPolicyWorkerBase pid=487915) 36608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 7d2 sendbuff 0x401913c00000 recvbuff 0x401913c00000 count 48236608 datatype 7 op (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 [repeated 10x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487914) opCount 784 sendbuff 0x4017d0200000 recvbuff 0x4017d0200000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) ot 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 (FSDPPolicyWorkerBase pid=487915) m 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 7df sendbuff 0x401bb1800000 recvbu (FSDPPolicyWorkerBase pid=487913) jpbo-0 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401740000000 recvbuff 0x401cc0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400df16b (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) : opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatte (FSDPPolicyWorkerBase pid=487915) ff 0x401bb1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 (FSDPPolicyWorkerBase pid=487915) jpbo- (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) pCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4021c2a45800 recvbuff (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 7c1 sendbuff 0x401c00000000 recvbuff 0x401c00000 [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) INFO 06-12 05:47:27 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 304.5 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 28.2%, Prefix cache hit rate: 94.4% [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401ec22fa800 recvbuff 0x401e00000000 count 48236608 datatype (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401a3ed98000 recvbuff 0x401300000000 count 48236608 datatype 7 op (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) r: opCount 0 sendbuff 0x401a40000000 recvbuff 0x401b80000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 7ed sendbuff 0x401c80000000 recvbuff 0x401c80000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:4 (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 33x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype 9 op 0 root 0 (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 roo (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 0e192bd380 (FSDPPolicyWorkerBase pid=487914) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40179cc05280 recvbuff 0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: op (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10  [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487913) 0 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount 7f4 sendbuff 0x4017b0a00000 recvbuff 0x4017b0a00000 coun (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) t 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount 7f4 sendbuff 0x401732a00000 recvbuff 0x401732a00000 c (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INF (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff  [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllG (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 984290 [nranks=2] stream 0x400e427d8400 [repeated 8x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 [repeated 761x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) [repeated 35x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 36x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823: (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 com (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) Count 0 sendbuff 0x401677004200 recvbuff 0x401677004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=487912) fa4050 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) t 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [ (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) ount 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:33403 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 7de sendbuff 0x401 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487914) O AllGather: opCount 0 sendbuff 0x401785c01080 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401d06000000 recvbuff 0x401bd1800000 count 48236608 datatype 7 op 4 root 0 c (FSDPPolicyWorkerBase pid=487915) ather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401d20000000 count 48236608 datat (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401d05e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 [repeated 422x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 804 sendbuff 0x401d05e00000 recvbuff 0x401d05e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 [repeated 419x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 27c7f40 (FSDPPolicyWorkerBase pid=487913) m 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 8 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount 80f sendbuff 0x401dc5e00000 recvbuff 0x401dc5e00000 count 48236608 datatype 7 (FSDPPolicyWorkerBase pid=487914) omm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllGather: opCount 0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount 0 sendb (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) ype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO A (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:3 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 0000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=487913) 09 sendbuff 0x401c11800000 recvbuff 0x401c11800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (FSDPPolicyWorkerBase pid=487913) jpbo-006- (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) op 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:3 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-1 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401677004200 recvbuff 0x401677004200 co (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) llGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 data (FSDPPolicyWorkerBase pid=487915) uff 0x4017a5180000 recvbuff 0x401780000000 count 155582464 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017a2806300 recvbuff 0x401797004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [ (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401aa0000000 recvbuff 0x401752a00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e156cd8d0 [nrank (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) type 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) unt 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceSc (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGa (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) ] stream 0x400df13e9b50 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 rec (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401740000000 recvbuff 0x4017c0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400df16bdc50 [nranks=4 [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) atter: opCount 0 sendbuff 0x4019e6000000 recvbuff 0x4019ca600000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 82d sendbuff 0x401b91800000 recvbuff 0x401b91800000 count 48236608 datatype 7 op 4 root 0 comm 0x40 (FSDPPolicyWorkerBase pid=487913) jp (FSDPPolicyWorkerBase pid=487912) : opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) vbuff 0x401800000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 83b s (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 0x400e192bd380 (FSDPPolicyWorkerBase pid=487912) ther: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:4 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915: (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401ae0000000 recvbuff 0x401c20000000 cou (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbu (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount 83a sendbuff 0x401ce0000000 recvbuff 0x401ce0000000 count 48236 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 83a sendbuff 0x401cd1800000 recvbuff 0x401cd1800000 count 48236608 da (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:47:31 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:47:31 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:47:31 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:47:31 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:47:31 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:47:31 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:47:31 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:47:31 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:47:31 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:47:31 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:47:31 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:47:31 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:47:31 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:47:31 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:47:31 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:47:31 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:47:31 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:47:31 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:47:31 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:47:31 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:47:31 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:47:31 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:47:31 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:47:31 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:47:31 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) ERROR 06-12 05:47:31 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (FSDPPolicyWorkerBase pid=487913) nt 48236608 datatype 7 op 4 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 (FSDPPolicyWorkerBase pid=487912) ff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401677004200 recvbuff 0x401677004200 count 48236608 datatype 9 o (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 9c80 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff [repeated 6x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) endbuff 0x401da0000000 recvbuff 0x401da0000000 count 155583488 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401677004200 recvbuff 0x401677004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [n (FSDPPolicyWorkerBase pid=487914) 608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] (FSDPPolicyWorkerBase pid=487915) tatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487913) 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401788407380 recvbuff 0x401777004200 count 48236608 datatype 9 op 0 r (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) INFO 06-12 05:47:31 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 363.4 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 32.7%, Prefix cache hit rate: 93.4% [repeated 47x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO R (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401c2e600000 count 48236608 d (FSDPPolicyWorkerBase pid=487912) p 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (FSDPPolicyWorkerBase pid=487915) INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401b00000000 recvbuff 0x40194a400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [ (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) educeScatter: opCount 0 sendbuff 0x401740000000 recvbuff 0x4017d1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df13e9b50 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount 84f sendbuff 0x401ce0000000 recvbuff 0x401ce0000000 count 48236608 datatype 7 op 4 root 0 co (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x4016 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) O ReduceScatter: opCount 0 sendbuff 0x4019aa612000 recvbuff 0x4017aa400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9c80 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount 84f sendbuff 0x401b51800000 recvbuff 0x401b51800000 count 48236608 datatype 7 op 4 root 0 (FSDPPolicyWorkerBase pid=487913) oot 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCou (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40179cc05280 recvbuff 0x401797004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e616c (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) atatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL IN (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-1 (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INF (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40177 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) ranks=4] stream 0x400e40fa4c00 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 [repeated 1340x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) [repeated 71x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 72x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) mm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatter: opCou (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 862 sendbuff 0x401da0000000 recvbuff 0x401da0000000 cou (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) FO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401e40000000 count 48236608 datatype 7 op 4 root 0 (FSDPPolicyWorkerBase pid=487913) nt 849 sendbuff 0x401b2e800000 recvbuff 0x401b2e800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 81e sendbuff 0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401cb8888000 recvbuff 0x401c31800000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 [repeated 703x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 85b sendbuff 0x401c20000000 recvbuff 0x401c20000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 [repeated 700x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) cc05280 recvbuff 0x401777004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 [repeated 6x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount 861 sendbuff 0x401e00000000 recvbuff 0x401e00000000 count 155583488 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] s (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount 80f sendbuff 0x401c40000000 recvbuff 0x401c40000000 count 48236608 datatype (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream ( (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401822806300 recvbuff 0x401817004200 count 48236608 datatype 9 op 0 root 0 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [ (FSDPPolicyWorkerBase pid=487912) 0146a400000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllG (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 0 sendbuff 0x401740000000 recvbuff 0x401df1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df13e9b50 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487914) tream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 (FSDPPolicyWorkerBase pid=487915) 0x400e193f14b0 (FSDPPolicyWorkerBase pid=487913) nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401771403180 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [ (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) ather: opCount 0 sendbuff 0x401848407380 recvbuff 0x401837004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] str (FSDPPolicyWorkerBase pid=487912) a4050 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 86d sendbuff 0x401a20c00000 recvbuff 0x401a20c00000 count 48236608 datatype 7 op 4 root 0 comm (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401a80000000 recvbuff 0x4014a5200000 count 48236608 datatype 7 op 4 root 0 comm [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCC (FSDPPolicyWorkerBase pid=487913) 0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401ae0000000 recvbuff 0x401b40000000 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11: (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007- (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 s (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) eam 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 879 sendbuff 0x401d40000000 recv (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 se (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount 87a sendbuff 0x401c00000000 recvbuff 0x401c00000000 count 48236608 datatype 7 (FSDPPolicyWorkerBase pid=487915) L INFO ReduceScatter: opCount 0 sendbuff 0x401ce0000000 recvbuff 0x4019b2800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 87a sendbuff 0x401be5e00000 recvbuff 0x401be5e00000 count 48236608 datatype 7 op 4 r (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:4 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 85c sendbuff 0x401c31800000 recvbuff 0x401c31800000 count 4823 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) buff 0x401d40000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 (FSDPPolicyWorkerBase pid=487912) ndbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401e333e5c00 recvbuff 0x401c45e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nrank (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:33 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:33 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:33 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:33 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:33 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:33 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:33 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:33 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:33 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:33 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:33 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:33 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:33 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) nt 155582464 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 841 sendbuff 0x40146a400000 recvbuff 0x4 (FSDPPolicyWorkerBase pid=487914) op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INF (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendb (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllR (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) endbuff 0x4019e0000000 recvbuff 0x401bc0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 430 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (ni (FSDPPolicyWorkerBase pid=487912) 6608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:47:37 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 281.9 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 30.6%, Prefix cache hit rate: 94.7% [repeated 51x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:33 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: o (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 90 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) l) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 25x across cluster] (FSDPPolicyWorkerBase pid=487914) O AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401785c01080 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca (FSDPPolicyWorkerBase pid=487915) educe: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [n (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401848407380 recvbuff 0x401837 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=487915) oot 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 (FSDPPolicyWorkerBase pid=487913) uff 0x401917a40000 recvbuff 0x4018e0000000 count 155582464 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: o (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-0 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000  [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487912) pCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-00 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 5b2840 [nranks=2] stream 0x400e153fcbf0 [repeated 8x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 [repeated 1117x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 437 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 70x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 70x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) L INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401d51800000 count 48236608 datatype 7 op 4 roo (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401785c01080 rec (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) nt 0 sendbuff 0x401953236000 recvbuff 0x401c60000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9c80 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount 8aa sendbuff 0x401c71800000 recvbuff 0x401c71800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stre (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) t 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllRedu (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401b00000000 recvbuff 0x401f00000000 count 155582464 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 [repeated 529x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 8ad sendbuff 0x401f00000000 recvbuff 0x401f00000000 count 155583488 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 [repeated 530x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0x401d792f3400 recvbuff 0x401b80000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9c80 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 89c sendbuff 0x401b51800000 recvbuff 0x401b51800000 count (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487914) vbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487915) jpbo (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 8a4 sendbuff 0x (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40177cc05280 recvbuff 0x401777004200 count 48236608 datatype 9 op 0 root (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [ (FSDPPolicyWorkerBase pid=487915) 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 [repeated 7x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] (FSDPPolicyWorkerBase pid=487914) 0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 r (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount 88f sendbuff 0x40180a400000 recvbuff 0x40180a400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e15 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 0x400df13eca20 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) am 0x400e153fcbf0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 roo (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCC (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) ce: opCount 8ae sendbuff 0x401f00000000 recvbuff 0x401f00000000 count 155582464 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:3340 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 c (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401788407380 recvbuff 0x401777004200 count 48236608 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype 9 op (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 0 sendbuff 0x401420000000 recvbuff 0x401a8a400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 8c8 sendbuff 0x401bd1800000 recvbuff 0x401bd1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nr (FSDPPolicyWorkerBase pid=487914) oot 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] (FSDPPolicyWorkerBase pid=487915) comm 0x400e192daf20 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL I (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:47:42 [loggers.py:259] Engine 000: Avg prompt throughput: 9.2 tokens/s, Avg generation throughput: 210.5 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 23.3%, Prefix cache hit rate: 94.7% [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 36608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduc (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllRe (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 18x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401848407380 recvbuff 0x401837004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e19 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401e6d444000 recvbuff 0x40236460 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487913) pCount 889 sendbuff 0x401480000000 recvbuff 0x401480000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 6bcf10 [nranks=4] stream 0x400e192bd380 [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x402231403180 recvbuff 0x402220000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 [repeated 631x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 440 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) [repeated 35x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 34x across cluster] (FSDPPolicyWorkerBase pid=487912) ount 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:43 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:43 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:43 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:43 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:43 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:43 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:43 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:43 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:43 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:43 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:43 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:43 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:43 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:43 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:43 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:43 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:43 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) ERROR 06-12 05:47:43 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) t 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401a80000000 recvbuff 0x40152a400000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 [repeated 258x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 8b5 sendbuff 0x40152a400000 recvbuff 0x40152a400000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 [repeated 253x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) recvbuff 0x401c60000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 [repeated 4x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount 8ba sendbuff 0x401a2f000000 recvbuff 0x401a2f000000 count 48236608 datatype 7 op 4 root 0 co (FSDPPolicyWorkerBase pid=487915) NFO AllGather: opCount 0 sendbuff 0x4017a2806300 recvbuff 0x401797004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 8b9 sendbuff 0x401c60000000  [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487913) 0000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:4879 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 8ba sendbuff 0x401a23400000 recvbuff 0x401a23400000 count 48236608 datatype 7 op 4 root 0 comm 0x4 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4019750cc000 recvbuff 0x401b60000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [n (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487914) mm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 send (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487912) 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:47:47 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 233.6 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 23.6%, Prefix cache hit rate: 94.7% [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=487913) datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401ae0000000 recvbuff 0x401c91800000 count 48236608 da (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 45x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:47:47] OK: 383 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:47:47] OK: RSS 3.48 GiB | node mem 382.4/858.0 GiB used (44.6%), avail 475.5 GiB (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000 (FSDPPolicyWorkerBase pid=487912) ranks=4] stream 0x400dfa7c68a0 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 [repeated 140x across cluster] (FSDPPolicyWorkerBase pid=487914) 0 sendbuff 0x4017e0000000 recvbuff 0x401c20000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllRedu (FSDPPolicyWorkerBase pid=487915) buff 0x4017e0000000 recvbuff 0x401c11800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: op (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 441 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 442 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 443 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 444 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 445 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 446 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 447 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count (FSDPPolicyWorkerBase pid=487913) tatype 7 op 4 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x402248407380 recvbuff 0x402237004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] str (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:47:49 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:47:49 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:47:49 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:47:49 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:47:49 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:47:49 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:47:49 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:47:49 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:47:49 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:47:49 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:47:49 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:47:49 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:47:49 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:47:49 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:47:49 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:47:49 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:47:49 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:47:49 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:47:49 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:47:49 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:47:49 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:47:49 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:47:49 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:47:49 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:47:49 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) ERROR 06-12 05:47:49 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401d60000000 recvbuff 0x4018e0000000 count 155582464 datatype 7 op 4 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 [repeated 161x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 8d4 sendbuff 0x4018e0000000 recvbuff 0x4018e0000000 count 155582464 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 [repeated 164x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) e: opCount 8d1 sendbuff 0x401e11800000 recvbuff 0x401e11800000 count 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype 9 op 0 root 0 c (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) duce: opCount 8d1 sendbuff 0x401c85e00000 recvbuff 0x401c85e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root (FSDPPolicyWorkerBase pid=487912) 000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [ (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) =8] stream (nil) (FSDPPolicyWorkerBase pid=487914) ce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40179cc05280 recvbuff 0x4 (FSDPPolicyWorkerBase pid=487915) Count fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) omm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401e95ff2400 re (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401d82344a00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401c71800000 count 48236608 datatype 7 op 4 (FSDPPolicyWorkerBase pid=487913) eam 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e19 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) cvbuff 0x401b05200000 count 48236608 datatype 7 op 4 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df13e9b50 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 8dc sendbuff 0x40156b000000 recvbuff 0x40156b000000 co (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:33403 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401e95ff2400 recvbuff 0x401e25e00000 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) bf0 (FSDPPolicyWorkerBase pid=487914) 01797004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45: (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 8e4 sendbuf (FSDPPolicyWorkerBase pid=487912) unt 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 dataty (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487913) f 0x401e0ba00000 recvbuff 0x401e0ba00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stre (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 8f7 sendbuff 0x401c71800000 recvbuff 0x401c71800000 count 48236608 datatype (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:3328 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 com (FSDPPolicyWorkerBase pid=487915) pe 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x40 (FSDPPolicyWorkerBase pid=487913) am 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11: (FSDPPolicyWorkerBase pid=487912) 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) INFO 06-12 05:47:52 [loggers.py:259] Engine 000: Avg prompt throughput: 182.4 tokens/s, Avg generation throughput: 594.8 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 46.3%, Prefix cache hit rate: 93.6% [repeated 47x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opC (FSDPPolicyWorkerBase pid=487914) m 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [n (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 23x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 co (FSDPPolicyWorkerBase pid=487913) AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45: (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) ount 0 sendbuff 0x401a4c000000 recvbuff 0x401300000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 908 sendbuff 0x401aee600000 recvbuff 0x401aee600000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4019c0000000 recvbuff 0x40194a400000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) mm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount (FSDPPolicyWorkerBase pid=487913) 6c3c90 [nranks=4] stream 0xaaaafb324d40 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 [repeated 1032x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) d380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 914 sendbuff 0x401d80000000 recvbuff 0x401d800 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401ae0000000 recvbuff 0x401e17400000 count 4823660 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 63x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 452 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) [repeated 64x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:47:53 [serving.py:315] Error in preprocessing prompt inputs [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:47:53 [serving.py:315] Traceback (most recent call last): [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:47:53 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:47:53 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:47:53 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:47:53 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:47:53 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:47:53 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:47:53 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:47:53 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:47:53 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:47:53 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:47:53 [serving.py:315] return await asyncio.gather( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:47:53 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:47:53 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:47:53 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:47:53 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:47:53 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:47:53 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:47:53 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:47:53 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:47:53 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:47:53 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:47:53 [serving.py:315] raise VLLMValidationError( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) ERROR 06-12 05:47:53 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 911 sendbuff 0x401ba0000000 recvbuff 0x401ba0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401ca0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9c80 [repeated 483x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount 91e sendbuff 0x401ca0000000 recvbuff 0x401ca0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 [repeated 483x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401b20000000 recvbuff 0x401ee0000000 count 155583488 datatype 7 op 4 root 0 comm 0x400df16bdc50 [ (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 31403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487915) 0e192daf20 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017a2806300 recvbuff 0x4017970 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4019 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401785c01080 recvbuff 0x401780000000 count 48236608 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype 9 op 0 root 0 comm  [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487914) 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487913) 8 datatype 7 op 4 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487915) 193f14b0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGat (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 d (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400e (FSDPPolicyWorkerBase pid=487913) INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401ca0000 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) b9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401857a41800 recvbuff 0x401820000000 count 155583488 datatype (FSDPPolicyWorkerBase pid=487912) x400dfa7c68a0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) atatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 c (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 (FSDPPolicyWorkerBase pid=487915) her: opCount 0 sendbuff 0x401925180000 recvbuff 0x401900000000 count 155582464 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NC (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:3 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-1 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGath (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 924 sen (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) ount 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:3 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401677004200 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:48 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:47:57 [loggers.py:259] Engine 000: Avg prompt throughput: 686.1 tokens/s, Avg generation throughput: 344.1 tokens/s, Running: 10 reqs, Waiting: 0 reqs, GPU KV cache usage: 37.4%, Prefix cache hit rate: 94.7% [repeated 49x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:47:57] OK: 391 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:47:57] OK: RSS 3.10 GiB | node mem 387.9/858.0 GiB used (45.2%), avail 470.1 GiB (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401785c01080 recvbuff 0x401780000000 count 48236608 datatype 9 op (FSDPPolicyWorkerBase pid=487915) CL INFO AllReduce: opCount 921 sendbuff 0x4014a0000000 recvbuff 0x4014a0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) er: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 13x across cluster] (FSDPPolicyWorkerBase pid=487913) dbuff 0x401c31800000 recvbuff 0x401c31800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [ (FSDPPolicyWorkerBase pid=487914) AllGather: opCount 0 sendbuff 0x4019128c0000 recvbuff 0x401900000000 count 155582464 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:33403 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 937 sendbuff 0x401b85e00000 recvbuff 0x401b85e00000 count 48236608 datat (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401d91800000 count 48236608 datatype 7 op 4 root 0 comm 0x400 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount 911 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0 [nranks=4] stream 0x400e153f9c80 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 [repeated 719x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487914) 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduc (FSDPPolicyWorkerBase pid=487915) 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opC (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401aa0000000 recvbuff 0x401bb1800000 count 48236 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) [repeated 36x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 45b sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) [repeated 35x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 93f sendbuff 0x402334600000 r (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) sendbuff 0x401d31800000 recvbuff 0x401d31800000 count 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401d00000000 recvbuff 0x402334600000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 [repeated 360x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 93e sendbuff 0x402328c00000 recvbuff 0x402328c00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 [repeated 360x across cluster] (FSDPPolicyWorkerBase pid=487912) ype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL IN (FSDPPolicyWorkerBase pid=487912) recvbuff 0x401677004200 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] strea (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jp (FSDPPolicyWorkerBase pid=487914) e: opCount 93c sendbuff 0x401c45e00000 recvbuff 0x401c45e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487915) ount 93c sendbuff 0x401c31800000 recvbuff 0x401c31800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017a2806300 recvbuff 0x401797004200 cou (FSDPPolicyWorkerBase pid=487913) ecvbuff 0x402334600000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401822806300 recvbuff 0x401817004200 count 48236608 datatype 9 op 0 root 0 co (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) m 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 948 sendbuff 0x401311800000 recvbuff 0x401311800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e4298 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x4 [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount 951 sendbuff 0 (FSDPPolicyWorkerBase pid=487912) FO AllReduce: opCount 945 sendbuff 0x401da0000000 recvbuff 0x401da0000000 count 155583488 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401677004200 recvbuff 0x401677004200 count 48236608 datatype 9 op 0 ro (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) mm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 954 sendbuff 0x401c91800000 recvbuff 0x401 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 608 datatype 7 op 4 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9c80 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4017e0000000 re (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40179cc05280 recvbuff 0x4017970042 (FSDPPolicyWorkerBase pid=487914) 00 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487915) nt 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4018a2000000 recvbuff (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:3 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype (FSDPPolicyWorkerBase pid=487912) ot 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487914) cvbuff 0x4014b1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (FSDPPolicyWorkerBase pid=487915) jpbo- (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401960000000 recvbuff 0x401ba0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] strea (FSDPPolicyWorkerBase pid=487914) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:48:02 [loggers.py:259] Engine 000: Avg prompt throughput: 25.9 tokens/s, Avg generation throughput: 370.2 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 30.7%, Prefix cache hit rate: 94.7% [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401d00000000 recvbuff 0x402340000000 count 48236608 datatype 7 op 4 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40177cc05280 recvbuff 0x401777004200 count 48236608 datatype 9 op 0 root 0 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:02 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:02 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:02 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:02 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:02 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:02 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:02 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:02 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:02 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:02 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:02 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:02 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:02 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:02 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:02 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:02 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:02 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:02 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:02 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:02 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:02 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:02 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:02 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:02 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:02 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:02 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 28x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 00df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e88c00 recvbuff 0x400e79e88 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40179cc05280 recvbuff 0x401797004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nrank (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017a2806300 recvbuff 0x401797004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] s (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 0x400e427d8400 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO A (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:4 (FSDPPolicyWorkerBase pid=487915) tream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks (FSDPPolicyWorkerBase pid=487914) s=4] stream 0xaaaaef0966d0 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 [repeated 857x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 977 sendbuff 0x401a12000000 recvbuff 0x401a12000000 count 48236608 d (FSDPPolicyWorkerBase pid=487913) root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) c00 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:3328 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:3 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4018a2000000 recvbuff 0x401c45e00000 count  [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount 97c (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 97c sendb (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO All (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) [repeated 70x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 46d sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) [repeated 72x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) llGather: opCount 0 sendbuff 0x401677004200 recvbuff 0x401677004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:3327 (FSDPPolicyWorkerBase pid=487912) atatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO (FSDPPolicyWorkerBase pid=487915) uff 0x401b4e600000 recvbuff 0x401b4e600000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] st (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4018a2000000 recvbuff 0x401c45e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 [repeated 464x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount 989 sendbuff 0x401c45e00000 recvbuff 0x401c45e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 [repeated 464x across cluster] (FSDPPolicyWorkerBase pid=487913) 00 recvbuff 0x4020f0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 [repeated 6x across cluster] (FSDPPolicyWorkerBase pid=487912) m 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401cc0000000 recvbuff 0x401a2b000000 count 48236608 datatype 7 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 97f sendbuff 0x4020f00000 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) Gather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401760000000 recvbuff 0x401ca5e00000 count 48236608 datatype 7 op 4 root 0 comm 0 (FSDPPolicyWorkerBase pid=487915) 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 data (FSDPPolicyWorkerBase pid=487914) count 48236608 datatype 7 op 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4017e0000000 recvbuff 0x401c60000000 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] [05:48:05] OK: 420 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=593552, ip=10.128.34.1) [fd-monitor] [05:48:05] OK: RSS 3.29 GiB | node mem 399.8/858.0 GiB used (46.6%), avail 458.2 GiB (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 4 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df13e9b50 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) op 4 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9c80 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGat (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 send (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) type 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INF (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487912) ReduceScatter: opCount 0 sendbuff 0x401966000000 recvbuff 0x401bb1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 46f sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) : opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) her: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) buff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 4 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) INFO 06-12 05:48:07 [loggers.py:259] Engine 000: Avg prompt throughput: 246.9 tokens/s, Avg generation throughput: 578.5 tokens/s, Running: 12 reqs, Waiting: 0 reqs, GPU KV cache usage: 58.1%, Prefix cache hit rate: 93.6% [repeated 47x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401862000000 recvbuff 0x401c60000000 count 48236608 datatype 7 op (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) O AllReduce: opCount 98a sendbuff 0x401c20000000 recvbuff 0x401c20000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401677004200 recvbuff 0x401 (FSDPPolicyWorkerBase pid=487915) ream 0xaaaabca967e0 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b80210 (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:48:05 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:48:05 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:48:05 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:48:05 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:48:05 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:48:05 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:48:05 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:48:05 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:48:05 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:48:05 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:48:05 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:48:05 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:48:05 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:48:05 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:48:05 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:48:05 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:48:05 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:48:05 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:48:05 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:48:05 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:48:05 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:48:05 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:48:05 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:48:05 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:48:05 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 13x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuf (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 994 sendbuff 0x401670c00000 recvbuff 0 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487913) d40 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:08 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:08 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:08 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:08 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:08 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:08 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:08 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:08 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:08 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:08 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:08 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:08 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:08 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:08 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:08 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:08 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:08 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jp (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) f 0x401420000000 recvbuff 0x4014a5200000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 9a3 sendbuff 0x401b20000000 recvbuff 0x401b20000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] st (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007- (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401785c01080 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks (FSDPPolicyWorkerBase pid=487914) =4] stream 0xaaaaef0966d0 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 [repeated 674x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount 9ac sendbuff 0x401d65e00000 recvbuff 0x401d65e00000 c (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 9af sendbuff 0x401db1800000 recvbuff 0x401db1800000 count 48 (FSDPPolicyWorkerBase pid=487912) b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401420000000 recvb (FSDPPolicyWorkerBase pid=487914) sendbuff 0x401b60000000 recvbuff 0x401b60000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) [repeated 36x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 476 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) [repeated 35x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401420000000 recvbuff 0x401c20000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 [repeated 305x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 9b0 sendbuff 0x401c20000000 recvbuff 0x401c20000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 [repeated 305x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 0 recvbuff 0x401800000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487914) AllGather: opCount 0 sendbuff 0x401785c01080 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:4 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915: (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 8236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) x401670c00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487912) 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401a20000000 recvbuff 0x401a72000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] s (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4017e0000000 recvbuff 0x401b71800000 count 48236608 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGat (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4018a2000000 recvbuff 0x401b60000000 count 48236608 datat (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401d17c10000 recvbuff 0x4022b4800000 count 48236608 datatype 7 o (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487915) her: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:48:12 [loggers.py:259] Engine 000: Avg prompt throughput: 75.0 tokens/s, Avg generation throughput: 323.0 tokens/s, Running: 7 reqs, Waiting: 0 reqs, GPU KV cache usage: 29.3%, Prefix cache hit rate: 94.7% [repeated 49x across cluster] (FSDPPolicyWorkerBase pid=487912) tream 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbu (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) ream 0x400e427d8400 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0 [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) ERROR 06-12 05:48:09 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 37x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount 9ac sendbuff 0x401be0000000 recvbuff 0x401be000000 (FSDPPolicyWorkerBase pid=487914) datatype 7 op 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:48:12 [serving.py:315] Error in preprocessing prompt inputs [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:48:12 [serving.py:315] Traceback (most recent call last): [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:48:12 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:48:12 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:48:12 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:48:12 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:48:12 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:48:12 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:48:12 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:48:12 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:48:12 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:48:12 [serving.py:315] return await asyncio.gather( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:48:12 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:48:12 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:48:12 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:48:12 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:48:12 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:48:12 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:48:12 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:48:12 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:48:12 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:48:12 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:48:12 [serving.py:315] raise VLLMValidationError( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:48:12 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487915) ype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO A (FSDPPolicyWorkerBase pid=487913) p 4 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGathe (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 [repeated 294x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487912) ff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 477 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 478 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 479 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 47a sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCou (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 3x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) ount 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e88c00 recvbuff 0x400e79e88c00 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e88c00 recvbuff 0x400e79e88c00 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e88c00 recvbuff 0x400e79e88c00 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e88c00 recvbuff 0x400e79e88c00 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e88c00 recvbuff 0x400e79e88c00 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e88c00 recvbuff 0x400e79e88c00 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e88c00 recvbuff 0x400e79e88c00 count (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 co (FSDPPolicyWorkerBase pid=487914) INFO AllGather: opCount 0 sendbuff 0x401785c01080 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401785c01080 recvbuff 0x40178000000 (FSDPPolicyWorkerBase pid=487915) llGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 coun (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) uff 0x401c31800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401677004200 recvbuff 0x401677004200 count 48236608 datatype 9 op 0 root 0 comm 0x400 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401a20000000 recvbuff 0x401f00000000 count 155582464 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 [repeated 183x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 9b8 sendbuff 0x401f00000000 recvbuff 0x401f00000000 count 155582464 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 [repeated 188x across cluster] (FSDPPolicyWorkerBase pid=487913) x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) r: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datat (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] [05:48:14] OK: 408 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=2062840, ip=10.128.17.224) [fd-monitor] [05:48:14] OK: RSS 2.99 GiB | node mem 390.6/858.0 GiB used (45.5%), avail 467.4 GiB (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401825181000 recvbuff 0x401800000000 count 155583488 datatype 9 op 0 root 0 co (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) unt 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017728c0800 recvbuff 0x401760000000 count 155583488 datatype 9 op 0 root 0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401752a00000 count 48236608 datatype 7 op 4 root 0 co (FSDPPolicyWorkerBase pid=487912) nt 47b sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount 9bc sendbuff 0x40 (FSDPPolicyWorkerBase pid=487915) t 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGather: opCount 0 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) mm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather: opCount 0 se (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487914) 172e800000 recvbuff 0x40172e800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) ndbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401862000000 recvbuff 0x401d80000000 count 48236608 datatype 7 op 4 root 0 comm 0x400df16bdc50 [nrank (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401d79057a00 recvbuff 0x401bf1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e156cd8d0 [nr (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) INFO AllReduce: opCount 9ca sendbuff 0x401b31800000 recvbuff 0x401b31800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401dc5e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nra (FSDPPolicyWorkerBase pid=487914) 0 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 [repeated 2x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487912) INFO ReduceScatter: opCount 0 sendbuff 0x401cb9e9c000 recvbuff 0x401ace600000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount 9d2 sendbuff 0x401bc0000000 recvbuff 0x401bc0000000 count 48236608 datatype 7 op 4 roo (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487914) aaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount 9d7 sendbuff 0x401c80000000 recvbuff 0x (FSDPPolicyWorkerBase pid=487915) 967e0 (FSDPPolicyWorkerBase pid=487913) b324d40 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487912) t 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 486 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [ (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:33302 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NC (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] st (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:3328 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:3 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) CL INFO AllGather: opCount 0 sendbuff 0x401848407380 recvbuff 0x401837004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo- (FSDPPolicyWorkerBase pid=487912) comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 155583488 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nrank (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount 9e3 sendbuff 0x401540000000 recvbuff 0x401540000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401d0f40e000 recvbuff 0x401772a00000 count 48236608 datatype 7 op (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) INFO 06-12 05:48:17 [loggers.py:259] Engine 000: Avg prompt throughput: 115.4 tokens/s, Avg generation throughput: 701.1 tokens/s, Running: 13 reqs, Waiting: 0 reqs, GPU KV cache usage: 51.7%, Prefix cache hit rate: 93.6% [repeated 47x across cluster] (FSDPPolicyWorkerBase pid=487915) ream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401cef1ec000 recvbuff 0x401752a00000 count 48236608 datatype 7 op 4 roo (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount 9ec sendbuff 0x401c80000000 recvbuff 0x401c80000000 count 48236608 (FSDPPolicyWorkerBase pid=487913) aaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4019a0000000 recvbuff 0x401c80000000 count 48236608 datatype (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) ] stream 0x400e427d8400 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401420000000 r (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount 9ef sendbuff 0x401cd1800000 recvbuff 0x401cd1800000 coun (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:48:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 29x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 re (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INF (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL  [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:48:13 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:48:13 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:48:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:48:13 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:48:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:48:13 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:48:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:48:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:48:13 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:48:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:48:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:48:13 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:48:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:48:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:48:13 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:48:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:48:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:48:13 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:48:13 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:48:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:48:13 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:48:13 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:48:13 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) ERROR 06-12 05:48:13 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (FSDPPolicyWorkerBase pid=487913) ype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487914) 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllGather (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCo (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) ecvbuff 0x401b45e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 (FSDPPolicyWorkerBase pid=487913) 7 op 4 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllG (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 [repeated 1154x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487912) cvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487914) : opCount 0 sendbuff 0x401785c01080 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401b00000000 recvbuff 0x401c91800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e6 (FSDPPolicyWorkerBase pid=487915) unt 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487913) ather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) O ReduceScatter: opCount 0 sendbuff 0x401740000000 recvbuff 0x401d91800000 count 48236608 datatype 7 op 4 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df13e9b50 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbuf (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 488 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) [repeated 66x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recv (FSDPPolicyWorkerBase pid=487912) a0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO Al (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) [repeated 82x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401848407380 recvb (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401785c01080 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4 (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:48:19 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:48:19 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:48:19 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:48:19 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:48:19 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:48:19 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:48:19 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:48:19 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:48:19 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:48:19 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:48:19 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:48:19 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:48:19 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:48:19 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:48:19 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:48:19 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:48:19 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:48:19 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:48:19 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:48:19 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:48:19 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:48:19 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:48:19 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) ERROR 06-12 05:48:19 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401d00000000 recvbuff 0x401da0000000 count 155582464 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 [repeated 584x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount a04 sendbuff 0x401da0000000 recvbuff 0x401da0000000 count 155582464 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 [repeated 583x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO A (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGath (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) f 0x401800000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) buff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=487915) 0 [nranks=4] stream 0x400e193ee7c0 [repeated 8x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 4 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) uff 0x401837004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCoun (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGa (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 9bf sendbuff 0x401c71 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) t 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) mm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 ro (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) t 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=487912) lGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 155583488 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] N (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 1765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 [repeated 5x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487913) 57400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 400df13e9b50 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0x400e153f9c80 (FSDPPolicyWorkerBase pid=487914) llGather: opCount 0 sendbuff 0x40179cc05280 recvbuff 0x401797004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jp (FSDPPolicyWorkerBase pid=487915) er: opCount 0 sendbuff 0x4017a2806300 recvbuff 0x401797004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) ot 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: o (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) INFO ReduceScatter: opCount 0 sendbuff 0x401940000000 recvbuff 0x401c00000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9c80 (skyrl_entrypoint pid=487747) [fd-monitor] [05:48:20] OK: 565 / 131,072 FDs open (0.4% of soft limit, hard limit: 131,072) (skyrl_entrypoint pid=487747) [fd-monitor] [05:48:20] OK: RSS 14.64 GiB | node mem 483.2/858.0 GiB used (56.3%), avail 374.8 GiB (FSDPPolicyWorkerBase pid=487913) ther: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount 9da sendbuff 0x402357400000 recvbuff 0x4023 [repeated 2x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487912) CCL INFO ReduceScatter: opCount 0 sendbuff 0x401431f34000 recvbuff 0x401060000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount a12 sendbuff 0x401ae0000000 recvbuff 0x401ae0000000 count 48236608 datatype 7 op 4 (FSDPPolicyWorkerBase pid=487915) t 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [ (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount a17 sendbuff 0x401ba0000000 recvbuff 0x401ba0000000 c (FSDPPolicyWorkerBase pid=487913) aaafb324d40 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:48:22 [loggers.py:259] Engine 000: Avg prompt throughput: 166.0 tokens/s, Avg generation throughput: 433.8 tokens/s, Running: 12 reqs, Waiting: 0 reqs, GPU KV cache usage: 49.8%, Prefix cache hit rate: 94.8% [repeated 49x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount a17 sendbuff 0x401b85e00000 recvbuff 0x401b85e00000 count 4 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:48:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 8x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (FSDPPolicyWorkerBase pid=487914) ount 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:48926 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 08 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL (FSDPPolicyWorkerBase pid=487915) 8236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] (FSDPPolicyWorkerBase pid=487912) root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487913) jpbo-006- (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401822806300 recvbuff 0x401817004200 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 [repeated 664x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017a2806300 recvbuff 0x401797004200 count (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401848407380 recvbuff 0x401837004200 count 48236608 datatype 9 op 0 root 0 c (FSDPPolicyWorkerBase pid=487912) : opCount 0 sendbuff 0x4014e2000000 recvbuff 0x401be5e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401ce0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10  [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) pCount a25 sendbuff 0x401c51800000 recvbuff 0x401c51800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatyp (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllGather: opCount 0 se (FSDPPolicyWorkerBase pid=487915) 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 49a sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 72x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 54x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 6x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 6x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount 0 sendbuff [repeated 5x across cluster] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:48:21 [serving.py:315] Error in preprocessing prompt inputs [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:48:21 [serving.py:315] Traceback (most recent call last): [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:48:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:48:21 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:48:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:48:21 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:48:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:48:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:48:21 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:48:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:48:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:48:21 [serving.py:315] return await asyncio.gather( [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:48:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:48:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:48:21 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:48:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:48:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:48:21 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:48:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:48:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:48:21 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:48:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:48:21 [serving.py:315] raise VLLMValidationError( [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) ERROR 06-12 05:48:21 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401420000000 recvbuff 0x401af1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa [repeated 352x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount a38 sendbuff 0x401ae0000000 recvbuff 0x401ae0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 [repeated 352x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount a2c sendbuff 0x4014b1800000 recvbuff 0x4014b1800000 count 48236608 datatype 7 (FSDPPolicyWorkerBase pid=487914) ndbuff 0x401785c01080 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4017e0000000 recvbuff 0x401bb1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e616ca9c0 [nrank (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401b00000000 recvbuff 0x401ba0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] s (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) omm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=487913) AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401ae0000000 recvbuff 0x4022d7600000 count 48236608 datatype 7 op 4 root 0 com (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0e156cd8d0 [nranks=4] stream 0x400e152ce4c0 [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=487912) 7c68a0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) e 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4014200000 (FSDPPolicyWorkerBase pid=487915) tream 0x400e193ee7c0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbu (FSDPPolicyWorkerBase pid=487913) m 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCoun (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatte (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceSca (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:3340 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 00 recvbuff 0x401565e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) r: opCount 0 sendbuff 0x401862000000 recvbuff 0x401ca5e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df13e9b50 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount a47 sendbuff 0x401da0000000 recvbuff 0x401da0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400df1 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) tter: opCount 0 sendbuff 0x401e1fc44000 recvbuff 0x401b20000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9c80 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount a4a sendbuff 0x401df1800000 recvbuff 0x401df1800000 count 48236608 datatyp (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount a2f sendbuff 0x4018e0000000 recvbuff 0x4018e0000000  [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) d8400 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4014380a4000 recvbuff 0x401c65e0 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount a2c sendbuff 0x4016f0000000 recvbuff 0x4016f0000000 count 48236608 datatype 7 op (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:3330 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-0 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) j (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40179cc05280 recvbuff 0x401797004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks= (FSDPPolicyWorkerBase pid=487915) ff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017a2806300 recvbuff 0x401797004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] str (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) INFO 06-12 05:48:27 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 420.0 tokens/s, Running: 10 reqs, Waiting: 0 reqs, GPU KV cache usage: 44.2%, Prefix cache hit rate: 94.8% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 0e40fa4c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op (FSDPPolicyWorkerBase pid=487913) t 0 sendbuff 0x401ec0000000 recvbuff 0x4018e0000000 count 155583488 datatype 7 op 4 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:26 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 6x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 44x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 coun (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 c (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduc (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 [repeated 955x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) t 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:3340 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401ab8aae800 recvbuff 0x4018f1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bc (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) ount 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:3 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) e: opCount a65 sendbuff 0x401b71800000 recvbuff 0x401b71800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 (FSDPPolicyWorkerBase pid=487915) eam 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount a57 sendbuff 0x4017ca400000 recvbuff 0x4017ca400000 count 48236608 dataty (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) [repeated 36x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 36x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount a57 sendbuff 0x401830200000 recvbuff 0x401830200000 count 48236608 (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 3x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401 [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:26 [serving.py:315] Error in preprocessing prompt inputs [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:26 [serving.py:315] Traceback (most recent call last): [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:26 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:26 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:26 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:26 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:26 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:26 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:26 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:26 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:26 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:26 [serving.py:315] return await asyncio.gather( [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:26 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:26 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:26 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:26 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:26 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:26 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:26 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:26 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:26 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:26 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:26 [serving.py:315] raise VLLMValidationError( [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) ERROR 06-12 05:48:26 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401a80000000 recvbuff 0x401830200000 count 48236608 datatype 7 op 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 [repeated 402x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount a56 sendbuff 0x4017d0200000 recvbuff 0x4017d0200000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 [repeated 399x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount a52 sendbuff 0x40106ba00000 recvbuff 0x40106ba00000 count 48236608 datatype 7 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount a5a sendbuff 0x401ca5e00000 recvbuf (FSDPPolicyWorkerBase pid=487913) c3c90 [nranks=4] stream 0xaaaafb324d40 [repeated 7x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) e 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=487912) op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceSca (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 0000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7f40 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount a47 sendbuff 0x401c11800000 recvbuff 0x401c11800000 count 48236608 datatype 7 op 4 root 0 comm 0x400 (FSDPPolicyWorkerBase pid=487915) pe 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO Redu (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487914) datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INF (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487913) f 0x401ca5e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (FSDPPolicyWorkerBase pid=487913) jpbo-0 (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) INFO 06-12 05:48:32 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 252.4 tokens/s, Running: 4 reqs, Waiting: 0 reqs, GPU KV cache usage: 16.0%, Prefix cache hit rate: 94.6% [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=487912) tter: opCount 0 sendbuff 0x401420000000 recvbuff 0x401b00000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount a6d sendbuff 0x401c00000000 recvbuff 0x401c00000000 count 48236608 datatype 7 op 4 root 0 comm 0x400 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 26x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487915) ceScatter: opCount 0 sendbuff 0x4018a2000000 recvbuff 0x401bb1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 (FSDPPolicyWorkerBase pid=487914) O ReduceScatter: opCount 0 sendbuff 0x401dc3a70000 recvbuff 0x401bc0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount a72 sendbuff 0x401cc0000000 recvbuff 0x401cc0000000 count 48236608 datatype 7 op 4 root 0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401d60000000 recvbuff 0x4018e000000 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e42 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017f7a40000 recvbuff 0x4017c0000000 count 155582464 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 [repeated 234x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:48 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401822806300 recvbuff (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40177cc05280 recvb (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGath (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 0x400e193e8890 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sen (FSDPPolicyWorkerBase pid=487914) comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO Al (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGathe (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 s (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) [repeated 27x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) uff 0x401777004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatter: opCount (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913: (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 36x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) er: opCount 0 sendbuff 0x401677004200 recvbuff 0x401677004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e40fa4c00 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) dbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823: (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401420000000 recvbuff 0x401300000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x40 [repeated 180x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount a78 sendbuff 0x40106ba00000 recvbuff 0x40106ba00000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 [repeated 182x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) endbuff 0x401862000000 recvbuff 0x4017b0a00000 count 48236608 datatype 7 op 4 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df13e9b50 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount a87 sendbuff 0x401cc0000000 recvbuff 0x401cc0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0 sendbuff 0x401aa0000000 recvbuff 0x401732a00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9c80 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount a87 sendbuff 0x401b31800000 recvbuff 0x401b31800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nra (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:48:34 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:48:34 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:48:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:48:34 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:48:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:48:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:48:34 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:48:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:48:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:48:34 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:48:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:48:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:48:34 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:48:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:48:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:48:34 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:48:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:48:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:48:34 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:48:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:48:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:48:34 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:48:34 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:48:34 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:48:34 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) ERROR 06-12 05:48:34 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) e427d8400 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount a8a sendbuff 0x401d05e00000 recvbuff 0x401d05e00000 count 48236608 dat (FSDPPolicyWorkerBase pid=487914) lGather: opCount 0 sendbuff 0x401785c01080 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487915) r: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487915) 0x400e195a79c0 [nranks=2] stream 0x400e193f14b0 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4019cfdf0000 recvbuff 0x401cc0000000 count 48236608 datatype 7 op 4 root 0 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401862000000 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11: (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) atype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO R (FSDPPolicyWorkerBase pid=487912) 0dfa7c68a0 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487914) 400e613f6c90 (FSDPPolicyWorkerBase pid=487915) 3ee7c0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount (FSDPPolicyWorkerBase pid=487912) jpbo-006-45 (FSDPPolicyWorkerBase pid=487914) jpbo-006- (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487 (FSDPPolicyWorkerBase pid=487915) 000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 [repeated 6x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) recvbuff 0x401dc5e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df13e9b50 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401822806300 recvbuff 0x401817004200 count 48236608 datatype 9 op 0 root 0 comm 0x400df1 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) educeScatter: opCount 0 sendbuff 0x401eb8ee4a00 recvbuff 0x401e11800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 (FSDPPolicyWorkerBase pid=487913) [nranks=8] stream (nil) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO Reduc (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatte (FSDPPolicyWorkerBase pid=487915) 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opC (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 dat (FSDPPolicyWorkerBase pid=487913) jp (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllR (FSDPPolicyWorkerBase pid=487912) eScatter: opCount 0 sendbuff 0x401420000000 recvbuff 0x401311800000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount aaa sendbuff 0x401ae0000000 recvbuff 0x401ae0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount aab sendbuff 0x401af1800000 recvbuff 0x401af1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount aac sendbuff 0x401b00000000 recvbuff 0x401b00000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount aad sendbuff 0x401b11800000 recvbuff 0x401b11800000 count 48236608 datatype 7 op 4 root 0 comm 0 (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) INFO 06-12 05:48:37 [loggers.py:259] Engine 000: Avg prompt throughput: 94.0 tokens/s, Avg generation throughput: 214.7 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 21.1%, Prefix cache hit rate: 94.6% [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [repeated 2x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount aad sendbuff 0x401b71800000 recvbuff 0x401b71800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount aae sendbuff 0x401b85e00000 recvbuff 0x401b85e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount aaf sendbuff 0x401ba0000000 recvbuff 0x401ba0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487915) ount 0 sendbuff 0x401978dd0000 recvbuff 0x40186b000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount ab5 sendbuff 0x4022eea00000 recvbuff 0x4022eea00000 c (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 30x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL IN (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount aba sendbuff 0x401c11800000 recvbuff 0x401c11800000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 se (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) educe: opCount aa5 sendbuff 0x40196a400000 recvbuff 0x40196a400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401882000000 recvbuff 0x401d20000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] str (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount abb sendbuff 0x401c80000000 recvbuff 0x401c80000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount abc sendbuff 0x401c91800000 recvbuff 0x401c91800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount abd sendbuff 0x401ca5e00000 recvbuff 0x401ca5e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount abe sendbuff 0x401cc0000000 recvbuff 0x401cc0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount abf sendbuff 0x401cd1800000 recvbuff 0x401cd1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4019e0000000 recvbuff 0x401b (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017a5180000 recvbuff 0x401780000000 count 155582464 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 [repeated 1141x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) FO AllReduce: opCount aae sendbuff 0x401cd1800000 recvbuff 0x401cd1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff 0x401800000000 count 48236608 datatype 9 op 0 root 0 co (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) INFO AllReduce: opCount aae sendbuff 0x401b40000000 recvbuff 0x401b40000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401677004200 recvbuff 0x401677004 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) eam 0x400e193e8890 (FSDPPolicyWorkerBase pid=487913) comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount ac0 se (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 4b5 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) [repeated 44x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 36x across cluster] (FSDPPolicyWorkerBase pid=487913) ount 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401d60000000 recvbuff 0x4018e0000000 count 155582464 datatype 7 op 4 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 [repeated 540x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount ac2 sendbuff 0x4018e0000000 recvbuff 0x4018e0000000 count 155582464 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 [repeated 462x across cluster] (FSDPPolicyWorkerBase pid=487913) 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:37 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:37 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:37 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:37 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:37 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:37 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:37 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:37 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:37 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:37 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:37 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:37 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:37 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:37 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:37 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:37 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) ERROR 06-12 05:48:37 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (FSDPPolicyWorkerBase pid=487912) ndbuff 0x4014e2000000 recvbuff 0x401c20000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401677004200 recvbuff 0x401677004200 count 48236608 dataty (FSDPPolicyWorkerBase pid=487914) r: opCount 0 sendbuff 0x4019649de000 recvbuff 0x401870c00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) mm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [ (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:33302 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 0e42a9f600 [nranks=4] stream 0x400e40fa4c00 [repeated 8x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401 [repeated 3x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) atype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stre (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:48:39 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:48:39 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:48:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:48:39 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:48:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:48:39 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:48:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:48:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:48:39 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:48:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:48:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:48:39 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:48:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:48:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:48:39 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:48:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:48:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:48:39 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:48:39 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:48:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:48:39 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:48:39 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:48:39 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) ERROR 06-12 05:48:39 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] st (FSDPPolicyWorkerBase pid=487913) f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017e8407380 recvbuff 0x4017d7004200 count 48236608 datatype 9 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487912) pe 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487914) ] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount aca sendbuff 0x401865200000 recvbuff 0x401865200000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount acb sendbuff 0x401870c00000 recvbuff 0x401870c00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks=2] stream 0x400e613f9c00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401785c01080 recvbuff 0x401780000000 coun (FSDPPolicyWorkerBase pid=487915) am 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 sendbuff 0 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487913) op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount acc sendbuff 0x401ca5e00000 recvbuff 0x401ca5e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount acd sendbuff 0x401cc0000000 recvbuff 0x401cc0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount ace sendbuff 0x401cd1800000 recvbuff 0x401cd1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opC (FSDPPolicyWorkerBase pid=487915) 8a2000000 recvbuff 0x401cd1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 [repeated 6x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401677004200 recvbuff 0x401677004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e42a9f [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) ream 0x400e192bd380 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount aca sendbuff 0x401bce000000 recvbuff 0x401bce000000 count 48236608 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40178b802100 recvbuff 0x401780000000 count 4823 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount acf sendbuff 0x401ace600000 recvbuff 0x401ace600000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401420000000 recvbuff 0x (FSDPPolicyWorkerBase pid=487914) t 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:4892 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount ada sendbuff 0x4022e3000000 recvbuff 0x4022e3000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount adb sendbuff 0x4022eea00000 recvbuff 0x4022eea00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401d00000000 recvbuff 0x4022fa400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c3c (FSDPPolicyWorkerBase pid=487915) 6608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 00df13eca20 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0x400e153fcbf0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount adc sendbuff 0x401bc0000000 recvbuff 0x401bc0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount add sendbuff 0x401bd1800000 recvbuff 0x401bd1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount ade sendbuff 0x401be5e00000 recvbuff 0x401be5e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount adf sendbuff 0x401c00000000 recvbuff 0x401c00000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount aaa sendbuff 0x401c91800000 recvbuff 0x401c91800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount aab sendbuff 0x401ca5e00000 recvbuff 0x401ca5e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount aac sendbuff 0x401cc0000000 recvbuff 0x401cc0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 [repeated 7x across cluster] (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) INFO 06-12 05:48:42 [loggers.py:259] Engine 000: Avg prompt throughput: 23.0 tokens/s, Avg generation throughput: 267.8 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 19.1%, Prefix cache hit rate: 94.6% [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount ab2 sendbuff 0x401bc0000000 recvbuff 0x401bc0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401d60000000 recvbuff 0x401e00000000 count 155583488 datat (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401d40000000 recvbuff 0x401de0000000 count 155583488 datatype 7 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount aad sendbuff 0x401cd1800000 recvbuff 0x401cd1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 [repeated 6x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount aae sendbuff 0x401ce0000000 recvbuff 0x401ce0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount aaf sendbuff 0x401b51800000 recvbuff 0x401b51800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=487913) ount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-1 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL IN (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount aba sendbuff 0x402328c00000 recvbuff 0x402328c00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-00 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount abb sendbuff 0x401c20000000 recvbuff 0x401c20000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount abc sendbuff 0x401c51800000 recvbuff 0x401c51800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount abd sendbuff 0x401c60000000 recvbuff 0x401c60000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount abe sendbuff 0x401c71800000 recvbuff 0x401c71800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount abf sendbuff 0x401e25e00000 recvbuff 0x401e25e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) FO ReduceScatter: opCount 0 sendbuff 0x401eb73f1400 recvbuff 0x401d31800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount ae5 sendbuff 0x401e25e00000 recvbuff 0x401e25e00000 count 48236608 datatype 7 op 4 root (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recvbuff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 [repeated 638x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount ae7 sendbuff 0x401da0000000 recvbuff 0x401da0000000 count (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x4 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbu (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INF (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) [repeated 49x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) [repeated 54x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:33 (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401d00000000 recvbuff 0x401da0000000 count 155582464 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 [repeated 326x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount ae8 sendbuff 0x401da0000000 recvbuff 0x401da0000000 count 155582464 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 [repeated 222x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:48:43 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) ndbuff 0x401c91800000 recvbuff 0x401c91800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 (FSDPPolicyWorkerBase pid=487913) 90 [nranks=4] stream 0x400e193f12b0 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount aea sendbuff 0x4016f0000000 recvbuff 0x4016f0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount aeb sendbuff 0x401732a00000 recvbuff 0x401732a00000 count 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount aec sendbuff 0x4017b0a00000 recvbuff 0x4017b0a00000 count 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount aed sendbuff 0x4017c0000000 recvbuff 0x4017c0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) ff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e152ce4c0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduc (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount aee sendbuff 0x401551800000 recvbuff 0x401551800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount aef sendbuff 0x401565e00000 recvbuff 0x401565e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:48:43 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:48:43 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:48:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:48:43 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:48:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:48:43 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:48:43 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:48:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:48:43 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:48:43 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:48:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:48:43 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:48:43 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:48:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:48:43 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:48:43 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:48:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:48:43 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:48:43 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:48:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:48:43 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:48:43 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:48:43 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:48:43 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) O AllGather: opCount 0 sendbuff 0x401848407380 recvbuff 0x401837004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=487914) ype 7 op 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO All (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) opCount aee sendbuff 0x4017d1800000 recvbuff 0x4017d1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount afa sendbuff 0x401cd1800000 recvbuff 0x401cd1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount afb sendbuff 0x401ce0000000 recvbuff 0x401ce0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) e: opCount aee sendbuff 0x4017aa400000 recvbuff 0x4017aa400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 0x401800000000 count 48236608 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df12be390 [repeated 5x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount afc sendbuff 0x401d05e00000 recvbuff 0x401d05e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount afd sendbuff 0x401d20000000 recvbuff 0x401d20000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount afe sendbuff 0x401d31800000 recvbuff 0x401d31800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401831403180 recv (FSDPPolicyWorkerBase pid=487913) 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount aca sendbuff 0x401845200000 recvbuff 0x401845200000 count 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 [repeated 6x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount acb sendbuff 0x401c2e600000 recvbuff 0x401c2e600000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount aff sendbuff 0x401b91800000 recvbuff 0x401b91800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount b00 (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:48:45 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:48:45 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:48:45 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:48:45 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:48:45 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:48:45 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:48:45 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:48:45 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:48:45 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:48:45 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:48:45 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:48:45 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:48:45 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:48:45 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:48:45 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:48:45 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:48:45 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:48:45 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:48:45 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:48:45 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:48:45 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:48:45 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:48:45 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) ERROR 06-12 05:48:45 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487915) op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount acc sendbuff 0x4019ca600000 recvbuff 0x4019ca600000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount acd sendbuff 0x401a20c00000 recvbuff 0x401a20c00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount ace sendbuff 0x401a2c600000 recvbuff 0x401a2c600000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401765c01080 recvbuff 0x401760000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e15 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:334050 [0] NCCL INFO AllReduce: opCount b09 sendbuff 0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401660000000 recvbuff 0x401660000000 count 4823 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount acf sendbuff 0x401a8a400000 recvbuff 0x401a8a400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO AllReduce: opCount b09 sendbuf (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:334038 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401cc0000000 recvbuff 0 [repeated 2x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) sendbuff 0x401ba5e00000 recvbuff 0x401ba5e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) buff 0x401820000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e192bd380 (FSDPPolicyWorkerBase pid=487912) 00df8fa4050 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount aed sendbuff 0x40146a400000 recvbuff 0x40146a400000 count 48236608 datatype 7 op 4 root 0 co (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount ada sendbuff 0x401ba5e00000 recvbuff 0x401ba5e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount adb sendbuff 0x401bc0000000 recvbuff 0x401bc0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e42984290 [nranks=2] stream 0x400e427d8400 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401e60000000 recvbuff 0x401c2e600000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] (FSDPPolicyWorkerBase pid=487912) 6608 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400df8fa4050 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount adc sendbuff 0x4022fa400000 recvbuff 0x4022fa400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount add sendbuff 0x402305e00000 recvbuff 0x402305e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount ade sendbuff 0x402311800000 recvbuff 0x402311800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount adf sendbuff 0x40231d200000 recvbuff 0x40231d200000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=487914) Gather: opCount 0 sendbuff 0x40179cc05280 recvbuff 0x401797004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0xaaaaef0966d0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO AllReduce: opCount af2 sendbuff 0x4018ea400000 recvbuff 0x4018ea400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e615aff00 [nranks (FSDPPolicyWorkerBase pid=487915) : opCount 0 sendbuff 0x4017a2806300 recvbuff 0x401797004200 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0xaaaabca967e0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO AllReduce: opCount af2 sendbuff 0x40194a400000 recvbuff 0x40194a400000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a79c0 [nranks=2] st (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) INFO 06-12 05:48:47 [loggers.py:259] Engine 000: Avg prompt throughput: 45.9 tokens/s, Avg generation throughput: 258.3 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 19.7%, Prefix cache hit rate: 94.6% [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:334037 [0] NCCL INFO AllReduce: opCount b0e sendbuff 0x401da0000000 recvbuff 0x401da0000000 count 155582464 datatype 7 op 4 root 0 comm 0x400e42984 [repeated 2x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 10x across cluster] (FSDPPolicyWorkerBase pid=487913) 324d40 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount af5 sendbuff 0x401ce0000000 recvbuff 0x401ce00000 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017d1403180 recvbuff 0x4017c0000000 count 48236608 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 [repeated 523x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40180b802100 recvbuff (FSDPPolicyWorkerBase pid=487912) mm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) [repeated 21x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 18x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 2x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x4014e2000000 recvbuff 0x401b25e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 [repeated 192x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount af9 sendbuff 0x401b11800000 recvbuff 0x401b11800000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 [repeated 114x across cluster] (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:48:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:489268 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401960000000 (FSDPPolicyWorkerBase pid=487915) ream 0x400e193f14b0 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:489269 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401b13050000 recvb (FSDPPolicyWorkerBase pid=487914) =2] stream 0x400e613f9c00 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount aea sendbuff 0x401c11800000 recvbuff 0x401c11800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount aeb sendbuff 0x401c20000000 recvbuff 0x401c20000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount aec sendbuff 0x401c31800000 recvbuff 0x401c31800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount aed sendbuff 0x401c45e00000 recvbuff 0x401c45e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 [repeated 6x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount aee sendbuff 0x4014a5200000 recvbuff 0x4014a5200000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount aef sendbuff 0x40152a400000 recvbuff 0x40152a400000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:4 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount afa sendbuff 0x401eac200000 recvbuff 0x401eac200000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount afb sendbuff 0x4020f0000000 recvbuff 0x4020f0000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 [repeated 6x across cluster] (FSDPPolicyWorkerBase pid=487913) 00 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount afc sendbuff 0x4022b4800000 recvbuff 0x4022b4800000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 [repeated 6x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount afd sendbuff 0x4022c0200000 recvbuff 0x4022c0200000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 [repeated 6x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount afe sendbuff 0x4022cbc00000 recvbuff 0x4022cbc00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 [repeated 6x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllReduce: opCount aff sendbuff 0x4022d7600000 recvbuff 0x4022d7600000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a95c0 [nranks=2] stream 0x400e193f4220 [repeated 6x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount b08 sendbuff 0x401c31800000 recvbuff 0x401c31800000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nra (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:48:46 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:48:46 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:48:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:48:46 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:48:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:48:46 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:48:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:48:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:48:46 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:48:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:48:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:48:46 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:48:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:48:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:48:46 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:48:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:48:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:48:46 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:48:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:48:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:48:46 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:48:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:48:46 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) ERROR 06-12 05:48:46 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487912) 0 sendbuff 0x4014e2000000 recvbuff 0x401b40000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c68a0 (FSDPPolicyWorkerBase pid=487914) recvbuff 0x401c00000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6c90 (FSDPPolicyWorkerBase pid=487915) uff 0x401be5e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee7c0 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) INFO 06-12 05:48:52 [loggers.py:259] Engine 000: Avg prompt throughput: 157.6 tokens/s, Avg generation throughput: 260.6 tokens/s, Running: 4 reqs, Waiting: 0 reqs, GPU KV cache usage: 9.6%, Prefix cache hit rate: 94.6% [repeated 48x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) Provider List: https://docs.litellm.ai/docs/providers [repeated 13x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] N (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) f 0x401c60000000 recvbuff 0x401c60000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e155b2840 [nranks=2] stream 0x400e153fcbf0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913: (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4017f7a40000 recvbuff 0x4017c0000000 count 155582464 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0xaaaafb324d40 [repeated 112x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) [repeated 26x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) [repeated 50x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:489267 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401d60000000 recvbuff 0x4018e0000000 count 155582464 datatype 7 op 4 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f12b0 [repeated 108x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 0 sendbuff 0x400eb9feca00 recvbuff 0x400eb9feca00 count 1 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f11f0 [repeated 101x across cluster] (FSDPPolicyWorkerBase pid=487912) nks=2] stream 0x400dfa7c98e0 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount afb sendbuff 0x401b40000000 recvbuff 0x401b40000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount afc sendbuff 0x401b51800000 recvbuff 0x401b51800000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount afd sendbuff 0x401b60000000 recvbuff 0x401b60000000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount afe sendbuff 0x401b71800000 recvbuff 0x401b71800000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:489266 [0] NCCL INFO AllReduce: opCount aff sendbuff 0x401b85e00000 recvbuff 0x401b85e00000 count 48236608 datatype 7 op 4 root 0 comm 0x400dfa982e60 [nranks=2] stream 0x400dfa7c98e0 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) x401df1800000 recvbuff 0x401df1800000 count 48236608 datatype 7 op 4 root 0 comm 0x400df15a2c40 [nranks=2] stream 0x400df13eca20 (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) INFO 06-12 05:48:57 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 190.4 tokens/s, Running: 4 reqs, Waiting: 0 reqs, GPU KV cache usage: 9.9%, Prefix cache hit rate: 94.6% [repeated 48x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 21x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) (EngineCore_DP0 pid=388616) WARNING 06-12 05:48:58 [block_pool.py:435] Failed to reset prefix cache because some blocks (4611) are not freed yet (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount b0c sendbuff 0x401e40000000 recvbuff 0x401e40000000 count 48236608 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401b00000000 recvbuff 0x401f00000000 count 155583488 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount b0d sendbuff 0x401f00000000 recvbuff 0x401f00000000 count 155583488 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO ReduceScatter: opCount 0 sendbuff 0x401b00000000 recvbuff 0x401f00000000 count 155582464 datatype 7 op 4 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e8890 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:334045 [0] NCCL INFO AllReduce: opCount b0e sendbuff 0x401f00000000 recvbuff 0x401f00000000 count 155582464 datatype 7 op 4 root 0 comm 0x400e195a2000 [nranks=2] stream 0x400e193eb800 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 0 sendbuff 0x401c2e577400 recvbuff 0x401c2e577400 count 1 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e87d0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 0 sendbuff 0x401c2e577200 recvbuff 0x401c2e577200 count 1 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e87d0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 0 sendbuff 0x401c2e577200 recvbuff 0x401c2e577200 count 1 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e87d0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 0 sendbuff 0x401c2e577400 recvbuff 0x401c2e577400 count 1 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e87d0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x401c2e400200 recvbuff 0x401c2e400200 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4012c0000000 recvbuff 0x401be0000000 count 155582464 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e87d0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4010f1200000 recvbuff 0x4018c7da2000 count 4194304 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e87d0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4010e5a00000 recvbuff 0x401718c00000 count 1048576 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e87d0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 290 [nranks=2] stream 0x400e427d8400 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 18f sendbuff 0x401aa0000000 recvbuff 0x401aa0000000 count 622329856 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 190 sendbuff 0x4014e2000000 recvbuff 0x4014e2000000 count 16777216 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 191 sendbuff 0x401389580000 recvbuff 0x401389580000 count 4194304 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 192 sendbuff 0x4014e2000000 recvbuff 0x4014e2000000 count 4194304 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllR (FSDPPolicyWorkerBase pid=487914) 00e613f6c90 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9f93600 recvbuff 0x400eb9f93600 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400eb9e2ca00 recvbuff 0x400eb9e (FSDPPolicyWorkerBase pid=487915) ee7c0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 r (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) educe: opCount 4e7 sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 193 sendbuff 0x4014e2800000 recvbuff 0x4014e2800000 count 16777216 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 194 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 128 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 195 sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 128 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 196 sendbuff 0x4014e2000000 recvbuff 0x4014e2000000 count 50331648 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 197 sendbuff 0x4014e8000000 recvbuff 0x4014e8000000 count 50331648 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 198 sendbuff 0x4014e2000000 recvbuff 0x4014e2000000 count 50331648 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 199 sendbuff 0x400e9fa89000 recvbuff 0x400e9fa89000 count 4096 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 19a sendbuff 0x400e9fa8b000 recvbuff 0x400e9fa8b000 count 4096 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 19b sendbuff 0x4014e2000000 recvbuff 0x4014e2000000 count 16777216 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 19c sendbuff 0x401389580000 recvbuff 0x401389580000 count 4194304 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 19d sendbuff 0x4014e2000000 recvbuff 0x4014e2000000 count 4194304 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 19e sendbuff 0x4014e2800000 recvbuff 0x4014e2800000 count 16777216 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 19f sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 128 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 1a0 sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 128 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 4f5 sendbuff 0x400e9fa89000 recvbuff 0x400e9fa89000 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (ni (FSDPPolicyWorkerBase pid=487912) 4de sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) CCL INFO AllReduce: opCount fe sendbuff 0x400e79e88c00 recvbuff 0x400e79e88c00 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-00 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) ] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) jpbo-010-46:2080788:2080788 [0] NCCL INFO Bro (FSDPPolicyWorkerBase pid=487914) 88c00 count 32 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6bd0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe se (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) 0 comm 0xaaab403c3100 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) jpbo-010-48:2062599:2062599 [0] NCCL INFO Broadcast: opCount fa sendbuff 0x4003bdd82400 recvbuff 0x4003bdd82400 count 128 datatype 9 op 0 root 0 comm 0xaaab403c3100 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) jpbo-010-48:2062599:2062599 [0] NCCL INFO Broadcast: opCount fa sendbuff 0x4003bdd82400 recvbuff 0x4003bdd82400 count 128 datatype 9 op 0 root 0 comm 0xaaab403c3100 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) jpbo-010-48:2062599:2062599 [0] NCCL INFO Broadcast: opCount fa sendbuff 0x4017a0000000 recvbuff 0x4017a0000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab403c3100 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) jpbo-010-48:2062599:2062599 [0] NCCL INFO Broadcast: opCount fb sendbuff 0x4017a0000000 recvbuff 0x4017a0000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab403c3100 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) jpbo-010-48:2062599:2062599 [0] NCCL INFO Broadcast: opCount fc sendbuff 0x4017a0000000 recvbuff 0x4017a0000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab403c3100 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) jpbo-010-48:2062599:2062599 [0] NCCL INFO Broadcast: opCount fd sendbuff 0x40075203c400 recvbuff 0x40075203c400 count 4096 datatype 9 op 0 root 0 comm 0xaaab403c3100 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) jpbo-010-48:2062599:2062599 [0] NCCL INFO Broadcast: opCount fd sendbuff 0x40075203c400 recvbuff 0x40075203c400 count 4096 datatype 9 op 0 root 0 comm 0xaaab403c3100 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) jpbo-010-48:2062599:2062599 [0] NCCL INFO Broadcast: opCount fd sendbuff 0x40075203c400 recvbuff 0x40075203c400 count 4096 datatype 9 op 0 root 0 comm 0xaaab403c3100 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) jpbo-010-48:2062599:2062599 [0] NCCL INFO Broadcast: opCount fd sendbuff 0x4017a0000000 recvbuff 0x4017a0000000 count 622329856 datatype 9 op 0 root 0 comm 0xaaab403c3100 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) jpbo-010-48:2062599:2062599 [0] NCCL INFO Broadcast: opCount fe sendbuff 0x4017a0000000 recvbuff 0x4017a0000000 count 622329856 datatype 9 op 0 root 0 comm 0xaaab403c3100 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) jpbo-010-48:2062599:2062599 [0] NCCL INFO Broadcast: opCount ff sendbuff 0x401739800000 recvbuff 0x401739800000 count 16777216 datatype 9 op 0 root 0 comm 0xaaab403c3100 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) jpb (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) INFO AllGather: opCount 0 sendbuff 0x4010e5c00000 recvbuff 0x4018c7da2000 count 1048576 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e87d0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4012d3c00000 recvbuff 0x401718c00000 count 1048576 datatype 9 op 0 root 0 comm 0x400e196bcf10 [n (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) jpbo-007-31:389428:389428 [0] NCCL INFO Broadcast: opCount 1ab sendbuff 0x4003ddd82400 recvbuff 0x4003ddd82400 count 128 d (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) jpbo-051-38:386894:386894 [0] NCCL INFO Broadcast: opCount 111 sendbuff 0x401720000000 (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) jpbo-051-40:388616:388616 [0] NCCL INFO Broadcast: opCount 1ab sendbuff 0x40039dd82400 recvbuff 0x40039dd82400 count 128 dINFO 06-12 05:49:00 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 228.7 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 20.7%, Prefix cache hit rate: 93.5% (FSDPPolicyWorkerBase pid=487913) ecvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL I (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) l) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400e9fa5de00 recvbuff 0x400e9fa89000 coun (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 50b sendbuff 0x400e9fa89000 recvbuff 0x400e9fa89000 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b25 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40107e800000 recvbuff 0x401868000000 count 12582912 datatype 9 op 0 root 0 comm 0x400df16b (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) t 1024 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7e80 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadca (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88e00 recvbuff 0x400e (FSDPPolicyWorkerBase pid=487914) fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [ (FSDPPolicyWorkerBase pid=487915) ndbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCC (FSDPPolicyWorkerBase pid=487913) NFO AllGather: opCount 0 sendbuff 0x4010f2000000 recvbuff 0x401882000000 count 12582912 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f11f0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4010fc000000 recvbuff 0x401888000000 count 12582912 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nr (FSDPPolicyWorkerBase pid=487912) 60 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) st: opCount 1bd sendbuff 0x401389580000 recvbuff 0x401389580000 count 4194304 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 1cb sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 128 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=4 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e88c00 recvbuff 0x400e79e88c00 count 1 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 coun (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) adcast: opCount 1a7 sendbuff 0x401799200000 recvbuff 0x401799200000 count 4194304 datatype 9 op 0 root 0 comm 0xaaab07142a30 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) jpbo-010-48:2062599:2062599 [0] NCCL INFO (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) b9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401304000000 recvbuff 0x4018a2000000 count 1048576 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stre (FSDPPolicyWorkerBase pid=487915) L INFO AllGather: opCount 0 sendbuff 0x40105e400000 recvbuff 0x4018d38f6000 count 4194304 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee700 (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) atatype 9 op 0 root 0 comm 0xaaab1e9d50e0 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) recvbuff 0x4003e0000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab20ae6560 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) jpbo-051-40 (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) jpbo-007-21 (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) jpbo-007-21 (AsyncVLLMInferenceEngine pid=389198, ip=10.128.17.63) jpbo-007-31 (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) jpbo-007-31 (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) jpbo-007-02 (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) jpbo-007-02 (AsyncVLLMInferenceEngine pid=386652, ip=10.128.34.6) jpbo-051-38 (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) jpbo-051-38 (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) jpbo-051-38 (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) jpbo-051-37 (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) jpbo-051-37 (AsyncVLLMInferenceEngine pid=379670, ip=10.128.34.2) jpbo-051-34 (AsyncVLLMInferenceEngine pid=379668, ip=10.128.34.2) jpbo-051-34 (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) jpbo-007-23 (AsyncVLLMInferenceEngine pid=308847, ip=10.128.17.55) jpbo-007-23 (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) jpbo-051-33 (AsyncVLLMInferenceEngine pid=388390, ip=10.128.34.8) jpbo-051-40 (AsyncVLLMInferenceEngine pid=593056, ip=10.128.34.1) jpbo-051-33 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4012d1400000 re (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) INFO 06-12 05:49:02 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 183.5 tokens/s, Running: 4 reqs, Waiting: 0 reqs, GPU KV cache usage: 10.1%, Prefix cache hit rate: 94.6% [repeated 47x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) t 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40130fc00000 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) : opCount 0 sendbuff 0x4010de800000 recvbuff 0x4018c7da2000 count 4194304 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e87d0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400eb9e37200 recvbuff 0x400eb9e8b000 count 1024 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] strea (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCoun (FSDPPolicyWorkerBase pid=487914) am 0x400e613f6bd0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 ro (FSDPPolicyWorkerBase pid=487915) 00e193ee700 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 c (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: o (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 40x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) t 0 sendbuff 0x401853400000 recvbuff 0x4014e2000000 count 4194304 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7e80 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40185e400000 recvbuff 0x4014e2800000 count 4194304 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] s (FSDPPolicyWorkerBase pid=487912) INFO AllReduce: opCount 536 sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 54b sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [ (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) cvbuff 0x401862000000 count 12582912 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df13e9a90 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO Al (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) recvbuff 0x4018b3182000 count 12582912 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9bc0 (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) jpbo-010-48:2062603:2062603 [0] NCCL INFO Broadcast: opCount 1fb sendbuff 0x401720000000 recvbuff 0x401720000000 coun (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) Broadcast: opCount 129 sendbuff 0x401739800000 recvbuff 0x401739800000 count 16777216 datatype 9 op 0 root 0 comm 0xaaab19362b60 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) m 0x400e193e87d0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 co (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) jpbo-007-21:336339:336339 [0] NCCL INFO Broadcast: opCount 200 sendbuff 0x401739200000 recvbuff 0x401739200000 count 4194304 datatype 9 op 0 root 0 comm 0xaaaaf5247510 [n (FSDPPolicyWorkerBase pid=487914) ot 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487915) omm 0x400e192daf20 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401316e00000 recvbuff 0x401726400000 co (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) jpbo-051-38:386886:386886 [0] NCCL INFO Broadcast: opCount 148 sendbuff 0x4016b9800000 recvbuff 0x4016b9800000 count 16777216 data (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) tream 0x400e427c7e80 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 204 sendbuff 0x4014e2000000 recvbuff 0x4014e2000000 count (FSDPPolicyWorkerBase pid=487913) pCount 0 sendbuff 0x4012eb800000 recvbuff 0x401882000000 count 12582912 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f11f0 (AsyncVLLMInferenceEngine pid=388389, ip=10.128.34.8) (EngineCore_DP0 pid=388620) WARNING 06-12 05:48:58 [block_pool.py:435] Failed to reset prefix cache because some blocks (5399) are not freed yet [repeated 47x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401847400000 recvbuff 0x401514808000 c (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) lReduce: opCount fe sendbuff 0x400e79e89000 recvbuff 0x400e79e89000 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) AllReduce: opCount fe sendbuff 0x400eb9e89000 recvbuff 0x400eb9e89000 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllRedu (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) unt 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount f (FSDPPolicyWorkerBase pid=487915) unt 1048576 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee700 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe send (FSDPPolicyWorkerBase pid=487913) x400e193f11f0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 cou (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount 0 sendbuff 0x401698c9c400 recvbuff 0x401698c9c400 count 1 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee700 [repeated 14x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) [repeated 568x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400eb9e3ca00 recvbuff 0x400eb9e88c00 count 32 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f11f0 [repeated 1024x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL (FSDPPolicyWorkerBase pid=487913) anks=4] stream 0x400e193f11f0 [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 212 sendbuff 0x400e9fa89000 recvbuff 0x400e9fa89000 count 4096 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) [repeated 5849x across cluster] (FSDPPolicyWorkerBase pid=487912) ount 4194304 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllR (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 539x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401315600000 recvbuff 0x401724c00 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) ce: opCount 567 sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 575 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] s (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) jpbo-010-47:2073774:2073774 [0] NCCL INFO Broadcast: opCount 15f sendbuff 0x401720000000 recvbuff 0x401720000000 c (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) jpbo-007-31:389428:389428 [0] NCCL INFO Broadcast: opCount 22b sendbuff 0x401759200000 recvbuff 0x4 (FSDPPolicyWorkerBase pid=487914) e sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] (FSDPPolicyWorkerBase pid=487915) buff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) type 9 op 0 root 0 comm 0xaaab25666d20 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=379417, ip=10.128.34.5) jpbo-051-37:379644:379644 [0] NCCL INFO Broadcast: opCount (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) tream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400e9fa6cc00 recvbuff 0x400e9fa88e00 count 32 (FSDPPolicyWorkerBase pid=487913) nt 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCou (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) jpbo-010-47:2073782:2073782 [0] NCCL INFO Bro [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=487914) 000 count 1048576 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6bd0 [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) 0 comm 0xaaab215f1670 [nranks=49] stream (nil) [repeated 24x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) jpbo-051-38:386887:386887 [0] NCCL INFO Broadcast: opCount fa sendbuff 0x401740000000 recvbuff 0x401740000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab215f1670 [nranks=49] stream (nil) [repeated 72x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) jpbo-051-38:386887:386887 [0] NCCL INFO Broadcast: opCount fb sendbuff 0x401740000000 recvbuff 0x401740000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab215f1670 [nranks=49] stream (nil) [repeated 24x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) jpbo-051-38:386887:386887 [0] NCCL INFO Broadcast: opCount fc sendbuff 0x401740000000 recvbuff 0x401740000000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab215f1670 [nranks=49] stream (nil) [repeated 24x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) jpbo-051-38:386887:386887 [0] NCCL INFO Broadcast: opCount fd sendbuff 0x401740000000 recvbuff 0x401740000000 count 622329856 datatype 9 op 0 root 0 comm 0xaaab215f1670 [nranks=49] stream (nil) [repeated 96x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) jpbo-051-38:386887:386887 [0] NCCL INFO Broadcast: opCount fe sendbuff 0x401740000000 recvbuff 0x401740000000 count 622329856 datatype 9 op 0 root 0 comm 0xaaab215f1670 [nranks=49] stream (nil) [repeated 24x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) jpbo-051-38:386887:386887 [0] NCCL INFO Broadcast: opCount ff sendbuff 0x4016d9800000 recvbuff 0x4016d9800000 count 16777216 datatype 9 op 0 root 0 comm 0xaaab215f1670 [nranks=49] stream (nil) [repeated 24x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpb [repeated 6x across cluster] (FSDPPolicyWorkerBase pid=487912) educe: opCount 576 sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 58b sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stre (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) jpbo-007-21:336339:336339 [0] NCCL INFO Broadcast: opCount 1ab sendbuff 0x4003bdd82400 recvbuff 0x4003bdd82400 count 128 d [repeated 15x across cluster] (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) jpbo-010-47:2073795:2073795 [0] NCCL INFO Broadcast: opCount 145 sendbuff  [repeated 30x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e88c00 recvbuff 0x400e79e88c00 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7e80 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broad (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 193e87d0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 dataty (FSDPPolicyWorkerBase pid=487915) INFO AllGather: opCount 0 sendbuff 0x401333400000 recvbuff 0x4018d30f6000 count 4194304 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee700 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400eb9e3de00 recvbuff 0x400eb9e88c00 count 32 datatype 9 op 0 root 0 comm 0x400e [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=487913) nt 0 sendbuff 0x400eb9e3f200 recvbuff 0x400eb9e88c00 count 32 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f11f0 (FSDPPolicyWorkerBase pid=487912) am (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40185f400000 recvbuff 0x401389180000 count 1048576 d (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) jpbo-051-37:379657:379657 [0] NCCL INFO Broadca [repeated 17x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 52e sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00  [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) cast: opCount 23d sendbuff 0x4014e2000000 recvbuff 0x4014e2000000 count 50331648 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) jpbo-010-48:2062603:2062603 [0] NCCL INFO Broadcast: opCount 24f sendbuff 0x4003bdd87600 recvbuff 0x4003bdd87600 count 128 datatype 9 op 0 root 0 comm 0xaaaaeca090b0 [nranks=49] stream (n (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) ount 50331648 datatype 9 op 0 root 0 comm 0xaaab19362b60 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) jpbo-010-46:2080799:2080799 [0] NCCL INFO Broadcast: opCount 17a sendbuff 0x4003bdd8ca00 recvbuff 0x4003bdd8ca00 count 128 datatype 9 op 0 root 0 com (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) pe 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) 0x400e613f6bd0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 r (FSDPPolicyWorkerBase pid=487915) e193ee700 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) jpbo-051-40:388616:388616 [0] (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) jpbo-007-23:309070:309070 [0] (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) 163 sendbuff 0x401739200000 recvbuff 0x401739200000 count 4194304 datatype 9 op 0 root 0 comm 0xaaab38751670 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=308848, ip=10.128.17.55) jpbo-007-23:309096:309096 [0] NCCL INFO Broadcast: opCount 17d sendbuff 0x4016d9800000 recvbuff 0x4016d9800000 count 16777216 datatype 9 op 0 root 0 comm 0xaaab38751670 [nranks=49] s (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) jpbo-007-23:309090:309090 [0] (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) jpbo-051-33:593278:593278 [0] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) jpbo-051-33:593291:593291 [0] (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) jpbo-007-21:336355:336355 [0] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) xaaab2ebd6ee0 [nranks=49] stream (nil) [repeated 25x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) jpbo-007-21:336339:336339 [0] (AsyncVLLMInferenceEngine pid=389069, ip=10.128.17.63) jpbo-007-31:389428:389428 [0] (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) st: opCount 1d6 sendbuff 0x4003ddd82400 recvbuff 0x4003ddd82400 count 128 datatype 9 op 0 root 0 comm 0xaaab3461efd0 [nranks=49] stream (nil) [repeated 17x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 24b sendbuff 0x4014e2000000 recvbuff 0x4014e2000000 count 16777216 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=4 [repeated 7x across cluster] (AsyncVLLMInferenceEngine pid=389197, ip=10.128.17.63) jpbo-007-31:389444:389444 [0] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) jpbo-007-02:632066:632066 [0] (AsyncVLLMInferenceEngine pid=631841, ip=10.128.17.34) jpbo-007-02:632077:632077 [0] (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) jpbo-007-02:632073:632073 [0] NCCL INFO Broadcast: opCount 17d sendbuff 0x401759800000 recvbuff 0x401759800000 count 16777216 datatype 9 op 0 root 0 comm 0xaaab28dd2810 [nranks=49] sINFO 06-12 05:49:06 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 193.0 tokens/s, Running: 4 reqs, Waiting: 0 reqs, GPU KV cache usage: 16.7%, Prefix cache hit rate: 94.0% (FSDPPolicyWorkerBase pid=487913) e193f11f0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 data (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) jpbo-051-38:386885:386885 [0] (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) jpbo-051-37:379664:379664 [0] (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) jpbo-051-37:379657:379657 [0] (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) jpbo-051-34:379901:379901 [0] (AsyncVLLMInferenceEngine pid=379669, ip=10.128.34.2) jpbo-051-34:379917:379917 [0] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) jpbo-051-40:388625:388625 [0] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) adcast: opCount 1a7 sendbuff 0x401719200000 recvbuff 0x401719200000 count 4194304 datatype 9 op 0 root 0 comm 0xaaab0bce14a0 [nranks=49] stream (nil) [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=388262, ip=10.128.34.8) 01759200000 count 4194304 datatype 9 op 0 root 0 comm 0xaaaaebe55a60 [nranks=49] stream (nil) [repeated 21x across cluster] (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) atatype 9 op 0 root 0 comm 0xaaab1d13b270 [nranks=49] stream (nil) [repeated 16x across cluster] (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) 0x40077203c400 recvbuff 0x40077203c400 count 4096 datatype 9 op 0 root 0 comm 0xaaab26735d60 [nranks=49] stream (nil) [repeated 30x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] N (FSDPPolicyWorkerBase pid=487912) atatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-00 (FSDPPolicyWorkerBase pid=487914) oot 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 50331648 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) INFO 06-12 05:49:07 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 502.3 tokens/s, Running: 12 reqs, Waiting: 0 reqs, GPU KV cache usage: 65.0%, Prefix cache hit rate: 61.6% [repeated 48x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) t 50331648 datatype 9 op 0 root 0 comm 0xaaab2ebd6ee0 [nranks=49] stream (nil) [repeated 6x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCoun [repeated 2x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 20x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) t 0 sendbuff 0x400eb9e3ba00 recvbuff 0x400eb9e89000 count 1024 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e87d0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 0131e400000 recvbuff 0x401888000000 count 12582912 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e87d0 [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) jpbo-010-46:2080778:2080778 [0] NCCL INFO Broadcast: opCount 1fb sendbuff 0x401760000000 recvbuff 0x401760000000 coun [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) t 0 sendbuff 0x40189b400000 recvbuff 0x4014e8000000 count 12582912 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7e80 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) Broadcast: opCount 129 sendbuff 0x4016d9800000 recvbuff 0x4016d9800000 count 16777216 datatype 9 op 0 root 0 comm 0xaaab26735d60 [nranks=49] stream (nil) [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo- (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) INFO AllReduce: opCount fe sendbuff 0x400e79e88e00 recvbuff 0x400e79e88e00 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) CCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) jpbo-007-21:336355:336355 [0] NCCL INFO Broadcast: opCount 200 sendbuff 0x401739200000 recvbuff 0x401739200000 count 4194304 datatype 9 op 0 root 0 comm 0xaaab3461efd0 [n [repeated 16x across cluster] (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) il) (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) jpbo-010-46:2080788:2080788 [0] (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) m 0xaaab19362b60 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) jpbo-010-46:2080778:2080778 [0] (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) jpbo-010-48:2062603:2062603 [0] (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) jpbo-007-02:632092:632092 [0] NCCL INFO Broadcast: opCount 148 sendbuff 0x4016f9800000 recvbuff 0x4016f9800000 count 16777216 data [repeated 18x across cluster] (AsyncVLLMInferenceEngine pid=2062356, ip=10.128.17.224) jpbo-010-48:2062583:2062583 [0] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e89000 recvbuff 0x400eb9e89000 count 1 datatype 7 op 0 root 0 (AsyncVLLMInferenceEngine pid=2073418, ip=10.128.17.223) jpbo-010-47:2073787:2073787 [0] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) jpbo-010-47:2073782:2073782 [0] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCou (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) jpbo-051-33:593278:593278 [0] NCCL INFO Broadcast: opCount 280 sendbuff 0x40077203c400 recvbuff 0x40077203c400 count 4096 datatype 9 op 0 root 0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 0x400e427c7e80 (FSDPPolicyWorkerBase pid=487913) jpbo (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400e9fa77600 recvbuff 0x400e9fa8b000 count 1024 datatype 9 op 0 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) [repeated 467x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400eb9e4a600 recvbuff 0x400eb9e88c00 count 32 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e87d0 [repeated 932x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCo (FSDPPolicyWorkerBase pid=487914) nt fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NC (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 156cd8d0 [nranks=4] stream 0x400e153f9bc0 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 292 sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 128 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) [repeated 6124x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 ro (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 504x across cluster] (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400eb9e45a00 recvbuff 0x400eb9e8900 [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) jpbo-010-47:2073795:2073795 [0] NCCL INFO Broadcast: opCount 15f sendbuff 0x401740000000 recvbuff 0x401740000000 c [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) unt 5e7 sendbuff 0x400e9fa89000 recvbuff 0x400e9fa89000 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 5f5 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (FSDPPolicyWorkerBase pid=487912) root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 5f6 sendbu (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400e79e88e00 recvbuff 0x400e79e88e00 co (AsyncVLLMInferenceEngine pid=2080424, ip=10.128.17.222) jpbo-010-46:2080788:2080788 [0] NCCL INFO Broadcast: opCount 2a4 send (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 284 sendbuff 0x4014e2000000 recvbuff 0x4014e200000 [repeated 36x across cluster] (AsyncVLLMInferenceEngine pid=2080552, ip=10.128.17.222) jpbo-010-46:2080799:2080799 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL (FSDPPolicyWorkerBase pid=487913) type 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) [repeated 19x across cluster] (AsyncVLLMInferenceEngine pid=379416, ip=10.128.34.5) jpbo-051-37:379652:379652 [0] NCCL INFO Broadcast: opCount  [repeated 18x across cluster] (AsyncVLLMInferenceEngine pid=2062354, ip=10.128.17.224) jpbo-010-48:2062599:2062599 (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) jpbo-010-48:2062587:2062587 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:33282 (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) jpbo-010-47:2073774:2073774 (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) jpbo-010-47:2073795:2073795 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) jpbo-010-46:2080782:2080782 (AsyncVLLMInferenceEngine pid=2080553, ip=10.128.17.222) tream (nil) [repeated 25x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401378800000 recvbuff 0x4018a8000000 count 12582912 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] str (FSDPPolicyWorkerBase pid=487915) CL INFO AllGather: opCount 0 sendbuff 0x401371400000 recvbuff 0x4018d30f6000 count 12582912 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee700 (FSDPPolicyWorkerBase pid=487915) 0 count 1024 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee700 [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=631843, ip=10.128.17.34) jpbo-007-02:632073:632073 [0] NCCL (AsyncVLLMInferenceEngine pid=379540, ip=10.128.34.2) comm 0xaaaae96a6e60 [nranks=49] stream (nil) [repeated 12x across cluster] (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpb (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) jpbo-051-33:593291:593291 [0] NCCL INFO Broadcast: opCount 2ab sendbuff 0x4INFO 06-12 05:49:10 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 96.3 tokens/s, Running: 2 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.3%, Prefix cache hit rate: 94.6% (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) ] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4018bc800000 recvbuff 0x401389580000 count 104 (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) jpbo-007-23:309090:309090 [0] NCCL INFO Broadcast: opCount 2ab sendbuff 0x4 [repeated 16x across cluster] (FSDPPolicyWorkerBase pid=487913) ot 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487912) ff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:4 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) unt 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 8576 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7e80 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broad (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 com (FSDPPolicyWorkerBase pid=487914) eam 0x400e613f6bd0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 ro (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40135f600000 recvbuff 0x401704c00000 count 1048576 datatype 9 op 0 root 0 comm 0x4 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=487915) 400e193ee700 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 c (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 5ae sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) jpbo-010-46:2080778:2080778 [0] NCCL INFO Broadcast: opCount 24f sendbuff 0x4003ddd82400 recvbuff 0x4003ddd82400 count 128 datatype 9 op 0 root 0 comm 0xaaab2ebd6ee0 [nranks=49] stream (n [repeated 5x across cluster] (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) ount 50331648 datatype 9 op 0 root 0 comm 0xaaab26735d60 [nranks=49] stream (nil) [repeated 5x across cluster] (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) jpbo-010-47:2073795:2073795 [0] NCCL INFO Broadcast: opCount 17a sendbuff 0x4003bdd8ca00 recvbuff 0x4003bdd8ca00 count 128 datatype 9 op 0 root 0 com [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4018ae800000 recvbuff 0x401514808000 count 12582912 datatype 9 op 0 root 0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) cast: opCount 2bd sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 128 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 2cb sendbuff 0x4014e8000000 recvbuff 0x4014e8000000 count 50331648 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks= (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) buff 0x401719200000 recvbuff 0x401719200000 count 4194304 datatype 9 op 0 root 0 comm 0xaaaaeca090b0 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL I (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCC (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) jpbo-010-47:2073774:2073774 [0] NCCL INFO Broadcast: opCount 1ca s (FSDPPolicyWorkerBase pid=487915) sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 20x across cluster] (AsyncVLLMInferenceEngine pid=2073545, ip=10.128.17.223) jpbo-010-47:2073795:2073795 [0] NCCL INFO Broadcast: opCount 195 sendbuff 0x4016d9800000 recvbuff 0x4016d9800000 count 16777216 datatype 9 op 0 root 0 comm 0xaaab26735d60 [nranks=49] s [repeated 23x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 9] stream (nil) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401361800000 recvbuf (FSDPPolicyWorkerBase pid=487915) omm 0x400e192daf20 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401391400000 recvbuff 0x4018d30f6000 cou (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) [repeated 22x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) INFO Broadcast: opCount 1b5 sendbuff 0x40079203c400 recvbuff 0x40079203c400 count 4096 datatype 9 op 0 root 0 comm 0xaaab0bda6fa0 [nranks=49] stream (nil) (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) jpbo-007-02:632092:632092 [0] NCCL INFO Broadcast: opCount 1cf sendbuff 0x401760000000 recvbuff 0x401760000000 count 50331648 datatype 9 op 0 root 0 comm 0xaa (AsyncVLLMInferenceEngine pid=379545, ip=10.128.34.5) 0092a400000 recvbuff 0x40092a400000 count 50331648 datatype 9 op 0 root 0 comm 0xaaab1d13b270 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllReduce: opCount 62e sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 coun (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) INFO 06-12 05:49:12 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 471.1 tokens/s, Running: 11 reqs, Waiting: 0 reqs, GPU KV cache usage: 59.0%, Prefix cache hit rate: 61.6% [repeated 47x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40137c [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=487912) comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 636 sendbuff 0x400e9fa (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) NFO AllReduce: opCount fe sendbuff 0x400e79e88c00 recvbuff 0x400e79e88c00 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jp (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) L INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) t 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 25x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) f 0x401882000000 count 4194304 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e87d0 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 (FSDPPolicyWorkerBase pid=487915) nt 12582912 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee700 (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendb (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 000000 recvbuff 0x4018b3182000 count 4194304 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9bc0 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) opCount 0 sendbuff 0x4018de400000 recvbuff 0x4014e2800000 count 4194304 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7e80 (AsyncVLLMInferenceEngine pid=2062225, ip=10.128.17.224) jpbo-010-48:2062603:2062603 [0] NCCL INFO Broadcast: opCount 2f8 sendbuff 0x401720000000 recvbuff 0x401720000000 count 50331648 datatype 9 op (AsyncVLLMInferenceEngine pid=379289, ip=10.128.34.5) jpbo- [repeated 17x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) endbuff 0x4016b9800000 recvbuff 0x4016b9800000 count 16777216 datatype 9 op 0 root 0 comm 0xaaab46c65e00 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) il) [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) m 0x400e192d5230 [nranks=8] stream (nil) [repeated 6x across cluster] (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) jpbo-007-21:336355:336355 [0] NCCL INFO Broadcast: opCount 280 sendbuff 0x40077203c400 recvbuff 0x40077203c400 count 4096 datatype 9 op 0 root 0  [repeated 16x across cluster] (FSDPPolicyWorkerBase pid=487915) uff 0x400eb9e89000 recvbuff 0x400eb9e89000 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] N (AsyncVLLMInferenceEngine pid=308714, ip=10.128.17.55) jpbo-007-23:309070:309070 [0] NCCL INFO Broadcast: opCount 300 sendbuff 0x40039dd82400 recvbuff 0x40039dd82400 count 128 dat (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) jpbo-007-21:336335:336335 [0] NCCL INFO Broadcast: opCount 1ea sendbuff 0x401740000000 r (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 0e427c7e80 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 660 sendbuff 0x400e9fa89000 recvbuff 0x400e9fa89000 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) [repeated 482x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4013ae800000 recvbuff 0x4018b3182000 count 12582912 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9bc0 [repeated 987x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllRed (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400eb9e59600 recvbuff 0x400eb9e88c00 count 32 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] st (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 400e156cd8d0 [nranks=4] stream 0x400e153f9bc0 [repeated 4x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 312 sendbuff 0x401389580000 recvbuff 0x401389580000 count 4194304 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) [repeated 6136x across cluster] (FSDPPolicyWorkerBase pid=487915) CCL INFO AllGather: opCount 0 sendbuff 0x400eb9e57000 recvbuff 0x400eb9e88e00 count 32 datatype 9 op 0 root 0 comm 0x400e196c2490 [nranks=4] stream 0x400e193ee700 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192daf20 [nranks=8] stream (nil) [repeated 504x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401379a00000 recvbuff 0x401718c000 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192d52 [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2073676, ip=10.128.17.223) jpbo-010-47:2073782:2073782 [0] NCCL INFO Broadcast: opCount 2a4 send [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 304 sendbuff 0x400e9fa89000 recvbuff 0x400e9fa8900 [repeated 13x across cluster] (FSDPPolicyWorkerBase pid=487914) 00 count 12582912 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stream 0x400e613f6bd0 (AsyncVLLMInferenceEngine pid=631842, ip=10.128.17.34) jpbo-007-02:632092:632092 [0] NCCL [repeated 18x across cluster] (AsyncVLLMInferenceEngine pid=593055, ip=10.128.34.1) comm 0xaaab10396860 [nranks=49] stream (nil) [repeated 7x across cluster] (FSDPPolicyWorkerBase pid=487914) ot 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e8 [repeated 2x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=336104, ip=10.128.17.53) (EngineCore_DP0 pid=336335) WARNING 06-12 05:49:16 [block_pool.py:435] Failed to reset prefix cache because some blocks (7239) are not freed yet (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4018ce400000 recvbuff 0x401514808000 count 12582912 datatype 9 op 0 root 0 comm 0x400dfaa9d [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:16 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:16 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:16 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:16 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:16 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:16 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:16 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:16 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:16 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:16 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:16 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:16 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:16 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:16 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:16 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:16 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:16 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:16 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:16 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:16 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:16 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:16 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:16 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:16 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:16 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:16 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) buff 0x401759200000 recvbuff 0x401759200000 count 4194304 datatype 9 op 0 root 0 comm 0xaaab2ebd6ee0 [nranks=49] stream (nil) [repeated 5x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) jpbo-010-48:2062587:2062587 [0] NCCL INFO Broadcast: opCount 1ca s [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=487914) sendbuff 0x400eb9e89000 recvbuff 0x400eb9e89000 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400eb9e54a00 recvbuf (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 0 count 4096 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) [repeated 7x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) INFO Broadcast: opCount 1b5 sendbuff 0x40075203c400 recvbuff 0x40075203c400 count 4096 datatype 9 op 0 root 0 comm 0xaaab215f1670 [nranks=49] stream (nil) [repeated 18x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) jpbo-051-38:386887:386887 [0] NCCL INFO Broadcast: opCount 1cf sendbuff 0x401740000000 recvbuff 0x401740000000 count 50331648 datatype 9 op 0 root 0 comm 0xaa [repeated 18x across cluster] (FSDPPolicyWorkerBase pid=487912) 88e00 recvbuff 0x400e9fa88e00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) [repeated 17x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 30 [nranks=8] stream (nil) [repeated 21x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) INFO 06-12 05:49:17 [loggers.py:259] Engine 000: Avg prompt throughput: 525.3 tokens/s, Avg generation throughput: 468.1 tokens/s, Running: 16 reqs, Waiting: 0 reqs, GPU KV cache usage: 69.3%, Prefix cache hit rate: 62.0% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 28x across cluster] (FSDPPolicyWorkerBase pid=487913) f 0x400eb9e89000 count 1024 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f11f0 (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) [rank-0]: Saving model to /e/data1/datasets/playground/ot-baf/explore-tis-minp/explore-tis-minp/checkpoints/global_step_82/policy/model_world_size_8_rank_0.pt (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) jpbo-010-46:2080778:2080778 [0] NCCL INFO Broadcast: opCount 2f8 sendbuff 0x401760000000 recvbuff 0x401760000000 count 50331648 datatype 9 op [repeated 5x across cluster] (AsyncVLLMInferenceEngine pid=2062355, ip=10.128.17.224) endbuff 0x4016f9800000 recvbuff 0x4016f9800000 count 16777216 datatype 9 op 0 root 0 comm 0xaaab0645a9d0 [nranks=49] stream (nil) [repeated 5x across cluster] (AsyncVLLMInferenceEngine pid=388391, ip=10.128.34.8) jpbo-051-40:388616:388616 [0] NCCL INFO Broadcast: opCount 300 sendbuff 0x40039dd87600 recvbuff 0x40039dd87600 count 128 dat [repeated 16x across cluster] (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) jpbo-007-21:336331:336331 [0] NCCL INFO Broadcast: opCount 1ea sendbuff 0x401740000000 r [repeated 18x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f0 [nranks=8] stream (nil) [repeated 56x across cluster] (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40138a000000 recvbuff 0x401888000000 count 12582912 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f11f0 [repeated 98x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e192dd9f (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) [rank-0]: Saving optim to /e/data1/datasets/playground/ot-baf/explore-tis-minp/explore-tis-minp/checkpoints/global_step_82/policy/optim_world_size_8_rank_0.pt (AsyncVLLMInferenceEngine pid=335976, ip=10.128.17.53) (EngineCore_DP0 pid=336355) WARNING 06-12 05:49:16 [block_pool.py:435] Failed to reset prefix cache because some blocks (5886) are not freed yet [repeated 47x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:49:21 [serving.py:315] Error in preprocessing prompt inputs [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:49:21 [serving.py:315] Traceback (most recent call last): [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:49:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:49:21 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:49:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 12x across cluster] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:49:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:49:21 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:49:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:49:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:49:21 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:49:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:49:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:49:21 [serving.py:315] return await asyncio.gather( [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:49:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:49:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:49:21 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:49:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:49:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:49:21 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:49:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:49:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:49:21 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:49:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:49:21 [serving.py:315] raise VLLMValidationError( [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=389196, ip=10.128.17.63) ERROR 06-12 05:49:21 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 6x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 6x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 6x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) INFO 06-12 05:49:22 [loggers.py:259] Engine 000: Avg prompt throughput: 32.6 tokens/s, Avg generation throughput: 725.5 tokens/s, Running: 16 reqs, Waiting: 0 reqs, GPU KV cache usage: 70.3%, Prefix cache hit rate: 62.0% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 30x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) [rank-0]: Saving extra_state to /e/data1/datasets/playground/ot-baf/explore-tis-minp/explore-tis-minp/checkpoints/global_step_82/policy/extra_state_world_size_8_rank_0.pt (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) [rank-0]: Checkpoint saved to /e/data1/datasets/playground/ot-baf/explore-tis-minp/explore-tis-minp/checkpoints/global_step_82/policy (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:49:21 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:49:21 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:49:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:49:21 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:49:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:49:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:49:21 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:49:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:49:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:49:21 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:49:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:49:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:49:21 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:49:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:49:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:49:21 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:49:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:49:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:49:21 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:49:21 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:49:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:49:21 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:49:21 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:49:21 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=2080551, ip=10.128.17.222) ERROR 06-12 05:49:21 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) INFO 06-12 05:49:27 [loggers.py:259] Engine 000: Avg prompt throughput: 124.3 tokens/s, Avg generation throughput: 657.3 tokens/s, Running: 15 reqs, Waiting: 0 reqs, GPU KV cache usage: 65.4%, Prefix cache hit rate: 62.1% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 43x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:49:29 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:49:29 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:49:29 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:49:29 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:49:29 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:49:29 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:49:29 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:49:29 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:49:29 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:49:29 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:49:29 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:49:29 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:49:29 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:49:29 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:49:29 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:49:29 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:49:29 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:49:29 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:49:29 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:49:29 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:49:29 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:49:29 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:49:29 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=593057, ip=10.128.34.1) ERROR 06-12 05:49:29 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:49:29 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) INFO 06-12 05:49:32 [loggers.py:259] Engine 000: Avg prompt throughput: 69.2 tokens/s, Avg generation throughput: 634.4 tokens/s, Running: 13 reqs, Waiting: 0 reqs, GPU KV cache usage: 55.9%, Prefix cache hit rate: 62.3% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) Provider List: https://docs.litellm.ai/docs/providers [repeated 33x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:49:29 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:49:29 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:49:29 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:49:29 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:49:29 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:49:29 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:49:29 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:49:29 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:49:29 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:49:29 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:49:29 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:49:29 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:49:29 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:49:29 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:49:29 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:49:29 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:49:29 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:49:29 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:49:29 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:49:29 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:49:29 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:49:29 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:49:29 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=336105, ip=10.128.17.53) ERROR 06-12 05:49:29 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:49:37 [serving.py:315] Error in preprocessing prompt inputs (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:49:37 [serving.py:315] Traceback (most recent call last): (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:49:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:49:37 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:49:37 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:49:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:49:37 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:49:37 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:49:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:49:37 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:49:37 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:49:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:49:37 [serving.py:315] return await asyncio.gather( (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:49:37 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:49:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:49:37 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:49:37 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:49:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:49:37 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:49:37 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:49:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:49:37 [serving.py:315] tokens = validator(tokenizer, tokens) (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:49:37 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:49:37 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:49:37 [serving.py:315] raise VLLMValidationError( (AsyncVLLMInferenceEngine pid=631712, ip=10.128.17.34) ERROR 06-12 05:49:37 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. (RolloutCoordinator pid=309396, ip=10.128.17.55) (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) INFO 06-12 05:49:37 [loggers.py:259] Engine 000: Avg prompt throughput: 86.7 tokens/s, Avg generation throughput: 565.7 tokens/s, Running: 12 reqs, Waiting: 0 reqs, GPU KV cache usage: 53.5%, Prefix cache hit rate: 62.3% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 35x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) [rank-0]: Saving model to /e/data1/datasets/playground/ot-baf/explore-tis-minp/explore-tis-minp/checkpoints/global_step_83/policy/model_world_size_8_rank_0.pt (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) [rank-0]: Saving optim to /e/data1/datasets/playground/ot-baf/explore-tis-minp/explore-tis-minp/checkpoints/global_step_83/policy/optim_world_size_8_rank_0.pt (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:49:41 [serving.py:315] Error in preprocessing prompt inputs [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:49:41 [serving.py:315] Traceback (most recent call last): [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:49:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:49:41 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:49:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 8x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:49:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:49:41 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:49:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:49:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:49:41 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:49:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:49:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:49:41 [serving.py:315] return await asyncio.gather( [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:49:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:49:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:49:41 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:49:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:49:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:49:41 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:49:41 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:49:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:49:41 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:49:41 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:49:41 [serving.py:315] raise VLLMValidationError( [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=336106, ip=10.128.17.53) ERROR 06-12 05:49:41 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 4x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 6x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 6x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) INFO 06-12 05:49:42 [loggers.py:259] Engine 000: Avg prompt throughput: 96.0 tokens/s, Avg generation throughput: 546.5 tokens/s, Running: 13 reqs, Waiting: 0 reqs, GPU KV cache usage: 50.6%, Prefix cache hit rate: 62.3% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 15x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) [rank-0]: Saving extra_state to /e/data1/datasets/playground/ot-baf/explore-tis-minp/explore-tis-minp/checkpoints/global_step_83/policy/extra_state_world_size_8_rank_0.pt (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) [rank-0]: Checkpoint saved to /e/data1/datasets/playground/ot-baf/explore-tis-minp/explore-tis-minp/checkpoints/global_step_83/policy (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:49:46 [serving.py:315] Error in preprocessing prompt inputs [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:49:46 [serving.py:315] Traceback (most recent call last): [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:49:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:49:46 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:49:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 4x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:49:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:49:46 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:49:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:49:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:49:46 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:49:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:49:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:49:46 [serving.py:315] return await asyncio.gather( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:49:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:49:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:49:46 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:49:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:49:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:49:46 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:49:46 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:49:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:49:46 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:49:46 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:49:46 [serving.py:315] raise VLLMValidationError( [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386520, ip=10.128.34.6) ERROR 06-12 05:49:46 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 2x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 3x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 3x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) INFO 06-12 05:49:47 [loggers.py:259] Engine 000: Avg prompt throughput: 78.7 tokens/s, Avg generation throughput: 613.3 tokens/s, Running: 15 reqs, Waiting: 0 reqs, GPU KV cache usage: 66.2%, Prefix cache hit rate: 62.7% [repeated 48x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) [rank-0]: Created output directory: /e/data1/datasets/playground/ot-baf/explore-tis-minp/explore-tis-minp/exports/global_step_83/policy (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) [rank-0]: Detected FSDP version: 2 (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:49:47] OK: 430 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=336583, ip=10.128.17.53) [fd-monitor] [05:49:47] OK: RSS 3.62 GiB | node mem 382.6/858.0 GiB used (44.6%), avail 475.3 GiB (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487912) 5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 661 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4018cfc00000 recvbuff 0x40151a808000 count 12582912 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 662 sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4018d1400000 recvbuff 0x401514808000 count 12582912 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 663 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400e9fa85e00 recvbuff 0x400e9fa89000 count 1024 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 664 sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400e9fa86600 recvbuff 0x400e9fa8b000 count 1024 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 665 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4018cd800000 recvbuff 0x401514808000 count 4194304 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 666 sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4018c8400000 recvbuff 0x401389180000 count 1048576 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 667 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4018c8600000 recvbuff 0x401389980000 count 1048576 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 668 sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4018d2c00000 recvbuff 0x401514808000 count 4194304 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 669 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400e9fa86e00 recvbuff 0x400e9fa88c00 count 32 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 66a sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400e9fa87000 recvbuff 0x400e9fa88e00 count 32 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 66b sendbuff 0x400e9fa89000 recvbuff 0x400e9fa89000 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4018d4000000 recvbuff 0x401514808000 count 12582912 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 66c sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4018d5800000 recvbuff 0x40151a808000 count 12582912 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 66d sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4018d7000000 recvbuff 0x401514808000 count 12582912 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 66e sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400e9fa87200 recvbuff 0x400e9fa89000 count 1024 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 66f sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400e9fa87a00 recvbuff 0x400e9fa8b000 count 1024 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 670 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400e9fa88200 recvbuff 0x400e9fa89000 count 1024 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 671 sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4018e0000000 recvbuff 0x402220000000 count 155582464 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 672 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 673 sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 674 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 675 sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 676 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllReduce: opCount 677 sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 1 datatype 7 op 0 root 0 comm 0x400dfa6b2560 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40164c000000 recvbuff 0x402220000000 count 155582464 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401645400000 recvbuff 0x401514808000 count 4194304 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401645c00000 recvbuff 0x401389180000 cou (FSDPPolicyWorkerBase pid=487914) ream 0x400e613f6bd0 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e89000 recvbuff 0x400eb9e89000 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e612e3300 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487914) j (FSDPPolicyWorkerBase pid=487912) nt 1048576 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401832c (FSDPPolicyWorkerBase pid=487913) 0 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400eb9e39200 recvbuff 0x400eb9e88e00 count 1024 datatype 9 o (FSDPPolicyWorkerBase pid=487915) x400e193ee700 (FSDPPolicyWorkerBase pid=487915) jpbo-00 (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401345c00000 recvbuff 0x4018a2000000 count 12582912 datatype 9 op 0 root 0 comm 0x400e616ca9c0 (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 23x across cluster] (FSDPPolicyWorkerBase pid=487912) 00000 recvbuff 0x401514808000 count 12582912 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400e9fa72600 recvbuff 0x400e9fa88e00 count 1024 dat (FSDPPolicyWorkerBase pid=487913) p 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f11f0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400eb9e42e00 recvbuff 0x (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:488176 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401348c00000 recvbuff 0x4018d30f6000 count 12582912 datatype 9 op 0 root 0 comm 0x400e196c2490 [nran (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487 (FSDPPolicyWorkerBase pid=487912) atype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400e9fa7bc00 recv (FSDPPolicyWorkerBase pid=487913) 400eb9e88e00 count 32 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f11f0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sen (FSDPPolicyWorkerBase pid=487915) jpbo-006-45:487915:48 (FSDPPolicyWorkerBase pid=487913) dbuff 0x400eb9e4c600 recvbuff 0x400eb9e88e00 count 1024 datatype 9 op 0 root 0 comm 0x400e196c3c90 [nranks=4] stream 0x400e193f11f0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401361800000 recvbuff 0x401882000000 count 12582912 datatype 9 op 0 root 0 c (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) uce: opCount 667 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 1 datatype 7 op 0 root 0 comm 0x400e426b3c30 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 313 sendbuff 0x4014e2000000 recvbuff 0x4014e2000000 count 4194304 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 314 sendbuff 0x4014e2800000 recvbuff 0x4014e2800000 count 16777216 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 315 sendbuff 0x400e9fa88c00 recvbuff 0x400e9fa88c00 count 128 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 316 sendbuff 0x400e9fa88e00 recvbuff 0x400e9fa88e00 count 128 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 317 sendbuff 0x4014e2000000 recvbuff 0x4014e2000000 count 50331648 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 318 sendbuff 0x4014e8000000 recvbuff 0x4014e8000000 count 50331648 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 319 sendbuff 0x4014e2000000 recvbuff 0x4014e2000000 count 50331648 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 31a sendbuff 0x400e9fa89000 recvbuff 0x400e9fa89000 count 4096 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 31b sendbuff 0x400e9fa8b000 recvbuff 0x400e9fa8b000 count 4096 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 31c sendbuff 0x400e9fa89000 recvbuff 0x400e9fa89000 count 4096 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:333163 [0] NCCL INFO Broadcast: opCount 31d sendbuff 0x401aa0000000 recvbuff 0x401aa0000000 count 622329856 datatype 9 op 0 root 0 comm 0x400fa44ecc80 [nranks=49] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 0 count 1 datatype 7 op 0 root 0 comm 0x400df12d6240 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 send (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 s (FSDPPolicyWorkerBase pid=487914) jpbo-006-45:487914:488174 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4013b1400000 recvbuff 0x4018a2000000 count 12582912 datatype 9 op 0 root 0 comm 0x400e616ca9c0 [nranks=4] stre (FSDPPolicyWorkerBase pid=487912) buff 0x400e9fa88e00 count 32 datatype 9 op 0 root 0 comm 0x400dfaa9d5a0 [nranks=4] stream 0x400dfa7c67e0 (FSDPPolicyWorkerBase pid=487912) jpbo-006-45:487912:488175 [0] NCCL INFO AllGather: opCo (FSDPPolicyWorkerBase pid=487913) omm 0x400e196c3c90 [nranks=4] stream 0x400e193f11f0 (FSDPPolicyWorkerBase pid=487913) jpbo-006-45:487913:488177 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88c00 recvbuff 0x400eb9e88c00 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 427c7e80 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) buff 0x4012b3e00000 recvbuff 0x401862000000 count 12582912 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df13e9a90 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40103e400000 recvbuff 0x4018a7c6e000 count 4194304 datatype 9 op 0 r (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) endbuff 0x4012f3e00000 recvbuff 0x4018b3182000 count 12582912 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9bc0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40105c600000 recvbuff 0x4018b3182000 count 4194304 datatype 9 op (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 400e193e87d0 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) oot 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df13e9a90 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9bc0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:3330 (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllG (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) ather: opCount 0 sendbuff 0x40186b800000 recvbuff 0x4014e2000000 count 12582912 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7e80 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401881800000 recvbuff 0x4014e2000000 count 4194 (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) INFO AllGather: opCount 0 sendbuff 0x400eb9e3cc00 recvbuff 0x400eb9e88e00 count 32 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e87d0 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) 304 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7e80 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo- (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) 3e9a90 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) e153f9bc0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendb (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) uff 0x400eb9e45600 recvbuff 0x400eb9e88e00 count 32 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9bc0 (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpbo-007-11:332821:333039 [0] NCCL IN (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400eb9e4f800 recvbuff 0x400eb9e88e00 count 32 datatype 9 op 0 root 0 comm 0x400e196bcf10 [n (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGathe (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) FO AllGather: opCount 0 sendbuff 0x400e79e4ee00 recvbuff 0x400e79e88e00 count 1024 datatype 9 op 0 root 0 comm 0x400df16bdc50 [nranks=4] stream 0x400df13e9a90 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) r: opCount 0 sendbuff 0x400e9fa80a00 recvbuff 0x400e9fa88e00 count 32 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nranks=4] stream 0x400e427c7e80 (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4018eb400000 recvbuff 0x401389580000 count 1048576 datatype 9 (FSDPPolicyWorkerBase pid=332821, ip=10.128.17.43) jpb (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40137ec00000 recvbuff 0x4018c7da2000 count 4194304 dataty (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:51 [serving.py:315] Error in preprocessing prompt inputs [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:51 [serving.py:315] Traceback (most recent call last): [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/chat_completion/serving.py", line 299, in render_chat_request [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:51 [serving.py:315] conversation, engine_prompts = await self._preprocess_chat( [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:51 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 6x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/entrypoints/openai/engine/serving.py", line 1018, in _preprocess_chat [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:51 [serving.py:315] (conversation,), (engine_prompt,) = await renderer.render_chat_async( [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:51 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 381, in render_chat_async [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:51 [serving.py:315] tok_prompts = await self.tokenize_prompts_async(dict_prompts, tok_params) [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:51 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 278, in tokenize_prompts_async [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:51 [serving.py:315] return await asyncio.gather( [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:51 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/protocol.py", line 271, in tokenize_prompt_async [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:51 [serving.py:315] return params.apply_post_tokenization(self.tokenizer, prompt) # type: ignore[arg-type] [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:51 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 373, in apply_post_tokenization [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:51 [serving.py:315] prompt["prompt_token_ids"] = self._validate_tokens( # type: ignore[typeddict-unknown-key] [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:51 [serving.py:315] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 357, in _validate_tokens [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:51 [serving.py:315] tokens = validator(tokenizer, tokens) [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:51 [serving.py:315] File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/renderers/params.py", line 337, in _token_len_check [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:51 [serving.py:315] raise VLLMValidationError( [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=308849, ip=10.128.17.55) ERROR 06-12 05:49:51 [serving.py:315] vllm.exceptions.VLLMValidationError: You passed 32769 input tokens and requested 0 output tokens. However, the model's context length is only 32768 tokens, resulting in a maximum input length of 32768 tokens. Please reduce the length of the input prompt. (parameter=input_tokens, value=32769) [repeated 3x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 3x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 3x across cluster] (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) INFO 06-12 05:49:52 [loggers.py:259] Engine 000: Avg prompt throughput: 938.4 tokens/s, Avg generation throughput: 674.4 tokens/s, Running: 16 reqs, Waiting: 0 reqs, GPU KV cache usage: 69.7%, Prefix cache hit rate: 62.7% [repeated 48x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) ranks=4] stream 0x400e193e87d0 [repeated 8x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllReduce: opCount 18f sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e192d5230 [nranks=8] stream (nil) [repeated 43x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400eb9e58e00 recvbuff 0x400eb9e88e00 count 1024 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e87d0 [repeated 3114x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllReduce: opCount fe sendbuff 0x400eb9e88e00 recvbuff 0x400eb9e88e00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) [repeated 54x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) j (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401343000 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x40135f600000 recvbuff 0x401704c00000 count 1048576 datatype 9 op 0 root 0 comm  [repeated 4x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) Provider List: https://docs.litellm.ai/docs/providers [repeated 27x across cluster] (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) 000 recvbuff 0x401882000000 count 12582912 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e87d0 [repeated 5x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x400eb9e3ba00 recvbuff 0x400eb9 [repeated 2x across cluster] (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) jpbo-007-11:332746:332816 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x4018c2c00000 recvbuff 0x4014e2000000 count 12582912 datatype 9 op 0 root 0 comm 0x400e42a9f600 [nran (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) e88e00 count 1024 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9bc0 [repeated 3x across cluster] (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 8c00 count 1 datatype 7 op 0 root 0 comm 0x400e152e6370 [nranks=8] stream (nil) (FSDPPolicyWorkerBase pid=332746, ip=10.128.17.43) [rank-0]: Successfully saved model to /e/data1/datasets/playground/ot-baf/explore-tis-minp/explore-tis-minp/exports/global_step_83/policy (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) jpbo-007-11:332823:333047 [0] NCCL INFO AllG (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL (FSDPPolicyWorkerBase pid=332823, ip=10.128.17.43) ather: opCount 0 sendbuff 0x40136d000000 recvbuff 0x4018c7da2000 count 4194304 datatype 9 op 0 root 0 comm 0x400e196bcf10 [nranks=4] stream 0x400e193e87d0 (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) jpbo-007-11:332822:333029 [0] NCCL INFO AllGather: opCount 0 sendbuff 0x401395800000 recvbuff 0x4018b3182000 count 125 [repeated 3x across cluster] (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) INFO AllGather: opCount 0 sendbuff 0x400eb9e4ee00 recvbuff 0x400eb9e88e00 count 1024 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9bc0 (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (FSDPPolicyWorkerBase pid=332822, ip=10.128.17.43) 82912 datatype 9 op 0 root 0 comm 0x400e156cd8d0 [nranks=4] stream 0x400e153f9bc0 [repeated 3x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 148x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 148x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=386650, ip=10.128.34.6) INFO 06-12 05:49:57 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 696.5 tokens/s, Running: 14 reqs, Waiting: 0 reqs, GPU KV cache usage: 69.9%, Prefix cache hit rate: 62.7% [repeated 48x across cluster] (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:49:58] OK: 346 / 131,072 FDs open (0.3% of soft limit, hard limit: 131,072) (RolloutCoordinator pid=309396, ip=10.128.17.55) [fd-monitor] [05:49:58] OK: RSS 3.20 GiB | node mem 388.0/858.0 GiB used (45.2%), avail 470.0 GiB (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 3x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=336583, ip=10.128.17.53) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=593552, ip=10.128.34.1) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (RolloutCoordinator pid=2062840, ip=10.128.17.224) (AsyncVLLMInferenceEngine pid=592922, ip=10.128.34.1) ERROR 06-12 05:50:02 [core_client.py:616] Engine core proc EngineCore_DP0 died unexpectedly, shutting down client. (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) (RolloutCoordinator pid=309396, ip=10.128.17.55) Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new [repeated 140x across cluster] (RolloutCoordinator pid=309396, ip=10.128.17.55) LiteLLM.Info: If you need to debug this error, use `litellm._turn_on_debug()'. [repeated 140x across cluster] (AsyncVLLMInferenceEngine pid=2073546, ip=10.128.17.223) INFO 06-12 05:50:02 [loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 277.8 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 26.1%, Prefix cache hit rate: 93.9% [repeated 41x across cluster] (RolloutCoordinator pid=336583, ip=10.128.17.53) Provider List: https://docs.litellm.ai/docs/providers [repeated 2x across cluster] (AsyncVLLMInferenceEngine pid=386651, ip=10.128.34.6) ERROR 06-12 05:50:02 [core_client.py:616] Engine core proc EngineCore_DP0 died unexpectedly, shutting down client. [repeated 34x across cluster] Stopping Ray cluster... Ray cluster stopped [RLJobRunner] Launching trace upload (training exit code: 0): repo_id: DCAgent/explore-tis-minp job_dir: /e/data1/datasets/playground/ot-baf/explore-tis-minp/explore-tis-minp episodes: last log: /e/data1/datasets/playground/ot-baf/explore-tis-minp/logs/explore-tis-minp_trace_upload.log [RLJobRunner] Waiting for trace upload to complete... [RLJobRunner] Trace upload failed with exit code 1. Preserving Ray logs to /e/data1/datasets/playground/ot-baf/explore-tis-minp/ray_logs/ Collecting Ray logs from worker jpbo-007-02... Collecting Ray logs from worker jpbo-007-11... Collecting Ray logs from worker jpbo-007-21... Collecting Ray logs from worker jpbo-007-23... Collecting Ray logs from worker jpbo-007-31... Collecting Ray logs from worker jpbo-010-46... Collecting Ray logs from worker jpbo-010-47... Collecting Ray logs from worker jpbo-010-48... Collecting Ray logs from worker jpbo-051-33... Collecting Ray logs from worker jpbo-051-34... Collecting Ray logs from worker jpbo-051-37... Collecting Ray logs from worker jpbo-051-38... Collecting Ray logs from worker jpbo-051-40... Ray log preservation complete