The following have been reloaded with a version change: 1) GCCcore/.14.3.0 => GCCcore/14.3.0 Lmod is automatically replacing "GCC/14.3.0" with "nvidia-compilers/25.9-CUDA-13". Deactivating conda environment: /e/scratch/jureap59/feuer1/miniforge3/envs/otagent Activating RL environment: /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl Python executable: /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/bin/python Python path check: /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/bin/python [ray] RAY_TMPDIR=/tmp/ray/ray_629884 [triton_cache] Triton cache: /tmp/triton_cache_feuer1_629884 [triton_cache] TorchInductor cache: /tmp/torchinductor_cache_feuer1_629884 [proxy] ✓ Found proxychains binary at /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 [proxy] Setting up SSH tunnel to jpbl-s01-01 [proxy] SSH key: /e/home/jusers/feuer1/jupiter/.ssh/authorized_keys/id_ed25519_jsc [proxy] Tunnel port: 7003 [proxy] Node IP: 10.128.33.2 (workers will connect here) [proxy] ✓ SSH tunnel started successfully [proxy] ✓ Generated proxychains config at /e/home/jusers/feuer1/jupiter/.proxychains/proxychains_629884.conf [proxy] - Internal traffic (10.x.x.x, 172.x.x.x, 169.254.x.x) → DIRECT [proxy] - External traffic (internet) → PROXY via tunnel [proxy] ✓ Daytona timeout settings configured [proxy] Testing proxy connectivity... [proxychains] config file found: /e/home/jusers/feuer1/jupiter/.proxychains/proxychains_629884.conf [proxychains] preloading /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/lib/libproxychains4.so [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e [proxy] ✓ Proxy connectivity test passed (huggingface.co reachable via wrapped binary) [proxy] ⚠ Tunnel not accessible at 10.128.33.2:7003 (workers may fail) [proxy] ✓ Proxy setup complete (using wrapped binary for Ray workers) [container_runtime] Using cloud backend: daytona (no local container setup) === Universal RL Training Runner === Config: /e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqnorm/configs/ablation-pymethods2test-seqnorm_rl_config.json Working directory: /e/scratch/jureap59/feuer1/OpenThoughts-Agent Python: /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/bin/python Python version: Python 3.12.12 UV_USE_IO_URING: 0 Proxy: DISABLED (direct internet or not configured) ======================================== === RLJobRunner: ablation-pymethods2test-seqnorm === [wandb_utils] Fixing permissions on: /e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqnorm/wandb [wandb_utils] WandB directory ready: /e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqnorm/wandb HF_TOKEN=****pDbg HF_HUB_CACHE=/e/data1/datasets/playground/ot-baf/hf_hub SUPABASE_URL=https://rpzmyuapoqilpghynmza.s... (direct Supabase config) Environment configured: TENSOR_PARALLEL_SIZE=1 NUM_INFERENCE_ENGINES=56 POLICY_NUM_NODES=14 WANDB_DIR=/e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqnorm/wandb Starting Ray cluster with 14 nodes, 4 GPUs/node Cleaning up existing Ray instances... === Starting Ray Cluster === Nodes: 14 GPUs per node: 4 CPUs per node: 288 Head node: jpbo-046-18 (10.128.33.2) Ray port: 6379 ============================ Starting Ray head on jpbo-046-18 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_head_jpbo-046-18.log)... Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.2 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-046-18 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --head --node-ip-address=10.128.33.2 --port=6379 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 Started Ray head on jpbo-046-18 Starting Ray worker on jpbo-046-19 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-046-19.log)... Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.3 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-046-19 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.33.2:6379 --node-ip-address=10.128.33.3 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 Started Ray worker 1 on jpbo-046-19 Starting Ray worker on jpbo-046-21 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-046-21.log)... Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.5 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-046-21 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.33.2:6379 --node-ip-address=10.128.33.5 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 Started Ray worker 2 on jpbo-046-21 Starting Ray worker on jpbo-046-22 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-046-22.log)... Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.6 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-046-22 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.33.2:6379 --node-ip-address=10.128.33.6 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 Started Ray worker 3 on jpbo-046-22 Starting Ray worker on jpbo-046-23 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-046-23.log)... Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.7 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-046-23 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.33.2:6379 --node-ip-address=10.128.33.7 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 Started Ray worker 4 on jpbo-046-23 Starting Ray worker on jpbo-046-24 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-046-24.log)... Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.8 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-046-24 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.33.2:6379 --node-ip-address=10.128.33.8 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 Started Ray worker 5 on jpbo-046-24 Starting Ray worker on jpbo-046-25 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-046-25.log)... Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.9 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-046-25 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.33.2:6379 --node-ip-address=10.128.33.9 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 Started Ray worker 6 on jpbo-046-25 Starting Ray worker on jpbo-046-26 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-046-26.log)... Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.10 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-046-26 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.33.2:6379 --node-ip-address=10.128.33.10 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 Started Ray worker 7 on jpbo-046-26 Starting Ray worker on jpbo-046-29 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-046-29.log)... Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.13 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-046-29 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.33.2:6379 --node-ip-address=10.128.33.13 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 Started Ray worker 8 on jpbo-046-29 Starting Ray worker on jpbo-046-31 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-046-31.log)... Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.15 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-046-31 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.33.2:6379 --node-ip-address=10.128.33.15 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 Started Ray worker 9 on jpbo-046-31 Starting Ray worker on jpbo-047-40 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-047-40.log)... Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.72 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-047-40 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.33.2:6379 --node-ip-address=10.128.33.72 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 Started Ray worker 10 on jpbo-047-40 Starting Ray worker on jpbo-047-41 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-047-41.log)... Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.73 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-047-41 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.33.2:6379 --node-ip-address=10.128.33.73 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 Started Ray worker 11 on jpbo-047-41 Starting Ray worker on jpbo-047-42 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-047-42.log)... Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.74 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-047-42 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.33.2:6379 --node-ip-address=10.128.33.74 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 Started Ray worker 12 on jpbo-047-42 Starting Ray worker on jpbo-047-44 (logging to /e/data1/datasets/playground/ot-baf/experiments/_ray_logs/ray_worker_jpbo-047-44.log)... Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.76 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-047-44 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.33.2:6379 --node-ip-address=10.128.33.76 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 Started Ray worker 13 on jpbo-047-44 Waiting for cluster (56 GPUs, 14 nodes)... Connecting to Ray at 10.128.33.2:6379 (expecting 14 nodes, 56.0 GPUs) Ray connection established, polling for resources... [Ray wait] nodes=14/14 GPUs=56.0/56.0 resources={'accelerator_type:GH200': 14.0, 'memory': 10140326887424.0, 'CPU': 4032.0, 'GPU': 56.0, 'object_store_memory': 601295421440.0, 'node:10.128.33.9': 1.0, 'node:__internal_head__': 1.0, 'node:10.128.33.2': 1.0, 'node:10.128.33.73': 1.0, 'node:10.128.33.15': 1.0, 'node:10.128.33.10': 1.0, 'node:10.128.33.8': 1.0, 'node:10.128.33.74': 1.0, 'node:10.128.33.5': 1.0, 'node:10.128.33.76': 1.0, 'node:10.128.33.3': 1.0, 'node:10.128.33.13': 1.0, 'node:10.128.33.72': 1.0, 'node:10.128.33.6': 1.0, 'node:10.128.33.7': 1.0} ✓ Ray cluster ready === Ray Cluster Ready === Address: 10.128.33.2:6379 Total GPUs: 56 ========================= Ray cluster ready at 10.128.33.2:6379 Total GPUs available: 56 [RLJobRunner] Pinggy check: url=False, token=False, needs_tunnel=False (agent=terminus-2, env=daytona) [RLJobRunner] No Pinggy tunnel needed, using local vLLM Running SkyRL: Python: /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/bin/python Entrypoint: examples.terminal_bench.entrypoints.main_tbench Args: 120 Hydra arguments Working dir: /e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train Using proxychains binary: /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 Executing command with srun: /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f $PROXYCHAINS_CONF_FILE /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/bin/python -m examples.terminal_bench.entrypoints.main_tbench +terminal_bench_config=terminal_bench trainer.strategy=fsdp2 trainer.algorithm.advantage_estimator=rloo_n trainer.algorithm.use_kl_loss=false trainer.algorithm.kl_loss_coef=0.0 trainer.algorithm.eps_clip_low=0.2 trainer.algorithm.eps_clip_high=0.05 trainer.algorithm.loss_reduction=seq_mean_token_sum_norm_global trainer.epochs=2 trainer.max_steps=80 trainer.update_epochs_per_batch=1 trainer.train_batch_size=64 trainer.policy_mini_batch_size=64 trainer.eval_batch_size=64 trainer.micro_forward_batch_size_per_gpu=4 trainer.micro_train_batch_size_per_gpu=1 trainer.max_prompt_length=999999 trainer.eval_interval=999999 trainer.eval_before_train=false trainer.ckpt_interval=2 trainer.resume_mode=latest trainer.hf_save_interval=5 ++trainer.hf_hub_repo_id=laion/ablation-pymethods2test-seqnorm ++trainer.hf_hub_private=false ++trainer.hf_hub_revision=main ++trainer.enable_db_registration=false trainer.project_name=OpenThoughts-Agent trainer.log_level=INFO trainer.tracker_commit_each_step=true trainer.logger=console trainer.run_name=ablation-pymethods2test-seqnorm trainer.ckpt_path=/e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqnorm/ablation-pymethods2test-seqnorm/checkpoints trainer.export_path=/e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqnorm/ablation-pymethods2test-seqnorm/exports trainer.policy.optimizer_config.lr=8e-6 trainer.policy.optimizer_config.weight_decay=0.0 trainer.policy.optimizer_config.adam_betas=[0.9,0.999] trainer.policy.optimizer_config.max_grad_norm=0.9 trainer.policy.fsdp_config.cpu_offload=false trainer.policy.fsdp_config.reshard_after_forward=true trainer.policy.fsdp_config.fsdp_size=4 trainer.policy.model.path=/e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6 trainer.ref.fsdp_config.cpu_offload=false trainer.ref.fsdp_config.reshard_after_forward=true trainer.ref.fsdp_config.fsdp_size=4 trainer.placement.colocate_all=false trainer.placement.policy_num_nodes=2 trainer.placement.ref_num_nodes=2 trainer.placement.policy_num_gpus_per_node=4 trainer.placement.ref_num_gpus_per_node=4 trainer.fully_async.max_staleness_steps=16 trainer.fully_async.num_parallel_generation_workers=338 generator.backend=vllm generator.timeout_multiplier=1.0 generator.model_dtype=bfloat16 generator.inference_engine_tensor_parallel_size=1 generator.num_inference_engines=48 generator.n_samples_per_prompt=8 generator.eval_n_samples_per_prompt=8 generator.gpu_memory_utilization=0.75 generator.max_num_seqs=24 generator.max_num_batched_tokens=65536 generator.enable_prefix_caching=true generator.enable_chunked_prefill=true generator.run_engines_locally=true generator.weight_sync_backend=nccl generator.async_engine=true generator.batched=false generator.enable_http_endpoint=true generator.enable_ray_prometheus_stats=false generator.vllm_stats_interval=1 generator.append_eos_token_after_stop_str_in_multi_turn=true generator.max_turns=999999 generator.sampling_params.max_generate_length=4096 generator.sampling_params.temperature=0.7 generator.sampling_params.top_p=0.95 generator.sampling_params.top_k=20 ++generator.engine_init_kwargs.max_model_len=32768 ++generator.engine_init_kwargs.custom_chat_template_chat_completion_path=chat_templates/qwen3_thinking_acc.jinja2 ++generator.engine_init_kwargs.served_model_name=0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6 data.train_data=["/e/scratch/jureap59/feuer1/tasks/exp_rpt_pymethods2test-large"] data.val_data=[] +terminal_bench_config.trials_dir=/e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqnorm/ablation-pymethods2test-seqnorm/trace_jobs +terminal_bench_config.harbor.name=terminus-2 +terminal_bench_config.harbor.max_episodes=999999 +terminal_bench_config.harbor.enable_summarize=false +terminal_bench_config.harbor.store_all_messages=true +terminal_bench_config.harbor.trajectory_config.raw_content=true +terminal_bench_config.harbor.enable_episode_logging=false +terminal_bench_config.harbor.record_terminal_session=false +terminal_bench_config.harbor.enable_pane_logging=false +terminal_bench_config.harbor.strict_json_parser=true +terminal_bench_config.harbor.interleaved_thinking=true +terminal_bench_config.harbor.extra_body.chat_template_kwargs.enable_thinking=true +terminal_bench_config.harbor.override_timeout_sec=900 +terminal_bench_config.harbor.override_cpus=1 +terminal_bench_config.harbor.override_memory_mb=2048 +terminal_bench_config.harbor.override_storage_mb=2048 +terminal_bench_config.harbor.auto_snapshot=true +terminal_bench_config.harbor.verifier_override_timeout_sec=120 +terminal_bench_config.harbor.max_retries=3 +terminal_bench_config.harbor.min_wait_sec=60.0 +terminal_bench_config.harbor.max_wait_sec=600.0 +terminal_bench_config.harbor.wait_multiplier=2.0 +terminal_bench_config.harbor.exclude_exceptions=["VerifierTimeoutError","VerifierRuntimeError","RewardFileNotFoundError","RewardFileEmptyError","VerifierOutputParseError"] +terminal_bench_config.harbor.n_concurrent_trials=675 +terminal_bench_config.harbor.log_level=INFO +terminal_bench_config.harbor.enable_reward_shaping=false +terminal_bench_config.harbor.enable_error_classification=true +terminal_bench_config.harbor.mask_exceptions=["DaytonaError","EnvironmentStartTimeoutError","NetworkError","ConnectionError","RewardFileNotFoundError","RewardFileEmptyError","AgentEnvironmentTimeoutError","ContextLengthExceededError"] +terminal_bench_config.harbor.default_error_treatment=zero +terminal_bench_config.harbor.passthrough_exceptions=["AgentTimeoutError"] +terminal_bench_config.harbor.zero_exceptions=[] +terminal_bench_config.model_info.max_input_tokens=32000 +terminal_bench_config.model_info.max_output_tokens=4096 +terminal_bench_config.archiving.enabled=false +terminal_bench_config.trace_upload.enabled=true +terminal_bench_config.trace_upload.repo_org=DCAgent +terminal_bench_config.trace_upload.episodes=last +terminal_bench_config.trace_upload.dataset_type=SFT +terminal_bench_config.trace_upload.cleanup=true [proxychains] config file found: /e/home/jusers/feuer1/jupiter/.proxychains/proxychains_629884.conf [proxychains] preloading /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/lib/libproxychains4.so [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e Error executing job with overrides: ['+terminal_bench_config=terminal_bench', 'trainer.strategy=fsdp2', 'trainer.algorithm.advantage_estimator=rloo_n', 'trainer.algorithm.use_kl_loss=false', 'trainer.algorithm.kl_loss_coef=0.0', 'trainer.algorithm.eps_clip_low=0.2', 'trainer.algorithm.eps_clip_high=0.05', 'trainer.algorithm.loss_reduction=seq_mean_token_sum_norm_global', 'trainer.epochs=2', 'trainer.max_steps=80', 'trainer.update_epochs_per_batch=1', 'trainer.train_batch_size=64', 'trainer.policy_mini_batch_size=64', 'trainer.eval_batch_size=64', 'trainer.micro_forward_batch_size_per_gpu=4', 'trainer.micro_train_batch_size_per_gpu=1', 'trainer.max_prompt_length=999999', 'trainer.eval_interval=999999', 'trainer.eval_before_train=false', 'trainer.ckpt_interval=2', 'trainer.resume_mode=latest', 'trainer.hf_save_interval=5', '++trainer.hf_hub_repo_id=laion/ablation-pymethods2test-seqnorm', '++trainer.hf_hub_private=false', '++trainer.hf_hub_revision=main', '++trainer.enable_db_registration=false', 'trainer.project_name=OpenThoughts-Agent', 'trainer.log_level=INFO', 'trainer.tracker_commit_each_step=true', 'trainer.logger=console', 'trainer.run_name=ablation-pymethods2test-seqnorm', 'trainer.ckpt_path=/e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqnorm/ablation-pymethods2test-seqnorm/checkpoints', 'trainer.export_path=/e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqnorm/ablation-pymethods2test-seqnorm/exports', 'trainer.policy.optimizer_config.lr=8e-6', 'trainer.policy.optimizer_config.weight_decay=0.0', 'trainer.policy.optimizer_config.adam_betas=[0.9,0.999]', 'trainer.policy.optimizer_config.max_grad_norm=0.9', 'trainer.policy.fsdp_config.cpu_offload=false', 'trainer.policy.fsdp_config.reshard_after_forward=true', 'trainer.policy.fsdp_config.fsdp_size=4', 'trainer.policy.model.path=/e/data1/datasets/playground/ot-baf/hf_hub/models--laion--GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink/snapshots/0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', 'trainer.ref.fsdp_config.cpu_offload=false', 'trainer.ref.fsdp_config.reshard_after_forward=true', 'trainer.ref.fsdp_config.fsdp_size=4', 'trainer.placement.colocate_all=false', 'trainer.placement.policy_num_nodes=2', 'trainer.placement.ref_num_nodes=2', 'trainer.placement.policy_num_gpus_per_node=4', 'trainer.placement.ref_num_gpus_per_node=4', 'trainer.fully_async.max_staleness_steps=16', 'trainer.fully_async.num_parallel_generation_workers=338', 'generator.backend=vllm', 'generator.timeout_multiplier=1.0', 'generator.model_dtype=bfloat16', 'generator.inference_engine_tensor_parallel_size=1', 'generator.num_inference_engines=48', 'generator.n_samples_per_prompt=8', 'generator.eval_n_samples_per_prompt=8', 'generator.gpu_memory_utilization=0.75', 'generator.max_num_seqs=24', 'generator.max_num_batched_tokens=65536', 'generator.enable_prefix_caching=true', 'generator.enable_chunked_prefill=true', 'generator.run_engines_locally=true', 'generator.weight_sync_backend=nccl', 'generator.async_engine=true', 'generator.batched=false', 'generator.enable_http_endpoint=true', 'generator.enable_ray_prometheus_stats=false', 'generator.vllm_stats_interval=1', 'generator.append_eos_token_after_stop_str_in_multi_turn=true', 'generator.max_turns=999999', 'generator.sampling_params.max_generate_length=4096', 'generator.sampling_params.temperature=0.7', 'generator.sampling_params.top_p=0.95', 'generator.sampling_params.top_k=20', '++generator.engine_init_kwargs.max_model_len=32768', '++generator.engine_init_kwargs.custom_chat_template_chat_completion_path=chat_templates/qwen3_thinking_acc.jinja2', '++generator.engine_init_kwargs.served_model_name=0e3bff0c4e51f6b9ec0713b98b9eec36efb91cc6', 'data.train_data=["/e/scratch/jureap59/feuer1/tasks/exp_rpt_pymethods2test-large"]', 'data.val_data=[]', '+terminal_bench_config.trials_dir=/e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqnorm/ablation-pymethods2test-seqnorm/trace_jobs', '+terminal_bench_config.harbor.name=terminus-2', '+terminal_bench_config.harbor.max_episodes=999999', '+terminal_bench_config.harbor.enable_summarize=false', '+terminal_bench_config.harbor.store_all_messages=true', '+terminal_bench_config.harbor.trajectory_config.raw_content=true', '+terminal_bench_config.harbor.enable_episode_logging=false', '+terminal_bench_config.harbor.record_terminal_session=false', '+terminal_bench_config.harbor.enable_pane_logging=false', '+terminal_bench_config.harbor.strict_json_parser=true', '+terminal_bench_config.harbor.interleaved_thinking=true', '+terminal_bench_config.harbor.extra_body.chat_template_kwargs.enable_thinking=true', '+terminal_bench_config.harbor.override_timeout_sec=900', '+terminal_bench_config.harbor.override_cpus=1', '+terminal_bench_config.harbor.override_memory_mb=2048', '+terminal_bench_config.harbor.override_storage_mb=2048', '+terminal_bench_config.harbor.auto_snapshot=true', '+terminal_bench_config.harbor.verifier_override_timeout_sec=120', '+terminal_bench_config.harbor.max_retries=3', '+terminal_bench_config.harbor.min_wait_sec=60.0', '+terminal_bench_config.harbor.max_wait_sec=600.0', '+terminal_bench_config.harbor.wait_multiplier=2.0', '+terminal_bench_config.harbor.exclude_exceptions=["VerifierTimeoutError","VerifierRuntimeError","RewardFileNotFoundError","RewardFileEmptyError","VerifierOutputParseError"]', '+terminal_bench_config.harbor.n_concurrent_trials=675', '+terminal_bench_config.harbor.log_level=INFO', '+terminal_bench_config.harbor.enable_reward_shaping=false', '+terminal_bench_config.harbor.enable_error_classification=true', '+terminal_bench_config.harbor.mask_exceptions=["DaytonaError","EnvironmentStartTimeoutError","NetworkError","ConnectionError","RewardFileNotFoundError","RewardFileEmptyError","AgentEnvironmentTimeoutError","ContextLengthExceededError"]', '+terminal_bench_config.harbor.default_error_treatment=zero', '+terminal_bench_config.harbor.passthrough_exceptions=["AgentTimeoutError"]', '+terminal_bench_config.harbor.zero_exceptions=[]', '+terminal_bench_config.model_info.max_input_tokens=32000', '+terminal_bench_config.model_info.max_output_tokens=4096', '+terminal_bench_config.archiving.enabled=false', '+terminal_bench_config.trace_upload.enabled=true', '+terminal_bench_config.trace_upload.repo_org=DCAgent', '+terminal_bench_config.trace_upload.episodes=last', '+terminal_bench_config.trace_upload.dataset_type=SFT', '+terminal_bench_config.trace_upload.cleanup=true'] Traceback (most recent call last): File "", line 198, in _run_module_as_main File "", line 88, in _run_code File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train/examples/terminal_bench/entrypoints/main_tbench.py", line 142, in main() File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/hydra/main.py", line 94, in decorated_main _run_hydra( File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/hydra/_internal/utils.py", line 394, in _run_hydra _run_app( File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/hydra/_internal/utils.py", line 457, in _run_app run_and_report( File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/hydra/_internal/utils.py", line 223, in run_and_report raise ex File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/hydra/_internal/utils.py", line 220, in run_and_report return func() ^^^^^^ File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/hydra/_internal/utils.py", line 458, in lambda: hydra.run( ^^^^^^^^^^ File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/hydra/_internal/hydra.py", line 132, in run _ = ret.return_value ^^^^^^^^^^^^^^^^ File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/hydra/core/utils.py", line 260, in return_value raise self._return_value File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/hydra/core/utils.py", line 186, in run_job ret.return_value = task_function(task_cfg) ^^^^^^^^^^^^^^^^^^^^^^^ File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train/examples/terminal_bench/entrypoints/main_tbench.py", line 111, in main validate_cfg(cfg) File "/e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train/skyrl_train/utils/utils.py", line 278, in validate_cfg assert cfg.trainer.algorithm.loss_reduction in ( ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ AssertionError: invalid loss_reduction: seq_mean_token_sum_norm_global. Must be one of `['token_mean', 'sequence_mean', 'seq_mean_token_sum_norm']` Stopping Ray cluster... Ray cluster stopped [RLJobRunner] Crash detected (exit!=0) — preserving Ray logs to /e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqnorm/ablation-pymethods2test-seqnorm/ray_logs BEFORE trace upload (so a wall-clock kill can't lose crash evidence)... Collecting Ray logs from worker jpbo-046-19... Collecting Ray logs from worker jpbo-046-21... Collecting Ray logs from worker jpbo-046-22... Collecting Ray logs from worker jpbo-046-23... Collecting Ray logs from worker jpbo-046-24... Collecting Ray logs from worker jpbo-046-25... Collecting Ray logs from worker jpbo-046-26... Collecting Ray logs from worker jpbo-046-29... Collecting Ray logs from worker jpbo-046-31... Collecting Ray logs from worker jpbo-047-40... Collecting Ray logs from worker jpbo-047-41... Collecting Ray logs from worker jpbo-047-42... Collecting Ray logs from worker jpbo-047-44... [RLJobRunner] Crash-time Ray log preservation complete. [RLJobRunner] No trace_jobs directory found at /e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqnorm/ablation-pymethods2test-seqnorm/trace_jobs, skipping upload. Preserving Ray logs to /e/data1/datasets/playground/ot-baf/ablation-pymethods2test-seqnorm/ray_logs/ Collecting Ray logs from worker jpbo-046-19... Collecting Ray logs from worker jpbo-046-21... Collecting Ray logs from worker jpbo-046-22... Collecting Ray logs from worker jpbo-046-23... Collecting Ray logs from worker jpbo-046-24... Collecting Ray logs from worker jpbo-046-25... Collecting Ray logs from worker jpbo-046-26... Collecting Ray logs from worker jpbo-046-29... Collecting Ray logs from worker jpbo-046-31... Collecting Ray logs from worker jpbo-047-40... Collecting Ray logs from worker jpbo-047-41... Collecting Ray logs from worker jpbo-047-42... Collecting Ray logs from worker jpbo-047-44... Ray log preservation complete