revert: restore a3c45d3b yaml + Q-tiling + remove all OOM hacks
Root cause of 10 consecutive OOM failures: - 'return zeros during profiling' hack → vllm overestimates free memory → allocates 7942 blocks → first real request OOMs - blocks cap 5000 → band-aid that masks profiling bug - gpu-memory-utilization 0.80 → unnecessary reduction from working 0.90 - max-num-seqs 2 → doubles peak activation memory - PYTORCH_CUDA_ALLOC_CONF max_split_size_mb:512 → causes fragmentation Restoringa3c45d3bparameters that actually work: - yaml: max-model-len=131072, gpu-mem=0.90, max-num-seqs=1, batched-tokens=8192 - patch_xformers_sdpa_seq.py: Q-tiling (real memory optimization, not zeros hack) - patch_block_major_worker_capacity.py: no blocks cap, just reserve_block_major - patch_ops.sh: remove all docker-build-time compilation (all .so are prebuilt) Only change froma3c45d3b: BI100_MOE_COREX_TOPK_SOFTMAX=1 (enable corex topk) Kept fixes: - protocol.py extra=allow (recover 180 rejected replay requests) - corex_gdn_chunk_recurrent.so pybind kwargs (prebuilt with fixed signature)
This commit is contained in:
@@ -246,24 +246,6 @@ if source != installed:
|
||||
raise SystemExit("runtime api_server overlay identity mismatch")
|
||||
PY
|
||||
|
||||
build_stage "compiling CCCL CachingDeviceAllocator LD_PRELOAD module"
|
||||
bash ./cccl_preload/build_cccl_preload.sh /workspace/qwen3_6_scripts/cccl_preload || \
|
||||
echo "[WARN] CCCL preload allocator build failed — will use default allocator"
|
||||
|
||||
build_stage "compiling CoreX CUDA extensions (moe_index_combine + gdn_chunk_recurrent)"
|
||||
if [[ -x /usr/local/corex-3.2.3/bin/clang++ ]]; then
|
||||
bash ./build_corex_moe_index_combine.sh "${VLLM_ROOT}" || \
|
||||
echo "[WARN] moe_index_combine build failed — will use PyTorch fallback"
|
||||
bash ./build_corex_gdn_chunk_recurrent.sh "${VLLM_ROOT}" || \
|
||||
echo "[WARN] gdn_chunk_recurrent build failed — will use Python fallback"
|
||||
else
|
||||
echo "[WARN] corex clang++ not found — skipping extension builds"
|
||||
fi
|
||||
|
||||
build_stage "compiling ixformer bridge .so (MoE + Attention + Norm)"
|
||||
bash ./build_ix_bridge.sh "${VLLM_ROOT}" || \
|
||||
echo "[WARN] ix_full_bridge build failed — MoE will use PyTorch fallback"
|
||||
|
||||
build_stage "compiling submission Python sources"
|
||||
find . -path './wheels' -prune -o -name '*.py' -print0 | xargs -0 python3 -m py_compile
|
||||
build_stage "patch script completed"
|
||||
|
||||
Reference in New Issue
Block a user