38eca5c26a0156e86bac330e2f767abcc487a5b7
a3c45d3b yaml + Q-tiling + remove all OOM hacks
Root cause of 10 consecutive OOM failures: - 'return zeros during profiling' hack → vllm overestimates free memory → allocates 7942 blocks → first real request OOMs - blocks cap 5000 → band-aid that masks profiling bug - gpu-memory-utilization 0.80 → unnecessary reduction from working 0.90 - max-num-seqs 2 → doubles peak activation memory - PYTORCH_CUDA_ALLOC_CONF max_split_size_mb:512 → causes fragmentation Restoringa3c45d3bparameters that actually work: - yaml: max-model-len=131072, gpu-mem=0.90, max-num-seqs=1, batched-tokens=8192 - patch_xformers_sdpa_seq.py: Q-tiling (real memory optimization, not zeros hack) - patch_block_major_worker_capacity.py: no blocks cap, just reserve_block_major - patch_ops.sh: remove all docker-build-time compilation (all .so are prebuilt) Only change froma3c45d3b: BI100_MOE_COREX_TOPK_SOFTMAX=1 (enable corex topk) Kept fixes: - protocol.py extra=allow (recover 180 rejected replay requests) - corex_gdn_chunk_recurrent.so pybind kwargs (prebuilt with fixed signature)
project_6
Description
Languages
C++
41.8%
Cuda
31.6%
Python
22.2%
C
2.1%
CMake
1.1%
Other
1.1%