456380eed02fff8c51f2f0e229b233fef2199570
Profiling zeros-out attention → vllm overestimates free memory → 7942 blocks → first real request OOMs. Cap at 5000 (80K tokens / 16 block_size). patch_block_major_worker_capacity.py reads BI100_MAX_GPU_BLOCKS from env, caps num_gpu_blocks after reserve_block_major_gpu_blocks.
project_6
Description
Languages
C++
41.8%
Cuda
31.6%
Python
22.2%
C
2.1%
CMake
1.1%
Other
1.1%