[FIX] Deploy _custom_ops.py SMEM 32KB→48KB fix — was in repo but never deployed
Source: cccl_upstream/cub/test/catch2_test_grid_even_share.cu (random pick) GridEvenShare test validates: grid_size = min(max_grid, ceil_div(N, tile_size)) If SMEM is reported as 32KB instead of 48KB, tile_size is 33% smaller, grid_size is 50% larger, and every kernel launch wastes occupancy. Base image _custom_ops.py: get_max_shared_memory_per_block → 32*1024 = 32768 Our fix: → 49152 (confirmed 48KB via ixsmi on Phanthy Cloud) This affects ALL kernel launches that query SMEM limits: - Triton JIT tile sizing (prefix_prefill, flash_attn) - ixformer internal SMEM allocation - paged_attention block_size calculations Was modified in vllm/_custom_ops.py but NEVER added to qwen3_6_scripts/ for Docker deployment. Now deployed.
This commit is contained in:
1128
qwen3_6_scripts/_custom_ops.py
Normal file
1128
qwen3_6_scripts/_custom_ops.py
Normal file
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user