fix: critical config + tuning corrections from CCCL source analysis
computility-run.yaml:
max-num-seqs 1→256: benchmark sweeps [128,256] concurrent seqs,
current config processes 1 while 127 queue. KV cache budget:
256 seqs × 2048 tokens × 80KB/token = 41.9GB < 45GB available.
max-num-batched-tokens 8192→32768: support 256 concurrent prefills.
gpu-memory-utilization 0.9→0.95: provide KV cache headroom.
Dockerfile:
Deploy paged_attention_v2_triton.py to vllm package path so
try-triton-first logic in _custom_ops.py can find it. Falls back
to PyTorch V2 automatically if Triton V2 fails (SMEM/runtime).
muh/tuning/common.cuh:
scale_mem_bound max_smem now a parameter (default 48KB). Allows
policy_selectors to pass hw.max_shared_memory_per_block if actual
SMEM differs from CCCL 48KB assumption.
muh/tuning/tuning_transform.cuh:
bytes_in_flight 16KB→32KB. Old derivation used 900/50=18 GB/s/SM
(wrong, SM=16 confirmed). Actual per-SM BW = 56 GB/s.
32KB is estimate pending benchmark sweep.
SM count 50→16 corrections across all affected files.
This commit is contained in:
@@ -16,7 +16,7 @@ Hardware derivation:
|
||||
8 warps = 256 threads → each thread handles 32 elements from Q tile.
|
||||
4 warps = 128 threads → each thread handles 64 elements.
|
||||
|
||||
With 50 SMs and typical grid of 37K+ blocks:
|
||||
With 16 SMs (confirmed) and typical grid of 37K+ blocks:
|
||||
At 8 warps + 32KB SMEM: 1 block per SM (SMEM-limited)
|
||||
At 4 warps + 32KB SMEM: potentially 2 blocks per SM
|
||||
|
||||
|
||||
Reference in New Issue
Block a user