fix: critical config + tuning corrections from CCCL source analysis

computility-run.yaml:
  max-num-seqs 1→256: benchmark sweeps [128,256] concurrent seqs,
    current config processes 1 while 127 queue. KV cache budget:
    256 seqs × 2048 tokens × 80KB/token = 41.9GB < 45GB available.
  max-num-batched-tokens 8192→32768: support 256 concurrent prefills.
  gpu-memory-utilization 0.9→0.95: provide KV cache headroom.

Dockerfile:
  Deploy paged_attention_v2_triton.py to vllm package path so
  try-triton-first logic in _custom_ops.py can find it. Falls back
  to PyTorch V2 automatically if Triton V2 fails (SMEM/runtime).

muh/tuning/common.cuh:
  scale_mem_bound max_smem now a parameter (default 48KB). Allows
  policy_selectors to pass hw.max_shared_memory_per_block if actual
  SMEM differs from CCCL 48KB assumption.

muh/tuning/tuning_transform.cuh:
  bytes_in_flight 16KB→32KB. Old derivation used 900/50=18 GB/s/SM
  (wrong, SM=16 confirmed). Actual per-SM BW = 56 GB/s.
  32KB is estimate pending benchmark sweep.

SM count 50→16 corrections across all affected files.
This commit is contained in:
Claude
2026-08-03 06:45:54 +00:00
parent 071fa361a3
commit cdc01bbc6a
7 changed files with 33 additions and 14 deletions

View File

@@ -18,6 +18,15 @@ RUN python3 /workspace/qwen3_6_scripts/patch_ixformer_native.py
# 1. PagedAttention V2 — fills the NotImplementedError hole
# Enables partitioned attention for long sequences (>8192 tokens)
# Deploy BOTH PyTorch and Triton V2 to vllm package — _custom_ops.py
# tries Triton first, falls back to PyTorch if import/runtime fails.
# Triton V2 risk: SMEM=32KB zero margin at head_dim=256 BLOCK_N=32.
# If Triton V2 crashes, PyTorch V2 (batched bmm, no intermediate tensor
# savings but correct) takes over automatically via try/except.
RUN cp /workspace/paged_attention_v2_triton.py \
/usr/local/corex/lib/python3/dist-packages/vllm/paged_attention_v2_triton.py 2>/dev/null || \
cp /workspace/paged_attention_v2_triton.py \
/usr/local/corex/lib64/python3/dist-packages/vllm/paged_attention_v2_triton.py 2>/dev/null || true
RUN python3 /workspace/qwen3_6_scripts/patch_paged_attention_v2.py
# 2. Triton kernel tuning: BLOCK=64, NUM_WARPS=4