8e9c22f6c1457f79c660d8ad066d3557f77afb70
paged_attn.py: - Remove use_v1=True hardcode that forced all decode through ixf_F V1 - Wire up paged_attention_v2_triton.py as Tier 2 decode path for seq_len > 8192 - 3-tier dispatch: V1 (short) → Triton V2 (long) → PyTorch (fallback) - Triton V2 uses CCCL compound-reduce pattern (summary_statistics.cu) with GQA broadcast (6x KV read reduction for Qwen3.6) - This is the single highest-impact change: Output TPS is 83% of score prefix_prefill.py: - CCCL scan-tuning-informed block sizes for BI-V100 (SM=16, 48KB SMEM) - BI-V100 path: BLOCK=64 NUM_WARPS=4 (vs BLOCK=128 NUM_WARPS=8 on A100+) - Matches muh/tuning/tuning_scan.cuh bi100_lookback_4B_o4 pattern - Fewer warps = less register pressure = higher occupancy on 16 SMs computility-run.yaml: - Add --num-scheduler-steps=8: batch 8 decode iterations per Python call (cuts scheduler overhead ~8x, directly improves Output TPS) - Add --preemption-mode=recompute (cheaper than swap on BI-V100 HBM) - Add TRITON_CACHE_DIR for JIT warmup persistence - Add TRITON_PRINT_AUTOTUNING=0 (use hardcoded CCCL configs, skip autotune) Competition impact estimate: - Tier 2 Triton V2 replaces PyTorch fallback for 8K-100K contexts → ~5-10x decode speedup - Multi-step scheduling → ~20-30% Output TPS improvement - SM=16 block tuning → ~10-15% Input TPS improvement
project_6
Description
Languages
C++
41.8%
Cuda
31.6%
Python
22.2%
C
2.1%
CMake
1.1%
Other
1.1%