6d8de852ad1934a7763fbd1c7f4b8a3b4d467e39
CRITICAL DISCOVERY: Qwen3.6-35B-A3B uses head_dim=256 (not 128). text_cfg.head_dim=256, num_heads=24, num_kv_heads=4, GQA=6 This means ALL previous SMEM calculations were wrong: BLOCK=64 + head_dim=256: 64×256×2×2 = 64KB > 48KB → OVERFLOW BLOCK=64 + head_dim=128: 64×128×2×2 = 32KB ≤ 48KB → OK (but wrong model) Fix: head_dim-dependent BLOCK selection in prefix_prefill.py: head_dim ≤ 128: BLOCK=64, NUM_WARPS=4 (32KB SMEM) head_dim = 256: BLOCK=32, NUM_WARPS=4 (32KB SMEM) head_dim > 256: BLOCK=16, NUM_WARPS=2 (16KB SMEM) Also: _Q_CHUNK in _run_sdpa_fallback reduced 256→128 for head_dim=256 to avoid OOM on long sequences (256×100K×24×4=2.3GB vs 128×100K×24×4=1.2GB). Without this patch, Triton prefill CANNOT work for Qwen3.6. patch_enable_triton.py's try/fallback would always fall back to PyTorch.
project_6
Description
Languages
C++
41.8%
Cuda
31.6%
Python
22.2%
C
2.1%
CMake
1.1%
Other
1.1%