fe64650681b59ca1d615259059f7c8fe1d1d9852
8 steps in sequence, no user interaction needed: 1. Hardware diagnostics (SM count, SMEM, VRAM per GPU) 2. SMEM 32KB vs 48KB definitive answer from torch.cuda.get_device_properties 3. Triton availability check 4. prefix_prefill kernel import test 5. Triton compilation smoke test (compile+run trivial kernel) 6. Actual prefill kernel benchmark: 16 variants × 4 ctx_lens 7. Show current computility-run.yaml 8. fused_moe BLOCK_SIZE_M dispatch table for Qwen3.6 dimensions
project_6
Description
Languages
C++
41.8%
Cuda
31.6%
Python
22.2%
C
2.1%
CMake
1.1%
Other
1.1%