9f93d695a9bfa8fdfbd3d9bb1e8da88e4faecf6f
muh_dispatch.py: - Fix missing os/sys imports (was crashing on import) - Fix SM count 50→16 (confirmed via ixsmi, matches hardware.cuh) - Fix C++ struct name lookup to match actual tuning_reduce.cuh names: bi100_plus_float32_o4, bi100_plus_float64_o4, bi100_plus_accum2_o4 (was: bi100_float32_plus_o4 — wrong name, would always fall through to default) Dockerfile: - Add COPY for prefix_prefill.py and muh_dispatch.py - Deploy CCCL-tuned prefix_prefill.py into vllm attention ops (BLOCK=64, NUM_WARPS=4 for BI-V100 SM=16) - Deploy muh_dispatch.py into vllm package for type-dispatched kernel configs - These files were written but never deployed — dead code until now Impact: prefix_prefill.py deployment means the CCCL-derived block sizes actually take effect at runtime. Previously the base image's original prefix_prefill.py (BLOCK=128 for cc>=80, or 64 for cc<80) was used, which is correct for BI-V100 but our version adds explicit SM=16 documentation and the path for future tuning.
project_6
Description
Languages
C++
41.8%
Cuda
31.6%
Python
22.2%
C
2.1%
CMake
1.1%
Other
1.1%