9605415404595af2e5acae232971b46bb7908e29
Tests scale_mem_bound CCCL parity (8/8), register pressure for all 14 bi100_* structs, summary_statistics.cu 28-byte AccumT safety, and vectorization alignment. All pass. Key finding: BI-V100 float32 tile is 1.5x SM100's (12288 vs 8192) because 16 SMs need larger tiles to compensate for fewer CTAs. float64 tile is 0.6x SM100's (6144 vs 10240) because threads=640 was reduced to 384 (clean warp count) and vec=2 added.
project_6
Description
Languages
C++
41.8%
Cuda
31.6%
Python
22.2%
C
2.1%
CMake
1.1%
Other
1.1%