88db0ed89c9ad9278e3a849fa4243f13c3f7756b
tuning_reduce.cuh (201→311 lines): - Add accum_size=1/2/16 branches (int8, bfloat16, int128) - Add min/max op dispatch (same params as plus for BI-V100) - SM=16 tile maximization: det_float32 tile 11648→49152 (23%→100% SMEM) - SM=16 tile maximization: det_float64 tile 11264→49152 (23%→100% SMEM) - Add float32_o8, int64_o4/o8 variants with vec_size dispatch - Increase float32 items 16→24 (32768→49152, fill SMEM for fewer CTAs) tuning_scan.cuh: - Fix 1B tile from 9216→16384 (19%→33% SMEM, scan needs 2x buffer) - Fix 2B tile from 13312→24576 (27%→100% SMEM with double buffer) - Fix 8B_o4 tile: threads 416→384 for warp alignment, items 14→16 - Update header comments with confirmed SM=16 hardware profile - Document lookback delay heuristic for L2=6MB tuning_transform.cuh (128→168 lines): - CRITICAL: bytes_in_flight 16KB→32KB (was based on 900/50=18 GB/s, actual is 900/16=56 GB/s — 3× error) - Add full PrefetchPolicy struct matching CCCL upstream - Add AsyncCopyPolicy with BI-V100 fallback (no cp.async support) - Document CCCL cc_to_min_bytes_in_flight reference values - Add vec_size calculation from element size (16-byte vector loads) - Cap items_per_thread at 32 to prevent register pressure hardware.cuh: - Add SMEM 48KB vs 32KB disambiguation note
project_6
Description
Languages
C++
41.8%
Cuda
31.6%
Python
22.2%
C
2.1%
CMake
1.1%
Other
1.1%