Key findings from reading dispatch_scan.cuh: 1. Lookahead scan requires PTX ISA >= 860 (NVIDIA SM100+), completely unavailable on BI-V100. Our lookback-only strategy is correct. 2. Lookback scan passes 0 dynamic SMEM — SMEM is all static via __shared__. Different from lookahead which uses dynamic stages. 3. Scan launches exactly num_tiles blocks (not sm_count * subscription), one CTA per tile. For 100K tokens: ~12 tiles all fit in one wave on 16 SMs, explaining why no_delay (dcid=0) is optimal. 4. Lookahead's num_stages auto-tuning is irrelevant for BI-V100 but reveals NVIDIA's pipeline depth selection strategy.
6.2 KiB
CCCL Reduce Architecture Notes
Source:
dispatch_reduce.cuh,kernel_reduce.cuh,agent_reduce.cuh,tuning_reduce.cuh,util_arch.cuhRead: 2026-08-04 by Claude from CCCL upstream in project_6/cccl_upstream/
Key Architecture
Two-pass dispatch (dispatch_reduce.cuh)
num_items <= single_tile.threads * single_tile.items
→ SingleTile: one CTA, one kernel launch
→ DeviceReduceSingleTileKernel(d_in, d_out, num_items, ...)
num_items > single_tile threshold
→ Pass 1: DeviceReduceKernel — N CTAs each reduce their share → d_block_reductions[N]
→ Pass 2: DeviceReduceSingleTileKernel — 1 CTA reduces d_block_reductions[N] → d_out
Grid size for Pass 1: max_blocks = sm_occupancy * sm_count * subscription_factor(5)
For BI-V100: 2 * 16 * 5 = 160 blocks max.
Each block processes ceil(num_items / 160) elements.
Tile consumption (agent_reduce.cuh)
Critical: tile data is in registers, NOT SMEM.
AccumT items[ITEMS_PER_THREAD]; // <-- register array, per-thread
// ... load from global memory ...
thread_aggregate = ThreadReduce(items, reduction_op); // per-thread reduction
// Only SMEM used:
BlockReduce(temp_storage.reduce).Reduce(thread_aggregate, reduction_op);
TempStorage = BlockReduce::TempStorage ≈ threads * sizeof(AccumT) bytes.
NOT threads * items * sizeof(AccumT).
Vectorized loads
ATTEMPT_VECTORIZATION = (vec_size > 1) && (ITEMS_PER_THREAD % vec_size == 0)
&& is_pointer<InputIteratorT>
&& (is_primitive<InputT> || is_trivially_relocatable<InputT>)
&& sizeof(InputT) <= 8;
For fp32 scores: vec_size=2 → loads 8 bytes (2 floats) per instruction. For fp16 KV cache: vec_size=4 → loads 8 bytes (4 halfs) per instruction.
scale_mem_bound vs scale_reg_bound (util_arch.cuh)
Two scaling functions with different constraints:
scale_mem_bound (memory-bound algorithms: reduce, transform):
- items = clamp(nominal * 4 / type_size, 1, nominal * 2) ← allows 2x expansion
- threads = min(nominal, round_up(48KB / (type_size * items), 32))
scale_reg_bound (register-bound algorithms: scan with complex state):
- items = max(1, nominal * 4 / max(4, type_size)) ← no expansion past nominal
- threads = min(nominal, ceil_div(48KB / (type_size * items), 32) * 32)
Key difference: scale_reg_bound uses max(4, type_size) preventing items from exceeding nominal for small types, and uses ceil_div instead of round_up for thread count. Both use 48KB as the cap, but this limits REGISTER PRESSURE (spill to local memory), not actual SMEM usage.
Impact on muh tuning
Our SMEM model was wrong for reduce
test_smem_safety.py and check_smem() in muh_kernel_map.py compute
tile_bytes = threads * items * type_size and check against 49152.
This is the scale_mem_bound cap, NOT the actual SMEM usage. The actual SMEM
for reduce is approximately threads * max(sizeof(AccumT), 4) bytes — about
2-8 KB, not 32-49 KB.
CCCL's SM100 float64 tuning uses threads=640, items=16 → scale_mem_bound
"tile" = 640168 = 81920 > 49152. But this doesn't overflow SMEM — it only
means scale_mem_bound will cap threads down. The actual kernel SMEM usage
with threads=640 is only ~5120 bytes.
Our float64/int64 tuning may be too conservative
We use threads=384 items=16 for float64, capped by scale_mem_bound. CCCL uses threads=640 items=16 on SM100. The question is whether BI-V100's register file (255 regs/thread) can hold 16 float64 items without spilling.
16 * 8 = 128 bytes = 32 registers per thread for tile data alone. With overhead (thread_aggregate, loop variables, etc.), ~40 registers/thread. 255 max registers → no spill risk. threads=640 may be safe on BI-V100.
TODO: Benchmark threads=640 items=16 for float64 on BI-V100.
paged_attn.py forces V1
Line 99: use_v1 = True overrides V1/V2 heuristic. V2 is completely disabled.
For 100K token sequences, V1 makes one CTA iterate over all KV blocks — bad
for latency. V2 would partition the work and reduce across partitions, which
is exactly CCCL's two-pass pattern.
TODO: Re-enable V2 for max_seq_len > 8192. Use muh's partition_size tuning.
_PARTITION_SIZE = 512 is hardcoded
Not controlled by muh. Should be tunable: larger partition = fewer blocks = less overhead but more work per block. Optimal value depends on SM count. For 16 SMs: partition_size=1024 may be better (fewer partitions to reduce).
CCCL Scan Architecture (dispatch_scan.cuh)
Added: 2026-08-04
Two algorithm paths
Lookback (all GPUs including BI-V100):
- Each CTA processes one tile, uses
ScanTileStatein global memory for inter-CTA communication - Lookback delay policy controls how aggressively CTAs poll predecessors
- SMEM: static only (
__shared__), passed as0dynamic SMEM - BI-V100 optimal:
no_delay(dcid=0) because 16 SMs → ~32 CTAs → tile_status fits in 6MB L2
Lookahead (SM100+ only, PTX ISA >= 860):
- Pipeline-based with
__pipeline_memcpy_asyncand bulk copy - Uses dynamic SMEM with auto-selected
num_stages - Not available on BI-V100 — requires NVIDIA PTX ISA 860+ instructions
- All lookahead structs in our tuning_scan.cuh can remain empty shells
ScanTileState allocation
Scan requires d_temp_storage for tile status descriptors:
tile_size = threads * items
num_tiles = ceil(num_items / tile_size)
temp_bytes = tile_state.AllocationSize(num_tiles)
For BI-V100 with 100K tokens and tile_size=384*22=8448: num_tiles = ceil(100000/8448) = 12 tiles → negligible temp storage.
Grid size for scan
Lookback scan launches num_tiles blocks (one per tile), NOT sm_count * subscription_factor.
This is different from reduce, which uses GridEvenShare.
For scan, every CTA processes exactly one tile and communicates with neighbors.
With 12 tiles on 16 SMs: all tiles fit in one wave, zero lookback contention.
This is why no_delay works on BI-V100 — the entire scan completes in a single wave.
Lookahead num_stages optimization (SM100 only)
CCCL dynamically selects pipeline depth:
max_stages = ceil(num_items / (sm_count * tile_size)) + 1
while (smem_for_stages(num_stages+1) <= max_dynamic_smem) num_stages++
For BI-V100 this is irrelevant (no pipeline support), but the formula shows NVIDIA's strategy: match pipeline depth to problem size / SM count ratio.