Files
project_6/docs/CCCL_REDUCE_ARCHITECTURE_NOTES.md
project_6 3cc97c1d4e [docs] CCCL scan architecture — lookback vs lookahead, tile_state allocation, grid sizing
Key findings from reading dispatch_scan.cuh:
1. Lookahead scan requires PTX ISA >= 860 (NVIDIA SM100+), completely
   unavailable on BI-V100. Our lookback-only strategy is correct.
2. Lookback scan passes 0 dynamic SMEM — SMEM is all static via
   __shared__. Different from lookahead which uses dynamic stages.
3. Scan launches exactly num_tiles blocks (not sm_count * subscription),
   one CTA per tile. For 100K tokens: ~12 tiles all fit in one wave
   on 16 SMs, explaining why no_delay (dcid=0) is optimal.
4. Lookahead's num_stages auto-tuning is irrelevant for BI-V100 but
   reveals NVIDIA's pipeline depth selection strategy.
2026-08-05 03:21:51 +00:00

6.2 KiB

CCCL Reduce Architecture Notes

Source: dispatch_reduce.cuh, kernel_reduce.cuh, agent_reduce.cuh, tuning_reduce.cuh, util_arch.cuh Read: 2026-08-04 by Claude from CCCL upstream in project_6/cccl_upstream/

Key Architecture

Two-pass dispatch (dispatch_reduce.cuh)

num_items <= single_tile.threads * single_tile.items
  → SingleTile: one CTA, one kernel launch
  → DeviceReduceSingleTileKernel(d_in, d_out, num_items, ...)

num_items > single_tile threshold
  → Pass 1: DeviceReduceKernel — N CTAs each reduce their share → d_block_reductions[N]
  → Pass 2: DeviceReduceSingleTileKernel — 1 CTA reduces d_block_reductions[N] → d_out

Grid size for Pass 1: max_blocks = sm_occupancy * sm_count * subscription_factor(5) For BI-V100: 2 * 16 * 5 = 160 blocks max. Each block processes ceil(num_items / 160) elements.

Tile consumption (agent_reduce.cuh)

Critical: tile data is in registers, NOT SMEM.

AccumT items[ITEMS_PER_THREAD];  // <-- register array, per-thread
// ... load from global memory ...
thread_aggregate = ThreadReduce(items, reduction_op);  // per-thread reduction

// Only SMEM used:
BlockReduce(temp_storage.reduce).Reduce(thread_aggregate, reduction_op);

TempStorage = BlockReduce::TempStorage ≈ threads * sizeof(AccumT) bytes. NOT threads * items * sizeof(AccumT).

Vectorized loads

ATTEMPT_VECTORIZATION = (vec_size > 1) && (ITEMS_PER_THREAD % vec_size == 0)
    && is_pointer<InputIteratorT>
    && (is_primitive<InputT> || is_trivially_relocatable<InputT>)
    && sizeof(InputT) <= 8;

For fp32 scores: vec_size=2 → loads 8 bytes (2 floats) per instruction. For fp16 KV cache: vec_size=4 → loads 8 bytes (4 halfs) per instruction.

scale_mem_bound vs scale_reg_bound (util_arch.cuh)

Two scaling functions with different constraints:

scale_mem_bound (memory-bound algorithms: reduce, transform):

  • items = clamp(nominal * 4 / type_size, 1, nominal * 2) ← allows 2x expansion
  • threads = min(nominal, round_up(48KB / (type_size * items), 32))

scale_reg_bound (register-bound algorithms: scan with complex state):

  • items = max(1, nominal * 4 / max(4, type_size)) ← no expansion past nominal
  • threads = min(nominal, ceil_div(48KB / (type_size * items), 32) * 32)

Key difference: scale_reg_bound uses max(4, type_size) preventing items from exceeding nominal for small types, and uses ceil_div instead of round_up for thread count. Both use 48KB as the cap, but this limits REGISTER PRESSURE (spill to local memory), not actual SMEM usage.

Impact on muh tuning

Our SMEM model was wrong for reduce

test_smem_safety.py and check_smem() in muh_kernel_map.py compute tile_bytes = threads * items * type_size and check against 49152.

This is the scale_mem_bound cap, NOT the actual SMEM usage. The actual SMEM for reduce is approximately threads * max(sizeof(AccumT), 4) bytes — about 2-8 KB, not 32-49 KB.

CCCL's SM100 float64 tuning uses threads=640, items=16 → scale_mem_bound "tile" = 640168 = 81920 > 49152. But this doesn't overflow SMEM — it only means scale_mem_bound will cap threads down. The actual kernel SMEM usage with threads=640 is only ~5120 bytes.

Our float64/int64 tuning may be too conservative

We use threads=384 items=16 for float64, capped by scale_mem_bound. CCCL uses threads=640 items=16 on SM100. The question is whether BI-V100's register file (255 regs/thread) can hold 16 float64 items without spilling.

16 * 8 = 128 bytes = 32 registers per thread for tile data alone. With overhead (thread_aggregate, loop variables, etc.), ~40 registers/thread. 255 max registers → no spill risk. threads=640 may be safe on BI-V100.

TODO: Benchmark threads=640 items=16 for float64 on BI-V100.

paged_attn.py forces V1

Line 99: use_v1 = True overrides V1/V2 heuristic. V2 is completely disabled. For 100K token sequences, V1 makes one CTA iterate over all KV blocks — bad for latency. V2 would partition the work and reduce across partitions, which is exactly CCCL's two-pass pattern.

TODO: Re-enable V2 for max_seq_len > 8192. Use muh's partition_size tuning.

_PARTITION_SIZE = 512 is hardcoded

Not controlled by muh. Should be tunable: larger partition = fewer blocks = less overhead but more work per block. Optimal value depends on SM count. For 16 SMs: partition_size=1024 may be better (fewer partitions to reduce).


CCCL Scan Architecture (dispatch_scan.cuh)

Added: 2026-08-04

Two algorithm paths

Lookback (all GPUs including BI-V100):

  • Each CTA processes one tile, uses ScanTileState in global memory for inter-CTA communication
  • Lookback delay policy controls how aggressively CTAs poll predecessors
  • SMEM: static only (__shared__), passed as 0 dynamic SMEM
  • BI-V100 optimal: no_delay (dcid=0) because 16 SMs → ~32 CTAs → tile_status fits in 6MB L2

Lookahead (SM100+ only, PTX ISA >= 860):

  • Pipeline-based with __pipeline_memcpy_async and bulk copy
  • Uses dynamic SMEM with auto-selected num_stages
  • Not available on BI-V100 — requires NVIDIA PTX ISA 860+ instructions
  • All lookahead structs in our tuning_scan.cuh can remain empty shells

ScanTileState allocation

Scan requires d_temp_storage for tile status descriptors:

tile_size = threads * items
num_tiles = ceil(num_items / tile_size)
temp_bytes = tile_state.AllocationSize(num_tiles)

For BI-V100 with 100K tokens and tile_size=384*22=8448: num_tiles = ceil(100000/8448) = 12 tiles → negligible temp storage.

Grid size for scan

Lookback scan launches num_tiles blocks (one per tile), NOT sm_count * subscription_factor. This is different from reduce, which uses GridEvenShare. For scan, every CTA processes exactly one tile and communicates with neighbors.

With 12 tiles on 16 SMs: all tiles fit in one wave, zero lookback contention. This is why no_delay works on BI-V100 — the entire scan completes in a single wave.

Lookahead num_stages optimization (SM100 only)

CCCL dynamically selects pipeline depth:

max_stages = ceil(num_items / (sm_count * tile_size)) + 1
while (smem_for_stages(num_stages+1) <= max_dynamic_smem) num_stages++

For BI-V100 this is irrelevant (no pipeline support), but the formula shows NVIDIA's strategy: match pipeline depth to problem size / SM count ratio.