diff --git a/docs/SPECIALIZATION_ANALYSIS.md b/docs/SPECIALIZATION_ANALYSIS.md new file mode 100644 index 00000000..ad0bbf8b --- /dev/null +++ b/docs/SPECIALIZATION_ANALYSIS.md @@ -0,0 +1,85 @@ +# muh vs CCCL SM100: Type Specialization Parity Analysis + +Generated: 2026-07-31 + +## Summary + +| Algorithm | CCCL SM100 branches | muh BI-V100 branches | Status | +|-----------|--------------------:|---------------------:|--------| +| reduce | 4+2 det = 6 | 4+2 det+1 default = 7 | ✓ PARITY+ | +| scan (lookback) | 7 | 7 (after 35ef79c5) | ✓ PARITY | +| scan (lookahead) | 6 | 6 | ✓ PARITY | +| topk | 1 (dynamic by key_size) | 1 (dynamic by key_size) | ✓ PARITY | +| transform | 1 (dynamic by elem_size) | 1 (dynamic by elem_size) | ✓ PARITY | +| batch_memcpy | 1 (uniform) | 1 (uniform) | ✓ PARITY | +| for | 1 (uniform) | 1 (uniform) | ✓ PARITY | + +## Detailed Breakdown + +### reduce (tuning_reduce.cuh) + +CCCL SM100 specializes by `(accum_type × offset_size)`: +- `int64 + o4`: ipt=15, tpb=512, ipv=2 +- `int64 + o8`: ipt=15, tpb=512, ipv=1 +- `float32 + o4`: ipt=16, tpb=512, ipv=2 +- `float64 + o4`: ipt=16, tpb=640, ipv=1 + +muh BI-V100 maps these with SMEM-derived corrections: +- `bi100_float32_plus_o4`: tpb=512, ipt=16, ipv=2 (direct match) +- `bi100_float64_plus_o4`: tpb=512, ipt=12, ipv=1 (SM100 ipt=16 → SMEM overflow at 8B, reduced) +- `bi100_int64_plus_o4`: tpb=384, ipt=16, ipv=2 (SM100 tpb=512 → SMEM overflow, reduced threads) +- `bi100_int64_plus_o8`: tpb=384, ipt=16, ipv=1 (same, vec=1 for 8B offset) +- `bi100_det_float32`: tpb=224, ipt=13 (deterministic path, RAKING) +- `bi100_det_float64`: tpb=128, ipt=11 (deterministic path, RAKING) +- `bi100_default`: tpb=256, ipt=16, ipv=4 (fallback) + +### scan (tuning_scan.cuh) + +CCCL SM100 lookback specializes by `(input_value_size × offset_size)`: +``` +offset=4: 1B→(512,18) 2B→(512,13) 4B→(384,22) 8B→(416,23) +offset=8: 1B→(384,14) [2B=skip] 4B→(416,19) 8B→(320,22) +``` + +muh BI-V100 after commit 35ef79c5: +``` +offset=4: 1B→(512,18) 2B→(512,13) 4B→(384,22) 8B→(416,14*) +offset=8: 1B→(384,14) 4B→(416,19) 8B→(320,19*) +``` +*items reduced to fit 49152B SMEM + +All delay parameters halved (ns×0.5, l2w×0.6) to account for +BI-V100 L2=6MB vs SM100 L2=50MB. + +### SMEM Constraint Validation + +Every muh bi100_* struct satisfies: `nominal_tile = tpb × ipt × 4 ≤ 49152` + +| Struct | tpb | ipt | nominal_tile | Status | +|--------|----:|----:|-------------:|--------| +| bi100_lookback_1B_o4 | 512 | 18 | 36864 | ✓ | +| bi100_lookback_2B_o4 | 512 | 13 | 26624 | ✓ | +| bi100_lookback_4B_o4 | 384 | 22 | 33792 | ✓ | +| bi100_lookback_4B_o8 | 416 | 19 | 31616 | ✓ | +| bi100_lookback_8B_o4 | 416 | 14 | 23296 | ✓ | +| bi100_lookback_8B_o8 | 320 | 19 | 24320 | ✓ | +| bi100_lookback_1B_o8 | 384 | 14 | 21504 | ✓ | +| bi100_float32_plus_o4 | 512 | 16 | 32768 | ✓ | +| bi100_float64_plus_o4 | 512 | 12 | 24576 | ✓ | +| bi100_int64_plus_o4 | 384 | 16 | 24576 | ✓ | +| bi100_int64_plus_o8 | 384 | 16 | 24576 | ✓ | + +## Non-Hot-Path Algorithms (20 missing) + +These 20 CCCL algorithms have muh/schema/*.yaml but no tuning header. +They are NOT on the vllm inference hot path for Qwen3.6 decode. +If any competition test case triggers them, they will use CCCL defaults +which may cause SMEM overflow on BI-V100 for large types. + +Priority to add (by SMEM overflow risk): +1. `radix_sort` (89KB tuning, 161 type dispatches) — HIGH risk +2. `reduce_by_key` (72KB, 134 dispatches) — HIGH risk +3. `select_if` (107KB, 2729 lines) — MEDIUM risk +4. `scan_by_key` (88KB) — MEDIUM risk +5. `unique_by_key` (61KB) — LOW risk +6. Others: LOW risk (small tile sizes, unlikely SMEM overflow)