From 9907c9b8ee8297c0b651f91a8def1dec721fb81b Mon Sep 17 00:00:00 2001 From: dylanyunlon Date: Fri, 7 Aug 2026 03:18:25 +0000 Subject: [PATCH] docs: comprehensive CCCL vs muh gap analysis with competition priority P0 (reduce/scan/transform): core params done, need real benchmarks P1 (topk/select_if/radix_sort): partial, select_if missing bi100 struct P2 (14 others): theoretical coverage only, no competition impact Key insight: only 5 of 26 algorithms affect competition score. Pipeline fixed: gen_config.py replaces broken gen_patch.py. --- MUH_TUNING_GAP_ANALYSIS.md | 85 ++++++++++++++++++++++++++++++++++++++ 1 file changed, 85 insertions(+) create mode 100644 MUH_TUNING_GAP_ANALYSIS.md diff --git a/MUH_TUNING_GAP_ANALYSIS.md b/MUH_TUNING_GAP_ANALYSIS.md new file mode 100644 index 00000000..610827eb --- /dev/null +++ b/MUH_TUNING_GAP_ANALYSIS.md @@ -0,0 +1,85 @@ +# muh Tuning Gap Analysis — CCCL vs BI-V100 适配 +## 2026-08-07 + +### 方法论 + +直接读取 CCCL 源码(26 个 tuning_*.cuh),提取竞赛相关的 benchmark annotations, +对比 muh 已有的 BI-V100 struct 值。每个算法的优先级由竞赛评分公式决定: + +``` +Score = Output_TPS × 16.796 + Input_TPS × 2.799 + Cache_TPS × 0.56 +``` + +Output TPS = 83%, Input TPS = 14%, Cache TPS = 3% + +--- + +### P0: 直接影响竞赛评分的算法 + +#### 1. REDUCE (Output TPS 83%) — ★★★★★ +- **竞赛路径**: paged_attention score reduction, float32, plus +- **CCCL SM100**: `ipt_16.tpb_512.ipv_2 → 1.061/1.000/1.065/1.167` +- **muh BI-V100**: `bi100_plus_float32_o4 {512, 24, 2}` — tile=12288 (1.5× SM100) +- **状态**: ✅ 完成 (62% 行覆盖) +- **待定**: SM=16 items 适配 (P0 BUG)、LOAD_LDG vs LOAD_DEFAULT benchmark + +#### 2. SCAN (Output TPS 83%) — ★★★★☆ +- **竞赛路径**: softmax denominator prefix sum, float32, plus +- **CCCL SM100**: `ipt_22.tpb_384.ns_1904.dcid_6.l2w_830 → 1.148/0.997/1.140/1.463` +- **muh BI-V100**: `bi100_lookback_4B_o4 {384, 22}` — 与 SM100 同 tile +- **状态**: ✅ 核心完成 (39% 行覆盖,lookback + SM90 fallback) +- **待定**: Lookback delay 参数需实测校准、8B structs 99% SMEM 需验证 + +#### 3. TRANSFORM (Input TPS 14% + all activations) — ★★★★☆ +- **竞赛路径**: SiLU/GeLU/RMSNorm, bfloat16 +- **CCCL**: bytes_in_flight 是核心参数, B200=64KB, H100=48KB +- **muh BI-V100**: bytes_in_flight=64KB (confirmed by babelstream bench) +- **状态**: ✅ 核心完成 +- **待定**: Vectorized vs prefetch algorithm 选择需实测 + +--- + +### P1: 间接影响性能的算法 + +#### 4. TOPK (sampling, Output TPS) — ★★★☆☆ +- **竞赛路径**: logit sampling, float32 keys +- **CCCL**: bits_per_pass, thread count, BLOCK_SCAN_WARP_SCANS +- **muh BI-V100**: 有 inline tuning (threads=512, bits_per_pass=11) +- **状态**: ✅ 基本完成 +- **待定**: Onesweep vs multi-sweep 选择 + +#### 5. SELECT_IF (MoE routing) — ★★☆☆☆ +- **竞赛路径**: expert selection, float32, not_flagged, no_rejects, offset_4 +- **CCCL SM80**: `{threads=256, items=18, WARP_TRANSPOSE, no_delay=1130}` +- **muh BI-V100**: 零 bi100 structs, 用 get_sm100_adapted() inline 计算 +- **状态**: ⚠️ 只需 1/77 个 specialization, 但完全缺失 +- **待定**: 需添加 bi100_select_float32_nf_nr_o4 struct + +#### 6. RADIX_SORT (topk helper) — ★★☆☆☆ +- **竞赛路径**: float32 key sort for sampling +- **CCCL**: 2381 行, onesweep + histogram, SM100 有复杂分支 +- **muh BI-V100**: 222 行 (9% 覆盖) +- **状态**: ⚠️ 需要 onesweep 路径 +- **待定**: bits_per_pass 和 histogram SMEM + +--- + +### P2: 理论覆盖但不直接影响评分 + +| 算法 | CCCL 行数 | muh 行数 | 覆盖率 | 竞赛影响 | +|------|----------|---------|-------|---------| +| reduce_by_key | 1735 | 217 | 13% | 低 | +| scan_by_key | 2008 | 161 | 8% | 低 | +| unique_by_key | 1510 | 179 | 12% | 低 | +| three_way_partition | 708 | 67 | 9% | 低 | +| segmented_reduce | 471 | 112 | 24% | 低 | +| 其余 14 个 | ~4000 | ~800 | ~20% | 无 | + +--- + +### 关键差距总结 + +1. **gen_patch.py 管道断裂** — 产出零 patch。已被 gen_config.py 替代。 +2. **muh headers 20% 完成** — 但竞赛相关的 5 个算法 (reduce/scan/transform/topk/select_if) 核心参数已就位。 +3. **缺 benchmark 验证** — 所有 BI-V100 speedup 标 TBD,需要在 Phanthy Cloud 上跑。 +4. **Python layer 是真正的注入点** — 已在 triton_flash_attention.py 添加 8 个 BI-V100 configs, prefix_prefill.py 修 BLOCK=64, _custom_ops.py 修 SMEM=48KB。gen_config.py 又发现 19 个新候选 configs。