muh-bot
b50bd2dfd5
[gen_patch] fix critical struct selection: dispatch by kernel data type
...
CCCL policy_selector dispatches by (accum_size, type_t, offset_size).
gen_patch was picking first non-default struct → bi100_plus_accum1_o4
(int8 items=32) for reduce. paged_attention uses float32 scores →
correct struct is bi100_plus_float32_o4 (items=24).
Before: 512*32*4=65536 > 49152 SMEM → crash
After: 512*24*4=49152 = 100% SMEM → correct
SCAN_BLOCK_SIZE: 512→384 (bench_bi100.py ipt=22 tpb=384 dcid=0)
2026-08-05 03:21:31 +00:00
project_6
db0df78580
[gen_patch] fix 3 critical bugs: reduce struct selection, topk bits_per_pass injection, transform/batch_memcpy extraction
...
Bug 1: reduce struct selection — gen_patch selected bi100_plus_accum1_o4 (int8
path, items=32) instead of bi100_plus_float32_o4 (fp32 score accumulator,
items=24). vllm paged_attention always uses fp32 for score accumulation, so
the wrong struct was injecting items=32 into the 83%-weight hot path.
Fix: preference-ordered struct selection — float32 > accum2 > first non-default.
Now correctly selects bi100_plus_float32_o4 → NUM_ITEMS_PER_THREAD=24.
Bug 2: topk bits_per_pass not injected — gen_patch only extracted threads=512
from topk inline policy_selector, missing calc_bits_per_pass(key_size).
For Qwen3.6 float32 logits (key_size=4), bits_per_pass=11 (not 8).
Fix: topk-specific extraction that parses calc_bits_per_pass and returns
bits_per_pass=11. Now generates RADIX_BITS=11 patch for sampling_kernels.cu.
Bug 3: transform/batch_memcpy extraction failed — these headers use
policy struct naming, not bi100_* naming, so extract_bi100_structs was empty.
Fix: algorithm-specific fallback extraction for transform (reads
bi100_bytes_in_flight constexpr) and batch_memcpy (reads threads from
policy_selector return).
Validation: gen_patch 7 patches (was 6), test_smem_safety 191/191 safe
2026-08-05 03:15:12 +00:00
dylanyunlon
2c5e77f370
feat(gen_patch): add TUNING_REGISTRY for all 26 algorithms
...
Registers all 26 CUB algorithms with metadata:
- 6 'injection' mode: have VLLM_INJECTION_POINTS (reduce/scan/topk/transform/batch_memcpy/for)
- 20 'library' mode: used via CCCL device API, no direct #define injection
- struct_mode: 'named' (bi100_* structs) vs 'inline' (policy_selector returns)
Also adds coverage reporting to generate_patches().
2026-08-01 12:33:14 +08:00
Claude
c7a63bc2c8
[MUH] Fix 7 structural discrepancies vs CCCL — read source, not grep
...
Fixes found by reading all 17 muh files + 6 CCCL counterpart
policy_selectors as full source code input:
1. topk: BLOCK_LOAD_DIRECT → BLOCK_LOAD_VECTORIZE (CCCL SM90+ uses
VECTORIZE). bits_per_pass was wrong (muh: ks<=4→9, CCCL: ks>=2→11).
items now computed dynamically (4*4/key_size) not hardcoded.
2. reduce: added determinism dispatch — three modes matching CCCL:
gpu_to_gpu (BLOCK_REDUCE_RAKING, vec_size=1, LOAD_DEFAULT),
run_to_run (WARP_REDUCTIONS, LOAD_LDG, default),
not_guaranteed (WARP_REDUCTIONS_NONDETERMINISTIC).
Added bi100_det_float32 and bi100_det_float64 tuning structs
with SM90 benchmark reference values.
3. batch_memcpy: flat single-tier → SmallBuffer+LargeBuffer two-tier
matching CCCL structure (128 threads small, 256 threads large,
warp_threshold=128, block_threshold=8192).
4. transform: single BulkPolicy → three-policy structure
(VectorizedPolicy + AsyncCopyPolicy + PrefetchPolicy) matching CCCL.
items_per_thread computed from bytes_in_flight / (threads * elem_size).
5. compile_test: 17 checks → 33 checks. Now verifies exact values:
reduce determinism modes, topk VECTORIZE + bits=11, batch_memcpy
two-tier thresholds, transform three-policy structure.
6. gen_patch: added fallback extraction for inline policy_selector
values (topk now generates SAMPLING_BLOCK_SIZE patch).
7. MUH_PROJECT_CHECKPOINT.md: 'PRD设计阶段还没有代码' → actual status.
7 files changed, 413 insertions, 265 deletions.
2026-07-30 14:37:38 +00:00
dylanyunlon
57e222b99d
[MUH] Fix three-layer disconnect — C++ headers are now the single source of truth
...
Problems fixed:
1. gen_patch.py was reading .muh YAML (all nulls) instead of C++ headers.
Now it parses bi100_* structs directly from tuning_*.cuh via regex,
extracts constexpr values, and maps them to vllm injection points.
Verified: 11 patches generated from 6 algorithms.
2. C++ headers had no build system or tests.
Added CMakeLists.txt (header-only library target) and compile_test.cpp.
Verified: g++ -std=c++17 compiles all headers, 17/17 runtime checks pass.
Also added cuda_compile_test.cu for when nvcc is available.
3. baseline.muh had a tuning section full of nulls duplicating C++ values.
Stripped to vllm launch config only. Tuning values live exclusively
in muh/include/muh/tuning/tuning_*.cuh bi100_* structs.
4. Fixed constexpr goto in tuning_scan.cuh (C++17 doesn't allow goto in
constexpr; replaced with early-return + default: break pattern).
Data flow is now:
tuning_*.cuh (bi100_* constexpr) ──→ gen_patch.py ──→ vllm patches
baseline.muh (launch config) ──→ gen_yaml.py ──→ computility-run.yaml
compile_test.cpp ──→ g++/nvcc ──→ verify values are real
2026-07-30 14:12:33 +00:00
dylanyunlon
9b21a13119
[MUH] Bootstrap muh toolchain — extract/parse/gen_yaml/gen_patch + baseline.muh
...
Pipeline:
1. extract.py: Parses all 26 CCCL tuning_*.cuh → 26 YAML schemas in muh/schema/
2. parse.py: .muh file parser with extends-inheritance + schema validation
3. gen_yaml.py: .muh → computility-run.yaml (verified: matches competition reference)
4. gen_patch.py: .muh → vllm kernel unified diff patches (6 algorithm mappings)
5. baseline.muh: Competition reference config, all tuning values pending BI-V100 benchmarks
Schemas extracted:
26 algorithms, 8-19 params each, SM75/80/90/100 reference tunings
Priority mapping: reduce→attention, topk→sampling, scan→paged_attention,
transform→activations, batch_memcpy→KV_cache, for→RoPE
Tested: extract→parse→validate→gen_yaml→gen_patch full pipeline passes
2026-07-30 10:39:06 +00:00