[MUH] Delete 20 dead-code batch-generated tuning headers

Audit results:
  - 20/20 files had IDENTICAL if-branch and fallback (dead code)
  - 787 lines total, 5% coverage of 15116 lines in CCCL originals
  - No type specializations, no offset_size branches, no benchmark data
  - 0 of 20 algorithms appear on vllm's Qwen3.6 inference hot path

The 6 headers that remain (reduce, topk, scan, transform, batch_memcpy, for)
are the only algorithms that execute during vllm decode/prefill/cache operations.
These 6 have real type specializations and CCCL SM100 reference values.

CCCL has 26 algorithms because it's a general-purpose library.
muh targets one workload: Qwen3.6-35B-A3B on 4× BI-V100.
Covering algorithms that don't execute is worse than not covering them —
it creates the illusion of completeness.
This commit is contained in:
dylanyunlon
2026-07-30 14:22:18 +00:00
parent 07b015f31e
commit e02134a3ce
21 changed files with 20 additions and 825 deletions

View File

@@ -1,38 +0,0 @@
// muh/include/muh/tuning/tuning_batched_topk.cuh — BI-V100 batched_topk tuning
//
// Mirrors: cccl_upstream/cub/cub/device/dispatch/tuning/tuning_batched_topk.cuh
// vllm impact: batched top-k across sequences
// Competition weight: Output TPS × 16.796
#pragma once
#include "muh/hardware.cuh"
#include "muh/tuning/common.cuh"
namespace muh::tuning::batched_topk {
struct BatchedTopkPolicy {
int threads_per_block;
int items_per_thread;
BlockLoadAlgorithm load_algorithm;
};
struct bi100_default {
static constexpr int threads = 256;
static constexpr int items = 16;
static constexpr int load_algo = BLOCK_LOAD_WARP_TRANSPOSE;
};
struct policy_selector {
int key_size;
constexpr BatchedTopkPolicy operator()(const hardware_capability& hw) const {
if (hw.at_least(hardware_capability::vendor_t::iluvatar, 100)) {
return {bi100_default::threads, bi100_default::items, BLOCK_LOAD_WARP_TRANSPOSE};
}
// Fallback
return {bi100_default::threads, bi100_default::items, BLOCK_LOAD_WARP_TRANSPOSE};
}
};
} // namespace muh::tuning::batched_topk