Commit Graph

2 Commits

Author SHA1 Message Date
muh-engine
0d810ff989 [ENGINE] muh_cc_dispatch + analysis: max_num_seqs=1 from computility-run.yaml
CRITICAL FINDING from reading computility-run.yaml:
  --max-num-seqs 1

This means the competition ALWAYS runs single-sequence inference.
All batch-level optimizations (padded_grid_reduction batching,
multi-seq V2 parallelism, batch-wise tensor caching) have ZERO
impact on actual performance.

The real bottleneck is single-sequence KV cache access:
  - decode: 1 seq × all heads × all KV blocks
  - prefill: 1 seq × chunked (max_num_batched_tokens=8192)
  - MoE: 1 seq × top_k=8 experts × 64 layers

Updated muh_cc_dispatch.py to record QWEN36_MAX_NUM_SEQS=1.

CCCL insight from padded_grid_reduction.cu: the padded grid batching
pattern is only beneficial when num_seqs > 1. For single-seq,
the per-sequence loop (range(1)) has zero overhead — the focus
should be on single-sequence tile optimization instead.

CCCL files: thrust/examples/padded_grid_reduction.cu,
cub/block/block_exchange.cuh
2026-08-06 01:02:21 +00:00
muh-engine
c0395ade14 [ENGINE] muh_cc_dispatch.py: CCCL cc_dispatch.cuh Python port
Unified kernel policy dispatch — single entry point for ALL kernel configs.

Architecture directly mirrors CCCL cc_dispatch.cuh:
  dispatch_compute_cap(policy_selector, cc, functor)
  → policy_getter<PolicySelector, CC>{}()
  → concrete policy struct

Our equivalent:
  dispatch_kernel_config('attention', hw=BI_V100)
  → pre-computed AttentionConfig (frozen dataclass)

Includes:
  - HardwareCapability (mirrors hardware.cuh bi_v100())
  - AttentionConfig (mirrors ReducePolicy for V1/V2 dispatch)
  - MoEConfig (mirrors TopkPolicy for fused_moe BLOCK_SIZE_M)
  - TransformConfig (mirrors transform bytes_in_flight)
  - CacheConfig (mirrors batch_memcpy threads)
  - grid_even_share() (Python port of GridEvenShare::DispatchInit)
  - check_smem() (SMEM constraint checker used by all policies)
  - Pre-computed configs for Qwen3.6 at import time
    (lowest_cc_resolver pattern: compute once, lookup always)

CCCL files read: cc_dispatch.cuh, dispatch_reduce.cuh,
dispatch_transform.cuh, dispatch_topk.cuh, dispatch_common.cuh,
grid_even_share.cuh, agent_reduce.cuh
2026-08-05 09:25:48 +00:00