Files
project_6/DLOPEN_DISPATCH_CHAIN.md
project6-dev 1be9449883 feat(EX): corex_gdn + corex_moe — dlopen dispatch chain from comp 168 log analysis
From 2d5232c5 docker log analysis:
  07-23 (168's docker): corex_gdn.py + corex_moe.py → full fused kernels
  08-07 (our docker): missing both → NaN GDN + PyTorch MoE fallback

corex_gdn.py: GDN fused kernel dispatch
  - FlashQLA .so loading (gdn_forward.cu pre-compiled)
  - PyTorch chunked delta rule with fp32 accum + clamp (no NaN)
  - Decode single-step recurrent with state clamping

corex_moe.py: MoE fused pipeline
  - topk_softmax: replaces MISSING ixf_F.vllm_moe_topk_softmax
  - Per-expert GEMM via torch.matmul (cublas under the hood)
  - ixformer.silu_and_mul for activation when available

DLOPEN_DISPATCH_CHAIN.md: complete .so loading chain map
deploy_corex_modules.sh: wire into VLLM/model_executor/models/
2026-08-10 03:37:15 +00:00

7.6 KiB
Raw Permalink Blame History

dlopen Dispatch Chain — BI-V100 Runtime .so Loading

Source: comp 168 docker log (2d5232c5)

Two runs in dockerrizhi.txt:

  • 07-23: Competitor 168's Docker (working, full fused kernels)
  • 08-07: Our Docker (broken MoE, NaN in GDN)

Competitor 168's Working AST Call Chain

HTTP Request → api_server.py → serving_chat.py
  → vLLM AsyncLLMEngine
    → model_runner.py:1074 (base image version, NOT our 1119)
      → qwen3_5.py (base image version with corex imports)
        │
        ├── Attention layers (32 of 36):
        │   → selector.py:115 → Using XFormers backend
        │   → ixf_F.vllm_single_query_cached_kv_attention  [ixformer .so — WORKS]
        │   → ixf_F.vllm_rotary_embedding_neox              [ixformer .so — WORKS]
        │
        ├── GDN layers (4 of 36):
        │   │
        │   ├── PREFILL:
        │   │   → corex_gdn.py:228 "Using fused CoreX GDN prefill operator"
        │   │   → corex_gdn.py:56  dlopen("/usr/local/corex/lib64/libcorex_gdn.so")
        │   │   → [chunked delta rule kernel — fp32 accumulate, NO NaN]
        │   │
        │   └── DECODE:
        │       → corex_gdn.py:138 "Using fused CoreX GDN decode operator"
        │       → [single-step recurrent kernel from libcorex_gdn.so]
        │
        ├── MoE layers (all 36):
        │   │
        │   ├── PREFILL (tokens=4096):
        │   │   → corex_moe.py:339 "Using CoreX fused MoE prefill: kernel=expert-grouped-wmma"
        │   │   → [topk routing — NOT via ixf_F, own implementation]
        │   │   → [expert GEMM via WMMA/cublas group_gemm]
        │   │   → ixf_F.silu_and_mul for activation
        │   │
        │   └── DECODE:
        │       → corex_moe.py:249 "Using CoreX fused MoE decode operator"
        │       → [same pipeline, fewer tokens]
        │
        └── Supporting ops (all via ixformer .so — confirmed working):
            → ixf_F.rms_norm
            → ixf_F.fused_add_rms_norm
            → ixf_F.vllm_cache_ops_reshape_and_cache
            → ixf_F.copy_blocks
            → ixf_F.swap_blocks

Our 08-07 Docker — What Broke

HTTP Request → api_server.py → serving_chat.py
  → vLLM AsyncLLMEngine
    → model_runner.py:1119 (OUR version, +45 lines from base)
      → qwen3_5.py (OUR version — 1500+ lines)
        │
        ├── GDN layers: ✗ NaN (99.98%)
        │   → No corex_gdn.py found
        │   → FlashQLA SM70 disabled (abs_mean=inf in test)
        │   → Falls to _torch_chunk_gated_delta_rule (our PyTorch)
        │   → qwen3_5.py:445 "NaN in prefill GatedDeltaNet layer N"
        │   → nan_to_num(0) → garbage output → quality collapse
        │
        └── MoE layers: ✗ fallback to pure PyTorch
            → No corex_moe.py found
            → Tries ixf_F.vllm_moe_topk_softmax → AttributeError (NOT IN ixformer!)
            → _custom_ops.py:58 "Error in calling custom op topk_softmax"
            → qwen3_5.py:913 "falling back to pure PyTorch experts permanently"
            → Python for-loop over 64 experts × 8 topk = ~50x slower

.so Files in Base Image

Available (confirmed by hardware probe):

/usr/local/corex/lib64/libcublas.so       ← used by torch.matmul
/usr/local/corex/lib64/libcublasLt.so     ← cublas lite
/usr/local/corex/lib64/libcuda.so         ← CUDA driver
/usr/local/corex/lib64/libcudart.so       ← CUDA runtime
/usr/local/corex/lib64/libcudnn.so        ← cuDNN
/usr/local/corex/lib64/libcutlass.so      ← CUTLASS
/usr/local/corex/lib64/libixattn.so       ← ixformer attention kernel
/usr/local/corex/lib64/libcuinfer.so      ← custom inference lib
/usr/local/corex/lib64/libixkninject.so   ← kernel injection

NOT available (must be built or bypassed):

/usr/local/corex/lib64/libcorex_gdn.so    ← GDN kernel (168 built this)
ixf_F.vllm_moe_topk_softmax              ← MoE routing (ABSENT from ixformer)
ixf_F.vllm_invoke_fused_moe_kernel       ← MoE GEMM (present but crashes)

What We Need to Build

Module 1: corex_gdn.py

Location: $VLLM/model_executor/models/corex_gdn.py Purpose: GDN fused kernel dispatch Dispatch:

  1. FlashQLA .so (gdn_forward.cu compiled on BI-V100) — needs inf fix
  2. PyTorch chunked delta rule with fp32 accumulation + clamping

Module 2: corex_moe.py

Location: $VLLM/model_executor/models/corex_moe.py Purpose: MoE fused pipeline (routing + expert GEMM + activation) Dispatch:

  1. PyTorch topk_softmax (replaces missing ixf_F.vllm_moe_topk_softmax)
  2. Per-expert torch.matmul (goes to cublas via libcublas.so)
  3. ixformer.silu_and_mul for activation (confirmed working)

Integration: patch_ops.sh additions

# Add to patch_ops.sh after line 10 (deploy corex modules):
cp /workspace/ex_engine/python/corex_gdn.py $VLLM/model_executor/models/
cp /workspace/ex_engine/python/corex_moe.py $VLLM/model_executor/models/

ixformer.functions — Confirmed API

WORKS (no errors in any log):

ixf_F.silu_and_mul(x, out)
ixf_F.gelu_and_mul(x, out)
ixf_F.gelu_tanh_and_mul(x, out)
ixf_F.rms_norm(input, weight, out, epsilon)
ixf_F.fused_add_rms_norm(input, residual, weight, epsilon)
ixf_F.vllm_single_query_cached_kv_attention(...)  → paged_attn v1
ixf_F.vllm_rotary_embedding_neox(positions, query, key, ...)
ixf_F.vllm_batched_rotary_embedding(...)
ixf_F.vllm_cache_ops_reshape_and_cache(key, value, ...)
ixf_F.reshape_and_cache_flash(...)
ixf_F.paged_attention_cache_appended(...)
ixf_F.copy_blocks(key_caches, value_caches, block_mapping)
ixf_F.swap_blocks(src, dst, block_mapping)
ixf_F.advance_step_flashattn(...)
ixf_F.w8a8(a, b, scale_a, scale_b, bias, ...)
ixf_F.w8a16(x, qweight, scales, ...)
ixf_F.static_scaled_int8_quant(output, input, scale)
ixf_F.dynamic_scaled_int8_quant(output, input, input_scales)
ixf_F.vllm_gptq_shuffle(q_weight, q_perm)
ixf_F.quantized_linear(input, qweight, scales, ...)
ixf_F.quantized_weight_dequant(...)

BROKEN/MISSING:

ixf_F.vllm_moe_topk_softmax        → AttributeError (doesn't exist)
ixf_F.vllm_invoke_fused_moe_kernel  → present but crashes (wrong BI-V100 config)
ixf_F.vllm_moe_align_block_size     → present, untested

Version Differences

Metric 168's Docker (07-23) Our Docker (08-07)
model_runner.py line :1074 :1119
Model weights 17.35 GB 16.23 GB
corex_gdn.py ✓ (built + deployed) ✗ (not found)
corex_moe.py ✓ (built + deployed) ✗ (not found)
GDN result clean (no NaN) 99.98% NaN
MoE result fused WMMA kernel PyTorch loop fallback
topk_softmax own implementation tries ixf_F (crashes)

CCCL Pattern Mapping

Kernel CCCL Algorithm .so Target
GDN prefill scan_by_key (chunked lookback) libcorex_gdn.so or PyTorch
GDN decode device_reduce (single-tile) libcorex_gdn.so or PyTorch
MoE topk device_select_if (softmax + argmax) PyTorch softmax + topk
MoE expert GEMM batch_memcpytransform (per-expert tile) cublas via torch.matmul
MoE activation transform (element-wise SiLU) ixformer.silu_and_mul
MoE scatter-add reduce_by_key (weighted accumulation) PyTorch scatter
Attention reduce (Q·K reduction) ixf_F.vllm_single_query_cached_kv_attention
Softmax scan (prefix sum for online softmax) XFormers SDPA backend
RoPE transform (element-wise rotation) ixf_F.vllm_rotary_embedding_neox
RMSNorm reduce + transform ixf_F.rms_norm