Files
project_6/upstream_ref
claude 051b02d3cd feat: ix_moe_bridge + ix_attn_bridge — dlopen bridges for full ixformer::infer API
Bridge architecture (from xllm/core/kernels/ilu/ixformer.h):

ix_moe_bridge.so (MoE 7-step fused pipeline):
  - topk_softmax → moe_compute_token_index_api → moe_expand_input
  - moe_w16a16_group_gemm (x2) → silu_and_mul → moe_output_reduce_sum
  - fused_moe_forward(): replaces entire Python expert loop
  - Fix: group_gemm format NT→TN (match xllm trans_b=true)

ix_attn_bridge.so (attention + linear):
  - ixinfer_flash_attn_unpad_with_block_tables (fused prefill)
  - xllm_paged_attention (fused paged decode)
  - ixformer_linear (matmul + activation)
  - residual_rms_norm (fused residual + norm)

Integration:
  - ix_fused_moe.py: Python loader (prebuilt .so → JIT → unavailable)
  - qwen3_5.py: Tier 0 dispatch in _pure_pytorch_experts()
  - patch_ops.sh: deploys ix_fused_moe.py + all prebuilt/*.so

Source: jd-opensource/xllm (fresh clone, all ILU kernels verified SAME)
Sync: upstream_ref/xllm_latest/models/llm/qwen3_next_hybrid_base.h (+32 lines)

Build on real machine:
  bash qwen3_6_scripts/build_ix_moe_bridge.sh
  bash qwen3_6_scripts/build_ix_attn_bridge.sh
2026-08-14 07:32:31 +00:00
..

Upstream Reference: Deep-Spark xllm + vllm (FULL TREE)

Source repos (cloned 2026-08-09, Apache 2.0):

  • Deep-Spark/xllm — Iluvatar official C++ LLM inference engine (1470 files)
  • Deep-Spark/vllm — Iluvatar official vllm fork (703 files, csrc + model layer)

What's here

xllm/ (complete source minus git/binaries/submodules)

天数智芯官方下一代推理引擎C++ 原生,多平台(CUDA/ILU/MLU/NPU)。 包含 kernels → layers → models → runtime → scheduler → api_service 完整栈。

Key subtrees:

  • xllm/core/kernels/ilu/ — ixformer API wrappers (ixformer.h是金矿)
  • xllm/core/kernels/cuda/moe/ — MoE CUDA kernels (topk_softmax, fused_topk)
  • xllm/core/kernels/cuda/ — activation, norm, rope, attention CUDA kernels
  • xllm/core/layers/ilu/ — Iluvatar FusedMoE完整pipeline
  • xllm/core/layers/npu_torch/ — GatedDeltaNet C++ implementation
  • xllm/models/llm/qwen3_5.h — Qwen3.5 model definition
  • xllm/compiler/tilelang/ — GDN kernel code generation

ds_vllm/ (csrc + model layers + fused_moe)

天数智芯官方vllm forkPython + CUDA torch extension。

  • csrc/ — ALL CUDA source (attention, moe, quantization, cache)
  • csrc/libtorch_stable/moe/topk_softmax_kernels.cu — vllm topk_softmax
  • vllm/_custom_ops.py — Python → torch.ops._moe_C bridge
  • vllm/model_executor/models/qwen3_5.py — ds_vllm的qwen3_5实现
  • vllm/model_executor/layers/fused_moe/ — vllm FusedMoE Python layer