Files
project_6/upstream_ref
claude 8d75652949 feat: import CUDA kernels from xllm/CCCL/FLA upstream repos
Sources cloned and tree'd (no --depth):
  - jd-opensource/xllm: ILU kernels, CUDA kernels, MoE kernels
  - NVIDIA/cccl: CUB tuning/dispatch headers (block-level primitives)
  - fla-org/flash-linear-attention: Triton GDN kernels
  - NVIDIA/cutlass: grouped GEMM reference (read, not copied)
  - Dao-AILab/flash-attention: attention kernel reference (SM80+, read only)

New CUDA kernels (from xllm, SM-agnostic, portable to BI-V100):
  ex_engine/xllm_kernels/cuda/activation.cu    (188 lines) — silu_and_mul, gelu
  ex_engine/xllm_kernels/cuda/norm.cu          (600 lines) — rms_norm, fused_add_rms_norm
  ex_engine/xllm_kernels/cuda/rope.cu          (258 lines) — rotary_embedding
  ex_engine/xllm_kernels/cuda/block_copy.cu    (209 lines) — copy_blocks, swap_blocks
  ex_engine/xllm_kernels/cuda/reshape_paged_cache.cu (101 lines) — KV cache ops
  ex_engine/xllm_kernels/cuda/headers/         (5 headers for compilation)

ILU bridge kernel sources (from xllm, verified SAME as upstream):
  ex_engine/xllm_kernels/ilu/    (10 files, 925 lines total)
  — activation.cpp, attention.cpp, fused_moe.cpp, group_gemm.cpp,
    matmul.cpp, norm.cpp, rope.cpp, ilu_ops_api.h, ixformer.h, utils.h

FLA Triton GDN kernels (for GatedDeltaNet without SM90+ FlashQLA):
  ex_engine/fla_kernels/gated_delta_rule/  (7 files, 2370 lines)
  — chunk_fwd.py (428), chunk.py (487), wy_fast.py (409),
    fused_recurrent.py (392), naive.py (161), gate.py (380)

CCCL sync (12 tuning + 14 dispatch headers updated from NVIDIA/cccl):
  cccl_upstream/cub/cub/device/dispatch/tuning/ — 12 changed files synced
  cccl_upstream/cub/cub/device/dispatch/ — 14 changed dispatch files synced

Compilation targets for real machine (ivcore10):
  1. CUDA kernels: --cuda-gpu-arch=ivcore10 via corex clang/16
  2. ILU bridges: torch.utils.cpp_extension linking ixformer .so
  3. FLA kernels: Triton JIT (if Triton works on BI-V100)
2026-08-14 07:48:52 +00:00
..

Upstream Reference: Deep-Spark xllm + vllm (FULL TREE)

Source repos (cloned 2026-08-09, Apache 2.0):

  • Deep-Spark/xllm — Iluvatar official C++ LLM inference engine (1470 files)
  • Deep-Spark/vllm — Iluvatar official vllm fork (703 files, csrc + model layer)

What's here

xllm/ (complete source minus git/binaries/submodules)

天数智芯官方下一代推理引擎C++ 原生,多平台(CUDA/ILU/MLU/NPU)。 包含 kernels → layers → models → runtime → scheduler → api_service 完整栈。

Key subtrees:

  • xllm/core/kernels/ilu/ — ixformer API wrappers (ixformer.h是金矿)
  • xllm/core/kernels/cuda/moe/ — MoE CUDA kernels (topk_softmax, fused_topk)
  • xllm/core/kernels/cuda/ — activation, norm, rope, attention CUDA kernels
  • xllm/core/layers/ilu/ — Iluvatar FusedMoE完整pipeline
  • xllm/core/layers/npu_torch/ — GatedDeltaNet C++ implementation
  • xllm/models/llm/qwen3_5.h — Qwen3.5 model definition
  • xllm/compiler/tilelang/ — GDN kernel code generation

ds_vllm/ (csrc + model layers + fused_moe)

天数智芯官方vllm forkPython + CUDA torch extension。

  • csrc/ — ALL CUDA source (attention, moe, quantization, cache)
  • csrc/libtorch_stable/moe/topk_softmax_kernels.cu — vllm topk_softmax
  • vllm/_custom_ops.py — Python → torch.ops._moe_C bridge
  • vllm/model_executor/models/qwen3_5.py — ds_vllm的qwen3_5实现
  • vllm/model_executor/layers/fused_moe/ — vllm FusedMoE Python layer