Files
project_6/upstream_ref/README.md
EX Engine 002f9879b2 ref(upstream): FULL TREE — Deep-Spark xllm (1470) + ds_vllm csrc/models (703)
Replaces cherry-picked upstream_ref with complete source trees.

xllm/ — Iluvatar official C++ inference engine (15MB, 1470 files)
  Complete: kernels → layers → models → runtime → scheduler → api
  Excluded: .git, binary images, third_party submodule checkouts

ds_vllm/ — Iluvatar official vllm fork (8MB, 703 files)
  Included: csrc/ (ALL CUDA kernels), fused_moe/, qwen3_5 model, _custom_ops
  Excluded: tests, benchmarks, docs, examples (not needed for reference)

Critical call chains now fully traceable:
  MoE: moe_topk_softmax_kernels.cuh → ixformer.h → fused_moe.cpp → layer
  GDN: qwen3_gated_delta_net_base.cpp → qwen3_5_gated_delta_net.cpp
  Attention: ixformer.h → xllm_paged_attention → attention.cpp
2026-08-10 02:54:03 +00:00

29 lines
1.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Upstream Reference: Deep-Spark xllm + vllm (FULL TREE)
Source repos (cloned 2026-08-09, Apache 2.0):
- `Deep-Spark/xllm` — Iluvatar official C++ LLM inference engine (1470 files)
- `Deep-Spark/vllm` — Iluvatar official vllm fork (703 files, csrc + model layer)
## What's here
### xllm/ (complete source minus git/binaries/submodules)
天数智芯官方下一代推理引擎C++ 原生,多平台(CUDA/ILU/MLU/NPU)。
包含 kernels → layers → models → runtime → scheduler → api_service 完整栈。
Key subtrees:
- `xllm/core/kernels/ilu/` — ixformer API wrappers (ixformer.h是金矿)
- `xllm/core/kernels/cuda/moe/` — MoE CUDA kernels (topk_softmax, fused_topk)
- `xllm/core/kernels/cuda/` — activation, norm, rope, attention CUDA kernels
- `xllm/core/layers/ilu/` — Iluvatar FusedMoE完整pipeline
- `xllm/core/layers/npu_torch/` — GatedDeltaNet C++ implementation
- `xllm/models/llm/qwen3_5.h` — Qwen3.5 model definition
- `xllm/compiler/tilelang/` — GDN kernel code generation
### ds_vllm/ (csrc + model layers + fused_moe)
天数智芯官方vllm forkPython + CUDA torch extension。
- `csrc/` — ALL CUDA source (attention, moe, quantization, cache)
- `csrc/libtorch_stable/moe/topk_softmax_kernels.cu` — vllm topk_softmax
- `vllm/_custom_ops.py` — Python → torch.ops._moe_C bridge
- `vllm/model_executor/models/qwen3_5.py` — ds_vllm的qwen3_5实现
- `vllm/model_executor/layers/fused_moe/` — vllm FusedMoE Python layer