Logo
Explore Help
Register Sign In
dylanyunlong/project_6_89d52222
1
0
Fork 0
You've already forked project_6_89d52222
Code Issues Pull Requests Actions Projects Releases Wiki Activity
Files
a875fa5d4c23f8ab8d87c49d84a59dc93dfe16ce
project_6_89d52222/upstream_ref/xllm/docs/en/features/basics.md

6 lines
428 B
Markdown
Raw Normal View History

ref(upstream): FULL TREE — Deep-Spark xllm (1470) + ds_vllm csrc/models (703) Replaces cherry-picked upstream_ref with complete source trees. xllm/ — Iluvatar official C++ inference engine (15MB, 1470 files) Complete: kernels → layers → models → runtime → scheduler → api Excluded: .git, binary images, third_party submodule checkouts ds_vllm/ — Iluvatar official vllm fork (8MB, 703 files) Included: csrc/ (ALL CUDA kernels), fused_moe/, qwen3_5 model, _custom_ops Excluded: tests, benchmarks, docs, examples (not needed for reference) Critical call chains now fully traceable: MoE: moe_topk_softmax_kernels.cuh → ixformer.h → fused_moe.cpp → layer GDN: qwen3_gated_delta_net_base.cpp → qwen3_5_gated_delta_net.cpp Attention: ixformer.h → xllm_paged_attention → attention.cpp
2026-08-10 02:53:54 +00:00
# Basics
- xLLM uses a one-device-per-process architecture. Across multiple devices, RPC is used for function calls, and data communication during model computation uses device collective communication libraries.
- HCCL/LCCL are high-performance collective communication frameworks that provide data-parallel and model-parallel collective communication for both single-node multi-device and multi-node multi-device scenarios.
Reference in New Issue Copy Permalink
Powered by Gitea Version: 1.24.3 Page: 202ms Template: 1ms
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API