ref(upstream): FULL TREE — Deep-Spark xllm (1470) + ds_vllm csrc/models (703)
Replaces cherry-picked upstream_ref with complete source trees. xllm/ — Iluvatar official C++ inference engine (15MB, 1470 files) Complete: kernels → layers → models → runtime → scheduler → api Excluded: .git, binary images, third_party submodule checkouts ds_vllm/ — Iluvatar official vllm fork (8MB, 703 files) Included: csrc/ (ALL CUDA kernels), fused_moe/, qwen3_5 model, _custom_ops Excluded: tests, benchmarks, docs, examples (not needed for reference) Critical call chains now fully traceable: MoE: moe_topk_softmax_kernels.cuh → ixformer.h → fused_moe.cpp → layer GDN: qwen3_gated_delta_net_base.cpp → qwen3_5_gated_delta_net.cpp Attention: ixformer.h → xllm_paged_attention → attention.cpp
This commit is contained in:
30
upstream_ref/xllm/docs/en/features/eplb.md
Normal file
30
upstream_ref/xllm/docs/en/features/eplb.md
Normal file
@@ -0,0 +1,30 @@
|
||||
# MoE Load Balancing (EPLB)
|
||||
|
||||
## Background
|
||||
|
||||
MoE models rely on dynamic token routing to distribute tokens among experts. However, in real-world deployments, uneven data distribution leads to expert load imbalance (some overloaded while others idle). Expert redundancy adjustment (e.g., adding/removing replicas) consumes additional GPU memory and may impact inference latency due to weight migration, posing significant implementation challenges. To address this, we employ an expert redundancy strategy (replicating hot experts) combined with hierarchical and global dynamic load balancing to achieve dynamic MoE load balancing.
|
||||
|
||||
## Features
|
||||
|
||||
The xLLM MoE Load Balancing (EPLB) functionality is implemented through three main modules:
|
||||
|
||||
- **EPLB Manager**: Responsible for monitoring expert loads, collecting and managing expert distribution updates. It uses a layer-by-layer update mechanism, determining whether to update each layer based on expert load changes.
|
||||
- **EPLB Executor**: The actual executor for expert distribution updates.
|
||||
- **EPLB Policy**: Strategy for generating new expert load tables.
|
||||
|
||||
Overall architecture diagram:
|
||||

|
||||
|
||||
## Usage
|
||||
|
||||
Simply add the following gflag parameters when launching xLLM:
|
||||
|
||||
(Replace with actual number of devices. `ep_size` must match the number of devices)
|
||||
|
||||
- xLLM provides the gflag parameter `enable_eplb` (default: false). Set to true in the xLLM service startup script to enable dynamic expert load balancing.
|
||||
- `expert_parallel_degree` and `ep_size` are MoE-related parameters. `expert_parallel_degree` should be set to `2`, and `ep_size` must match the actual number of NPU/GPU devices. See [moe_params](./moe_params.md)
|
||||
- `eplb_update_interval` sets the expert distribution update interval in seconds (default: 1000).
|
||||
- The expert distribution update uses a layer-by-layer mechanism based on expert load. When the similarity between consecutive loads for a layer is below `eplb_update_threshold`, that layer is updated (default: 1, range: 0-1).
|
||||
|
||||
```bash
|
||||
--enable_eplb=true --expert_parallel_degree=2 --ep_size=16 --eplb_update_interval=2000 --eplb_update_threshold=0.9
|
||||
Reference in New Issue
Block a user