Dynamic Expert Load Balance with Zero-like-overhead (#2956)

### Motivation
Currently dynamically experts balancing would stop-the-world.
Asynchronously expert load balancing would be better without flowing
problems:

Host-bound latency:
There are many cpu operations during EPLB such as
eplb-algorithm、creating p2p ops、and log2phy expert converting would
spend long cpu time, as ~1s.
Communication latency: The transfer time would cost much in the
situation without nvlink. As the weight of an expert maybe transfer to
multiple new positions, thus N times send/recv for one expert, with
result long latency. We had tested that batch_isend_irecv cost more
100ms for 16 experts weight transmission in A2 server of ascend.

SwiftBalancer would not stop-the-world anymore, in out test on NPU 1~2ms
cost for each layer while benefit 5ms-8ms decode latency with ep_size =
64.
The following updates have been made:
1、expert distribution recording with lower cost.
2、async cpu computing for eplb algo and other python operator.
3、new eplb algo with less expert rebalancing while almost the same
effect.
### Proposed Change
We will gradually migrate the EPLB logic to the VLLM community and
implement a generalized design. Relevant RFC:
https://github.com/vllm-project/vllm/issues/22246
The overall workflow involves:
<img width="801" height="302"
alt="474430541-23b06f58-23bc-44a3-a1be-00f268aeb15c"
src="https://github.com/user-attachments/assets/1d73a459-1b23-4b0a-812a-bf0a75debfed"
/>
1. Record experts distribution during forward. We using expert_token_num
after disptach instead of topk_ids, thus we got much smaller tensor
shape to reduce cost of hbm recording and add-operator.
2. Do all-gather for experts distribution. Using all-gather instead of
all-reduce as less traffic volume.
3. Wake up eplb worker process with experts distribution when
num_iterations comes. Run eplb algorithm in eplb worker.
4. Generate p2p send/recv ops and other operator such as log2phy would
cost long cpu time.
5. Lanch ibatch_send_recv in async_stream before forward.
6. After forward, wait for the ibatch_send_recv finish, then do uapte
expert map and expert weights.
### Co-author
Co-authored-by: raindaywhu raindaywhu@raindaywhu@ 163.con
Co-authored-by: njuyuan yuanjl19@smail.nju.edu.cn
Co-authored-by: qmkakaxi wjh1594260677@qq.com
Co-authored-by: Skywalker-EP 173723846@qq.com


- vLLM version: v0.10.2
- vLLM main:
567939953b

---------

Signed-off-by: offline0806 <z00858301@china.huawei.com>
Co-authored-by: offline0806 <z00858301@china.huawei.com>

This commit is contained in:

offline893

2025-09-17 10:36:43 +08:00

committed by

GitHub

parent ae758dda05

commit 76844eec78

30 changed files with 2891 additions and 47 deletions

									
										6

vllm_ascend/ops/moe/moe_comm_method.py
									
												View File
												
				@@ -88,7 +88,8 @@ class MoECommMethod(ABC):

				            # For load balance

				            log2phy: torch.Tensor = None,

				            global_redundant_expert_num: int = 0,

				            need_trans: bool = False) -> torch.Tensor:

				            need_trans: bool = False,

				            dynamic_eplb: bool = False):

				        # Check constraints

				        assert hidden_states.dtype in [

				            torch.float32, torch.float16, torch.bfloat16

				@@ -133,6 +134,9 @@ class MoECommMethod(ABC):

				        final_hidden_states = self.token_dispatcher.token_combine(

				            hidden_states=mlp_output)

				        if dynamic_eplb:

				            return (final_hidden_states, group_list_type, expert_tokens)

				        return final_hidden_states

				    @abstractmethod

Dynamic Expert Load Balance with Zero-like-overhead (#2956)

6 vllm_ascend/ops/moe/moe_comm_method.py Unescape Escape View File

6

vllm_ascend/ops/moe/moe_comm_method.py

View File