[Kernel] Add moe normal ops (#4810)

### What this PR does / why we need it? 1.Add the implementation of normal Aclnn operators: MoeCombineNormal, MoeDispatchNormal, NotifyDispatch，and DispatchLayout. - MoeCombineNormal: Implements the combine logic within MoE operations. - MoeDispatchNormal: Implements the dispatch logic within MoE operations. - NotifyDispatch: Exchanges topk_idx information among different ranks to calculate the device memory required for the dispatch stage. - DispatchLayout: Used to calculate information related to the device memory layout for the dispatch stage. 2.Provide PyTorch interfaces for normal operators—get_dispatch_layout, dispatch_prefill, and combine_prefill—to be used for MoE communication during the prefill stage in vLLM. - get_dispatch_layout: Calculates information related to the device memory layout for the dispatch operator, and is called before dispatch_prefill. - dispatch_prefill: Initiates the dispatch operation. - combine_prefill: Initiates the combine operation. ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? The functionality has already been validated using the local Qwen model. Test cases will be added after support for multi-NPU use cases in the CI pipeline is finalized. - vLLM version: v0.12.0 - vLLM main: ad32e3e19c Signed-off-by: shiro-zzzz <zhangdianhao@huawei.com>
2025-12-10 17:15:28 +08:00
parent c77dca54b2
commit bd8be2e759
39 changed files with 5365 additions and 4 deletions
--- a/csrc/dispatch_layout/op_kernel/dispatch_layout.cpp
+++ b/csrc/dispatch_layout/op_kernel/dispatch_layout.cpp
@@ -0,0 +1,17 @@
+#include "kernel_operator.h"
+#include "dispatch_layout.h"
+#include "dispatch_layout_tiling.h"
+
+
+extern "C" __global__ __aicore__ void dispatch_layout(GM_ADDR topkIdx, GM_ADDR numTokensPerRank, GM_ADDR numTokensPerExpert, 
+                                                      GM_ADDR isTokenInRank, GM_ADDR workspace, GM_ADDR tiling)
+{
+    REGISTER_TILING_DEFAULT(DispatchLayoutTilingData);
+    GET_TILING_DATA_WITH_STRUCT(DispatchLayoutTilingData, tilingData, tiling);
+
+    TPipe pipe;
+
+    DispatchLayout<int32_t> op;
+    op.Init(topkIdx, numTokensPerRank, numTokensPerExpert, isTokenInRank, workspace, &pipe, &tilingData);
+    op.Process();
+}