xc-llm-ascend

Files

linfeng-yuan 068ed706c8 [feat][torchair] support super kernel feat for quantized dsr1 (#3485 )

### What this PR does / why we need it?
Port #1916 and #2157 to master branch to fuse operators in deepseek moe
layers, which can reduce scheduling overhead on devices. Note that this
feature is valid only when `tp_size = 1` and
`multistream_overlap_shared_expert` is enabled with torchair graph mode.

### Does this PR introduce _any_ user-facing change?
Users can enable this feature with `--additional-config
'{"torchair_graph_config":{"enabled":true, "enable_super_kernel":true},
"multistream_overlap_shared_expert":true}'`.

### How was this patch tested?
E2E deepseek serving with 2P1D disaggregated prefill scenarios.


- vLLM version: v0.11.0rc3
- vLLM main: https://github.com/vllm-project/vllm/commit/v0.11.0

---------

Signed-off-by: linfeng-yuan <1102311262@qq.com>

2025-10-20 20:04:37 +08:00

configuration

[feat][torchair] support super kernel feat for quantized dsr1 (#3485 )

2025-10-20 20:04:37 +08:00

feature_guide

[Patch]patch of v1 executor when enable eplb. (#3511 )

2025-10-19 10:54:26 +08:00

support_matrix

[ReleaseNote] Release note of v0.10.0rc1 (#2225 )

2025-08-07 14:46:49 +08:00

release_notes.md

[Doc] Release note for v0.11.0rc0 (#3224 )

2025-09-30 03:26:18 +08:00