xc-llm-ascend

Author SHA1 Message Date

Author	SHA1	Message	Date
lilinsiman	cfe77e83ae	[Bugfix]Support Qwen3-MOE on aclgraph mode in sizes capture and add new ut (#2511 ) [Bugfix]Support Qwen3-MOE on aclgraph mode in sizes capture and add new ut What this PR does / why we need it? This PR solves the problem of sizes capture and stream error caused by using ACLgraph on the Qwen3-30B MOE model. Add new ut. Does this PR introduce any user-facing change? no How was this patch tested? ut - vLLM version: v0.10.1.1 - vLLM main: `6fad29b11b` Signed-off-by: lilinsiman <lilinsiman@gmail.com>	2025-08-26 12:39:21 +08:00
Ruri	e31b31f9c3	[main][Bugfix] Fix unable to load qwen3_moe quantized weights (#2219 ) ### What this PR does / why we need it? Fixes unable to load `qwen3_moe` quantized weights issue due to #1994 ### Does this PR introduce _any_ user-facing change? None ### How was this patch tested? Add a `qwen3_moe` W8A8 quantized model in `tests/e2e/multicard/test_qwen3_moe.py` - vLLM version: v0.10.0 - vLLM main: `c494f96fbc` --------- Signed-off-by: zhoux77899 <zhouxiang100@huawei.com>	2025-08-06 09:08:36 +08:00
wangxiyuan	458ab2db12	[BugFix] Fix the bug that qwen3 moe doesn't work with aclgraph (#2183 ) What's the PR does: 1. Move AscendSparseMoeBlock to qwen3 model, since it's only used by qwen3 model. 2. Disable AscendSparseMoeBlock if aclgraph is enabled, AscendSparseMoeBlock doesn't work with aclgraph currently. - vLLM version: v0.10.0 - vLLM main: `cdfd6871a5` Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>	2025-08-05 17:42:52 +08:00

lilinsiman

cfe77e83ae

[Bugfix]Support Qwen3-MOE on aclgraph mode in sizes capture and add new ut (#2511 )

[Bugfix]Support Qwen3-MOE on aclgraph mode in sizes capture and add new
ut

What this PR does / why we need it?
This PR solves the problem of sizes capture and stream error caused by
using ACLgraph on the Qwen3-30B MOE model.
Add new ut.

Does this PR introduce any user-facing change?
no

How was this patch tested?
ut

- vLLM version: v0.10.1.1
- vLLM main:
6fad29b11b

Signed-off-by: lilinsiman <lilinsiman@gmail.com>

2025-08-26 12:39:21 +08:00

Ruri

e31b31f9c3

[main][Bugfix] Fix unable to load qwen3_moe quantized weights (#2219 )

### What this PR does / why we need it?

Fixes unable to load `qwen3_moe` quantized weights issue due to #1994

### Does this PR introduce _any_ user-facing change?

None

### How was this patch tested?

Add a `qwen3_moe` W8A8 quantized model in
`tests/e2e/multicard/test_qwen3_moe.py`

- vLLM version: v0.10.0
- vLLM main:
c494f96fbc

---------

Signed-off-by: zhoux77899 <zhouxiang100@huawei.com>

2025-08-06 09:08:36 +08:00

wangxiyuan

458ab2db12

[BugFix] Fix the bug that qwen3 moe doesn't work with aclgraph (#2183 )

What's the PR does:
1. Move AscendSparseMoeBlock to qwen3 model, since it's only used by
qwen3 model.
2. Disable AscendSparseMoeBlock if aclgraph is enabled,
AscendSparseMoeBlock doesn't work with aclgraph currently.

- vLLM version: v0.10.0
- vLLM main:
cdfd6871a5

Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>

2025-08-05 17:42:52 +08:00

3 Commits