xc-llm-ascend

Files

wemaster 0ae9ee0f8a [BUGFIX] main-sd-bugfix && [UT] add mtp UT (#593 )

### What this PR does / why we need it?
The pr will fix some bug about spec decode / MTP
The pr add a mtp e2e UT `test_mtp_correctness.py`

**vllm_ascend/attention/attention.py**
1. add support `self.attn_mask_cache` only has 1 element to cover scene
in which both spec docode and chunked prefill are enabled.

**vllm_ascend/distributed/parallel_state.py**
1. remove 2 assert because spec decode worker would use init_worker
twice

**vllm_ascend/models/deepseek_mtp.py**
1. remove unused params;
2. add support w8a8 in `CustomDeepSeekMTP`

**vllm_ascend/quantization/quant_config.py**
1. use `AscendUnquantizedFusedMoEMethod` instead of
`UnquantizedFusedMoEMethod`

**other**
1. replace `from vllm.logger import init_logger` to `from vllm.logger
import logger` all of the vllm-ascend project



### Does this PR introduce _any_ user-facing change?


### How was this patch tested?

Signed-off-by: mengwei805 <mengwei25@huawei.com>

2025-04-21 19:25:51 +08:00

multicard

[SpecDecode] Add spec decode support (#500 )

2025-04-17 20:16:32 +08:00

ops

adopt rope in vllm-ascend (#530 )

2025-04-18 08:56:05 +08:00

scheduler

[BugFix] Fix scheduler problems in last PR. (#558 )

2025-04-18 08:49:48 +08:00

singlecard

[BUGFIX] main-sd-bugfix && [UT] add mtp UT (#593 )