xc-llm-ascend

Files

SparrowMu 54668e73c5 [Model] Support Minimax-m2.5 on NPU (#7105 )

### What this PR does / why we need it?

Initial version to support minimax-m2.5 on vllm-ascend. 
This commit coverting original fp8 weight to a quantilized bf16 to
support Minimax-m2.5 on NPU.

### Does this PR introduce _any_ user-facing change?

### How was this patch tested?

- vLLM version: v0.16.0
- vLLM main:
4034c3d32e

### Test Report
Self tested precision summary, where the official precision score of
AIME2025 is 86.3
<img width="426" height="84" alt="image"
src="https://github.com/user-attachments/assets/a3ce2452-92fa-4713-962e-862248e0b61a"
/>

---------

Signed-off-by: limuyuan <limuyuan3@huawei.com>
Signed-off-by: SparrowMu <52023119+SparrowMu@users.noreply.github.com>
Co-authored-by: limuyuan <limuyuan3@huawei.com>

2026-03-11 00:12:02 +08:00

__init__.py

[Model] Support Minimax-m2.5 on NPU (#7105 )

2026-03-11 00:12:02 +08:00

patch_balance_schedule.py

[Lint]Style: Convert vllm-ascend/ to ruff format(Batch #6 ) (#6001 )

2026-01-24 22:08:33 +08:00

patch_distributed.py

[Feature]: Support 310P device run qwen2.5/3 dense and qwen2.5vl models (#5776 )

2026-01-17 11:49:18 +08:00

patch_fusion_matcher_compat_ops.py

[Main2Main] Upgrade vLLM to 0303 (#6944 )