xc-llm-ascend

Go to file

weijinqian0 e9ada685ec [CI]Moe alltoall communication optimization (#1067 )

[CI]Moe alltoall communication optimization
The DeepSeek V3/R1 model has 256 routing experts. During parallel
inference, if the load of an EP rank is high, the overall communication
and computing time is slowed down, which becomes a weakness of parallel
inference because the load is unevenly distributed. However, the data
volume in the prefill phase is large, and the inter-card communication
time consumption/calculation time consumption and the data volume are
closely related to each other. Therefore, less non-linear precision loss
can be used to obtain a near-linear performance improvement.

During parallel inference, global synchronization occurs during
communication. As a result, the card with low load completes the
calculation first and waits for the card with the highest load to
complete the calculation. Therefore, if the load is unbalanced, the card
with high load slows down the overall time consumption. Significant
performance gains can be achieved by discarding a small number of
tokens, which is unacceptable in some precision-sensitive scenarios.
However, similar to quantification, it is a solution that uses an
acceptable precision loss in some scenarios for performance. In
addition, a trade-off between performance and precision can be achieved
by configuring a proportion of discarded tokens.

Perform the test on A3. The batch size is 8 (B), the prompt length is
3.5K tokens (S), and the parallel configuration is as follows: AttnDP=2,
AttnTP=8, MoeTP=1, and MoeEP=16. In this sence, we got a 10%-15%
performance gain.

Plus, the next version, we'll have an alltoallv moe.

---------

Signed-off-by: weijinqian_v1 <weijinqian@huawei.com>
Co-authored-by: weijinqian_v1 <weijinqian@huawei.com>

2025-06-07 10:15:56 +08:00

.github

[Worker][V1] Support sleep mode for v1 (#1084 )

2025-06-06 21:54:02 +08:00

benchmarks

[CI][Benchmark] Optimize performance benchmark workflow (#1039 )

2025-06-03 23:38:34 +08:00

cmake

[core] Support custom ascendc kernels in vllm-ascend (#233 )

2025-04-03 14:52:34 +08:00

csrc

[Performance]: Custom AscendC Kernel of Multi-Step Prepare Input (#814 )

2025-05-20 09:31:30 +08:00

docs

[Doc] Add graph mode user doc (#1083 )

2025-06-06 21:14:34 +08:00

examples

[ModelRunner] Support embedding inputs (#916 )

2025-06-06 20:21:13 +08:00

tests

[Worker][V1] Support sleep mode for v1 (#1084 )

2025-06-06 21:54:02 +08:00

tools

add workflow to build and release wheel (#775 )

2025-05-26 14:18:26 +08:00

vllm_ascend

[CI]Moe alltoall communication optimization (#1067 )

2025-06-07 10:15:56 +08:00

.gitignore

[Misc] version control by setuptools_scm (#21 )

2025-02-10 09:36:09 +08:00

.readthedocs.yaml

[Doc] Add sphinx build for vllm-ascend (#55 )

2025-02-13 18:44:17 +08:00

CMakeLists.txt

[Performance]: Custom AscendC Kernel of Multi-Step Prepare Input (#814 )

2025-05-20 09:31:30 +08:00

CODE_OF_CONDUCT.md

[Core] Init vllm-ascend (#3 )

2025-02-05 10:53:12 +08:00

collect_env.py

[CI]Add model basic accuracy test(Qwen2.5-0.5B-Instruct) (#460 )

2025-04-17 14:59:56 +08:00

DCO

[Core] Init vllm-ascend (#3 )

2025-02-05 10:53:12 +08:00

Dockerfile

[CI] upgrade to vllm 0.9.0 (#959 )

2025-05-28 21:18:41 +08:00

Dockerfile.openEuler

[CI] upgrade to vllm 0.9.0 (#959 )

2025-05-28 21:18:41 +08:00

format.sh

[CI] Refactor CI (#952 )

2025-05-28 06:31:35 +08:00

LICENSE

Initial commit

2025-01-29 02:44:13 -08:00

mypy.ini

[CI]Add model basic accuracy test(Qwen2.5-0.5B-Instruct) (#460 )

2025-04-17 14:59:56 +08:00

packages.txt

[CI/UT][PD Disaggreate] Initialize PD Disaggreate UT (#889 )

2025-05-29 10:17:12 +08:00

pyproject.toml

[CI] add codespell CI and fix format.sh (#827 )

2025-05-12 22:04:48 +08:00

pytest.ini

[CI/UT] Ignore vllm/tests/test_vllm_port.py (#887 )

2025-05-16 18:52:59 +08:00

README.md

[doc] add 0.7.3.post1 release note (#1008 )

2025-05-29 17:38:34 +08:00

README.zh.md

Upgrade CANN version to 8.1.rc1 (#747 )

2025-05-06 05:44:18 +08:00

requirements-dev.txt

[CI/UT][PD Disaggreate] Initialize PD Disaggreate UT (#889 )

2025-05-29 10:17:12 +08:00

requirements-lint.txt

[CI] Fix mypy CI (#443 )

2025-04-01 09:25:33 +08:00

requirements.txt

[CI] add codespell CI and fix format.sh (#827 )

2025-05-12 22:04:48 +08:00

setup.py

[Bug fix] fix a typo in setup.py (#762 )

2025-05-06 17:01:26 +08:00

README.md

vLLM Ascend Plugin

English | 中文

Latest News 🔥

[2025/03] We hosted the vLLM Beijing Meetup with vLLM team! Please find the meetup slides here.
[2025/02] vLLM community officially created vllm-project/vllm-ascend repo for running vLLM seamlessly on the Ascend NPU.
[2024/12] We are working with the vLLM community to support [RFC]: Hardware pluggable.

Overview

vLLM Ascend (vllm-ascend) is a community maintained hardware plugin for running vLLM seamlessly on the Ascend NPU.

It is the recommended approach for supporting the Ascend backend within the vLLM community. It adheres to the principles outlined in the [RFC]: Hardware pluggable, providing a hardware-pluggable interface that decouples the integration of the Ascend NPU with vLLM.

By using vLLM Ascend plugin, popular open-source models, including Transformer-like, Mixture-of-Expert, Embedding, Multi-modal LLMs can run seamlessly on the Ascend NPU.

Prerequisites

Hardware: Atlas 800I A2 Inference series, Atlas A2 Training series
OS: Linux
Software:
- Python >= 3.9, < 3.12
- CANN >= 8.1.RC1
- PyTorch >= 2.5.1, torch-npu >= 2.5.1
- vLLM (the same version as vllm-ascend)

Getting Started

Please refer to QuickStart and Installation for more details.

Contributing

See CONTRIBUTING for more details, which is a step-by-step guide to help you set up development environment, build and test.

We welcome and value any contributions and collaborations:

Please let us know if you encounter a bug by filing an issue
Please use User forum for usage questions and help.

Branch

vllm-ascend has main branch and dev branch.

main: main branch，corresponds to the vLLM main branch, and is continuously monitored for quality through Ascend CI.
vX.Y.Z-dev: development branch, created with part of new releases of vLLM. For example, v0.7.3-dev is the dev branch for vLLM v0.7.3 version.

Below is maintained branches:

Branch	Status	Note
main	Maintained	CI commitment for vLLM main branch and vLLM 0.9.x branch
v0.7.1-dev	Unmaintained	Only doc fixed is allowed
v0.7.3-dev	Maintained	CI commitment for vLLM 0.7.3 version

Please refer to Versioning policy for more details.

Weekly Meeting

vLLM Ascend Weekly Meeting: https://tinyurl.com/vllm-ascend-meeting
Wednesday, 15:00 - 16:00 (UTC+8, Convert to your timezone)

License

Apache License 2.0, as found in the LICENSE file.

Languages

C++ 51.3%

Python 46.3%

CMake 1%

Shell 0.7%

C 0.5%

README.md Unescape Escape

vLLM Ascend Plugin

Overview

Prerequisites

Getting Started

Contributing

Branch

Weekly Meeting

License

README.md