44
docs/source/user_guide/support_matrix/feature_matrix.md
Normal file
44
docs/source/user_guide/support_matrix/feature_matrix.md
Normal file
@@ -0,0 +1,44 @@
|
||||
# Feature Matrix
|
||||
|
||||
The table below shows mutually exclusive features and the support on Ascend hardware, extended from the [vLLM table](https://docs.vllm.ai/en/latest/features/#feature-x-feature).
|
||||
|
||||
The symbols used have the following meanings:
|
||||
|
||||
- ✅ = Full compatibility
|
||||
- 🟠 = Partial compatibility
|
||||
- ❌ = No compatibility
|
||||
- ❔ = Unknown or TBD
|
||||
|
||||
| Feature | [ACLGraph Full_Decode_Only](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/Design_Documents/ACL_Graph.html) | [ACLGraph Piecewise](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/Design_Documents/ACL_Graph.html) | Async Scheduling | [<abbr title="Automatic Prefix Caching">APC</abbr>](https://docs.vllm.ai/en/latest/features/automatic_prefix_caching/) | [Chunked Prefill](https://docs.vllm.ai/en/stable/configuration/optimization/#chunked-prefill) | [Context Parallel](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/Design_Documents/context_parallel.html) | [Cpu Binding](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/Design_Documents/cpu_binding.html) | [<abbr title="Data Parallel">DP</abbr>](https://docs.vllm.ai/en/latest/serving/data_parallel_deployment/) | [Disaggregated Prefill](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/Design_Documents/disaggregated_prefill.html) | [Eagle3](https://docs.vllm.ai/en/latest/features/speculative_decoding/eagle/) | [Eplb](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/Design_Documents/eplb_swift_balancer.html) | [<abbr title="Expert-Parallel">EP</abbr>](https://docs.vllm.ai/en/latest/serving/expert_parallel_deployment/) | Flashcomm1 | [KV Cache Pool](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/Design_Documents/KV_Cache_Pool_Guide.html) | Layer Sharding | Lmhead TP | Mlapo | [<abbr title="Multimodal Inputs">mm</abbr>](https://docs.vllm.ai/en/latest/features/multimodal_inputs/) | Multistream Moe | Shared Expert DP | [Quantization W4A4](https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/quantization.html) | [Quantization W4A8](https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/quantization.html) | [Quantization W8A8](https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/quantization.html) | <abbr title="Tensor Parallel">TP</abbr> | Weight nz |
|
||||
| - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - |
|
||||
| [ACLGraph Full_Decode_Only](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/Design_Documents/ACL_Graph.html) | ✅ | | | | | | | | | | | | | | | | | | | | | | | | |
|
||||
| [ACLGraph Piecewise](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/Design_Documents/ACL_Graph.html) | ❌ | ✅ | | | | | | | | | | | | | | | | | | | | | | | |
|
||||
| Async Scheduling | ✅ | ✅ | ✅ | | | | | | | | | | | | | | | | | | | | | | |
|
||||
| [<abbr title="Automatic Prefix Caching">APC</abbr>](https://docs.vllm.ai/en/latest/features/automatic_prefix_caching/) | ✅ | ✅ | ✅ | ✅ | | | | | | | | | | | | | | | | | | | | | |
|
||||
| [Chunked Prefill](https://docs.vllm.ai/en/stable/configuration/optimization/#chunked-prefill) | ✅ | ✅ | ✅ | ✅ | ✅ | | | | | | | | | | | | | | | | | | | | |
|
||||
| [Context Parallel](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/Design_Documents/context_parallel.html) | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | | | | | | | | | | | | | | | | | | | |
|
||||
| [Cpu Binding](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/Design_Documents/cpu_binding.html) | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | | | | | | | | | | | | | | | | | | |
|
||||
| [<abbr title="Data Parallel">DP</abbr>](https://docs.vllm.ai/en/latest/serving/data_parallel_deployment/) | ✅ | ✅ | ✅ | ✅ | ✅ | 🟠<sup>1</sup> | ✅ | ✅ | | | | | | | | | | | | | | | | | |
|
||||
| [Disaggregated Prefill](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/Design_Documents/disaggregated_prefill.html) | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | | | | | | | | | | | | | | | | |
|
||||
| [Eagle3](https://docs.vllm.ai/en/latest/features/speculative_decoding/eagle/) | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | | | | | | | | | | | | | | | |
|
||||
| [Eplb](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/Design_Documents/eplb_swift_balancer.html) | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | | | | | | | | | | | | | | |
|
||||
| [<abbr title="Expert-Parallel">EP</abbr>](https://docs.vllm.ai/en/latest/serving/expert_parallel_deployment/) | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | | | | | | | | | | | | | |
|
||||
| Flashcomm1 | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 🟠<sup>2</sup> | ✅ | ✅ | ✅ | ✅ | | | | | | | | | | | | |
|
||||
| [KV Cache Pool](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/Design_Documents/KV_Cache_Pool_Guide.html) | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | | | | | | | | | | | |
|
||||
| Layer Sharding | ✅ | ✅ | ✅ | ✅ | ✅ | 🟠 | ✅ | ✅ | 🟠<sup>3</sup> | ✅ | ✅ | ✅ | ✅ | ❔ | ✅ | | | | | | | | | | |
|
||||
| Lmhead TP | ✅ | ✅ | ✅ | ✅ | ✅ | ❔ | ✅ | 🟠<sup>4</sup> | ✅ | ✅ | ✅ | ✅ | ❌ | ❔ | ✅ | ✅ | | | | | | | | | |
|
||||
| Mlapo | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 🟠<sup>5</sup> | ✅ | ✅ | ✅ | ❌ | ❔ | ❌ | ✅ | ✅ | | | | | | | | |
|
||||
| [<abbr title="Multimodal Inputs">mm</abbr>](https://docs.vllm.ai/en/latest/features/multimodal_inputs/) | ✅ | ✅ | ✅ | ✅ | ✅ | 🟠 | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | ✅ | ✅ | | | | | | | |
|
||||
| Multistream Moe | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ❔ | ✅ | ✅ | ✅ | ✅ | | | | | | |
|
||||
| Shared Expert DP | ✅ | ✅ | ✅ | ✅ | ✅ | 🟠<sup>1</sup> | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ❔ | ✅ | ✅ | ❔ | ✅ | | | | | |
|
||||
| [Quantization W4A4](https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/quantization.html) | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | ❔ | ❔ | ✅ | ❔ | ✅ | ❔ | ❌ | ❔ | ❔ | ✅ | | | | |
|
||||
| [Quantization W4A8](https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/quantization.html) | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | ✅ | ❔ | ✅ | ❔ | ❌ | ✅ | ✅ | ❔ | ✅ | | | |
|
||||
| [Quantization W8A8](https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/quantization.html) | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ❔ | ✅ | ✅ | | |
|
||||
| <abbr title="Tensor Parallel">TP</abbr> | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | |
|
||||
| Weight nz | ✅ | ✅ | ✅ | ✅ | ✅ | ❔ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | 🟠 | ✅ | ✅ | ✅ |
|
||||
|
||||
- <sup>1</sup> Only dcp supports dp while pcp does not support dp.
|
||||
- <sup>2</sup> Flashcomm is only enabled on the prefill stage.
|
||||
- <sup>3</sup> Layer sharding is only enabled on the prefill stage.
|
||||
- <sup>4</sup> Lmhead TP is only enabled in the pure dp scenarios.
|
||||
- <sup>5</sup> MLAPO is only supported on the decode stage.
|
||||
@@ -1,10 +1,11 @@
|
||||
# Features and models
|
||||
# Features and Models
|
||||
|
||||
This section provides a detailed supported matrix by vLLM Ascend.
|
||||
This section provides a detailed matrix supported by vLLM Ascend.
|
||||
|
||||
:::{toctree}
|
||||
:caption: Support Matrix
|
||||
:maxdepth: 1
|
||||
supported_models
|
||||
supported_features
|
||||
feature_matrix
|
||||
:::
|
||||
|
||||
@@ -1,45 +1,48 @@
|
||||
# Feature Support
|
||||
# Supported Features
|
||||
|
||||
The feature support principle of vLLM Ascend is: **aligned with the vLLM**. We are also actively collaborating with the community to accelerate support.
|
||||
The feature support principle of vLLM Ascend is: **aligned with vLLM**. We are also actively collaborating with the community to accelerate support.
|
||||
|
||||
Functional call: <https://docs.vllm.ai/en/latest/features/tool_calling/>
|
||||
|
||||
You can check the [support status of vLLM V1 Engine][v1_user_guide]. Below is the feature support status of vLLM Ascend:
|
||||
|
||||
| Feature | Status | Next Step |
|
||||
|-------------------------------|----------------|------------------------------------------------------------------------|
|
||||
| Chunked Prefill | 🟢 Functional | Functional, see detail note: [Chunked Prefill][cp] |
|
||||
| Automatic Prefix Caching | 🟢 Functional | Functional, see detail note: [vllm-ascend#732][apc] |
|
||||
| LoRA | 🟢 Functional | [vllm-ascend#396][multilora], [vllm-ascend#893][v1 multilora] |
|
||||
| Speculative decoding | 🟢 Functional | Basic support |
|
||||
| Pooling | 🟢 Functional | CI needed and adapting more models; V1 support rely on vLLM support. |
|
||||
| Enc-dec | 🟡 Planned | vLLM should support this feature first. |
|
||||
| Multi Modality | 🟢 Functional | [Tutorial][multimodal], optimizing and adapting more models |
|
||||
| LogProbs | 🟢 Functional | CI needed |
|
||||
| Prompt logProbs | 🟢 Functional | CI needed |
|
||||
| Async output | 🟢 Functional | CI needed |
|
||||
| Beam search | 🟢 Functional | CI needed |
|
||||
| Guided Decoding | 🟢 Functional | [vllm-ascend#177][guided_decoding] |
|
||||
| Tensor Parallel | 🟢 Functional | Make TP >4 work with graph mode |
|
||||
| Pipeline Parallel | 🟢 Functional | Write official guide and tutorial. |
|
||||
| Expert Parallel | 🟢 Functional | Dynamic EPLB support. |
|
||||
| Data Parallel | 🟢 Functional | Data Parallel support for Qwen3 MoE. |
|
||||
| Prefill Decode Disaggregation | 🟢 Functional | Functional, xPyD is supported. |
|
||||
| Quantization | 🟢 Functional | W8A8 available; working on more quantization method support(W4A8, etc) |
|
||||
| Graph Mode | 🔵 Experimental| Experimental, see detail note: [vllm-ascend#767][graph_mode] |
|
||||
| Sleep Mode | 🟢 Functional | |
|
||||
| Chunked Prefill | 🟢 Functional | Functional, see detailed note: [Chunked Prefill][cp] |
|
||||
| Automatic Prefix Caching | 🟢 Functional | Functional, see detailed note: [vllm-ascend#732][apc] |
|
||||
| LoRA | 🔵 Experimental | Functional, see detailed note: [LoRA][LoRA] |
|
||||
| Speculative decoding | 🟢 Functional | Basic support |
|
||||
| Pooling | 🔵 Experimental | CI needed to adapt to more models; V1 support relies on vLLM support. |
|
||||
| Enc-dec | 🟡 Planned | vLLM should support this feature first. |
|
||||
| Multi Modality | 🟢 Functional | [Multi Modality][multimodal], optimizing and adapting more models |
|
||||
| LogProbs | 🟢 Functional | CI needed |
|
||||
| Prompt logProbs | 🟢 Functional | CI needed |
|
||||
| Async output | 🟢 Functional | CI needed |
|
||||
| Beam search | 🔵 Experimental | CI needed |
|
||||
| Guided Decoding | 🟢 Functional | [vllm-ascend#177][guided_decoding] |
|
||||
| Tensor Parallel | 🟢 Functional | Make TP >4 work with graph mode. |
|
||||
| Pipeline Parallel | 🟢 Functional | Write official guide and tutorial. |
|
||||
| Expert Parallel | 🟢 Functional | Support dynamic EPLB. |
|
||||
| Data Parallel | 🟢 Functional | Data Parallel support for Qwen3 MoE. |
|
||||
| Prefill Decode Disaggregation | 🟢 Functional | Functional, xPyD is supported. |
|
||||
| Quantization | 🟢 Functional | W8A8 available; working on more quantization method support (W4A8, etc) |
|
||||
| Graph Mode | 🟢 Functional | Functional, see detailed note: [Graph Mode][graph_mode] |
|
||||
| Sleep Mode | 🟢 Functional | Functional, see detailed note: [Sleep Mode][sleep_mode] |
|
||||
| Context Parallel | 🟢 Functional | Functional, see detailed note: [Context Parallel][context_parallel] |
|
||||
|
||||
- 🟢 Functional: Fully operational, with ongoing optimizations.
|
||||
- 🔵 Experimental: Experimental support, interfaces and functions may change.
|
||||
- 🚧 WIP: Under active development, will be supported soon.
|
||||
- 🟡 Planned: Scheduled for future implementation (some may have open PRs/RFCs).
|
||||
- 🔴 NO plan / Deprecated: No plan or deprecated by vLLM.
|
||||
- 🔴 NO plan/Deprecated: No plan or deprecated by vLLM.
|
||||
|
||||
[v1_user_guide]: https://docs.vllm.ai/en/latest/getting_started/v1_user_guide.html
|
||||
[multimodal]: https://vllm-ascend.readthedocs.io/en/latest/tutorials/single_npu_multimodal.html
|
||||
[v1_user_guide]: https://docs.vllm.ai/en/latest/usage/v1_guide/
|
||||
[multimodal]: https://docs.vllm.ai/projects/ascend/en/latest/tutorials/models/Qwen-VL-Dense.html
|
||||
[guided_decoding]: https://github.com/vllm-project/vllm-ascend/issues/177
|
||||
[multilora]: https://github.com/vllm-project/vllm-ascend/issues/396
|
||||
[v1 multilora]: https://github.com/vllm-project/vllm-ascend/pull/893
|
||||
[graph_mode]: https://github.com/vllm-project/vllm-ascend/issues/767
|
||||
[LoRA]: https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/lora.html
|
||||
[graph_mode]: https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/graph_mode.html
|
||||
[apc]: https://github.com/vllm-project/vllm-ascend/issues/732
|
||||
[cp]: https://docs.vllm.ai/en/stable/performance/optimization.html#chunked-prefill
|
||||
[cp]: https://docs.vllm.ai/en/stable/configuration/optimization/
|
||||
[1P1D]: https://github.com/vllm-project/vllm-ascend/pull/950
|
||||
[ray]: https://github.com/vllm-project/vllm-ascend/issues/1751
|
||||
[context_parallel]: https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/context_parallel.html
|
||||
[sleep_mode]: https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/sleep_mode.html
|
||||
|
||||
@@ -1,79 +1,194 @@
|
||||
# Model Support
|
||||
# Supported Models
|
||||
|
||||
Get the newest info here: https://github.com/vllm-project/vllm-ascend/issues/1608
|
||||
Get the latest info here: <https://github.com/vllm-project/vllm-ascend/issues/1608>
|
||||
|
||||
## Text-only Language Models
|
||||
**Legend Description**:
|
||||
|
||||
- ✅ = Supported model/feature
|
||||
- 🔵 = Experimental supported model/feature
|
||||
- ❌ = Not supported model/feature
|
||||
- 🟡 = Not tested or verified
|
||||
|
||||
## Text-Only Language Models
|
||||
|
||||
### Generative Models
|
||||
|
||||
| Model | Supported | Note |
|
||||
|-------------------------------|-----------|----------------------------------------------------------------------|
|
||||
| DeepSeek v3 | ✅ | |
|
||||
| DeepSeek R1 | ✅ | |
|
||||
| DeepSeek Distill (Qwen/LLama) | ✅ | |
|
||||
| Qwen3 | ✅ | |
|
||||
| Qwen3-based | ✅ | |
|
||||
| Qwen3-Coder | ✅ | |
|
||||
| Qwen3-Moe | ✅ | |
|
||||
| Qwen2.5 | ✅ | |
|
||||
| Qwen2 | ✅ | |
|
||||
| Qwen2-based | ✅ | |
|
||||
| QwQ-32B | ✅ | |
|
||||
| LLama2/3/3.1 | ✅ | |
|
||||
| Internlm | ✅ | [#1962](https://github.com/vllm-project/vllm-ascend/issues/1962) |
|
||||
| Baichuan | ✅ | |
|
||||
| Baichuan2 | ✅ | |
|
||||
| Phi-4-mini | ✅ | |
|
||||
| MiniCPM | ✅ | |
|
||||
| MiniCPM3 | ✅ | |
|
||||
| Ernie4.5 | ✅ | |
|
||||
| Ernie4.5-Moe | ✅ | |
|
||||
| Gemma-2 | ✅ | |
|
||||
| Gemma-3 | ✅ | |
|
||||
| Phi-3/4 | ✅ | |
|
||||
| Mistral/Mistral-Instruct | ✅ | |
|
||||
| GLM-4.5 | ✅ | |
|
||||
| GLM-4 | ❌ | [#2255](https://github.com/vllm-project/vllm-ascend/issues/2255) |
|
||||
| GLM-4-0414 | ❌ | [#2258](https://github.com/vllm-project/vllm-ascend/issues/2258) |
|
||||
| ChatGLM | ❌ | [#554](https://github.com/vllm-project/vllm-ascend/issues/554) |
|
||||
| DeepSeek v2.5 | 🟡 | Need test |
|
||||
| Mllama | 🟡 | Need test |
|
||||
| MiniMax-Text | 🟡 | Need test |
|
||||
#### Core Supported Models
|
||||
|
||||
:::::{tab-set}
|
||||
|
||||
::::{tab-item} Ascend 950 Products
|
||||
|
||||
| Model | Support | Note | BF16 | Supported Hardware | W8A8 | Chunked Prefill | Automatic Prefix Cache | LoRA | Speculative Decoding | Async Scheduling | Tensor Parallel | Pipeline Parallel | Expert Parallel | Data Parallel | Prefill-decode Disaggregation | Piecewise AclGraph | Fullgraph AclGraph | max-model-len | MLP Weight Prefetch | Doc |
|
||||
|-------|--------|--------|------|------|------|---------|-------|------|------|--------|-------|--------|--------|-------|-------|--------|----------|---------|----------|-----|
|
||||
|DeepSeek V4-Flash|✅|Native mixed MXFP8/MXFP4 weights||Ascend 950 Products|✅|✅|✅||✅|✅||✅|✅|✅|✅||✅|1M||[DeepSeek V4-Flash](../../tutorials/models/DeepSeek-V4-Flash.md)|
|
||||
|DeepSeek V4-Pro|✅|Native mixed MXFP8/MXFP4 weights||Ascend 950 Products|✅|✅|✅||✅|✅||✅|✅|✅|✅||✅|1M||[DeepSeek V4-Pro](../../tutorials/models/DeepSeek-V4-Pro.md)|
|
||||
|DeepSeek-V3.1|✅| |✅| Ascend 950 Products |✅|✅|✅||✅|✅|✅||✅|✅|✅|✅|✅|240k|| [DeepSeek-V3.1](../../tutorials/models/DeepSeek-V3.1.md) |
|
||||
|GLM-5.1|✅| |✅| Ascend 950 Products |✅|✅|✅||✅|✅|✅||✅|✅|✅||✅|200k||[GLM-5.1](../../tutorials/models/GLM5.md) |
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} A2/A3
|
||||
|
||||
| Model | Support | Note | BF16 | Supported Hardware | W8A8 | Chunked Prefill | Automatic Prefix Cache | LoRA | Speculative Decoding | Async Scheduling | Tensor Parallel | Pipeline Parallel | Expert Parallel | Data Parallel | Prefill-decode Disaggregation | Piecewise AclGraph | Fullgraph AclGraph | max-model-len | MLP Weight Prefetch | Doc |
|
||||
|-------------------------------|---------|-----------------------------------------------------|------|--------------------|------|-----------------|------------------------|------|----------------------|------------------|-----------------|-------------------|-----------------|---------------|-------------------------------|--------------------|--------------------|---------------|---------------------|-----|
|
||||
| DeepSeek V4-Flash | ✅ | | ✅ | A2/A3 | ✅ | ✅ |✅|| ✅ |✅| ✅ || ✅ | ✅ | ✅ || ✅ | 1M || [DeepSeek V4-Flash](../../tutorials/models/DeepSeek-V4-Flash.md) |
|
||||
| DeepSeek V4-Pro | ✅ | | ✅ | A2/A3 | ✅ | ✅ |✅|| ✅ |✅| ✅ || ✅ | ✅ | ✅ || ✅ | 1M || [DeepSeek-V4-Pro](../../tutorials/models/DeepSeek-V4-Pro.md) |
|
||||
| DeepSeek V3/3.1 | ✅ | | ✅ | A2/A3 | ✅ | ✅ | ✅ || ✅ || ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 240k || [DeepSeek-V3.1](../../tutorials/models/DeepSeek-V3.1.md) |
|
||||
| DeepSeek V3.2 | ✅ | | ✅ | A2/A3 | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 160k | ✅ | [DeepSeek-V3.2](../../tutorials/models/DeepSeek-V3.2.md) |
|
||||
| DeepSeek R1 | ✅ | | ✅ | A2/A3 | ✅ | ✅ | ✅ || ✅ || ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 128k || [DeepSeek-R1](../../tutorials/models/DeepSeek-R1.md) |
|
||||
| Qwen3-Dense | ✅ | | ✅ | A2/A3 | ✅ | ✅ | ✅ ||| ✅ | ✅ ||| ✅ || ✅ | ✅ | 128k | ✅ | [Qwen3-Dense](../../tutorials/models/Qwen3-Dense.md) |
|
||||
| Qwen3-30B-A3B | ✅ | | ✅ | A2/A3 | ✅ | ✅ | ✅ || ✅ | ✅ | ✅ || ✅ | ✅ || ✅ | ✅ ||| [Qwen3-30B-A3B](../../tutorials/models/Qwen3-30B-A3B.md) |
|
||||
| Qwen3-Coder-30B-A3B | 🔵 | | ✅ | A2/A3 | ✅ | ✅ | ✅ || ✅ | ✅ | ✅ || ✅ | ✅ || ✅ | ✅ ||| [Qwen3-Coder-30B-A3B](../../tutorials/models/Qwen3-Coder-30B-A3B.md) |
|
||||
| Qwen3-235B-A22B | ✅ | | ✅ | A2/A3 | ✅ | ✅ | ✅ ||| ✅ | ✅ || ✅ | ✅ | ✅ | ✅ | ✅ | 256k || [Qwen3-235B-A22B](../../tutorials/models/Qwen3-235B-A22B.md) |
|
||||
| Qwen3-Next | 🔵 | | ✅ | A2/A3 | ✅ |||||| ✅ ||| ✅ || ✅ | ✅ ||| [Qwen3-Next](../../tutorials/models/Qwen3-Next.md) |
|
||||
| GLM-4.x | ✅ | | | A2/A3 |✅|✅|✅||✅|✅|✅||✅|✅|✅|✅|✅|198k||[GLM-4.x](../../tutorials/models/GLM4.x.md)|
|
||||
| GLM-5/5.1 | ✅ | | ✅ | A2/A3 | ✅ | ✅ | ✅ || ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 200k || [GLM-5](../../tutorials/models/GLM5.md) |
|
||||
| GLM-5.2 | 🔵 | | ✅ | A2/A3 | ✅ | ✅ | ✅ || ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 200k || [GLM-5](../../tutorials/models/GLM5.2.md) |
|
||||
| Kimi-K2-Thinking | 🔵 | | | A2/A3 |||||||||||||||| [Kimi-K2-Thinking](../../tutorials/models/Kimi-K2-Thinking.md) |
|
||||
| DeepSeekOCR2 | ✅ | | ✅ | A2/A3 ||✅||||✅|||||||||| [DeepSeekOCR2](../../tutorials/models/DeepSeekOCR2.md) |
|
||||
| MiniMax-M2.5/2.7 | ✅ | | ✅ | A2/A3/Ascend950 (Ascend950 experimental) |✅|✅|✅|❌|✅|✅|✅|🟡|✅|✅|✅|🟡|✅|200k|🟡| [MiniMax-M2](../../tutorials/models/MiniMax-M2.md) |
|
||||
| Qwen2.5-Math-RM-72B | 🔵 | vllm-rm, tensor_parallel_size=4, max_model_len=4096 | ✅ | A2 | ✅ | 🟡 | 🟡 | ❌ | 🟡 | ✅ | ✅ | 🟡 | 🟡 | 🟡 | 🟡 | 🟡 | 🟡 | 4096 | 🟡 | [Qwen2.5-Math-RM-72B](../../tutorials/models/Qwen2.5-Math-RM-72B.md) |
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
|
||||
| Model | Support | Note | BF16 | Supported Hardware | W8A8 | Chunked Prefill | Automatic Prefix Cache | LoRA | Speculative Decoding | Async Scheduling | Tensor Parallel | Prefill-decode Disaggregation | Piecewise AclGraph | Fullgraph AclGraph | max-model-len | Doc |
|
||||
|---------------|---------|------|------|--------------------|------|-----------------|------------------------|------|----------------------|------------------|-----------------|-------------------------------|--------------------|--------------------|---------------|-----|
|
||||
| Qwen3-Dense | ✅ | | ❌ | Atlas 300I DUO | ✅ | ✅ | ✅ | ❌ | 🟡 | ✅ | ✅ | ❌ | ✅ | ✅ | 20k | [Qwen3-Dense](../../tutorials/models/Qwen3-Dense.md) |
|
||||
| Qwen3-30B-A3B | ✅ | | ❌ | Atlas 300I DUO | ✅ | ✅ | ✅ | ❌ | 🟡 | ✅ | ✅ | ❌ | ✅ | ✅ | 16k | [Qwen3-30B-A3B](../../tutorials/models/Qwen3-30B-A3B.md) |
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
#### Extended Compatible Models
|
||||
|
||||
| Model | Support | Note | Supported Hardware |
|
||||
|-------------------------------|-----------|----------------------------------------------------------------------|--------------------|
|
||||
| DeepSeek Distill (Qwen/Llama) | 🔵 | | A2/A3 |
|
||||
| Qwen3-based | 🔵 | | A2/A3 |
|
||||
| Qwen2 | 🔵 | | A2/A3 |
|
||||
| Qwen2.5 | ✅ | | A2/A3 |
|
||||
| Qwen2-based | 🔵 | | A2/A3 |
|
||||
| QwQ-32B | 🔵 | | A2/A3 |
|
||||
| Llama2/3/3.1/3.2 | 🔵 | | A2/A3 |
|
||||
| Internlm | 🔵 | [#1962](https://github.com/vllm-project/vllm-ascend/issues/1962) | A2/A3 |
|
||||
| Baichuan | 🔵 | | A2/A3 |
|
||||
| Baichuan2 | 🔵 | | A2/A3 |
|
||||
| Phi-4-mini | 🔵 | | A2/A3 |
|
||||
| MiniCPM | 🔵 | | A2/A3 |
|
||||
| MiniCPM3 | 🔵 | | A2/A3 |
|
||||
| Ernie4.5 | 🔵 | | A2/A3 |
|
||||
| Ernie4.5-Moe | 🔵 | | A2/A3 |
|
||||
| Gemma-2 | 🔵 | | A2/A3 |
|
||||
| Gemma-3 | 🔵 | | A2/A3 |
|
||||
| Phi-3/4 | 🔵 | | A2/A3 |
|
||||
| Mistral/Mistral-Instruct | 🔵 | | A2/A3 |
|
||||
| Hy3-preview | 🔵 | | A3 |
|
||||
| DeepSeek V2.5 | 🟡 | Need test | |
|
||||
| Mllama | 🟡 | Need test | |
|
||||
| MiniMax-Text | 🟡 | Need test | |
|
||||
|
||||
### Pooling Models
|
||||
|
||||
| Model | Supported | Note |
|
||||
|-------------------------------|-----------|----------------------------------------------------------------------|
|
||||
| Qwen3-Embedding | ✅ | |
|
||||
| Molmo | ✅ | [1942](https://github.com/vllm-project/vllm-ascend/issues/1942) |
|
||||
| XLM-RoBERTa-based | ❌ | [1960](https://github.com/vllm-project/vllm-ascend/issues/1960) |
|
||||
:::::{tab-set}
|
||||
::::{tab-item} A2/A3
|
||||
|
||||
| Model | Support | Note | Supported Hardware | W8A8 | Doc |
|
||||
|-------------------------------|-----------|----------------------------------------------------------------------|------------------------------|------|------|
|
||||
| Qwen3-Embedding | 🔵 | | A2/A3 |🟡| [Qwen3-Embedding](../../tutorials/models/Qwen3-Embedding.md)|
|
||||
| Qwen3-VL-Embedding | 🔵 | | A2/A3 |🔵| [Qwen3-VL-Embedding](../../tutorials/models/Qwen3-VL-Embedding.md)|
|
||||
| Qwen3-Reranker | 🔵 | | A2/A3 |🟡| [Qwen3-Reranker](../../tutorials/models/Qwen3-Reranker.md)|
|
||||
| Qwen3-VL-Reranker | 🔵 | | A2/A3 |🔵| [Qwen3-VL-Reranker](../../tutorials/models/Qwen3-VL-Reranker.md)|
|
||||
| Molmo | 🔵 | [1942](https://github.com/vllm-project/vllm-ascend/issues/1942) | A2/A3 |🟡| |
|
||||
| XLM-RoBERTa-based | 🔵 | | A2/A3 |🟡| |
|
||||
| Bert | 🔵 | | A2/A3 |🟡| |
|
||||
| Qwen2.5-Math-RM-72B | 🔵 | Reward Model, gsm8k_correctness accuracy=0.80 | A2 |🟡| [Qwen2.5-Math-RM-72B](../../tutorials/models/Qwen2.5-Math-RM-72B.md) |
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
|
||||
| Model | Support | Note | Supported Hardware | W8A8| Doc |
|
||||
|-------------------|---------|------|--------------------|-----|--------------------------------------------------------------------|
|
||||
| Qwen3-Embedding | 🔵 | FP16 | Atlas 300I DUO |🟡| [Qwen3-Embedding](../../tutorials/models/Qwen3-Embedding.md) |
|
||||
| Qwen3-VL-Embedding| 🔵 | FP16 | Atlas 300I DUO |🔵| [Qwen3-VL-Embedding](../../tutorials/models/Qwen3-VL-Embedding.md) |
|
||||
| Qwen3-Reranker | 🔵 | FP16 | Atlas 300I DUO |🟡| [Qwen3-Reranker](../../tutorials/models/Qwen3-Reranker.md) |
|
||||
| Qwen3-VL-Reranker | 🔵 | FP16 | Atlas 300I DUO |🔵| [Qwen3-VL-Reranker](../../tutorials/models/Qwen3-VL-Reranker.md) |
|
||||
| XLM-RoBERTa-based | 🔵 | FP16; embedding and scoring | Atlas 300I DUO |🟡| |
|
||||
| Qwen2.5-based | 🔵 | FP16 classification | Atlas 300I DUO |🟡| |
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
## Multimodal Language Models
|
||||
|
||||
### Generative Models
|
||||
|
||||
| Model | Supported | Note |
|
||||
|--------------------------------|---------------|----------------------------------------------------------------------|
|
||||
| Qwen2-VL | ✅ | |
|
||||
| Qwen2.5-VL | ✅ | |
|
||||
| Qwen2.5-Omni | ✅ | [1760](https://github.com/vllm-project/vllm-ascend/issues/1760) |
|
||||
| QVQ | ✅ | |
|
||||
| LLaVA 1.5/1.6 | ✅ | [1962](https://github.com/vllm-project/vllm-ascend/issues/1962) |
|
||||
| InternVL2 | ✅ | |
|
||||
| InternVL2.5 | ✅ | |
|
||||
| Qwen2-Audio | ✅ | |
|
||||
| Aria | ✅ | |
|
||||
| LLaVA-Next | ✅ | |
|
||||
| LLaVA-Next-Video | ✅ | |
|
||||
| MiniCPM-V | ✅ | |
|
||||
| Mistral3 | ✅ | |
|
||||
| Phi-3-Vison/Phi-3.5-Vison | ✅ | |
|
||||
| Gemma3 | ✅ | |
|
||||
| LLama4 | ❌ | [1972](https://github.com/vllm-project/vllm-ascend/issues/1972) |
|
||||
| LLama3.2 | ❌ | [1972](https://github.com/vllm-project/vllm-ascend/issues/1972) |
|
||||
| Keye-VL-8B-Preview | ❌ | [1963](https://github.com/vllm-project/vllm-ascend/issues/1963) |
|
||||
| Florence-2 | ❌ | [2259](https://github.com/vllm-project/vllm-ascend/issues/2259) |
|
||||
| GLM-4V | ❌ | [2260](https://github.com/vllm-project/vllm-ascend/issues/2260) |
|
||||
| InternVL2.0/2.5/3.0<br>InternVideo2.5/Mono-InternVL | ❌ | [2064](https://github.com/vllm-project/vllm-ascend/issues/2064) |
|
||||
| Whisper | ❌ | [2262](https://github.com/vllm-project/vllm-ascend/issues/2262) |
|
||||
| Ultravox | 🟡 Need test | |
|
||||
#### Core Supported Models
|
||||
|
||||
:::::{tab-set}
|
||||
|
||||
::::{tab-item} Ascend 950 Products
|
||||
|
||||
| Model | Support | Note | BF16 | Supported Hardware | W8A8 | Chunked Prefill | Automatic Prefix Cache | LoRA | Speculative Decoding | Async Scheduling | Tensor Parallel | Pipeline Parallel | Expert Parallel | Data Parallel | Prefill-decode Disaggregation | Piecewise AclGraph | Fullgraph AclGraph | max-model-len | MLP Weight Prefetch | Doc |
|
||||
|-----------------|----------|--------|------|------|------|---------|-------|------|------|--------|-------|--------|--------|-------|-------|--------|----------|---------|----------|-----|
|
||||
|Qwen3.5-397B-A17B|✅ | |✅ | Ascend 950DT |✅|✅|✅||✅|✅|✅||✅|✅|✅|✅|✅|1010000|| [Qwen3.5-397B-A17B](../../tutorials/models/Qwen3.5-397B-A17B.md) |
|
||||
|Qwen3.6-27B |✅ | |✅ | Ascend 950 Products |✅|✅|✅||✅|✅|✅||✅|✅|✅|✅|✅|262144|| [Qwen3.5-27B / Qwen3.6-27B](../../tutorials/models/Qwen3.5-27B-Qwen3.6-27B.md) |
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} A2/A3
|
||||
|
||||
| Model | Support | Note | BF16 | Supported Hardware | W8A8 | Chunked Prefill | Automatic Prefix Cache | LoRA | Speculative Decoding | Async Scheduling | Tensor Parallel | Pipeline Parallel | Expert Parallel | Data Parallel | Prefill-decode Disaggregation | Piecewise AclGraph | Fullgraph AclGraph | max-model-len | MLP Weight Prefetch | Doc |
|
||||
|-------------------------------------|---------------|------|------|--------------------|------|-----------------|------------------------|------|----------------------|------------------|-----------------|-------------------|-----------------|---------------|-------------------------------|--------------------|--------------------|---------------|---------------------|-----|
|
||||
| Qwen3-VL | ✅ | | |A2/A3|||||||✅|||||✅|✅||| [Qwen-VL-Dense](../../tutorials/models/Qwen-VL-Dense.md) |
|
||||
| Qwen3-VL-30B-A3B/Qwen3-VL-235B-A22B | ✅ | | ✅ | A2/A3 | ✅ | ✅ | ✅ | | | ✅ | ✅ | | ✅ | ✅ | ✅ | ✅ | ✅ | 262144 || [Qwen3-VL-30B-A3B](../../tutorials/models/Qwen3-VL-30B-A3B-Instruct.md)/[Qwen3-VL-235B-A22B](../../tutorials/models/Qwen3-VL-235B-A22B-Instruct.md) |
|
||||
| Qwen3.5-397B-A17B | ✅ | | ✅ | A2/A3 |✅|✅|✅||✅|✅|✅||✅|✅|✅|✅|✅|1010000|| [Qwen3.5-397B-A17B](../../tutorials/models/Qwen3.5-397B-A17B.md) |
|
||||
| Qwen3.5-27B / Qwen3.6-27B | ✅ | | ✅ | A2/A3 |✅|✅|✅||✅|✅|✅||✅|✅|✅|✅|✅|262144|| [Qwen3.5-27B / Qwen3.6-27B](../../tutorials/models/Qwen3.5-27B-Qwen3.6-27B.md) |
|
||||
| Qwen3.6-35B-A3B | ✅ | | ✅ | A2/A3 |✅|✅|✅||🔵|✅|✅||✅|✅|❌|✅|✅|262144|| [Qwen3.6-35B-A3B](../../tutorials/models/Qwen3.6-35B-A3B.md) |
|
||||
| Qwen3-Omni-30B-A3B-Thinking | ✅ | | |A2/A3|||||||✅||✅|||||||[Qwen3-Omni-30B-A3B-Thinking](../../tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md)|
|
||||
| Kimi-K2.5/Kimi-K2.6 | ✅ | | |A2/A3||✅|✅||✅|✅|✅||✅|✅|✅|✅|✅|262144||[Kimi-K2.5](../../tutorials/models/Kimi-K2.5.md)/[Kimi-K2.6](../../tutorials/models/Kimi-K2.6.md)|
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
|
||||
| Model | Support | Note | BF16 | Supported Hardware | W8A8 | Chunked Prefill | Automatic Prefix Cache | LoRA | Speculative Decoding | Async Scheduling | Tensor Parallel | Prefill-decode Disaggregation | Piecewise AclGraph | Fullgraph AclGraph | max-model-len | Doc |
|
||||
|-----------------|---------|------|------|--------------------|------|-----------------|------------------------|------|----------------------|------------------|-----------------|-------------------------------|--------------------|--------------------|---------------|-----|
|
||||
| Qwen3-VL | ✅ | | ❌ | Atlas 300I DUO | ✅ | ✅ | ✅ | ❌ | 🟡 | ✅ | ✅ | ❌ | ✅ | ✅ | 16k | [Qwen-VL-Dense](../../tutorials/models/Qwen-VL-Dense.md) |
|
||||
| Qwen3.5-Dense | ✅ | | ❌ | Atlas 300I DUO | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | 256k | [Qwen3.5-Dense](../../tutorials/models/Qwen3.5-Dense.md) |
|
||||
| Qwen3.5-35B-A3B | ✅ | | ❌ | Atlas 300I DUO | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | 256k | [Qwen3.5-35B-A3B](../../tutorials/models/Qwen3.6-35B-A3B.md) |
|
||||
| Qwen3.6-27B | ✅ | | ❌ | Atlas 300I DUO | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | 256k | [Qwen3.6-27B](../../tutorials/models/Qwen3.5-27B-Qwen3.6-27B.md) |
|
||||
| Qwen3.6-35B-A3B | ✅ | | ❌ | Atlas 300I DUO | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | 256k | [Qwen3.6-35B-A3B](../../tutorials/models/Qwen3.6-35B-A3B.md) |
|
||||
| PaddleOCR-VL | 🔵 | | ❌ | Atlas 300I DUO | ❌ | ✅ | ✅ | ❌ | ❌ | ✅ | ❌ | ❌ | ✅ | ✅ | 16k | [PaddleOCR-VL](../../tutorials/models/PaddleOCR-VL.md) |
|
||||
| Qwen3-ASR | 🔵 | | ❌ | Atlas 300I DUO | ❌ | ✅ | ✅ | ❌ | ❌ | ✅ | 🟡 | ❌ | ✅ | ✅ | 4096 | [Qwen3-ASR-1.7B](../../tutorials/models/Qwen3-ASR-1.7B.md) |
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
#### Extended Compatible Models
|
||||
|
||||
| Model | Support | Note | Supported Hardware |
|
||||
|--------------------------------|---------------|----------------------------------------------------------------------|--------------------|
|
||||
| Qwen2-VL | 🔵 | | A2/A3 |
|
||||
| Qwen3-Omni | 🔵 | | A2/A3 |
|
||||
| QVQ | 🔵 | | A2/A3 |
|
||||
| Qwen2-Audio | 🔵 | | A2/A3 |
|
||||
| Aria | 🔵 | | A2/A3 |
|
||||
| LLaVA-Next | 🔵 | | A2/A3 |
|
||||
| LLaVA-Next-Video | 🔵 | | A2/A3 |
|
||||
| MiniCPM-V | 🔵 | | A2/A3 |
|
||||
| Mistral3 | 🔵 | | A2/A3 |
|
||||
| Phi-3-Vision/Phi-3.5-Vision | 🔵 | | A2/A3 |
|
||||
| Gemma3 | 🔵 | | A2/A3 |
|
||||
| Llama3.2 | 🔵 | | A2/A3 |
|
||||
| PaddleOCR-VL | 🔵 | | A2/A3 |
|
||||
| Llama4 | ❌ | [1972](https://github.com/vllm-project/vllm-ascend/issues/1972) | |
|
||||
| Keye-VL-8B-Preview | ❌ | [1961](https://github.com/vllm-project/vllm-ascend/issues/1961) | |
|
||||
| Florence-2 | ❌ | [2259](https://github.com/vllm-project/vllm-ascend/issues/2259) | |
|
||||
| GLM-4V | ❌ | [2260](https://github.com/vllm-project/vllm-ascend/issues/2260) | |
|
||||
| InternVL2.0/2.5/3.0<br>InternVideo2.5/Mono-InternVL | ❌ | [2064](https://github.com/vllm-project/vllm-ascend/issues/2064) | |
|
||||
| Whisper | ❌ | [2262](https://github.com/vllm-project/vllm-ascend/issues/2262) | |
|
||||
| Ultravox | 🟡 | Need test | |
|
||||
|
||||
Reference in New Issue
Block a user