469
docs/source/tutorials/models/DeepSeek-R1.md
Normal file
469
docs/source/tutorials/models/DeepSeek-R1.md
Normal file
@@ -0,0 +1,469 @@
|
||||
# DeepSeek-R1
|
||||
|
||||
## 1 Introduction
|
||||
|
||||
DeepSeek-R1 is a high-performance Mixture-of-Experts (MoE) large language model developed by DeepSeek Company. It excels in complex logical reasoning, mathematical problem-solving, and code generation. By dynamically activating its expert networks, it delivers exceptional performance while maintaining computational efficiency. Building upon R1, DeepSeek-R1-W8A8 is a fully quantized version of the model. It employs 8-bit integer (INT8) quantization for both weights and activations, which significantly reduces the model's memory footprint and computational requirements, enabling more efficient deployment and application in resource-constrained environments.
|
||||
This article takes the `DeepSeek-R1-W8A8` version as an example to introduce the deployment of the R1 series models.
|
||||
|
||||
This document will show the main verification steps of the model, including supported features, feature configuration, environment preparation, single-node and multi-node deployment, accuracy and performance evaluation.
|
||||
|
||||
This document is validated and written based on **vLLM-Ascend v0.13.0**. The current model (DeepSeek-R1) is first supported in this version.
|
||||
|
||||
## 2 Supported Features
|
||||
|
||||
Refer to [supported features](../../user_guide/support_matrix/supported_models.md) to get the model's supported feature matrix.
|
||||
|
||||
Refer to [feature guide](../../user_guide/feature_guide/index.md) to get the feature's configuration.
|
||||
|
||||
## 3 Prerequisites
|
||||
|
||||
### 3.1 Model Weight
|
||||
|
||||
- `DeepSeek-R1-W8A8`(Quantized version): require 1 Atlas 800 A3 (64G × 16) nodes or 2 Atlas 800 A2 (64G × 8) nodes. [Download model weight](https://www.modelscope.cn/models/vllm-ascend/DeepSeek-R1-W8A8)
|
||||
|
||||
It is recommended to download the model weight to the shared directory of multiple nodes.
|
||||
|
||||
### 3.2 Verify Multi-node Communication (Optional)
|
||||
|
||||
If you want to deploy multi-node environment, you need to verify multi-node communication according to [verify multi-node communication environment](../../installation.md#verify-multi-node-communication).
|
||||
|
||||
## 4 Installation
|
||||
|
||||
### 4.1 Docker Image Installation
|
||||
|
||||
You can use our official docker image to run `DeepSeek-R1-W8A8` directly.
|
||||
|
||||
Select an image based on your machine type and start the docker image on your node, refer to [using docker](../../installation.md#set-up-using-docker).
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: install
|
||||
|
||||
::::{tab-item} A3 series
|
||||
:sync: A3
|
||||
|
||||
Start the docker image on each node.
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-a3
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--privileged=true \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci8 \
|
||||
--device /dev/davinci9 \
|
||||
--device /dev/davinci10 \
|
||||
--device /dev/davinci11 \
|
||||
--device /dev/davinci12 \
|
||||
--device /dev/davinci13 \
|
||||
--device /dev/davinci14 \
|
||||
--device /dev/davinci15 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} A2 series
|
||||
:sync: A2
|
||||
|
||||
Start the docker image on each node.
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--privileged=true \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
After a successful docker run, you can verify the running container service by executing the `docker ps` command.
|
||||
|
||||
### 4.2 Source Code Installation
|
||||
|
||||
If you don't want to use the docker image as above, you can also build all from source:
|
||||
|
||||
- Install `vllm-ascend` from source, refer to [installation](../../installation.md).
|
||||
|
||||
If you want to deploy multi-node environment, you need to set up environment on each node.
|
||||
|
||||
## 5 Online Service Deployment
|
||||
|
||||
### 5.1 Single-Node Online Deployment
|
||||
|
||||
Single-node deployment completes both Prefill and Decode within the same node. The quantized model `DeepSeek-R1-W8A8` can be deployed on 1 Atlas 800 A3 (64G × 16) or 2 Atlas 800 A2 (64G × 8).
|
||||
|
||||
Startup Command:
|
||||
|
||||
```shell
|
||||
#!/bin/sh
|
||||
|
||||
# this obtained through ifconfig
|
||||
# nic_name is the network interface name corresponding to local_ip of the current node
|
||||
nic_name="xxxx"
|
||||
local_ip="xxxx"
|
||||
|
||||
# AIV
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
export VLLM_ASCEND_BALANCE_SCHEDULING=1
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
|
||||
vllm serve vllm-ascend/DeepSeek-R1-W8A8 \
|
||||
--host 0.0.0.0 \
|
||||
--port 8000 \
|
||||
--data-parallel-size 4 \
|
||||
--tensor-parallel-size 4 \
|
||||
--quantization ascend \
|
||||
--seed 1024 \
|
||||
--served-model-name deepseek_r1 \
|
||||
--enable-expert-parallel \
|
||||
--max-num-seqs 16 \
|
||||
--max-model-len 16384 \
|
||||
--max-num-batched-tokens 4096 \
|
||||
--trust-remote-code \
|
||||
--gpu-memory-utilization 0.92 \
|
||||
--speculative-config '{"num_speculative_tokens":3,"method":"mtp"}' \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}'
|
||||
```
|
||||
|
||||
Key Parameter Descriptions:
|
||||
|
||||
- Setting the environment variable `VLLM_ASCEND_BALANCE_SCHEDULING=1` enables balance scheduling. This may help increase output throughput and reduce TPOT in v1 scheduler. However, TTFT may degrade in some scenarios. Furthermore, enabling this feature is not recommended in scenarios where PD is separated.
|
||||
- For single-node deployment, we recommend using `dp4tp4` instead of `dp2tp8`.
|
||||
- `--max-model-len` specifies the maximum context length - that is, the sum of input and output tokens for a single request. For performance testing with an input length of 3.5k and output length of 1.5k, a value of `16384` is sufficient, however, for precision testing, please set it to at least `35000`.
|
||||
- `--no-enable-prefix-caching` indicates that prefix caching is disabled. To enable it, remove this option.
|
||||
- If you use the w4a8 weight, more memory will be allocated to kvcache, and you can try to increase system throughput to achieve greater throughput.
|
||||
|
||||
Common Issues Tip: If you encounter issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html) for troubleshooting.
|
||||
|
||||
Service Verification:
|
||||
|
||||
```shell
|
||||
curl http://<node_ip>:8000/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "deepseek_r1",
|
||||
"messages": [{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "The future of AI is"
|
||||
}]
|
||||
}],
|
||||
"max_tokens": 1024,
|
||||
"temperature": 1.0,
|
||||
"top_p": 0.95
|
||||
}'
|
||||
```
|
||||
|
||||
Expected Result:
|
||||
|
||||
The service returns HTTP 200 OK with a JSON response containing the `choices` field. Example output:
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "chatcmpl-xxxxxxxxxxxxx",
|
||||
"object": "chat.completion",
|
||||
"model": "deepseek_r1",
|
||||
"choices": [
|
||||
{
|
||||
"index": 0,
|
||||
"message": {
|
||||
"role": "assistant",
|
||||
"content": "<think>\nOkay, the user wrote \"The future of AI is...",
|
||||
"finish_reason": "length"
|
||||
}
|
||||
}
|
||||
],
|
||||
"usage": {
|
||||
"prompt_tokens": 8,
|
||||
"total_tokens": 1032,
|
||||
"completion_tokens": 1024
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### 5.2 Multi-Node Data Parallel Deployment
|
||||
|
||||
Run the following scripts on two nodes respectively.
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: Deployment
|
||||
|
||||
::::{tab-item} Node 0
|
||||
:sync: Node 0
|
||||
|
||||
Startup Command:
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
#!/bin/sh
|
||||
|
||||
# this obtained through ifconfig
|
||||
# nic_name is the network interface name corresponding to local_ip of the current node
|
||||
nic_name="xxxx"
|
||||
local_ip="xxxx"
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export HCCL_BUFFSIZE=200
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export VLLM_ASCEND_BALANCE_SCHEDULING=1
|
||||
export HCCL_INTRA_PCIE_ENABLE=1
|
||||
export HCCL_INTRA_ROCE_ENABLE=0
|
||||
|
||||
vllm serve vllm-ascend/DeepSeek-R1-W8A8 \
|
||||
--host 0.0.0.0 \
|
||||
--port 8000 \
|
||||
--data-parallel-size 4 \
|
||||
--data-parallel-size-local 2 \
|
||||
--data-parallel-address $local_ip \
|
||||
--data-parallel-rpc-port 13389 \
|
||||
--tensor-parallel-size 4 \
|
||||
--quantization ascend \
|
||||
--seed 1024 \
|
||||
--served-model-name deepseek_r1 \
|
||||
--enable-expert-parallel \
|
||||
--max-num-seqs 16 \
|
||||
--max-model-len 16384 \
|
||||
--max-num-batched-tokens 4096 \
|
||||
--trust-remote-code \
|
||||
--gpu-memory-utilization 0.92 \
|
||||
--speculative-config '{"num_speculative_tokens":3,"method":"mtp"}' \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}'
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} Node 1
|
||||
:sync: Node 1
|
||||
|
||||
Startup Command:
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
#!/bin/sh
|
||||
|
||||
# this is obtained through ifconfig
|
||||
# nic_name is the network interface name corresponding to local_ip of the current node
|
||||
nic_name="xxxx"
|
||||
local_ip="xxxx"
|
||||
node0_ip="xxxx" # same as the local_IP address in node 0
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export HCCL_BUFFSIZE=200
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export VLLM_ASCEND_BALANCE_SCHEDULING=1
|
||||
export HCCL_INTRA_PCIE_ENABLE=1
|
||||
export HCCL_INTRA_ROCE_ENABLE=0
|
||||
|
||||
vllm serve vllm-ascend/DeepSeek-R1-W8A8 \
|
||||
--host 0.0.0.0 \
|
||||
--port 8000 \
|
||||
--headless \
|
||||
--data-parallel-size 4 \
|
||||
--data-parallel-size-local 2 \
|
||||
--data-parallel-start-rank 2 \
|
||||
--data-parallel-address $node0_ip \
|
||||
--data-parallel-rpc-port 13389 \
|
||||
--tensor-parallel-size 4 \
|
||||
--quantization ascend \
|
||||
--seed 1024 \
|
||||
--served-model-name deepseek_r1 \
|
||||
--enable-expert-parallel \
|
||||
--max-num-seqs 16 \
|
||||
--max-model-len 16384 \
|
||||
--max-num-batched-tokens 4096 \
|
||||
--trust-remote-code \
|
||||
--gpu-memory-utilization 0.92 \
|
||||
--speculative-config '{"num_speculative_tokens":3,"method":"mtp"}' \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}'
|
||||
```
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
Key Parameter Descriptions:
|
||||
|
||||
- `--data-parallel-size`: total number of data parallel ranks across all nodes. In this example, `4` means the model is split across 4 DP ranks total (2 per node).
|
||||
- `--data-parallel-size-local`: number of data parallel ranks running on the current node. In this example, each node runs 2 DP ranks.
|
||||
- `--data-parallel-start-rank`: starting rank offset for data parallel ranks on this node. Node 0 starts at rank 0 (default), Node 1 starts at rank 2. This ensures each node's DP ranks occupy distinct positions in the overall rank space.
|
||||
- `--data-parallel-address`: IP address of the data parallel master node (Node 0). This value must be consistent with `local_ip` set on Node 0.
|
||||
- `--data-parallel-rpc-port`: RPC port for data parallel master communication. Must be the same across all nodes.
|
||||
- `--headless`: indicates that this vLLM instance is not the master service node. Only set on non-master nodes (Node 1). The master node (Node 0) should NOT set this flag.
|
||||
- For single-node deployment, we recommend using `dp4 tp4` instead of `dp2 tp8`.
|
||||
|
||||
Common Issues Tip: If you encounter issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html) for troubleshooting.
|
||||
|
||||
Service Verification:
|
||||
|
||||
```shell
|
||||
curl http://<node_ip>:8000/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "deepseek_r1",
|
||||
"messages": [{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "The future of AI is"
|
||||
}]
|
||||
}],
|
||||
"max_tokens": 1024,
|
||||
"temperature": 1.0,
|
||||
"top_p": 0.95
|
||||
}'
|
||||
```
|
||||
|
||||
Expected Result:
|
||||
|
||||
The service returns HTTP 200 OK. The JSON response contains the `choices` field with the generated text.
|
||||
|
||||
### 5.3 Multi-Node PD Separation Deployment
|
||||
|
||||
We recommend using DeepSeek-V3.1 for deployment: [DeepSeek-V3.1](./DeepSeek-V3.1.md).
|
||||
|
||||
This solution has been tested and demonstrates excellent performance.
|
||||
|
||||
## 6 Functional Verification
|
||||
|
||||
Once your server is started, you can query the model with input prompts:
|
||||
|
||||
```shell
|
||||
curl http://<node0_ip>:<port>/v1/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "deepseek_r1",
|
||||
"prompt": "The future of AI is",
|
||||
"max_completion_tokens": 50,
|
||||
"temperature": 0
|
||||
}'
|
||||
```
|
||||
|
||||
## 7 Accuracy Evaluation
|
||||
|
||||
Here are two accuracy evaluation methods.
|
||||
|
||||
### Using AISBench
|
||||
|
||||
1. Refer to [Using AISBench](../../developer_guide/evaluation/using_ais_bench.md) for details.
|
||||
|
||||
2. After execution, you can get the result, here is the result of `DeepSeek-R1-W8A8` in `vllm-ascend:0.13.0` for reference only.
|
||||
|
||||
| dataset | version | metric | mode | vllm-api-general-chat |
|
||||
|----- | ----- | ----- | ----- | -----|
|
||||
| aime2024dataset | - | accuracy | gen | 80.00 |
|
||||
| gpqadataset | - | accuracy | gen | 72.22 |
|
||||
|
||||
### Using Language Model Evaluation Harness
|
||||
|
||||
As an example, take the `gsm8k` dataset as a test dataset, and run accuracy evaluation of `DeepSeek-R1-W8A8` in online mode.
|
||||
|
||||
1. Refer to [Using lm_eval](../../developer_guide/evaluation/using_lm_eval.md) for `lm_eval` installation.
|
||||
|
||||
2. Run `lm_eval` to execute the accuracy evaluation.
|
||||
|
||||
```shell
|
||||
lm_eval \
|
||||
--model local-completions \
|
||||
--model_args model=path/DeepSeek-R1-W8A8,base_url=http://<node0_ip>:<port>/v1/completions,tokenized_requests=False,trust_remote_code=True \
|
||||
--tasks gsm8k \
|
||||
--output_path ./
|
||||
```
|
||||
|
||||
3. After execution, you can get the result.
|
||||
|
||||
## 8 Performance Evaluation
|
||||
|
||||
### Using AISBench
|
||||
|
||||
Refer to [Using AISBench for performance evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation) for details.
|
||||
|
||||
### Using vLLM Benchmark
|
||||
|
||||
Run performance evaluation of `DeepSeek-R1-W8A8` as an example.
|
||||
|
||||
Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/) for more details.
|
||||
|
||||
There are three `vllm bench` subcommands:
|
||||
|
||||
- `latency`: Benchmark the latency of a single batch of requests.
|
||||
- `serve`: Benchmark the online serving throughput.
|
||||
- `throughput`: Benchmark offline inference throughput.
|
||||
|
||||
Take the `serve` as an example. Run the code as follows.
|
||||
|
||||
```shell
|
||||
vllm bench serve --model path/DeepSeek-R1-W8A8 --dataset-name random --random-input 200 --num-prompts 200 --request-rate 1 --save-result --result-dir ./
|
||||
```
|
||||
|
||||
After about several minutes, you can get the performance evaluation result.
|
||||
|
||||
## 9 Performance Tuning
|
||||
|
||||
We recommend using DeepSeek-V3.1 for deployment: [DeepSeek-V3.1](./DeepSeek-V3.1.md).
|
||||
|
||||
This solution has been tested and demonstrates excellent performance.
|
||||
|
||||
## 10 FAQ
|
||||
|
||||
For common environment, installation, and general parameter issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html).
|
||||
1307
docs/source/tutorials/models/DeepSeek-V3.1.md
Normal file
1307
docs/source/tutorials/models/DeepSeek-V3.1.md
Normal file
File diff suppressed because it is too large
Load Diff
889
docs/source/tutorials/models/DeepSeek-V3.2.md
Normal file
889
docs/source/tutorials/models/DeepSeek-V3.2.md
Normal file
@@ -0,0 +1,889 @@
|
||||
# DeepSeek-V3.2
|
||||
|
||||
## 1 Introduction
|
||||
|
||||
DeepSeek-V3.2 is a sparse attention model. The main architecture is similar to DeepSeek-V3.1, but with a sparse attention mechanism, which is designed to explore and validate optimizations for training and inference efficiency in long-context scenarios.
|
||||
|
||||
This document will show the main verification steps of the model, including supported features, feature configuration, environment preparation, single-node and multi-node deployment, accuracy and performance evaluation.
|
||||
|
||||
## 2 Supported Features
|
||||
|
||||
Refer to [supported features](../../user_guide/support_matrix/supported_models.md) to get the model's supported feature matrix.
|
||||
|
||||
Refer to [feature guide](../../user_guide/feature_guide/index.md) to get the feature's configuration.
|
||||
|
||||
## 3 Prerequisites
|
||||
|
||||
### 3.1 Model Weight
|
||||
|
||||
- `DeepSeek-V3.2-Exp-W8A8` (Quantized version): requires **1 Atlas 800 A3 (64G × 16) node** or **2 Atlas 800 A2 (64G × 8) nodes**. [Download model weight](https://www.modelscope.cn/models/vllm-ascend/DeepSeek-V3.2-Exp-W8A8)
|
||||
- `DeepSeek-V3.2-w8a8` (Quantized version): requires **1 Atlas 800 A3 (64G × 16) node** or **2 Atlas 800 A2 (64G × 8) nodes**. [Download model weight](https://www.modelscope.cn/models/vllm-ascend/DeepSeek-V3.2-W8A8/)
|
||||
|
||||
It is recommended to download the model weight to the shared directory of multiple nodes, such as `/root/.cache/`.
|
||||
|
||||
### 3.2 Verify Multi-node Communication (Optional)
|
||||
|
||||
If you want to deploy multi-node environment, you need to verify multi-node communication according to [verify multi-node communication environment](../../installation.md#verify-multi-node-communication).
|
||||
|
||||
## 4 Installation
|
||||
|
||||
### 4.1 Docker Image Installation
|
||||
|
||||
You can use our official docker image to run `DeepSeek-V3.2` directly.
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: install
|
||||
|
||||
::::{tab-item} A3 series
|
||||
:sync: A3
|
||||
|
||||
Start the docker image on each node.
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-a3
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--privileged=true \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci8 \
|
||||
--device /dev/davinci9 \
|
||||
--device /dev/davinci10 \
|
||||
--device /dev/davinci11 \
|
||||
--device /dev/davinci12 \
|
||||
--device /dev/davinci13 \
|
||||
--device /dev/davinci14 \
|
||||
--device /dev/davinci15 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} A2 series
|
||||
:sync: A2
|
||||
|
||||
Start the docker image on each node.
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--privileged=true \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
### 4.2 Source Code Installation
|
||||
|
||||
In addition, if you don't want to use the docker image as above, you can also build all from source:
|
||||
|
||||
- Install `vllm-ascend` from source, refer to [installation](../../installation.md).
|
||||
|
||||
If you want to deploy multi-node environment, you need to set up environment on each node.
|
||||
|
||||
## 5 Online Service Deployment
|
||||
|
||||
:::{note}
|
||||
In this tutorial, we suppose you downloaded the model weight to `/root/.cache/`. Feel free to change it to your own path.
|
||||
:::
|
||||
|
||||
### 5.1 Single-node Deployment
|
||||
|
||||
- Quantized model `DeepSeek-V3.2-w8a8` can be deployed on 1 Atlas 800 A3 (64G × 16).
|
||||
|
||||
Run the following script to execute online inference.
|
||||
|
||||
```shell
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=10
|
||||
export VLLM_USE_V1=1
|
||||
export HCCL_BUFFSIZE=200
|
||||
export VLLM_ASCEND_ENABLE_MLAPO=1
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export VLLM_ASCEND_ENABLE_FLASHCOMM1=1
|
||||
|
||||
vllm serve /root/.cache/modelscope/hub/models/vllm-ascend/DeepSeek-V3.2-W8A8 \
|
||||
--host 0.0.0.0 \
|
||||
--port 8000 \
|
||||
--data-parallel-size 2 \
|
||||
--tensor-parallel-size 8 \
|
||||
--quantization ascend \
|
||||
--seed 1024 \
|
||||
--served-model-name deepseek_v3_2 \
|
||||
--enable-expert-parallel \
|
||||
--max-num-seqs 16 \
|
||||
--max-model-len 8192 \
|
||||
--max-num-batched-tokens 4096 \
|
||||
--trust-remote-code \
|
||||
--no-enable-prefix-caching \
|
||||
--gpu-memory-utilization 0.92 \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
|
||||
--speculative-config '{"num_speculative_tokens": 3, "method": "deepseek_mtp"}'
|
||||
|
||||
```
|
||||
|
||||
### 5.2 Multi-node Deployment
|
||||
|
||||
- `DeepSeek-V3.2-w8a8`: require at least 2 Atlas 800 A2 (64G × 8).
|
||||
|
||||
Run the following scripts on two nodes respectively.
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: install
|
||||
|
||||
::::{tab-item} A3 series
|
||||
:sync: A3
|
||||
|
||||
**Node0**
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# this obtained through ifconfig
|
||||
# nic_name is the network interface name corresponding to local_ip of the current node
|
||||
nic_name="xxx"
|
||||
local_ip="xxx"
|
||||
|
||||
# The value of node0_ip must be consistent with the value of local_ip set in node0 (master node)
|
||||
node0_ip="xxxx"
|
||||
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=10
|
||||
export VLLM_USE_V1=1
|
||||
export HCCL_BUFFSIZE=200
|
||||
export VLLM_ASCEND_ENABLE_MLAPO=1
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export VLLM_ASCEND_ENABLE_FLASHCOMM1=1
|
||||
|
||||
vllm serve /root/.cache/modelscope/hub/models/vllm-ascend/DeepSeek-V3.2-W8A8 \
|
||||
--host 0.0.0.0 \
|
||||
--port 8077 \
|
||||
--data-parallel-size 2 \
|
||||
--data-parallel-size-local 1 \
|
||||
--data-parallel-address $node0_ip \
|
||||
--data-parallel-rpc-port 12890 \
|
||||
--tensor-parallel-size 16 \
|
||||
--quantization ascend \
|
||||
--seed 1024 \
|
||||
--served-model-name deepseek_v3_2 \
|
||||
--enable-expert-parallel \
|
||||
--max-num-seqs 16 \
|
||||
--max-model-len 8192 \
|
||||
--max-num-batched-tokens 4096 \
|
||||
--trust-remote-code \
|
||||
--no-enable-prefix-caching \
|
||||
--gpu-memory-utilization 0.92 \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
|
||||
--speculative-config '{"num_speculative_tokens": 3, "method": "deepseek_mtp"}'
|
||||
```
|
||||
|
||||
**Node1**
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# this obtained through ifconfig
|
||||
# nic_name is the network interface name corresponding to local_ip of the current node
|
||||
nic_name="xxx"
|
||||
local_ip="xxx"
|
||||
|
||||
# The value of node0_ip must be consistent with the value of local_ip set in node0 (master node)
|
||||
node0_ip="xxxx"
|
||||
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=10
|
||||
export VLLM_USE_V1=1
|
||||
export HCCL_BUFFSIZE=200
|
||||
export VLLM_ASCEND_ENABLE_MLAPO=1
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export VLLM_ASCEND_ENABLE_FLASHCOMM1=1
|
||||
|
||||
vllm serve /root/.cache/modelscope/hub/models/vllm-ascend/DeepSeek-V3.2-W8A8 \
|
||||
--host 0.0.0.0 \
|
||||
--port 8077 \
|
||||
--headless \
|
||||
--data-parallel-size 2 \
|
||||
--data-parallel-size-local 1 \
|
||||
--data-parallel-start-rank 1 \
|
||||
--data-parallel-address $node0_ip \
|
||||
--data-parallel-rpc-port 12890 \
|
||||
--tensor-parallel-size 16 \
|
||||
--quantization ascend \
|
||||
--seed 1024 \
|
||||
--served-model-name deepseek_v3_2 \
|
||||
--enable-expert-parallel \
|
||||
--max-num-seqs 16 \
|
||||
--max-model-len 8192 \
|
||||
--max-num-batched-tokens 4096 \
|
||||
--trust-remote-code \
|
||||
--no-enable-prefix-caching \
|
||||
--gpu-memory-utilization 0.92 \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
|
||||
--speculative-config '{"num_speculative_tokens": 3, "method": "deepseek_mtp"}'
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} A2 series
|
||||
:sync: A2
|
||||
|
||||
**Node0**
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# this obtained through ifconfig
|
||||
# nic_name is the network interface name corresponding to local_ip of the current node
|
||||
nic_name="xxx"
|
||||
local_ip="xxx"
|
||||
|
||||
# The value of node0_ip must be consistent with the value of local_ip set in node0 (master node)
|
||||
node0_ip="xxxx"
|
||||
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=100
|
||||
export VLLM_USE_V1=1
|
||||
export HCCL_BUFFSIZE=200
|
||||
export VLLM_ASCEND_ENABLE_MLAPO=1
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export VLLM_ASCEND_ENABLE_FLASHCOMM1=1
|
||||
export HCCL_CONNECT_TIMEOUT=120
|
||||
export HCCL_INTRA_PCIE_ENABLE=1
|
||||
export HCCL_INTRA_ROCE_ENABLE=0
|
||||
|
||||
vllm serve /root/.cache/modelscope/hub/models/vllm-ascend/DeepSeek-V3.2-W8A8 \
|
||||
--host 0.0.0.0 \
|
||||
--port 8077 \
|
||||
--data-parallel-size 2 \
|
||||
--data-parallel-size-local 1 \
|
||||
--data-parallel-address $node0_ip \
|
||||
--data-parallel-rpc-port 13389 \
|
||||
--tensor-parallel-size 8 \
|
||||
--quantization ascend \
|
||||
--seed 1024 \
|
||||
--served-model-name deepseek_v3_2 \
|
||||
--enable-expert-parallel \
|
||||
--max-num-seqs 16 \
|
||||
--max-model-len 8192 \
|
||||
--max-num-batched-tokens 4096 \
|
||||
--trust-remote-code \
|
||||
--no-enable-prefix-caching \
|
||||
--gpu-memory-utilization 0.92 \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY", "cudagraph_capture_sizes":[8, 16, 24, 32, 40, 48]}' \
|
||||
--speculative-config '{"num_speculative_tokens": 3, "method": "deepseek_mtp"}'
|
||||
|
||||
```
|
||||
|
||||
**Node1**
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# this obtained through ifconfig
|
||||
# nic_name is the network interface name corresponding to local_ip of the current node
|
||||
nic_name="xxx"
|
||||
local_ip="xxx"
|
||||
|
||||
# The value of node0_ip must be consistent with the value of local_ip set in node0 (master node)
|
||||
node0_ip="xxxx"
|
||||
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=100
|
||||
export VLLM_USE_V1=1
|
||||
export HCCL_BUFFSIZE=200
|
||||
export VLLM_ASCEND_ENABLE_MLAPO=1
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export VLLM_ASCEND_ENABLE_FLASHCOMM1=1
|
||||
export HCCL_CONNECT_TIMEOUT=120
|
||||
export HCCL_INTRA_PCIE_ENABLE=1
|
||||
export HCCL_INTRA_ROCE_ENABLE=0
|
||||
|
||||
vllm serve /root/.cache/modelscope/hub/models/vllm-ascend/DeepSeek-V3.2-W8A8 \
|
||||
--host 0.0.0.0 \
|
||||
--port 8077 \
|
||||
--headless \
|
||||
--data-parallel-size 2 \
|
||||
--data-parallel-size-local 1 \
|
||||
--data-parallel-start-rank 1 \
|
||||
--data-parallel-address $node0_ip \
|
||||
--data-parallel-rpc-port 13389 \
|
||||
--tensor-parallel-size 8 \
|
||||
--quantization ascend \
|
||||
--seed 1024 \
|
||||
--served-model-name deepseek_v3_2 \
|
||||
--enable-expert-parallel \
|
||||
--max-num-seqs 16 \
|
||||
--max-model-len 8192 \
|
||||
--max-num-batched-tokens 4096 \
|
||||
--trust-remote-code \
|
||||
--no-enable-prefix-caching \
|
||||
--gpu-memory-utilization 0.92 \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY", "cudagraph_capture_sizes":[8, 16, 24, 32, 40, 48]}' \
|
||||
--speculative-config '{"num_speculative_tokens": 3, "method": "deepseek_mtp"}'
|
||||
|
||||
```
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
### 5.3 Prefill-Decode Disaggregation
|
||||
|
||||
We recommend using Mooncake for deployment: [Mooncake](../features/pd_disaggregation_mooncake_multi_node.md).
|
||||
|
||||
```{warning}
|
||||
For DeepSeek-V3.2's Sparse Flash Attention (SFA) backend, when using Decode Context Parallelism (DCP) in PD disaggregation, enable it on both the prefiller and decoder nodes, or disable it on both. Enabling DCP on only one side can cause known accuracy issues.
|
||||
```
|
||||
|
||||
In the standard single-node deployment mode, Prefill (prompt processing) and Decode (token generation) tasks run on the same set of NPUs. PD (Prefill-Decode) separation addresses this by running Prefill and Decode on dedicated node groups, each configured independently:
|
||||
|
||||
- **Prefill nodes** focus on high-throughput prompt processing, optimized for compute and communication.
|
||||
- **Decode nodes** focus on low-latency token generation, optimized for memory bandwidth.
|
||||
|
||||
This architecture is recommended for production deployments with concurrent multi-user workloads, where stable latency and high throughput are both required.
|
||||
|
||||
We'd like to show the deployment guide of `DeepSeek-V3.2` on multi-node environment with 1P1D for better performance.
|
||||
|
||||
To run the vllm-ascend `Prefill-Decode Disaggregation` service, you need to deploy a `launch_online_dp.py` script and a `run_dp_template.sh` script on each node and deploy a `proxy.sh` script on prefill master node to forward requests.
|
||||
|
||||
[launch_online_dp.py](https://github.com/vllm-project/vllm-ascend/blob/main/examples/external_online_dp/launch_online_dp.py)
|
||||
|
||||
Parameter descriptions:
|
||||
|
||||
|Parameter|Type|Required|Default|Description|
|
||||
|---------|----|--------|-------|-----------|
|
||||
|`--dp-size`|int|Yes|-|Data parallel size (total number of DP ranks across all nodes).|
|
||||
|`--tp-size`|int|No|1|Tensor parallel size within each DP rank.|
|
||||
|`--dp-size-local`|int|No|(same as `--dp-size`)|Number of DP ranks on the current node. If not set, defaults to `--dp-size`.|
|
||||
|`--dp-rank-start`|int|No|0|Starting rank offset for data parallel ranks on this node.|
|
||||
|`--dp-address`|str|Yes|-|IP address of the data parallel master node (node 0).|
|
||||
|`--dp-rpc-port`|str|No|12345|RPC port for data parallel master communication.|
|
||||
|`--vllm-start-port`|int|No|9000|Starting port for each vLLM engine instance on this node. Each DP rank's engine port = `vllm_start_port` + local rank index.|
|
||||
|
||||
1. `run_dp_template.sh` script
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: pd-nodes
|
||||
|
||||
::::{tab-item} Node 0(Prefill)
|
||||
:sync: Node 0(Prefill)
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
nic_name="enp48s3u1u1" # change to your own nic name
|
||||
local_ip=141.61.39.105 # change to your own ip
|
||||
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=10
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export VLLM_USE_V1=1
|
||||
export HCCL_BUFFSIZE=256
|
||||
|
||||
export ASCEND_AGGREGATE_ENABLE=1
|
||||
export ASCEND_TRANSPORT_PRINT=1
|
||||
export ACL_OP_INIT_MODE=1
|
||||
export ASCEND_A3_ENABLE=1
|
||||
# Timeout (in seconds) for automatically releasing the prefiller’s KV cache for a particular request.
|
||||
export VLLM_MOONCAKE_ABORT_REQUEST_TIMEOUT=480
|
||||
|
||||
export ASCEND_RT_VISIBLE_DEVICES=$1
|
||||
|
||||
export VLLM_ASCEND_ENABLE_FLASHCOMM1=1
|
||||
|
||||
vllm serve /root/.cache/Eco-Tech/DeepSeek-V3.2-w8a8-mtp-QuaRot \
|
||||
--host 0.0.0.0 \
|
||||
--port $2 \
|
||||
--data-parallel-size $3 \
|
||||
--data-parallel-rank $4 \
|
||||
--data-parallel-address $5 \
|
||||
--data-parallel-rpc-port $6 \
|
||||
--tensor-parallel-size $7 \
|
||||
--enable-expert-parallel \
|
||||
--speculative-config '{"num_speculative_tokens": 1, "method":"deepseek_mtp"}' \
|
||||
--profiler-config \
|
||||
'{"profiler": "torch",
|
||||
"torch_profiler_dir": "./vllm_profile",
|
||||
"torch_profiler_with_stack": false}' \
|
||||
--seed 1024 \
|
||||
--served-model-name deepseek_v3.2 \
|
||||
--max-model-len 68000 \
|
||||
--max-num-batched-tokens 32560 \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 64 \
|
||||
--gpu-memory-utilization 0.82 \
|
||||
--quantization ascend \
|
||||
--enforce-eager \
|
||||
--no-enable-prefix-caching \
|
||||
--additional-config '{"layer_sharding": ["q_b_proj", "o_proj"], "enable_dsa_cp": true}' \
|
||||
--kv-transfer-config \
|
||||
'{"kv_connector": "MooncakeConnectorV1",
|
||||
"kv_role": "kv_producer",
|
||||
"kv_port": "30000",
|
||||
"kv_connector_extra_config": {
|
||||
"prefill": {
|
||||
"dp_size": 2,
|
||||
"tp_size": 16
|
||||
},
|
||||
"decode": {
|
||||
"dp_size": 8,
|
||||
"tp_size": 4
|
||||
}
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} Node 1(Prefill)
|
||||
:sync: Node 1(Prefill)
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
nic_name="enp48s3u1u1" # change to your own nic name
|
||||
local_ip=141.61.39.113 # change to your own ip
|
||||
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=10
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export VLLM_USE_V1=1
|
||||
export HCCL_BUFFSIZE=256
|
||||
|
||||
export ASCEND_AGGREGATE_ENABLE=1
|
||||
export ASCEND_TRANSPORT_PRINT=1
|
||||
export ACL_OP_INIT_MODE=1
|
||||
export ASCEND_A3_ENABLE=1
|
||||
# Timeout (in seconds) for automatically releasing the prefiller’s KV cache for a particular request.
|
||||
export VLLM_MOONCAKE_ABORT_REQUEST_TIMEOUT=480
|
||||
|
||||
export ASCEND_RT_VISIBLE_DEVICES=$1
|
||||
|
||||
export VLLM_ASCEND_ENABLE_FLASHCOMM1=1
|
||||
|
||||
vllm serve /root/.cache/Eco-Tech/DeepSeek-V3.2-w8a8-mtp-QuaRot \
|
||||
--host 0.0.0.0 \
|
||||
--port $2 \
|
||||
--data-parallel-size $3 \
|
||||
--data-parallel-rank $4 \
|
||||
--data-parallel-address $5 \
|
||||
--data-parallel-rpc-port $6 \
|
||||
--tensor-parallel-size $7 \
|
||||
--enable-expert-parallel \
|
||||
--speculative-config '{"num_speculative_tokens": 1, "method":"deepseek_mtp"}' \
|
||||
--profiler-config \
|
||||
'{"profiler": "torch",
|
||||
"torch_profiler_dir": "./vllm_profile",
|
||||
"torch_profiler_with_stack": false}' \
|
||||
--seed 1024 \
|
||||
--served-model-name deepseek_v3.2 \
|
||||
--max-model-len 68000 \
|
||||
--max-num-batched-tokens 32560 \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 64 \
|
||||
--gpu-memory-utilization 0.82 \
|
||||
--quantization ascend \
|
||||
--enforce-eager \
|
||||
--no-enable-prefix-caching \
|
||||
--additional-config '{"layer_sharding": ["q_b_proj", "o_proj"], "enable_dsa_cp": true}' \
|
||||
--kv-transfer-config \
|
||||
'{"kv_connector": "MooncakeConnectorV1",
|
||||
"kv_role": "kv_producer",
|
||||
"kv_port": "30000",
|
||||
"kv_connector_extra_config": {
|
||||
"prefill": {
|
||||
"dp_size": 2,
|
||||
"tp_size": 16
|
||||
},
|
||||
"decode": {
|
||||
"dp_size": 8,
|
||||
"tp_size": 4
|
||||
}
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Node 0(Decode)
|
||||
:sync: Node 0(Decode)
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
nic_name="enp48s3u1u1" # change to your own nic name
|
||||
local_ip=141.61.39.117 # change to your own ip
|
||||
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
|
||||
#Mooncake
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=10
|
||||
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export VLLM_USE_V1=1
|
||||
export HCCL_BUFFSIZE=256
|
||||
|
||||
export ASCEND_AGGREGATE_ENABLE=1
|
||||
export ASCEND_TRANSPORT_PRINT=1
|
||||
export ACL_OP_INIT_MODE=1
|
||||
export ASCEND_A3_ENABLE=1
|
||||
# Timeout (in seconds) for automatically releasing the prefiller’s KV cache for a particular request.
|
||||
export VLLM_MOONCAKE_ABORT_REQUEST_TIMEOUT=480
|
||||
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
|
||||
export ASCEND_RT_VISIBLE_DEVICES=$1
|
||||
|
||||
vllm serve /root/.cache/Eco-Tech/DeepSeek-V3.2-w8a8-mtp-QuaRot \
|
||||
--host 0.0.0.0 \
|
||||
--port $2 \
|
||||
--data-parallel-size $3 \
|
||||
--data-parallel-rank $4 \
|
||||
--data-parallel-address $5 \
|
||||
--data-parallel-rpc-port $6 \
|
||||
--tensor-parallel-size $7 \
|
||||
--enable-expert-parallel \
|
||||
--speculative-config '{"num_speculative_tokens": 2, "method":"deepseek_mtp"}' \
|
||||
--profiler-config \
|
||||
'{"profiler": "torch",
|
||||
"torch_profiler_dir": "./vllm_profile",
|
||||
"torch_profiler_with_stack": false}' \
|
||||
--seed 1024 \
|
||||
--served-model-name deepseek_v3.2 \
|
||||
--max-model-len 68000 \
|
||||
--max-num-batched-tokens 12 \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY", "cudagraph_capture_sizes":[3, 6, 9, 12]}' \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 4 \
|
||||
--gpu-memory-utilization 0.95 \
|
||||
--no-enable-prefix-caching \
|
||||
--quantization ascend \
|
||||
--kv-transfer-config \
|
||||
'{"kv_connector": "MooncakeConnectorV1",
|
||||
"kv_role": "kv_consumer",
|
||||
"kv_port": "30100",
|
||||
"kv_connector_extra_config": {
|
||||
"prefill": {
|
||||
"dp_size": 2,
|
||||
"tp_size": 16
|
||||
},
|
||||
"decode": {
|
||||
"dp_size": 8,
|
||||
"tp_size": 4
|
||||
}
|
||||
}
|
||||
}' \
|
||||
--additional-config '{"recompute_scheduler_enable" : true}'
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Node 1(Decode)
|
||||
:sync: Node 1(Decode)
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
nic_name="enp48s3u1u1" # change to your own nic name
|
||||
local_ip=141.61.39.181 # change to your own ip
|
||||
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
|
||||
#Mooncake
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=10
|
||||
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export VLLM_USE_V1=1
|
||||
export HCCL_BUFFSIZE=256
|
||||
|
||||
export ASCEND_AGGREGATE_ENABLE=1
|
||||
export ASCEND_TRANSPORT_PRINT=1
|
||||
export ACL_OP_INIT_MODE=1
|
||||
export ASCEND_A3_ENABLE=1
|
||||
# Timeout (in seconds) for automatically releasing the prefiller’s KV cache for a particular request.
|
||||
export VLLM_MOONCAKE_ABORT_REQUEST_TIMEOUT=480
|
||||
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
|
||||
export ASCEND_RT_VISIBLE_DEVICES=$1
|
||||
|
||||
vllm serve /root/.cache/Eco-Tech/DeepSeek-V3.2-w8a8-mtp-QuaRot \
|
||||
--host 0.0.0.0 \
|
||||
--port $2 \
|
||||
--data-parallel-size $3 \
|
||||
--data-parallel-rank $4 \
|
||||
--data-parallel-address $5 \
|
||||
--data-parallel-rpc-port $6 \
|
||||
--tensor-parallel-size $7 \
|
||||
--enable-expert-parallel \
|
||||
--speculative-config '{"num_speculative_tokens": 2, "method":"deepseek_mtp"}' \
|
||||
--profiler-config \
|
||||
'{"profiler": "torch",
|
||||
"torch_profiler_dir": "./vllm_profile",
|
||||
"torch_profiler_with_stack": false}' \
|
||||
--seed 1024 \
|
||||
--served-model-name deepseek_v3.2 \
|
||||
--max-model-len 68000 \
|
||||
--max-num-batched-tokens 12 \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY", "cudagraph_capture_sizes":[3, 6, 9, 12]}' \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 4 \
|
||||
--gpu-memory-utilization 0.95 \
|
||||
--no-enable-prefix-caching \
|
||||
--quantization ascend \
|
||||
--kv-transfer-config \
|
||||
'{"kv_connector": "MooncakeConnectorV1",
|
||||
"kv_role": "kv_consumer",
|
||||
"kv_port": "30100",
|
||||
"kv_connector_extra_config": {
|
||||
"prefill": {
|
||||
"dp_size": 2,
|
||||
"tp_size": 16
|
||||
},
|
||||
"decode": {
|
||||
"dp_size": 8,
|
||||
"tp_size": 4
|
||||
}
|
||||
}
|
||||
}' \
|
||||
--additional-config '{"recompute_scheduler_enable" : true}'
|
||||
```
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
Once the preparation is done, you can start the server with the following command on each node:
|
||||
Refer to [Distributed DP Server With Large-Scale Expert Parallelism](https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/large_scale_ep.html) to get the detailed boot method.
|
||||
|
||||
2. run server for each node:
|
||||
|
||||
```shell
|
||||
# p0
|
||||
python launch_online_dp.py --dp-size 2 --tp-size 16 --dp-size-local 1 --dp-rank-start 0 --dp-address 141.61.39.105 --dp-rpc-port 12890 --vllm-start-port 9100
|
||||
# p1
|
||||
python launch_online_dp.py --dp-size 2 --tp-size 16 --dp-size-local 1 --dp-rank-start 1 --dp-address 141.61.39.105 --dp-rpc-port 12890 --vllm-start-port 9100
|
||||
# d0
|
||||
python launch_online_dp.py --dp-size 8 --tp-size 4 --dp-size-local 4 --dp-rank-start 0 --dp-address 141.61.39.117 --dp-rpc-port 12777 --vllm-start-port 9100
|
||||
# d1
|
||||
python launch_online_dp.py --dp-size 8 --tp-size 4 --dp-size-local 4 --dp-rank-start 4 --dp-address 141.61.39.117 --dp-rpc-port 12777 --vllm-start-port 9100
|
||||
```
|
||||
|
||||
3. Run the `proxy.sh` script on the prefill master node
|
||||
|
||||
Run a proxy server on the same node with the prefiller service instance. You can get the proxy program in the repository's examples: [load_balance_proxy_server_example.py](https://github.com/vllm-project/vllm-ascend/blob/main/examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py)
|
||||
|
||||
```shell
|
||||
unset http_proxy
|
||||
unset https_proxy
|
||||
|
||||
python load_balance_proxy_server_example.py \
|
||||
--port 8000 \
|
||||
--host 141.61.39.105 \
|
||||
--prefiller-hosts \
|
||||
141.61.39.105 \
|
||||
141.61.39.113 \
|
||||
--prefiller-ports \
|
||||
9100 \
|
||||
9100 \
|
||||
--decoder-hosts \
|
||||
141.61.39.117 \
|
||||
141.61.39.117 \
|
||||
141.61.39.117 \
|
||||
141.61.39.117 \
|
||||
141.61.39.181 \
|
||||
141.61.39.181 \
|
||||
141.61.39.181 \
|
||||
141.61.39.181 \
|
||||
--decoder-ports \
|
||||
9100 9101 9102 9103 \
|
||||
9100 9101 9102 9103 \
|
||||
```
|
||||
|
||||
```shell
|
||||
cd vllm-ascend/examples/disaggregated_prefill_v1/
|
||||
bash proxy.sh
|
||||
```
|
||||
|
||||
Common Issues Tip: If you encounter issues with PD separation deployment, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html) for troubleshooting.
|
||||
|
||||
## 6 Functional Verification
|
||||
|
||||
Once your server is started, you can query the model with input prompts:
|
||||
|
||||
**Note**:
|
||||
|
||||
- `<node0_ip>`: The IP address of the node where the server is running (e.g., localhost). For PD-separated deployment, use the host IP of the node where the proxy script resides.
|
||||
- `<port>`: The port number specified in the server startup command (e.g., 8000). For PD-separated deployment, use the port configured in the proxy script.
|
||||
|
||||
```shell
|
||||
curl http://<node0_ip>:<port>/v1/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "deepseek_v3.2",
|
||||
"prompt": "The future of AI is",
|
||||
"max_completion_tokens": 50,
|
||||
"temperature": 0
|
||||
}'
|
||||
```
|
||||
|
||||
**Expected Result**:
|
||||
|
||||
```json
|
||||
{"id":"019eab54ead036b23e53f3a709e09289","object":"chat.completion","created":1780990929,"model":"deepseek_v3.2","choices":[{"index":0,"message":{"role":"assistant","content":"The future of AI is **not a single destination, but a complex, multi-faceted trajectory** that will reshape nearly every aspect of human society, technology, and our understanding of intelligence itself. It can be understood through several interconnected lenses:\n\n### "},"finish_reason":"length"}],"usage":{"prompt_tokens":9,"completion_tokens":50,"total_tokens":59,"completion_tokens_details":{"reasoning_tokens":0},"prompt_tokens_details":{"cached_tokens":0},"prompt_cache_hit_tokens":0,"prompt_cache_miss_tokens":9},"system_fingerprint":""}
|
||||
```
|
||||
|
||||
## 7 Accuracy Evaluation
|
||||
|
||||
Here are two accuracy evaluation methods.
|
||||
|
||||
### Using AISBench
|
||||
|
||||
1. Refer to [Using AISBench](../../developer_guide/evaluation/using_ais_bench.md) for details.
|
||||
|
||||
2. After execution, you can get the result.
|
||||
|
||||
### Using Language Model Evaluation Harness
|
||||
|
||||
As an example, take the `gsm8k` dataset as a test dataset, and run accuracy evaluation of `DeepSeek-V3.2-W8A8` in online mode.
|
||||
|
||||
1. Refer to [Using lm_eval](../../developer_guide/evaluation/using_lm_eval.md) for `lm_eval` installation.
|
||||
|
||||
2. Run `lm_eval` to execute the accuracy evaluation.
|
||||
|
||||
```shell
|
||||
lm_eval \
|
||||
--model local-completions \
|
||||
--model_args model=/root/.cache/Eco-Tech/DeepSeek-V3.2-w8a8-mtp-QuaRot,base_url=http://127.0.0.1:8000/v1/completions,tokenized_requests=False,trust_remote_code=True \
|
||||
--tasks gsm8k \
|
||||
--output_path ./
|
||||
```
|
||||
|
||||
3. After execution, you can get the result.
|
||||
|
||||
## 8 Performance Evaluation
|
||||
|
||||
### Using AISBench
|
||||
|
||||
Refer to [Using AISBench for performance evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation) for details.
|
||||
|
||||
The performance result is:
|
||||
|
||||
**Hardware**: A3-752T, 4 node
|
||||
|
||||
**Deployment**: 1P1D, Prefill node: DP2+TP16, Decode Node: DP8+TP4
|
||||
|
||||
**Input/Output**: 64k/3k
|
||||
|
||||
**Performance**: 533tps, TPOT 32ms
|
||||
|
||||
### Using vLLM Benchmark
|
||||
|
||||
Run performance evaluation of `DeepSeek-V3.2-W8A8` as an example.
|
||||
|
||||
Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/) for more details.
|
||||
|
||||
There are three `vllm bench` subcommands:
|
||||
|
||||
- `latency`: Benchmark the latency of a single batch of requests.
|
||||
- `serve`: Benchmark the online serving throughput.
|
||||
- `throughput`: Benchmark offline inference throughput.
|
||||
|
||||
Take the `serve` as an example. Run the code as follows.
|
||||
|
||||
```shell
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
vllm bench serve --model /root/.cache/Eco-Tech/DeepSeek-V3.2-w8a8-mtp-QuaRot --dataset-name random --random-input 200 --num-prompts 200 --request-rate 1 --save-result --result-dir ./
|
||||
```
|
||||
|
||||
## 9 Function Call
|
||||
|
||||
The function call feature is supported from v0.13.0rc1 on. Please use the latest version.
|
||||
|
||||
Refer to [DeepSeek-V3.2 Usage Guide](https://docs.vllm.ai/projects/recipes/en/latest/DeepSeek/DeepSeek-V3_2.html#tool-calling-example) for details.
|
||||
|
||||
## 10 FAQ
|
||||
|
||||
For common environment, installation, and general parameter issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html).
|
||||
1206
docs/source/tutorials/models/DeepSeek-V4-Flash.md
Normal file
1206
docs/source/tutorials/models/DeepSeek-V4-Flash.md
Normal file
File diff suppressed because it is too large
Load Diff
1754
docs/source/tutorials/models/DeepSeek-V4-Pro.md
Normal file
1754
docs/source/tutorials/models/DeepSeek-V4-Pro.md
Normal file
File diff suppressed because it is too large
Load Diff
258
docs/source/tutorials/models/DeepSeekOCR2.md
Normal file
258
docs/source/tutorials/models/DeepSeekOCR2.md
Normal file
@@ -0,0 +1,258 @@
|
||||
# DeepSeek-OCR-2
|
||||
|
||||
## 1 Introduction
|
||||
|
||||
DeepSeekOCR2 is a model to investigate the role of vision encoders from an LLM-centric viewpoint.
|
||||
|
||||
The `DeepSeek-OCR-2` model is first supported in `vllm-ascend:v0.16.0` and can stably run in v0.16.0 and later version.
|
||||
|
||||
This document will show the main verification steps of the model, including supported features, feature configuration, environment preparation, single-node deployment, accuracy and performance evaluation.
|
||||
|
||||
## 2 Supported Features
|
||||
|
||||
Refer to [Supported Features List](../../user_guide/support_matrix/supported_models.md) to get the model's supported feature matrix.
|
||||
|
||||
Refer to [Feature Guide](../../user_guide/feature_guide/index.md) to get the feature's configuration.
|
||||
|
||||
## 3 Prerequisites
|
||||
|
||||
### 3.1 Model Weight
|
||||
|
||||
- `DeepSeek-OCR-2`: [Download model weight](https://huggingface.co/deepseek-ai/DeepSeek-OCR-2).
|
||||
|
||||
It is recommended to download the model weight to the shared directory of multiple nodes, such as `/root/.cache/`.
|
||||
|
||||
### 3.2 Verify Multi-node Communication
|
||||
|
||||
If you want to deploy multi-node environment, you need to verify multi-node communication according to [verify multi-node communication environment](../../installation.md#verify-multi-node-communication).
|
||||
|
||||
## 4 Installation
|
||||
|
||||
### 4.1 Docker Image Installation
|
||||
|
||||
You can use our official docker image to run `DeepSeek-OCR-2` directly.
|
||||
|
||||
Select an image based on your machine type and start the docker image on your node, refer to [using docker](../../installation.md#set-up-using-docker).
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: install
|
||||
|
||||
::::{tab-item} A2 series
|
||||
:sync: A2
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=m.daocloud.io/quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
export NAME=vllm-ascend
|
||||
|
||||
# Run the container using the defined variables
|
||||
# Note: If you are running bridge network with docker, please expose available ports for multiple nodes communication in advance.
|
||||
docker run --rm \
|
||||
--name $NAME \
|
||||
--net=host \
|
||||
--shm-size=1g \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} A3 series
|
||||
:sync: A3
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=m.daocloud.io/quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
export NAME=vllm-ascend
|
||||
|
||||
# Run the container using the defined variables
|
||||
# Note: If you are running bridge network with docker, please expose available ports for multiple nodes communication in advance.
|
||||
docker run --rm \
|
||||
--name $NAME \
|
||||
--net=host \
|
||||
--shm-size=1g \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci8 \
|
||||
--device /dev/davinci9 \
|
||||
--device /dev/davinci10 \
|
||||
--device /dev/davinci11 \
|
||||
--device /dev/davinci12 \
|
||||
--device /dev/davinci13 \
|
||||
--device /dev/davinci14 \
|
||||
--device /dev/davinci15 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
If you want to deploy multi-node environment, you need to set up environment on each node.
|
||||
|
||||
## 5 Online Service Deployment
|
||||
|
||||
### 5.1 Single-Node Online Deployment
|
||||
|
||||
- `DeepSeek-OCR-2` can be deployed on 1 Atlas 800 A2.
|
||||
|
||||
Run the following script to execute online inference.
|
||||
|
||||
```shell
|
||||
#!/bin/sh
|
||||
|
||||
export VLLM_USE_V1=1
|
||||
export VLLM_ASCEND_ENABLE_NZ=0
|
||||
export TOKENIZERS_PARALLELISM=false
|
||||
export PYTORCH_NPU_ALLOC_CONF="expandable_segments:True"
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export TOKENIZERS_PARALLELISM=false
|
||||
|
||||
vllm serve /root/.cache/DeepSeek-OCR-2 \
|
||||
--served-model-name deepseekocr2 \
|
||||
--trust-remote-code \
|
||||
--tensor-parallel-size 1 \
|
||||
--port 1055 \
|
||||
--max_model_len 8192 \
|
||||
--no-enable-prefix-caching \
|
||||
--gpu-memory-utilization 0.8 \
|
||||
--allowed-local-media-path / \
|
||||
--additional-config '{
|
||||
"enable_cpu_binding": true,
|
||||
"multistream_overlap_shared_expert": true,
|
||||
"ascend_compilation_config": {"fuse_qknorm_rope": false}
|
||||
}' \
|
||||
--mm-processor-cache-gb 0
|
||||
```
|
||||
|
||||
**Notice:**
|
||||
The parameters are explained as follows:
|
||||
|
||||
- `--max-model-len` specifies the maximum context length - that is, the sum of input and output tokens for a single request.
|
||||
- `--no-enable-prefix-caching` indicates that prefix caching is disabled. To enable it, remove this option.
|
||||
- `--gpu-memory-utilization` represents the proportion of HBM that vLLM will use for actual inference. Its essential function is to calculate the available kv_cache size. During the warm-up phase (referred to as profile run in vLLM), vLLM records the peak GPU memory usage during an inference process with an input size of `--max-num-batched-tokens`. The available kv_cache size is then calculated as: `--gpu-memory-utilization` * HBM size - peak GPU memory usage. Therefore, the larger the value of `--gpu-memory-utilization`, the more kv_cache can be used. However, since the GPU memory usage during the warm-up phase may differ from that during actual inference (e.g., due to uneven EP load), setting `--gpu-memory-utilization` too high may lead to OOM (Out of Memory) issues during actual inference. The default value is `0.9`.
|
||||
|
||||
### 5.2 Multi-node Deployment
|
||||
|
||||
Single-node deployment is recommended.
|
||||
|
||||
### 5.3 Prefill-Decode Disaggregation
|
||||
|
||||
We don't need to Prefill-Decode disaggregation
|
||||
|
||||
## 6 Functional Verification
|
||||
|
||||
If your service start successfully, you can see the info shown below:
|
||||
|
||||
```bash
|
||||
INFO: Started server process [87471]
|
||||
INFO: Waiting for application startup.
|
||||
INFO: Application startup complete.
|
||||
```
|
||||
|
||||
Once your server is started, you can query the model with input prompts:
|
||||
|
||||
```shell
|
||||
curl http://<node0_ip>:<port>/v1/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "deepseekocr2",
|
||||
"prompt": "The future of AI is",
|
||||
"max_completion_tokens": 50,
|
||||
"temperature": 0
|
||||
}'
|
||||
```
|
||||
|
||||
## 7 Accuracy Evaluation
|
||||
|
||||
### Using AISBench
|
||||
|
||||
1. Refer to [Using AISBench](../../developer_guide/evaluation/using_ais_bench.md) for details.
|
||||
|
||||
2. After execution, you can get the result, here is the result of `DeepSeek-OCR-2` for reference only.
|
||||
|
||||
| dataset | version | metric | mode | vllm-api-general-chat | note |
|
||||
|----- | ----- | ----- | ----- | -----| ----- |
|
||||
| textvqa | - | accuracy | gen | 50.28 | 1 Atlas 800 A2 |
|
||||
| omnidocbench | - | accuracy | gen | 66.86 | 1 Atlas 800 A2 |
|
||||
|
||||
## 8 Performance Evaluation
|
||||
|
||||
### Using AISBench
|
||||
|
||||
Refer to [Using AISBench for performance evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation) for details.
|
||||
|
||||
The performance result is:
|
||||
|
||||
**Hardware**: A2-313T, 1 node
|
||||
|
||||
**Input/Output**: 1080P/256
|
||||
|
||||
**Performance**: TTFT = 2s, TPOT = 200ms, Average performance of each card is 864TPS (Token Per Second).
|
||||
|
||||
## 9 Performance Tuning
|
||||
|
||||
### 9.1 Recommended Configurations
|
||||
|
||||
> **Note**: The following configurations are validated in specific test environments and are for reference only. The optimal configuration depends on factors such as maximum input/output length, prefix cache hit rate, precision requirements, and deployment machine ratios. It is recommended to refer to Section 9.2 for tuning based on actual conditions.
|
||||
|
||||
#### Table 1: Scenario Overview
|
||||
|
||||
> `*Total NPUs` indicates the total number of NPUs used across all nodes. 1 node = 1 Atlas 800 A3 server (64GB × 16 NPUs).
|
||||
|
||||
|Scenario|Deployment Mode|*Total NPUs|Weight Version|Key Considerations|
|
||||
|--------|---------------|-----------|--------------|------------------|
|
||||
|Multimodal<br>(1080P)|Single-Node Mixed|16 (A3)|deepseekocr2|dp1 tp1 for high-resolution visual inputs|
|
||||
|
||||
### 9.2 Tuning Guidelines
|
||||
|
||||
#### 9.2.1 General Tuning Reference
|
||||
|
||||
Please refer to the [Public Performance Tuning Documentation](../../developer_guide/performance_and_debug/optimization_and_tuning.md) for tuning methods.
|
||||
|
||||
Please refer to the [Feature Guide](../../user_guide/support_matrix/feature_matrix.md) for detailed feature descriptions.
|
||||
|
||||
## 10 FAQ
|
||||
|
||||
For common environment, installation, and general parameter issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html).
|
||||
|
||||
- **Q: Startup fails with HCCL port conflicts (address already bound). What should I do?**
|
||||
|
||||
A: Clean up old processes and restart: `pkill -f vLLM*`.
|
||||
|
||||
- **Q: How to handle OOM or unstable startup?**
|
||||
|
||||
A: Reduce `--max-num-seqs` and `--max-model-len` first. If needed, reduce concurrency and load-testing pressure (e.g., `max-concurrency` / `num-prompts`).
|
||||
800
docs/source/tutorials/models/GLM4.x.md
Normal file
800
docs/source/tutorials/models/GLM4.x.md
Normal file
@@ -0,0 +1,800 @@
|
||||
# GLM-4.5/4.6/4.7
|
||||
|
||||
## 1 Introduction
|
||||
|
||||
GLM-4.x series models use a Mixture-of-Experts (MoE) architecture and are foundational models specifically designed for agent applications.
|
||||
|
||||
The `GLM-4.5` model is first supported in `vllm-ascend:v0.10.0rc1`.
|
||||
|
||||
This document will show the main verification steps of the model, including supported features, feature configuration, environment preparation, single-node and multi-node deployment, accuracy and performance evaluation.
|
||||
|
||||
## 2 Supported Features
|
||||
|
||||
Refer to [supported features](../../user_guide/support_matrix/supported_models.md) to get the model's supported feature matrix.
|
||||
|
||||
Refer to [feature guide](../../user_guide/feature_guide/index.md) to get the feature's configuration.
|
||||
|
||||
## 3 Prerequisites
|
||||
|
||||
### 3.1 Model Weight
|
||||
|
||||
- `GLM-4.5`(BF16 version): [Download model weight](https://www.modelscope.cn/models/ZhipuAI/GLM-4.5).
|
||||
- `GLM-4.6`(BF16 version): [Download model weight](https://www.modelscope.cn/models/ZhipuAI/GLM-4.6).
|
||||
- `GLM-4.7`(BF16 version): [Download model weight](https://www.modelscope.cn/models/ZhipuAI/GLM-4.7).
|
||||
- `GLM-4.5-w8a8-with-float-mtp`(Quantized version with mtp): [Download model weight](https://modelers.cn/models/Modelers_Park/GLM-4.5-w8a8).
|
||||
- `GLM-4.6-w8a8`(Quantized version without mtp): [Download model weight](https://modelers.cn/models/Modelers_Park/GLM-4.6-w8a8). Because vllm does not support GLM4.6 mtp in October, we do not provide an mtp version. Since it is now supported, you can use the following quantization scheme to add mtp weights to the quantized weights.
|
||||
- `GLM-4.7-w8a8-with-float-mtp`(Quantized version with mtp): [Download model weight](https://www.modelscope.cn/models/Eco-Tech/GLM-4.7-W8A8-floatmtp).
|
||||
- `Method of Quantization`: [quantization scheme](https://ai.gitcode.com/Ascend-SACT/GLM-4.5-w8a8). You can use these methods to quantize the model.
|
||||
|
||||
It is recommended to download the model weight to the shared directory of multiple nodes, such as `/root/.cache/`.
|
||||
|
||||
## 4 Installation
|
||||
|
||||
You can use our official docker image to run `GLM-4.x` directly.
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: install
|
||||
|
||||
::::{tab-item} A3 series
|
||||
:sync: A3
|
||||
|
||||
Start the docker image on each node.
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-a3
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci8 \
|
||||
--device /dev/davinci9 \
|
||||
--device /dev/davinci10 \
|
||||
--device /dev/davinci11 \
|
||||
--device /dev/davinci12 \
|
||||
--device /dev/davinci13 \
|
||||
--device /dev/davinci14 \
|
||||
--device /dev/davinci15 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} A2 series
|
||||
:sync: A2
|
||||
|
||||
Start the docker image on each node.
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
In addition, if you don't want to use the docker image as above, you can also build all from source:
|
||||
|
||||
- Install `vllm-ascend` from source, refer to [installation](../../installation.md).
|
||||
|
||||
If you want to deploy multi-node environment, you need to set up environment on each node.
|
||||
|
||||
## 5 Online Service Deployment
|
||||
|
||||
### 5.1 Single-node Deployment
|
||||
|
||||
- In low-latency scenarios, we recommend a single-machine deployment.
|
||||
- Quantized model `glm4.7_w8a8_with_float_mtp` can be deployed on 1 Atlas 800 A3 (64G × 16) or 1 Atlas 800 A2 (64G × 8).
|
||||
|
||||
Run the following script to execute online inference.
|
||||
|
||||
```shell
|
||||
#!/bin/sh
|
||||
export HCCL_BUFFSIZE=512
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export HCCL_OP_EXPANSION_MODE=AIV
|
||||
export VLLM_ASCEND_BALANCE_SCHEDULING=1
|
||||
export VLLM_ASCEND_ENABLE_TOPK_OPTIMIZE=1
|
||||
|
||||
vllm serve Eco-Tech/GLM-4.7-W8A8-floatmtp \
|
||||
--data-parallel-size 2 \
|
||||
--tensor-parallel-size 8 \
|
||||
--enable-expert-parallel \
|
||||
--seed 1024 \
|
||||
--served-model-name glm \
|
||||
--max-model-len 133000 \
|
||||
--max-num-batched-tokens 8192 \
|
||||
--max-num-seqs 16 \
|
||||
--quantization ascend \
|
||||
--trust-remote-code \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--speculative-config '{"num_speculative_tokens": 3, "method":"mtp", "enforce_eager": true}' \
|
||||
--compilation-config '{"cudagraph_capture_sizes": [1,2,4,8,16,32,64,128,256,512], "cudagraph_mode": "FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"enable_shared_expert_dp": true, "ascend_fusion_config": {"fusion_ops_gmmswigluquant": false}}'
|
||||
```
|
||||
|
||||
**Notice:**
|
||||
The parameters are explained as follows:
|
||||
|
||||
- `fusion_ops_gmmswigluquant` The performance of the GmmSwigluQuant fusion operator tends to degrade when the total number of NPUs is ≤ 16.
|
||||
- `VLLM_ASCEND_ENABLE_FLASHCOMM1` Due to the FD feature of the FIA operator being invalidated by padding data introduced by this feature, we recommend disabling the `flashcomm1` feature for long-sequence (≥16k) and low-concurrency (≤8 batch size) scenarios.For long-sequence and high-concurrency scenarios, you may enable this feature to achieve improved Prefill performance.
|
||||
|
||||
### 5.2 Multi-node Deployment
|
||||
|
||||
While the previous documentation advises against multi-node deployment on the Atlas 800 A2 (64G × 8) platform, this configuration can still be implemented for the GLM-4.x model if required. To proceed with a dual-node setup, execute the following scripts on each respective node.
|
||||
|
||||
**Node 0**
|
||||
|
||||
```shell
|
||||
#!/bin/sh
|
||||
|
||||
# this obtained through ifconfig
|
||||
# nic_name is the network interface name corresponding to local_ip of the current node
|
||||
nic_name="xxxx"
|
||||
local_ip="xxxx"
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_BUFFSIZE=512
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export HCCL_OP_EXPANSION_MODE=AIV
|
||||
export VLLM_ASCEND_BALANCE_SCHEDULING=1
|
||||
export VLLM_ASCEND_ENABLE_TOPK_OPTIMIZE=1
|
||||
|
||||
vllm serve Eco-Tech/GLM-4.7-W8A8-floatmtp \
|
||||
--host 0.0.0.0 \
|
||||
--port 8004 \
|
||||
--data-parallel-size 2 \
|
||||
--data-parallel-size-local 1 \
|
||||
--data-parallel-start-rank 0 \
|
||||
--data-parallel-address $local_ip \
|
||||
--data-parallel-rpc-port 13389 \
|
||||
--tensor-parallel-size 8 \
|
||||
--enable-expert-parallel \
|
||||
--seed 1024 \
|
||||
--max-model-len 140000 \
|
||||
--max-num-batched-tokens 8192 \
|
||||
--max-num-seqs 16 \
|
||||
--quantization ascend \
|
||||
--trust-remote-code \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--enable-auto-tool-choice \
|
||||
--reasoning-parser glm45 \
|
||||
--tool-call-parser glm47 \
|
||||
--served-model-name glm47 \
|
||||
--speculative-config '{"num_speculative_tokens": 3, "method":"mtp", "enforce_eager": true}' \
|
||||
--compilation-config '{"cudagraph_capture_sizes": [1,2,4,8,16,32,64,128,256,512], "cudagraph_mode": "FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"enable_shared_expert_dp": true, "ascend_fusion_config": {"fusion_ops_gmmswigluquant": false}}'
|
||||
```
|
||||
|
||||
**Node 1**
|
||||
|
||||
```shell
|
||||
#!/bin/sh
|
||||
|
||||
# this obtained through ifconfig
|
||||
# nic_name is the network interface name corresponding to local_ip of the current node
|
||||
nic_name="xxxx"
|
||||
local_ip="xxxx"
|
||||
node0_ip="xxxx" # same as the local_IP address in node 0
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_BUFFSIZE=512
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export HCCL_OP_EXPANSION_MODE=AIV
|
||||
export VLLM_ASCEND_BALANCE_SCHEDULING=1
|
||||
export VLLM_ASCEND_ENABLE_TOPK_OPTIMIZE=1
|
||||
|
||||
vllm serve Eco-Tech/GLM-4.7-W8A8-floatmtp \
|
||||
--host 0.0.0.0 \
|
||||
--port 8004 \
|
||||
--headless \
|
||||
--data-parallel-size 2 \
|
||||
--data-parallel-size-local 1 \
|
||||
--data-parallel-start-rank 1 \
|
||||
--data-parallel-address $node0_ip \
|
||||
--data-parallel-rpc-port 13389 \
|
||||
--tensor-parallel-size 8 \
|
||||
--enable-expert-parallel \
|
||||
--seed 1024 \
|
||||
--max-model-len 140000 \
|
||||
--max-num-batched-tokens 8192 \
|
||||
--max-num-seqs 16 \
|
||||
--quantization ascend \
|
||||
--trust-remote-code \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--enable-auto-tool-choice \
|
||||
--reasoning-parser glm45 \
|
||||
--tool-call-parser glm47 \
|
||||
--served-model-name glm47 \
|
||||
--speculative-config '{"num_speculative_tokens": 3, "method":"mtp", "enforce_eager": true}' \
|
||||
--compilation-config '{"cudagraph_capture_sizes": [1,2,4,8,16,32,64,128,256,512], "cudagraph_mode": "FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"enable_shared_expert_dp": true, "ascend_fusion_config": {"fusion_ops_gmmswigluquant": false}}'
|
||||
```
|
||||
|
||||
### 5.3 Prefill-Decode Disaggregation
|
||||
|
||||
We'd like to show the deployment guide of `GLM-4.7` in a multi-node environment with 2P1D for better performance.
|
||||
|
||||
Before you start, please
|
||||
|
||||
1. prepare the script `launch_online_dp.py` on each node:
|
||||
|
||||
```python
|
||||
import argparse
|
||||
import multiprocessing
|
||||
import os
|
||||
import subprocess
|
||||
import sys
|
||||
|
||||
def parse_args():
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument(
|
||||
"--dp-size",
|
||||
type=int,
|
||||
required=True,
|
||||
help="Data parallel size."
|
||||
)
|
||||
parser.add_argument(
|
||||
"--tp-size",
|
||||
type=int,
|
||||
default=1,
|
||||
help="Tensor parallel size."
|
||||
)
|
||||
parser.add_argument(
|
||||
"--dp-size-local",
|
||||
type=int,
|
||||
default=-1,
|
||||
help="Local data parallel size."
|
||||
)
|
||||
parser.add_argument(
|
||||
"--dp-rank-start",
|
||||
type=int,
|
||||
default=0,
|
||||
help="Starting rank for data parallel."
|
||||
)
|
||||
parser.add_argument(
|
||||
"--dp-address",
|
||||
type=str,
|
||||
required=True,
|
||||
help="IP address for data parallel master node."
|
||||
)
|
||||
parser.add_argument(
|
||||
"--dp-rpc-port",
|
||||
type=str,
|
||||
default=12345,
|
||||
help="Port for data parallel master node."
|
||||
)
|
||||
parser.add_argument(
|
||||
"--vllm-start-port",
|
||||
type=int,
|
||||
default=9000,
|
||||
help="Starting port for the engine."
|
||||
)
|
||||
return parser.parse_args()
|
||||
|
||||
args = parse_args()
|
||||
dp_size = args.dp_size
|
||||
tp_size = args.tp_size
|
||||
dp_size_local = args.dp_size_local
|
||||
if dp_size_local == -1:
|
||||
dp_size_local = dp_size
|
||||
dp_rank_start = args.dp_rank_start
|
||||
dp_address = args.dp_address
|
||||
dp_rpc_port = args.dp_rpc_port
|
||||
vllm_start_port = args.vllm_start_port
|
||||
|
||||
def run_command(visible_devices, dp_rank, vllm_engine_port):
|
||||
command = [
|
||||
"bash",
|
||||
"./run_dp_template.sh",
|
||||
visible_devices,
|
||||
str(vllm_engine_port),
|
||||
str(dp_size),
|
||||
str(dp_rank),
|
||||
dp_address,
|
||||
dp_rpc_port,
|
||||
str(tp_size),
|
||||
]
|
||||
subprocess.run(command, check=True)
|
||||
|
||||
if __name__ == "__main__":
|
||||
template_path = "./run_dp_template.sh"
|
||||
if not os.path.exists(template_path):
|
||||
print(f"Template file {template_path} does not exist.")
|
||||
sys.exit(1)
|
||||
|
||||
processes = []
|
||||
num_cards = dp_size_local * tp_size
|
||||
for i in range(dp_size_local):
|
||||
dp_rank = dp_rank_start + i
|
||||
vllm_engine_port = vllm_start_port + i
|
||||
visible_devices = ",".join(str(x) for x in range(i * tp_size, (i + 1) * tp_size))
|
||||
process = multiprocessing.Process(target=run_command,
|
||||
args=(visible_devices, dp_rank,
|
||||
vllm_engine_port))
|
||||
processes.append(process)
|
||||
process.start()
|
||||
|
||||
for process in processes:
|
||||
process.join()
|
||||
|
||||
```
|
||||
|
||||
2. prepare the script `run_dp_template.sh` on each node.
|
||||
|
||||
1. Prefill node 0
|
||||
|
||||
```shell
|
||||
nic_name="xxxx" # change to your own nic name
|
||||
local_ip="xxxx" # change to your own ip
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_BUFFSIZE=256
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export ASCEND_AGGREGATE_ENABLE=1
|
||||
export ASCEND_TRANSPORT_PRINT=1
|
||||
export ACL_OP_INIT_MODE=1
|
||||
export ASCEND_A3_ENABLE=1
|
||||
export VLLM_ASCEND_ENABLE_TOPK_OPTIMIZE=1
|
||||
export ASCEND_RT_VISIBLE_DEVICES=$1
|
||||
export LD_LIBRARY_PATH=/usr/local/Ascend/ascend-toolkit/latest/python/site-packages/mooncake:$LD_LIBRARY_PATH
|
||||
|
||||
vllm serve Eco-Tech/GLM-4.7-W8A8-floatmtp \
|
||||
--host 0.0.0.0 \
|
||||
--port $2 \
|
||||
--data-parallel-size $3 \
|
||||
--data-parallel-rank $4 \
|
||||
--data-parallel-address $5 \
|
||||
--data-parallel-rpc-port $6 \
|
||||
--tensor-parallel-size $7 \
|
||||
--enable-expert-parallel \
|
||||
--seed 1024 \
|
||||
--served-model-name glm \
|
||||
--max-model-len 133000 \
|
||||
--max-num-batched-tokens 8192 \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 64 \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--quantization ascend \
|
||||
--enforce-eager \
|
||||
--speculative-config '{"num_speculative_tokens": 1, "method":"mtp", "enforce_eager": true}' \
|
||||
--profiler-config '{"profiler": "torch", "torch_profiler_dir": "./vllm_profile", "torch_profiler_with_stack": false}' \
|
||||
--additional-config '{"enable_shared_expert_dp": true, "ascend_fusion_config": {"fusion_ops_gmmswigluquant": false}}' \
|
||||
--kv-transfer-config \
|
||||
'{"kv_connector": "MooncakeConnectorV1",
|
||||
"kv_role": "kv_producer",
|
||||
"kv_port": "30000",
|
||||
"kv_connector_extra_config": {
|
||||
"prefill": {
|
||||
"dp_size": 2,
|
||||
"tp_size": 8
|
||||
},
|
||||
"decode": {
|
||||
"dp_size": 8,
|
||||
"tp_size": 4
|
||||
}
|
||||
}
|
||||
}' 2>&1
|
||||
|
||||
```
|
||||
|
||||
2. Prefill node 1
|
||||
|
||||
```shell
|
||||
nic_name="xxxx" # change to your own nic name
|
||||
local_ip="xxxx" # change to your own ip
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_BUFFSIZE=256
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export ASCEND_AGGREGATE_ENABLE=1
|
||||
export ASCEND_TRANSPORT_PRINT=1
|
||||
export ACL_OP_INIT_MODE=1
|
||||
export ASCEND_A3_ENABLE=1
|
||||
export VLLM_ASCEND_ENABLE_TOPK_OPTIMIZE=1
|
||||
export ASCEND_RT_VISIBLE_DEVICES=$1
|
||||
export LD_LIBRARY_PATH=/usr/local/Ascend/ascend-toolkit/latest/python/site-packages/mooncake:$LD_LIBRARY_PATH
|
||||
|
||||
vllm serve Eco-Tech/GLM-4.7-W8A8-floatmtp \
|
||||
--host 0.0.0.0 \
|
||||
--port $2 \
|
||||
--data-parallel-size $3 \
|
||||
--data-parallel-rank $4 \
|
||||
--data-parallel-address $5 \
|
||||
--data-parallel-rpc-port $6 \
|
||||
--tensor-parallel-size $7 \
|
||||
--enable-expert-parallel \
|
||||
--seed 1024 \
|
||||
--served-model-name glm \
|
||||
--max-model-len 133000 \
|
||||
--max-num-batched-tokens 8192 \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 64 \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--quantization ascend \
|
||||
--enforce-eager \
|
||||
--speculative-config '{"num_speculative_tokens": 1, "method":"mtp", "enforce_eager": true}' \
|
||||
--profiler-config '{"profiler": "torch", "torch_profiler_dir": "./vllm_profile", "torch_profiler_with_stack": false}' \
|
||||
--additional-config '{"enable_shared_expert_dp": true, "ascend_fusion_config": {"fusion_ops_gmmswigluquant": false}}' \
|
||||
--kv-transfer-config \
|
||||
'{"kv_connector": "MooncakeConnectorV1",
|
||||
"kv_role": "kv_producer",
|
||||
"kv_port": "30100",
|
||||
"kv_connector_extra_config": {
|
||||
"prefill": {
|
||||
"dp_size": 2,
|
||||
"tp_size": 8
|
||||
},
|
||||
"decode": {
|
||||
"dp_size": 8,
|
||||
"tp_size": 4
|
||||
}
|
||||
}
|
||||
}' 2>&1
|
||||
```
|
||||
|
||||
3. Decode node 0
|
||||
|
||||
```shell
|
||||
nic_name="xxxx" # change to your own nic name
|
||||
local_ip="xxxx" # change to your own ip
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
|
||||
export HCCL_BUFFSIZE=512
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export ASCEND_AGGREGATE_ENABLE=1
|
||||
export ASCEND_TRANSPORT_PRINT=1
|
||||
export ACL_OP_INIT_MODE=1
|
||||
export ASCEND_A3_ENABLE=1
|
||||
# Timeout (in seconds) for automatically releasing the prefiller’s KV cache for a particular request.
|
||||
export VLLM_MOONCAKE_ABORT_REQUEST_TIMEOUT=480
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export LD_LIBRARY_PATH=/usr/local/Ascend/ascend-toolkit/latest/python/site-packages/mooncake:$LD_LIBRARY_PATH
|
||||
export VLLM_ASCEND_ENABLE_TOPK_OPTIMIZE=1
|
||||
export VLLM_ASCEND_ENABLE_FUSED_MC2=1
|
||||
export ASCEND_RT_VISIBLE_DEVICES=$1
|
||||
|
||||
vllm serve Eco-Tech/GLM-4.7-W8A8-floatmtp \
|
||||
--host 0.0.0.0 \
|
||||
--port $2 \
|
||||
--data-parallel-size $3 \
|
||||
--data-parallel-rank $4 \
|
||||
--data-parallel-address $5 \
|
||||
--data-parallel-rpc-port $6 \
|
||||
--tensor-parallel-size $7 \
|
||||
--enable-expert-parallel \
|
||||
--seed 1024 \
|
||||
--served-model-name glm \
|
||||
--max-model-len 133000 \
|
||||
--max-num-batched-tokens 128 \
|
||||
--max-num-seqs 4 \
|
||||
--trust-remote-code \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--quantization ascend \
|
||||
--speculative-config '{"num_speculative_tokens": 3, "method":"mtp", "enforce_eager": true}' \
|
||||
--profiler-config \
|
||||
'{"profiler": "torch",
|
||||
"torch_profiler_dir": "./vllm_profile",
|
||||
"torch_profiler_with_stack": false}' \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY", "cudagraph_capture_sizes":[1,2,4,6,8,10,12,14,16,18,20,24,26,28,30,32,64,128,256,512]}' \
|
||||
--additional-config '{"recompute_scheduler_enable": true, "enable_shared_expert_dp": true, "ascend_fusion_config": {"fusion_ops_gmmswigluquant": false}}' \
|
||||
--kv-transfer-config \
|
||||
'{"kv_connector": "MooncakeConnectorV1",
|
||||
"kv_role": "kv_consumer",
|
||||
"kv_port": "30200",
|
||||
"kv_connector_extra_config": {
|
||||
"prefill": {
|
||||
"dp_size": 2,
|
||||
"tp_size": 8
|
||||
},
|
||||
"decode": {
|
||||
"dp_size": 8,
|
||||
"tp_size": 4
|
||||
}
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
4. Decode node 1
|
||||
|
||||
```shell
|
||||
nic_name="xxxx" # change to your own nic name
|
||||
local_ip="xxxx" # change to your own ip
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
|
||||
export HCCL_BUFFSIZE=512
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export ASCEND_AGGREGATE_ENABLE=1
|
||||
export ASCEND_TRANSPORT_PRINT=1
|
||||
export ACL_OP_INIT_MODE=1
|
||||
export ASCEND_A3_ENABLE=1
|
||||
# Timeout (in seconds) for automatically releasing the prefiller’s KV cache for a particular request.
|
||||
export VLLM_MOONCAKE_ABORT_REQUEST_TIMEOUT=480
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export LD_LIBRARY_PATH=/usr/local/Ascend/ascend-toolkit/latest/python/site-packages/mooncake:$LD_LIBRARY_PATH
|
||||
export VLLM_ASCEND_ENABLE_TOPK_OPTIMIZE=1
|
||||
export VLLM_ASCEND_ENABLE_FUSED_MC2=1
|
||||
export ASCEND_RT_VISIBLE_DEVICES=$1
|
||||
|
||||
vllm serve Eco-Tech/GLM-4.7-W8A8-floatmtp \
|
||||
--host 0.0.0.0 \
|
||||
--port $2 \
|
||||
--data-parallel-size $3 \
|
||||
--data-parallel-rank $4 \
|
||||
--data-parallel-address $5 \
|
||||
--data-parallel-rpc-port $6 \
|
||||
--tensor-parallel-size $7 \
|
||||
--enable-expert-parallel \
|
||||
--seed 1024 \
|
||||
--served-model-name glm \
|
||||
--max-model-len 133000 \
|
||||
--max-num-batched-tokens 128 \
|
||||
--max-num-seqs 4 \
|
||||
--trust-remote-code \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--quantization ascend \
|
||||
--speculative-config '{"num_speculative_tokens": 3, "method":"mtp", "enforce_eager": true}' \
|
||||
--profiler-config \
|
||||
'{"profiler": "torch",
|
||||
"torch_profiler_dir": "./vllm_profile",
|
||||
"torch_profiler_with_stack": false}' \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY", "cudagraph_capture_sizes":[1,2,4,6,8,10,12,14,16,18,20,24,26,28,30,32,64,128,256,512]}' \
|
||||
--additional-config '{"recompute_scheduler_enable": true, "enable_shared_expert_dp": true, "ascend_fusion_config": {"fusion_ops_gmmswigluquant": false}}' \
|
||||
--kv-transfer-config \
|
||||
'{"kv_connector": "MooncakeConnectorV1",
|
||||
"kv_role": "kv_consumer",
|
||||
"kv_port": "30200",
|
||||
"kv_connector_extra_config": {
|
||||
"prefill": {
|
||||
"dp_size": 2,
|
||||
"tp_size": 8
|
||||
},
|
||||
"decode": {
|
||||
"dp_size": 8,
|
||||
"tp_size": 4
|
||||
}
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
Once the preparation is done, you can start the server with the following command on each node:
|
||||
|
||||
1. Prefill node 0
|
||||
|
||||
```shell
|
||||
# change ip to your own
|
||||
python launch_online_dp.py --dp-size 2 --tp-size 8 --dp-size-local 2 --dp-rank-start 0 --dp-address $node_p0_ip --dp-rpc-port 12880 --vllm-start-port 9300
|
||||
```
|
||||
|
||||
2. Prefill node 1
|
||||
|
||||
```shell
|
||||
# change ip to your own
|
||||
python launch_online_dp.py --dp-size 2 --tp-size 8 --dp-size-local 2 --dp-rank-start 0 --dp-address $node_p1_ip --dp-rpc-port 12880 --vllm-start-port 9300
|
||||
```
|
||||
|
||||
3. Decode node 0
|
||||
|
||||
```shell
|
||||
# change ip to your own
|
||||
python launch_online_dp.py --dp-size 8 --tp-size 4 --dp-size-local 4 --dp-rank-start 0 --dp-address $node_d0_ip --dp-rpc-port 12778 --vllm-start-port 9300
|
||||
```
|
||||
|
||||
4. Decode node 1
|
||||
|
||||
```shell
|
||||
# change ip to your own
|
||||
python launch_online_dp.py --dp-size 8 --tp-size 4 --dp-size-local 4 --dp-rank-start 4 --dp-address $node_d0_ip --dp-rpc-port 12778 --vllm-start-port 9300
|
||||
```
|
||||
|
||||
### 5.4 Request Forwarding
|
||||
|
||||
To set up request forwarding, run the following script on any machine. You can get the proxy program in the repository's examples: [load_balance_proxy_server_example.py](https://github.com/vllm-project/vllm-ascend/blob/main/examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py)
|
||||
|
||||
```shell
|
||||
unset http_proxy
|
||||
unset https_proxy
|
||||
|
||||
python load_balance_proxy_server_example.py \
|
||||
--port 8000 \
|
||||
--host 0.0.0.0 \
|
||||
--prefiller-hosts \
|
||||
$node_p0_ip $node_p0_ip \
|
||||
$node_p1_ip $node_p1_ip \
|
||||
--prefiller-ports \
|
||||
9300 9301 \
|
||||
9300 9301 \
|
||||
--decoder-hosts \
|
||||
$node_d0_ip \
|
||||
$node_d0_ip \
|
||||
$node_d0_ip \
|
||||
$node_d0_ip \
|
||||
$node_d1_ip \
|
||||
$node_d1_ip \
|
||||
$node_d1_ip \
|
||||
$node_d1_ip \
|
||||
--decoder-ports \
|
||||
9300 9301 9302 9303 \
|
||||
9300 9301 9302 9303
|
||||
```
|
||||
|
||||
## 6 Functional Verification
|
||||
|
||||
Once your server is started, you can query the model with input prompts:
|
||||
|
||||
```shell
|
||||
curl -H "Accept: application/json" \
|
||||
-H "Content-type: application/json" \
|
||||
-X POST \
|
||||
-d '{
|
||||
"model": "glm",
|
||||
"messages": [{
|
||||
"role": "user",
|
||||
"content": "The future of AI is"
|
||||
}],
|
||||
"stream": false,
|
||||
"ignore_eos": false,
|
||||
"temperature": 0,
|
||||
"max_tokens": 200
|
||||
}' http://<node0_ip>:<port>/v1/chat/completions
|
||||
```
|
||||
|
||||
## 7 Accuracy Evaluation
|
||||
|
||||
Here are two accuracy evaluation methods.
|
||||
|
||||
### Using AISBench
|
||||
|
||||
1. Refer to [Using AISBench](../../developer_guide/evaluation/using_ais_bench.md) for details.
|
||||
|
||||
2. After execution, you can get the result, here is the result of `GLM-4.7` in `vllm-ascend:main` (after `vllm-ascend:0.14.0rc1`) for reference only.
|
||||
|
||||
| dataset | version | metric | mode | vllm-api-general-chat | note |
|
||||
|----- | ----- | ----- | ----- | -----| ----- |
|
||||
| GPQA | - | accuracy | gen | 84.85 | 1 Atlas 800 A3 (64G × 16) |
|
||||
| MATH500 | - | accuracy | gen | 98.8 | 1 Atlas 800 A3 (64G × 16) |
|
||||
|
||||
### Using Language Model Evaluation Harness
|
||||
|
||||
Not tested yet.
|
||||
|
||||
## 8 Performance Evaluation
|
||||
|
||||
### Using AISBench
|
||||
|
||||
Refer to [Using AISBench for performance evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation) for details.
|
||||
|
||||
### Using vLLM Benchmark
|
||||
|
||||
Run performance evaluation of `GLM-4.x` as an example.
|
||||
|
||||
Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/) for more details.
|
||||
|
||||
There are three `vllm bench` subcommands:
|
||||
|
||||
- `latency`: Benchmark the latency of a single batch of requests.
|
||||
- `serve`: Benchmark the online serving throughput.
|
||||
- `throughput`: Benchmark offline inference throughput.
|
||||
|
||||
Take the `serve` as an example. Run the code as follows.
|
||||
|
||||
```shell
|
||||
vllm bench serve \
|
||||
--backend vllm \
|
||||
--dataset-name prefix_repetition \
|
||||
--prefix-repetition-prefix-len 22400 \
|
||||
--prefix-repetition-suffix-len 9600 \
|
||||
--prefix-repetition-output-len 1024 \
|
||||
--num-prompts 1 \
|
||||
--prefix-repetition-num-prefixes 1 \
|
||||
--ignore-eos \
|
||||
--model glm \
|
||||
--tokenizer Eco-Tech/GLM-4.7-W8A8-floatmtp \
|
||||
--seed 1000 \
|
||||
--host 0.0.0.0 \
|
||||
--port 8000 \
|
||||
--endpoint /v1/completions \
|
||||
--max-concurrency 1 \
|
||||
--request-rate 1
|
||||
```
|
||||
|
||||
After about several minutes, you can get the performance evaluation result.
|
||||
|
||||
## 9 Performance Tuning
|
||||
|
||||
In this chapter, we recommend best practices for three scenarios:
|
||||
|
||||
- Long-context: For long sequences with low concurrency (≤ 4): set `dp1 tp16`; For long sequences with high concurrency (> 4): set `dp2 tp8`
|
||||
- Low-latency: For short sequences with low latency: we recommend setting `dp2 tp8`
|
||||
- High-throughput: For short sequences with high throughput: we also recommend setting `dp2 tp8`
|
||||
|
||||
**Notice:**
|
||||
`max-model-len` and `max-num-seqs` need to be set according to the actual usage scenario. For other settings, please refer to the **[Deployment](#5-online-service-deployment)** chapter.
|
||||
|
||||
## 10 FAQ
|
||||
|
||||
- **Q: Startup fails with HCCL port conflicts (address already bound). What should I do?**
|
||||
|
||||
A: Clean up old processes and restart: `pkill -f VLLM*`.
|
||||
|
||||
- **Q: How to handle OOM or unstable startup?**
|
||||
|
||||
A: Reduce `--max-num-seqs` and `--max-model-len` first. If needed, reduce concurrency and load-testing pressure (e.g., `max-concurrency` / `num-prompts`).
|
||||
1685
docs/source/tutorials/models/GLM5.2.md
Normal file
1685
docs/source/tutorials/models/GLM5.2.md
Normal file
File diff suppressed because it is too large
Load Diff
2071
docs/source/tutorials/models/GLM5.md
Normal file
2071
docs/source/tutorials/models/GLM5.md
Normal file
File diff suppressed because it is too large
Load Diff
231
docs/source/tutorials/models/Hunyuan-A13B-Instruct.md
Normal file
231
docs/source/tutorials/models/Hunyuan-A13B-Instruct.md
Normal file
@@ -0,0 +1,231 @@
|
||||
# Hunyuan-A13B-Instruct
|
||||
|
||||
## Introduction
|
||||
|
||||
Hunyuan-A13B-Instruct is a fine-grained hybrid expert model (MoE) developed by Tencent. This model has a total of 80 billion parameters, 13 billion activation parameters, supports 256k ultra-long contexts, and possesses native thought chain (CoT) reasoning capabilities.
|
||||
|
||||
## Environment Preparation
|
||||
|
||||
### Model Weight
|
||||
|
||||
- `Hunyuan-A13B-Instruct`(BF16 version): [Download model weight](https://www.modelscope.cn/models/Tencent-Hunyuan/Hunyuan-A13B-Instruct).
|
||||
|
||||
It is recommended to download the model weight to the shared directory of multiple nodes, such as `/root/.cache/`
|
||||
|
||||
### Installation
|
||||
|
||||
Run docker container:
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# Update the vllm-ascend image
|
||||
# For Atlas A2 machines:
|
||||
# export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
# For Atlas A3 machines:
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-a3
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-p 8000:8000 \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
Build from source:
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# Install vLLM.
|
||||
git clone --depth 1 --branch |vllm_version| https://github.com/vllm-project/vllm
|
||||
cd vllm
|
||||
VLLM_TARGET_DEVICE=empty pip install -e .
|
||||
cd ..
|
||||
|
||||
# Install vLLM Ascend.
|
||||
git clone --depth 1 --branch |vllm_ascend_version| https://github.com/vllm-project/vllm-ascend.git
|
||||
cd vllm-ascend
|
||||
git submodule update --init --recursive
|
||||
pip install -e .
|
||||
cd ..
|
||||
```
|
||||
|
||||
### Software Stack Version Verification
|
||||
<!-- TODO: update to Python 3.12 after verification -->
|
||||
The environment is based on CANN built into the GiteeAI platform, and successfully runs vLLM |vllm_ascend_version|, and vLLM-Ascend:|vllm_ascend_version| through the Python 3.11.6 Conda environment.
|
||||
|
||||
## Deployment
|
||||
|
||||
### Single-node Deployment (4-NPU)
|
||||
|
||||
```bash
|
||||
export HCCL_INTRA_ROCE_ENABLE=1
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3
|
||||
export HF_HOME=/data
|
||||
export MODEL_PATH="Hunyuan-A13B-Instruct"
|
||||
|
||||
vllm serve ${MODEL_PATH} \
|
||||
--trust-remote-code \
|
||||
--host 0.0.0.0 \
|
||||
--port 8000 \
|
||||
--served-model-name Hunyuan \
|
||||
--tensor-parallel-size 4 \
|
||||
--max-model-len 32768 \
|
||||
--gpu-memory-utilization 0.90 \
|
||||
```
|
||||
|
||||
### Key Performance Indicators
|
||||
|
||||
Based on verified CANN 8.5.1 test logs:
|
||||
|
||||
- Memory usage for weights: each NPU has a static memory usage of approximately 37.46 GB.
|
||||
- Graph compilation (ACL Graph): with PIECEWISE mode enabled, the system automatically captures the graph in approximately 18 seconds, which can significantly accelerate subsequent inference.
|
||||
- KV cache capacity: the remaining NPU memory can provide concurrent cache space for approximately 529,152 tokens.
|
||||
|
||||
## Functional Verification
|
||||
|
||||
```bash
|
||||
curl http://localhost:8000/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "Hunyuan",
|
||||
"messages": [{"role": "user", "content": "Give me a short introduction to large language models."}],
|
||||
"max_tokens": 100,
|
||||
"temperature": 0.7
|
||||
}'
|
||||
```
|
||||
|
||||
Expected output:
|
||||
|
||||
```json
|
||||
{"id":"chatcmpl-9a60df2b23bb539f","object":"chat.completion","created":1774751760,"model":"Hunyuan","choices":[{"index":0,"message":{"role":"assistant","content":"<think>\nOkay, I need to write a short introduction to large language models. Let me start by recalling what I know. First, what are LLMs? They're machine learning models trained on vast amounts of text data. The key here is \"large\"—so they have a huge number of parameters. Maybe mention the scale, like billions or trillions of parameters.\n\nThen, how are they trained? They're trained on diverse text sources—books, websites, articles, etc. The","refusal":null,"annotations":null,"audio":null,"function_call":null,"tool_calls":[],"reasoning":null},"logprobs":null,"finish_reason":"length","stop_reason":null,"token_ids":null}],"service_tier":null,"system_fingerprint":null,"usage":{"prompt_tokens":12,"total_tokens":112,"completion_tokens":100,"prompt_tokens_details":null},"prompt_logprobs":null,"prompt_token_ids":null,"kv_transfer_params":null}
|
||||
```
|
||||
|
||||
## Accuracy Evaluation
|
||||
|
||||
On the GiteeAI platform, the model was tested and verified using the AISBench tool on the GSM8K benchmark set: Under the 7cd45e version configuration, the model achieved an accuracy of 94.77% in the accuracy generation mode.
|
||||
|
||||
```bash
|
||||
ais_bench --models vllm_api_general_chat --datasets gsm8k_gen_0_shot_cot_chat_prompt --summarizer example --debug
|
||||
```
|
||||
|
||||
output:
|
||||
|
||||
```bash
|
||||
03/29 03:20:03 - AISBench - INFO - Running 1-th replica of evaluation
|
||||
03/29 03:20:03 - AISBench - INFO - Task [vllm-api-general-chat/gsm8k]: {'accuracy': 94.76876421531463}
|
||||
03/29 03:20:03 - AISBench - INFO - time elapsed: 2.15s
|
||||
03/29 03:20:04 - AISBench - INFO - Evaluation tasks completed.
|
||||
03/29 03:20:04 - AISBench - INFO - Summarizing evaluation results...
|
||||
dataset version metric mode vllm-api-general-chat
|
||||
--------- --------- -------- ------ -----------------------
|
||||
gsm8k 7cd45e accuracy gen 94.77
|
||||
03/29 03:20:04 - AISBench - INFO - write summary to /data/outputs/default/20260329_025345/summary/summary_20260329_025345.txt
|
||||
03/29 03:20:04 - AISBench - INFO - write csv to /data/outputs/default/20260329_025345/summary/summary_20260329_025345.csv
|
||||
```
|
||||
|
||||
The markdown formatted result is as follows:
|
||||
|
||||
| dataset | version | metric | mode | vllm-api-general-chat |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| gsm8k | 7cd45e | accuracy | gen | 94.77 |
|
||||
|
||||
## Performance
|
||||
|
||||
### Using AISBench
|
||||
|
||||
```bash
|
||||
ais_bench --models vllm_api_stream_chat --datasets demo_gsm8k_gen_4_shot_cot_chat_prompt --summarizer default_perf --mode perf
|
||||
```
|
||||
|
||||
output:
|
||||
|
||||
```bash
|
||||
[2026-04-08 05:27:40,180] [ais_bench] [INFO] Performance Results of task [vllm-api-stream-chat/demo_gsm8k]:
|
||||
╒══════════════════════════╤═════════╤═════════════════╤═════════════════╤═════════════════╤═════════════════╤═════════════════╤═════════════════╤═════════════════╤═════╕
|
||||
│ Performance Parameters │ Stage │ Average │ Min │ Max │ Median │ P75 │ P90 │ P99 │ N │
|
||||
╞══════════════════════════╪═════════╪═════════════════╪═════════════════╪═════════════════╪═════════════════╪═════════════════╪═════════════════╪═════════════════╪═════╡
|
||||
│ E2EL │ total │ 29982.6 ms │ 16472.9 ms │ 41147.2 ms │ 30919.1 ms │ 33514.9 ms │ 39413.8 ms │ 40973.9 ms │ 8 │
|
||||
├──────────────────────────┼─────────┼─────────────────┼─────────────────┼─────────────────┼─────────────────┼─────────────────┼─────────────────┼─────────────────┼─────┤
|
||||
│ TTFT │ total │ 238.6 ms │ 107.9 ms │ 276.7 ms │ 254.0 ms │ 265.6 ms │ 272.4 ms │ 276.3 ms │ 8 │
|
||||
├──────────────────────────┼─────────┼─────────────────┼─────────────────┼─────────────────┼─────────────────┼─────────────────┼─────────────────┼─────────────────┼─────┤
|
||||
│ TPOT │ total │ 60.1 ms │ 57.7 ms │ 61.3 ms │ 60.4 ms │ 60.8 ms │ 61.2 ms │ 61.3 ms │ 8 │
|
||||
├──────────────────────────┼─────────┼─────────────────┼─────────────────┼─────────────────┼─────────────────┼─────────────────┼─────────────────┼─────────────────┼─────┤
|
||||
│ ITL │ total │ 59.7 ms │ 0.0 ms │ 219.7 ms │ 51.7 ms │ 64.1 ms │ 81.9 ms │ 146.2 ms │ 8 │
|
||||
├──────────────────────────┼─────────┼─────────────────┼─────────────────┼─────────────────┼─────────────────┼─────────────────┼─────────────────┼─────────────────┼─────┤
|
||||
│ InputTokens │ total │ 1457.5 │ 1426.0 │ 1511.0 │ 1456.5 │ 1465.25 │ 1481.6 │ 1508.06 │ 8 │
|
||||
├──────────────────────────┼─────────┼─────────────────┼─────────────────┼─────────────────┼─────────────────┼─────────────────┼─────────────────┼─────────────────┼─────┤
|
||||
│ OutputTokens │ total │ 497.5 │ 268.0 │ 710.0 │ 508.5 │ 555.75 │ 666.6 │ 705.66 │ 8 │
|
||||
├──────────────────────────┼─────────┼─────────────────┼─────────────────┼─────────────────┼─────────────────┼─────────────────┼─────────────────┼─────────────────┼─────┤
|
||||
│ OutputTokenThroughput │ total │ 16.5261 token/s │ 16.2402 token/s │ 17.2551 token/s │ 16.4461 token/s │ 16.5728 token/s │ 16.9063 token/s │ 17.2202 token/s │ 8 │
|
||||
╘══════════════════════════╧═════════╧═════════════════╧═════════════════╧═════════════════╧═════════════════╧═════════════════╧═════════════════╧═════════════════╧═════╛
|
||||
╒══════════════════════════╤═════════╤═══════════════════╕
|
||||
│ Common Metric │ Stage │ Value │
|
||||
╞══════════════════════════╪═════════╪═══════════════════╡
|
||||
│ Benchmark Duration │ total │ 41161.2934 ms │
|
||||
├──────────────────────────┼─────────┼───────────────────┤
|
||||
│ Total Requests │ total │ 8 │
|
||||
├──────────────────────────┼─────────┼───────────────────┤
|
||||
│ Failed Requests │ total │ 0 │
|
||||
├──────────────────────────┼─────────┼───────────────────┤
|
||||
│ Success Requests │ total │ 8 │
|
||||
├──────────────────────────┼─────────┼───────────────────┤
|
||||
│ Concurrency │ total │ 5.8273 │
|
||||
├──────────────────────────┼─────────┼───────────────────┤
|
||||
│ Max Concurrency │ total │ 16 │
|
||||
├──────────────────────────┼─────────┼───────────────────┤
|
||||
│ Request Throughput │ total │ 0.1944 req/s │
|
||||
├──────────────────────────┼─────────┼───────────────────┤
|
||||
│ Total Input Tokens │ total │ 11660 │
|
||||
├──────────────────────────┼─────────┼───────────────────┤
|
||||
│ Prefill Token Throughput │ total │ 6108.0184 token/s │
|
||||
├──────────────────────────┼─────────┼───────────────────┤
|
||||
│ Total Generated Tokens │ total │ 3980 │
|
||||
├──────────────────────────┼─────────┼───────────────────┤
|
||||
│ Input Token Throughput │ total │ 283.2758 token/s │
|
||||
├──────────────────────────┼─────────┼───────────────────┤
|
||||
│ Output Token Throughput │ total │ 96.6928 token/s │
|
||||
├──────────────────────────┼─────────┼───────────────────┤
|
||||
│ Total Token Throughput │ total │ 379.9686 token/s │
|
||||
╘══════════════════════════╧═════════╧═══════════════════╛
|
||||
```
|
||||
|
||||
### Using vLLM Benchmark
|
||||
|
||||
Run performance evaluation of `Hunyuan-A13B-Instruct` as an example.
|
||||
|
||||
Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/) for more details.
|
||||
|
||||
There are three `vllm bench` subcommands:
|
||||
|
||||
- `latency`: Benchmark the latency of a single batch of requests.
|
||||
- `serve`: Benchmark the online serving throughput.
|
||||
- `throughput`: Benchmark offline inference throughput.
|
||||
|
||||
Take the `serve` as an example. Run the code as follows.
|
||||
|
||||
```shell
|
||||
vllm bench serve \
|
||||
--model ./Hunyuan-A13B-Instruct/ \
|
||||
--port 8000 \
|
||||
--dataset-name random \
|
||||
--random-input 200 \
|
||||
--num-prompts 200 \
|
||||
--request-rate 1 \
|
||||
--save-result \
|
||||
--result-dir ./perf_results/ \
|
||||
--trust-remote-code
|
||||
```
|
||||
|
||||
After about several minutes, you can get the performance evaluation result.
|
||||
203
docs/source/tutorials/models/Hy3-preview.md
Normal file
203
docs/source/tutorials/models/Hy3-preview.md
Normal file
@@ -0,0 +1,203 @@
|
||||
# Hy3-preview
|
||||
|
||||
## Introduction
|
||||
|
||||
Hy3-preview is a Mixture-of-Experts model, with 295B total parameters, 21B active parameters and 3.8B MTP layer parameters, developed by the Tencent Hy Team. It is the first model trained on Tencent rebuilt infrastructure. It improves significantly on complex reasoning, instruction following, context learning, coding, and agent tasks.
|
||||
|
||||
This guide records the verified vLLM Ascend serving path for Hy3-preview on one Atlas A3 16-NPU node. The verified default path is TP16 + EP + MTP + ACLGraph.
|
||||
|
||||
## Supported Features
|
||||
|
||||
Refer to [supported features](../../user_guide/support_matrix/supported_models.md) to get the model's supported feature matrix.
|
||||
|
||||
Refer to [feature guide](../../user_guide/feature_guide/index.md) to get the feature's configuration.
|
||||
|
||||
## Environment Preparation
|
||||
|
||||
### Model Weight
|
||||
|
||||
- Hugging Face: [tencent/Hy3-preview](https://huggingface.co/tencent/Hy3-preview)
|
||||
- ModelScope: [Tencent-Hunyuan/Hy3-preview](https://www.modelscope.cn/models/Tencent-Hunyuan/Hy3-preview)
|
||||
- GitCode: [tencent_hunyuan/Hy3-preview](https://ai.gitcode.com/tencent_hunyuan/Hy3-preview)
|
||||
|
||||
Download or mount the checkpoint to a path shared by the runtime container, for example `/models/Hy3-preview`.
|
||||
|
||||
### Hardware
|
||||
|
||||
The verified configuration uses one Atlas A3 node with 16 NPUs and 64 GB HBM per NPU. The real-weight run used about 58 GB process memory per NPU after startup.
|
||||
|
||||
### Installation
|
||||
|
||||
You can use our official docker image to run Hy3-preview directly. For Atlas A3 machines, select the image variant with the `-a3` suffix. The official image already includes the vLLM and vLLM Ascend runtime needed for the verified serving path.
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-a3
|
||||
export NAME=vllm-ascend
|
||||
|
||||
docker run --rm \
|
||||
--name $NAME \
|
||||
--net=host \
|
||||
--shm-size=1g \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci8 \
|
||||
--device /dev/davinci9 \
|
||||
--device /dev/davinci10 \
|
||||
--device /dev/davinci11 \
|
||||
--device /dev/davinci12 \
|
||||
--device /dev/davinci13 \
|
||||
--device /dev/davinci14 \
|
||||
--device /dev/davinci15 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /models:/models \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
In addition, if you don't want to use the docker image as above, you can also build all from source:
|
||||
|
||||
- Install `vllm-ascend` from source, refer to [installation](https://github.com/vllm-project/vllm-ascend/blob/main/docs/source/installation.md).
|
||||
|
||||
## Deployment
|
||||
|
||||
### Single-node Deployment
|
||||
|
||||
Run `vllm serve` from `/workspace`.
|
||||
|
||||
```bash
|
||||
cd /workspace
|
||||
export MODEL_PATH=/models/Hy3-preview
|
||||
|
||||
HCCL_OP_EXPANSION_MODE=AIV \
|
||||
vllm serve ${MODEL_PATH} \
|
||||
--served-model-name hy3-preview \
|
||||
--tensor-parallel-size 16 \
|
||||
--speculative-config.method mtp \
|
||||
--speculative-config.num_speculative_tokens 1 \
|
||||
--enable-expert-parallel \
|
||||
--enable-ep-weight-filter \
|
||||
--tool-call-parser hy_v3 \
|
||||
--reasoning-parser hy_v3 \
|
||||
--enable-auto-tool-choice \
|
||||
--max-model-len 32768 \
|
||||
--max-num-seqs 8 \
|
||||
--host 0.0.0.0 \
|
||||
--port 8000
|
||||
```
|
||||
|
||||
- `--enable-ep-weight-filter` recommended. It skips expert weights that do not belong to the local EP rank during loading, reducing disk and host-memory pressure for very large MoE checkpoints. We recommend keeping it enabled.
|
||||
- Tool calling and reasoning are service interfaces declared in the Hy3 README, so it is recommended to pass the corresponding `hy_v3` parsers by default.
|
||||
|
||||
## Functional Verification
|
||||
|
||||
Check model readiness first:
|
||||
|
||||
```bash
|
||||
curl -sf http://127.0.0.1:8000/v1/models
|
||||
```
|
||||
|
||||
Run a text smoke request:
|
||||
|
||||
```bash
|
||||
curl -sS http://127.0.0.1:8000/v1/chat/completions \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{
|
||||
"model": "hy3-preview",
|
||||
"messages": [{"role": "user", "content": "Say hi in one word."}],
|
||||
"max_tokens": 16,
|
||||
"temperature": 0,
|
||||
"top_p": 1,
|
||||
"chat_template_kwargs": {"reasoning_effort": "no_think"}
|
||||
}'
|
||||
```
|
||||
|
||||
Expected result:
|
||||
|
||||
```json
|
||||
{
|
||||
"model": "hy3-preview",
|
||||
"choices": [
|
||||
{
|
||||
"message": {
|
||||
"role": "assistant",
|
||||
"content": "Hi"
|
||||
},
|
||||
"finish_reason": "stop"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
## Accuracy Evaluation
|
||||
|
||||
Here are two accuracy evaluation methods.
|
||||
|
||||
### Using AISBench
|
||||
|
||||
1. Refer to [Using AISBench](../../developer_guide/evaluation/using_ais_bench.md) for details.
|
||||
|
||||
2. After execution, you can get the result, here is the result for reference only.
|
||||
|
||||
| dataset | version | metric | mode | vllm-api-general-chat | note |
|
||||
|----- | ----- | ----- | ----- | -----| ----- |
|
||||
| GSM8K | - | accuracy | gen | 93.07 | 1 Atlas A3 (64G × 16) |
|
||||
| C-Eval | - | accuracy | gen | 87.64 | 1 Atlas A3 (64G × 16) |
|
||||
|
||||
### Using Language Model Evaluation Harness
|
||||
|
||||
Not tested yet.
|
||||
|
||||
## Performance Evaluation
|
||||
|
||||
### Using AISBench
|
||||
|
||||
Refer to [Using AISBench for performance evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation) for details.
|
||||
|
||||
### Using vLLM Benchmark
|
||||
|
||||
Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/) for more details.
|
||||
|
||||
### Lightweight Online Benchmark
|
||||
|
||||
The following numbers are from a real-weight smoke benchmark on one Atlas A3 16-NPU node, using TP16 + EP + MTP + ACLGraph. The benchmark used `vllm bench serve`, random prompts, output length 128, `--max-concurrency 1`, `--temperature 0`, and 4 requests per input length. These numbers are functional performance evidence, not tuned throughput limits.
|
||||
|
||||
```bash
|
||||
vllm bench serve \
|
||||
--backend openai-chat \
|
||||
--base-url http://127.0.0.1:8000 \
|
||||
--endpoint /v1/chat/completions \
|
||||
--model /models/Hy3-preview \
|
||||
--served-model-name hy3-preview \
|
||||
--dataset-name random \
|
||||
--random-input-len 1024 \
|
||||
--random-output-len 128 \
|
||||
--num-prompts 4 \
|
||||
--request-rate inf \
|
||||
--max-concurrency 1 \
|
||||
--temperature 0 \
|
||||
--top-p 1
|
||||
```
|
||||
|
||||
| Random input length | Success / total | Mean TTFT (ms) | Mean TPOT (ms) | Output throughput (tok/s) | Total token throughput (tok/s) |
|
||||
| --- | --- | ---: | ---: | ---: | ---: |
|
||||
| 1,024 | 4 / 4 | 484.64 | 30.10 | 29.71 | 270.90 |
|
||||
| 4,096 | 4 / 4 | 1379.41 | 30.24 | 24.52 | 811.99 |
|
||||
| 16,384 | 4 / 4 | 2604.58 | 30.43 | 19.79 | 2554.65 |
|
||||
|
||||
## Known Limitations
|
||||
|
||||
- The model config supports 262,144 tokens, but this guide only verifies 32,768-token serving. Larger contexts require a separate capacity validation.
|
||||
- Formal AISBench accuracy results are pending and should be added only after a real benchmark run.
|
||||
257
docs/source/tutorials/models/InternVL3.5.md
Normal file
257
docs/source/tutorials/models/InternVL3.5.md
Normal file
@@ -0,0 +1,257 @@
|
||||
# InternVL3.5(InternVL3_5-38B/241B-A28B)
|
||||
|
||||
## 1 Introduction
|
||||
|
||||
[InternVL3.5](https://huggingface.co/papers/2508.18265), a new family of open-source multimodal models that significantly advances versatility, reasoning capability, and inference efficiency along the InternVL series.
|
||||
|
||||
The `InternVL3.5` model is first supported in `vllm-ascend:v0.20.2`
|
||||
|
||||
This document will show the main verification steps of both `InternVL3_5-38B` and `InternVL3_5-241B-A28B` model, including supported features, feature configuration, environment preparation, single-node and multi-node deployment, accuracy and performance evaluation.
|
||||
|
||||
## 2 Supported Features
|
||||
|
||||
Refer to [supported features](../../user_guide/support_matrix/supported_models.md) to get the model's supported feature matrix.
|
||||
|
||||
Refer to [feature guide](../../user_guide/feature_guide/index.md) to get the feature's configuration.
|
||||
|
||||
## 3 Prerequisites
|
||||
|
||||
### 3.1 Model Weight
|
||||
|
||||
require 1 Atlas 800 A3 (64G × 16) node:
|
||||
|
||||
- `InternVL3_5-38B-w8a8`: requires 1 Atlas 800 A3 (64GB × 16) node [Download model weight](https://modelscope.cn/models/Eco-Tech/InternVL3_5-38B)
|
||||
- `InternVL3_5-241B-A28B-w8a8`: requires 1 Atlas 800 A3 (64GB × 16) node [Download model weight](https://huggingface.co/OpenGVLab/InternVL3_5-241B-A28B)
|
||||
|
||||
## 4 Installation
|
||||
|
||||
### 4.1 Docker Image Installation
|
||||
|
||||
You can use our official docker image to run InternVL3_5 directly.
|
||||
|
||||
``` bash
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:{{ vllm_ascend_version }}-a3
|
||||
export NAME=vllm-ascend
|
||||
|
||||
# Run the container using the defined variables
|
||||
# Note: If you are running bridge network with docker, please expose available ports for multiple nodes communication in advance
|
||||
docker run --rm \
|
||||
--name $NAME \
|
||||
--net=host \
|
||||
--shm-size=1g \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci8 \
|
||||
--device /dev/davinci9 \
|
||||
--device /dev/davinci10 \
|
||||
--device /dev/davinci11 \
|
||||
--device /dev/davinci12 \
|
||||
--device /dev/davinci13 \
|
||||
--device /dev/davinci14 \
|
||||
--device /dev/davinci15 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
To verify the successful installation of the environment, please refer to [installation](../../installation.md).
|
||||
|
||||
### 4.2 Source Code Installation
|
||||
|
||||
In addition, if you don't want to use the docker image as above, you can also build all from source:
|
||||
|
||||
- Install `vllm-ascend` from source, refer to [installation](../../installation.md).
|
||||
|
||||
If you want to deploy multi-node environment, you need to set up environment on each node.
|
||||
|
||||
## 5 Online Service Deployment
|
||||
|
||||
### 5.1 Single-Node Online Deployment
|
||||
|
||||
=== "InternVL3_5-38B"
|
||||
|
||||
- Quantized model `InternVL3_5-38B-w8a8` can be deployed on 1 Atlas 800 A3 (64G × 16) .
|
||||
|
||||
Run the following script to execute online inference.
|
||||
|
||||
Common Issues Tip: If you encounter issues, Refer to [FAQs](../../faqs.md).
|
||||
|
||||
```bash
|
||||
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
|
||||
sysctl -w vm.swappiness=0
|
||||
sysctl -w kernel.numa_balancing=0
|
||||
sysctl -w kernel.sched_migration_cost_ns=50000
|
||||
|
||||
export VLLM_ASCEND_ENABLE_FLASHCOMM1=1
|
||||
export VLLM_ASCEND_ENABLE_FUSED_MC2=1
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export VLLM_USE_V1=1
|
||||
export VLLM_TORCH_PROFILER_WITH_STACK=0
|
||||
export HCCL_BUFFSIZE=1536
|
||||
|
||||
vllm serve /root/.cache/modelscope/hub/models/vllm-ascend/InternVL3_5-38B-w8a8/ \
|
||||
--port 2002 \
|
||||
--served-model-name internvl3_5 \
|
||||
--trust-remote-code \
|
||||
--async-scheduling \
|
||||
--max-model-len 40960 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--tensor-parallel-size 4 \
|
||||
--max-num-seqs 32 \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--async-scheduling \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY", "cudagraph_capture_sizes":[4,32,64,128,192,256,512]}' \
|
||||
--additional-config '{"enable_weight_nz_layout": true, "enable_cpu_binding": true}' \
|
||||
--mm-processor-cache-gb 0 \
|
||||
--enable-chunked-prefill \
|
||||
--safetensors-load-strategy 'prefetch' \
|
||||
--allowed-local-media-path "/"
|
||||
|
||||
```
|
||||
|
||||
=== "InternVL3_5-241B-A28B"
|
||||
|
||||
- Quantized model `InternVL3_5-241B-A28B-w8a8` can be deployed on 1 Atlas 800 A3 (64G × 16) .
|
||||
|
||||
Run the following script to execute online inference.
|
||||
|
||||
Common Issues Tip: If you encounter issues, Refer to [FAQs](../../faqs.md).
|
||||
|
||||
```bash
|
||||
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
|
||||
sysctl -w vm.swappiness=0
|
||||
sysctl -w kernel.numa_balancing=0
|
||||
sysctl -w kernel.sched_migration_cost_ns=50000
|
||||
|
||||
export VLLM_ASCEND_ENABLE_FLASHCOMM1=1
|
||||
export VLLM_ASCEND_ENABLE_FUSED_MC2=1
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export VLLM_USE_V1=1
|
||||
export VLLM_TORCH_PROFILER_WITH_STACK=0
|
||||
export HCCL_BUFFSIZE=1536
|
||||
|
||||
vllm serve /root/.cache/modelscope/hub/models/vllm-ascend/InternVL3_5-241B-A28B-w8a8/ \
|
||||
--port 2001 \
|
||||
--served-model-name internvl3_5 \
|
||||
--trust-remote-code \
|
||||
--async-scheduling \
|
||||
--max-model-len 40960 \
|
||||
--max-num-batched-tokens 4096 \
|
||||
--tensor-parallel-size 4 \
|
||||
--data-parallel-size 2 \
|
||||
--max-num-seqs 70 \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--async-scheduling \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"enable_weight_nz_layout": true, "enable_cpu_binding": true}' \
|
||||
--mm-processor-cache-gb 0 \
|
||||
--enable-chunked-prefill \
|
||||
--enable-expert-parallel \
|
||||
--safetensors-load-strategy 'prefetch' \
|
||||
--allowed-local-media-path "/"
|
||||
```
|
||||
|
||||
**Notice:**
|
||||
|
||||
Some configurations for optimization are shown below:
|
||||
|
||||
- `VLLM_ASCEND_ENABLE_FLASHCOMM1`: Enable FlashComm optimization to reduce communication and computation overhead on prefill node. With FlashComm enabled, layer_sharding list cannot include o_proj as an element.
|
||||
- `VLLM_ASCEND_ENABLE_FUSED_MC2`: Enable the dispatch_ffn_combine/mega_moe fused operator.
|
||||
- The above parameters are validated in a specific test environment for reference only. Please adjust `--max-model-len`, `--max-num-seqs`, `--max-num-batched-tokens`, and `--gpu-memory-utilization` based on your actual input/output length, concurrency, and hardware configuration.
|
||||
- For Ascend-specific options passed through `--additional-config`, refer to [Additional Configuration](../../user_guide/configuration/additional_config.md). For Ascend-specific environment variables, refer to [Environment Variables](../../user_guide/configuration/env_vars.md).
|
||||
|
||||
### 5.2 Multi-Node PD Separation Deployment
|
||||
|
||||
Not support yet.
|
||||
|
||||
## 6 Functional Verification
|
||||
|
||||
Once your server is started, you can query the model with input prompts:
|
||||
|
||||
```bash
|
||||
curl http://localhost:8000/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "internvl3_5",
|
||||
"messages": [
|
||||
{"role": "system", "content": "You are a helpful assistant."},
|
||||
{"role": "user", "content": [
|
||||
{"type": "image_url", "image_url": {"url": "https://modelscope.oss-cn-beijing.aliyuncs.com/resource/tiger.jpeg"}},
|
||||
{"type": "text", "text": "What is the text in the illustration?"}
|
||||
]}
|
||||
]
|
||||
}'
|
||||
```
|
||||
|
||||
Expected Result:
|
||||
|
||||
```bash
|
||||
{"id":"chatcmpl-d3270d4a16cb4b98936f71ee3016451f","object":"chat.completion","created":1764924127,"model":"internvl3_5","choices":[{"index":0,"message":{"role":"assistant","content":"The text in the illustration is: **a tiger**","refusal":null,"annotations":null,"audio":null,"function_call":null,"tool_calls":[],"reasoning_content":null},"logprobs":null,"finish_reason":"stop","stop_reason":null,"token_ids":null}],"service_tier":null,"system_fingerprint":null,"usage":{"prompt_tokens":107,"total_tokens":123,"completion_tokens":16,"prompt_tokens_details":null},"prompt_logprobs":null,"prompt_token_ids":null,"kv_transfer_params":null}
|
||||
```
|
||||
|
||||
## 7 Accuracy Evaluation
|
||||
|
||||
### 7.1 Using AISBench
|
||||
|
||||
1. Refer to [Using AISBench](../../developer_guide/evaluation/using_ais_bench.md) for details.
|
||||
|
||||
2. After execution, you can get the result.
|
||||
|
||||
## 8 Performance Evaluation
|
||||
|
||||
### 8.1 Using AISBench
|
||||
|
||||
Refer to [Using AISBench for performance evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation) for details.
|
||||
|
||||
### 8.2 Using vLLM Benchmark
|
||||
|
||||
Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/) for more details.
|
||||
|
||||
## 9 Performance Tuning
|
||||
|
||||
### 9.1 Recommended Configurations
|
||||
|
||||
#### Table 1: Scenario Overview
|
||||
|
||||
| Scenario | Deployment Mode | *Total NPUs | Weight Version | Key Considerations |
|
||||
| ---------- | ---------------- | ------------- | ---------------- | ------------------------ |
|
||||
| InternVL3_5-241B-A28B-w8a8 High Throughput | Single node deployment | 8 (A3) | InternVL3_5-241B-A28B-w8a8 | For short-sequence high throughput, try tp4dp2 |
|
||||
| InternVL3_5-38B-w8a8 High Throughput | Single node deployment | 4 (A3) | InternVL3_5-38B-w8a8 | For short-sequence high throughput, try tp4 |
|
||||
|
||||
#### Table 2: Detailed Node Configuration
|
||||
|
||||
|Scenario|Configuration|NPUs|TP|DP|Max Num Seqs|Max Num Batched Tokens|Max Model Len|
|
||||
|--------|-------------|-----|--|--|------------|----------------------|--------------|
|
||||
|Single-Node (A3)|InternVL3_5-38B-w8a8 High Throughput|2|4|1|32|16384|135000|
|
||||
|Single-Node (A3)|InternVL3_5-241B-A28B-w8a8 High Throughput|4|4|2|32|4096|40960|
|
||||
|
||||
### 9.2 Tuning Guidelines
|
||||
|
||||
Please refer to the [Public Performance Tuning Documentation](../../developer_guide/performance_and_debug/optimization_and_tuning.md) for tuning methods.
|
||||
Please refer to the [Feature Guide](../../user_guide/support_matrix/feature_matrix.md) for detailed feature descriptions.
|
||||
|
||||
## 9 FAQ
|
||||
|
||||
- Common Issues Tip: If you encounter issues, Refer to [FAQs](../../faqs.md).
|
||||
371
docs/source/tutorials/models/Kimi-K2-Thinking.md
Normal file
371
docs/source/tutorials/models/Kimi-K2-Thinking.md
Normal file
@@ -0,0 +1,371 @@
|
||||
# Kimi-K2-Thinking
|
||||
|
||||
## 1 Introduction
|
||||
|
||||
Kimi-K2-Thinking is a large-scale Mixture-of-Experts (MoE) model developed by Moonshot AI. It features a hybrid thinking architecture that excels in complex reasoning and problem-solving tasks.
|
||||
|
||||
This document will demonstrate the main verification steps and references of the model, including supported features, environment preparation, installation, online service deployment, functional verification, accuracy evaluation, performance evaluation, performance tuning, and FAQ.
|
||||
|
||||
This document is validated and written based on **vLLM-Ascend v0.9.0rc1**. The current model (Kimi-K2-Thinking) is first supported in this version, and **v0.9.0rc1 and later versions** can run stably. It is recommended to use the latest release candidate or stable version alongside this document.
|
||||
|
||||
## 2 Supported Features
|
||||
|
||||
Refer to [supported features](../../user_guide/support_matrix/supported_models.md) to get the model's supported feature matrix.
|
||||
|
||||
Refer to [feature guide](../../user_guide/feature_guide/index.md) to get the feature's configuration.
|
||||
|
||||
## 3 Prerequisites
|
||||
|
||||
### 3.1 Model Weight
|
||||
|
||||
- `Kimi-K2-Thinking` (bfloat16): requires 1 Atlas 800 A3 (64G × 16) node. [Download model weight](https://huggingface.co/moonshotai/Kimi-K2-Thinking).
|
||||
|
||||
It is recommended to download the model weight to the shared directory, such as `/mnt/sfs_turbo/.cache/`.
|
||||
|
||||
After downloading the model weights, please edit the value of `"quantization_config.config_groups.group_0.targets"` from `["Linear"]` to `["MoE"]` in `config.json` of the original model to use the quantized model.
|
||||
|
||||
```json
|
||||
{
|
||||
"quantization_config": {
|
||||
"config_groups": {
|
||||
"group_0": {
|
||||
"targets": [
|
||||
"MoE"
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Your model files should look like:
|
||||
|
||||
```bash
|
||||
.
|
||||
|-- chat_template.jinja
|
||||
|-- config.json
|
||||
|-- configuration_deepseek.py
|
||||
|-- configuration.json
|
||||
|-- generation_config.json
|
||||
|-- model-00001-of-000062.safetensors
|
||||
|-- ...
|
||||
|-- model-00062-of-000062.safetensors
|
||||
|-- model.safetensors.index.json
|
||||
|-- modeling_deepseek.py
|
||||
|-- tiktoken.model
|
||||
|-- tokenization_kimi.py
|
||||
|-- tokenizer_config.json
|
||||
```
|
||||
|
||||
## 4 Installation
|
||||
|
||||
### 4.1 Docker Image Installation
|
||||
|
||||
You can use the official Docker image to run `Kimi-K2-Thinking` directly.
|
||||
|
||||
Select an image based on your machine type and start the Docker image on your node, refer to [using docker](../../installation.md#set-up-using-docker).
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# Update the vllm-ascend image according to your environment.
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-a3
|
||||
|
||||
# Run the container using the defined variables
|
||||
# Note: If you are running bridge network with docker, please expose available ports for multiple nodes communication in advance
|
||||
docker run --rm \
|
||||
--name $NAME \
|
||||
--net=host \
|
||||
--shm-size=1g \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci8 \
|
||||
--device /dev/davinci9 \
|
||||
--device /dev/davinci10 \
|
||||
--device /dev/davinci11 \
|
||||
--device /dev/davinci12 \
|
||||
--device /dev/davinci13 \
|
||||
--device /dev/davinci14 \
|
||||
--device /dev/davinci15 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /mnt/sfs_turbo/.cache:/home/cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
**Parameter Descriptions:**
|
||||
|
||||
- `IMAGE`: specifies the `vllm-ascend` image. The `-a3` suffix selects the Atlas A3 image.
|
||||
- `NAME`: specifies the container name.
|
||||
- `--net=host`: uses host networking, so the vLLM service port is exposed on the host directly.
|
||||
- `--shm-size=1g`: configures container shared memory.
|
||||
- `--device /dev/davinci[0-15]`: exposes 16 Ascend NPU devices to the container.
|
||||
- `--device /dev/davinci_manager`, `--device /dev/devmm_svm`, and `--device /dev/hisi_hdc`: expose required Ascend runtime device files.
|
||||
- `-v /usr/local/dcmi:/usr/local/dcmi`: mounts DCMI tools for device management.
|
||||
- `-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi`: mounts the NPU monitoring command.
|
||||
- `-v /usr/local/Ascend/driver/*`: mounts Ascend driver libraries and version files.
|
||||
- `-v /etc/ascend_install.info:/etc/ascend_install.info`: mounts Ascend installation metadata.
|
||||
- `-v /mnt/sfs_turbo/.cache:/home/cache`: mounts the shared model cache directory. Update it if you store model weights elsewhere.
|
||||
|
||||
After the container starts, run the following command on the host to verify the container status:
|
||||
|
||||
```bash
|
||||
docker ps --filter name=vllm-ascend --format "table {{.Names}}\t{{.Status}}"
|
||||
```
|
||||
|
||||
Expected Status:
|
||||
|
||||
- The container name is `vllm-ascend`.
|
||||
- The status is `Up ...`.
|
||||
- The container does not exit immediately.
|
||||
|
||||
Run the following command in the container to verify that Ascend devices are visible:
|
||||
|
||||
```bash
|
||||
npu-smi info
|
||||
```
|
||||
|
||||
Expected Status:
|
||||
|
||||
- The command exits successfully.
|
||||
- The output lists the expected NPU devices.
|
||||
- Device health status is normal.
|
||||
|
||||
### 4.2 Source Code Installation
|
||||
|
||||
If you do not want to use the Docker image, you can also build from source:
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# Install vLLM.
|
||||
git clone --depth 1 --branch |vllm_version| https://github.com/vllm-project/vllm
|
||||
cd vllm
|
||||
VLLM_TARGET_DEVICE=empty pip install -e .
|
||||
cd ..
|
||||
|
||||
# Install vLLM Ascend.
|
||||
git clone --depth 1 --branch |vllm_ascend_version| https://github.com/vllm-project/vllm-ascend.git
|
||||
cd vllm-ascend
|
||||
pip install -e .
|
||||
```
|
||||
|
||||
To verify the source installation, run:
|
||||
|
||||
```bash
|
||||
python -c "import vllm; import vllm_ascend; print('vllm and vllm_ascend import ok')"
|
||||
```
|
||||
|
||||
Expected Status:
|
||||
|
||||
- The command exits successfully.
|
||||
- `vllm and vllm_ascend import ok` is printed.
|
||||
|
||||
## 5 Online Service Deployment
|
||||
|
||||
### 5.1 Single-Node Online Deployment
|
||||
|
||||
Single-node deployment completes both Prefill and Decode within the same node, suitable for online inference scenarios with moderate concurrency requirements.
|
||||
|
||||
For an Atlas 800 A3 (64G × 16) node, `tensor-parallel-size` should be at least 16.
|
||||
|
||||
Run the following script to start the vLLM server:
|
||||
|
||||
```{model-code}
|
||||
:block_name: kimi_k2_thinking_single_node
|
||||
:converter_tag: single_node
|
||||
:test_case_path: tests/e2e/nightly/single_node/models/configs/Kimi-K2-Thinking.yaml
|
||||
```
|
||||
|
||||
**Parameter and Environment Variable Descriptions:**
|
||||
|
||||
The following table covers the generated `model`, all `envs`, and all `server_cmd` entries. Parameters are categorized by review priority: version-sensitive parameters, performance parameters, and Kimi-K2-Thinking-specific parameters.
|
||||
|
||||
| Parameter | Validated Value | Category | Description and Tuning Guidance |
|
||||
| --- | --- | --- | --- |
|
||||
| `moonshotai/Kimi-K2-Thinking` | model path | Model-specific | Specifies the model weight path passed to `vllm serve`. Because `--served-model-name` is not set in the script, API requests must use `moonshotai/Kimi-K2-Thinking` as the model name unless you add an explicit served-model-name override; see the FAQ in Chapter 10. |
|
||||
| `HCCL_BUFFSIZE` | `1024` | Performance | Configures the HCCL communication buffer used by distributed NPU communication. This document validates `1024`; other values need separate throughput, TTFT, TPOT, and HCCL stability validation. |
|
||||
| `TASK_QUEUE_ENABLE` | `1` | Version-sensitive / Performance | Enables task queue scheduling on Ascend. This document validates `1`; other values or version changes need startup and first-request validation. |
|
||||
| `OMP_PROC_BIND` | `false` | Performance | Avoids overly strict OpenMP CPU binding. This document validates `false`; other values need separate CPU affinity, NPU health, and HCCL stability validation. |
|
||||
| `HCCL_OP_EXPANSION_MODE` | `AIV` | Performance | Enables the AIV communication path. This document validates `AIV`; other values need separate throughput and latency validation. |
|
||||
| `PYTORCH_NPU_ALLOC_CONF` | `expandable_segments:True` | Memory / Performance | Reduces NPU memory fragmentation. This document validates `expandable_segments:True`; other allocator settings need separate startup, memory, and runtime stability validation. |
|
||||
| `SERVER_PORT` and `--port` | `8000` | Service | Sets the OpenAI-compatible service port. The documentation generator maps `DEFAULT_PORT` in the YAML to `8000`; update the curl examples if you change this value. |
|
||||
| `--tensor-parallel-size` | `16` | Model-specific / Performance | Uses all 16 NPUs on one Atlas 800 A3 node. This document validates `tp16`; other topologies need separate memory, accuracy, and communication validation. |
|
||||
| `--max-model-len` | `8192` | Performance | Sets the maximum input plus output tokens for one request and determines KV cache reservation. This document validates `8192`; larger values need separate NPU memory, accuracy, and performance validation. Keep it close to the real maximum input and output length for your workload. |
|
||||
| `--max-num-batched-tokens` | `8192` | Performance | Limits tokens processed in one scheduler step. This document validates `8192`; other values need separate memory, TTFT, TPOT, and throughput validation. |
|
||||
| `--max-num-seqs` | `12` | Performance | Limits active sequences scheduled at the same time. This document validates `12`; higher values need separate tail-latency and throughput validation. The reference sweep in Chapter 8 shows concurrency 16 causes severe TTFT growth; validate tail latency before raising it in production. |
|
||||
| `--gpu-memory-utilization` | `0.9` | Memory / Performance | Controls the fraction of NPU HBM used by vLLM for KV cache planning. This document validates `0.9`; other values need separate startup, OOM, and runtime stability validation. |
|
||||
| `--trust-remote-code` | enabled | Model-specific | Required because the model package contains model-specific configuration, modeling, tokenizer, and chat-template files. Disable it only after replacing the remote-code dependency with a validated native implementation. |
|
||||
| `--enable-expert-parallel` | enabled | Model-specific / Performance | Enables expert parallelism for Kimi-K2-Thinking MoE layers so experts can be distributed across NPUs. This document validates it as enabled; disabling it is not validated in this tutorial. |
|
||||
| `--no-enable-prefix-caching` | enabled | Performance | Disables prefix caching for the validated baseline and random-prompt benchmarks. Prefix caching is not validated in this tutorial. |
|
||||
|
||||
**Common Issues Tip:** For common environment, installation, and general parameter issues during deployment, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html). If the service runs under high concurrency, verify NPU health and HCCL status before increasing the request rate.
|
||||
|
||||
**Service Verification:**
|
||||
|
||||
After the service starts, you should see logs similar to:
|
||||
|
||||
```bash
|
||||
INFO: Started server process [...]
|
||||
INFO: Waiting for application startup.
|
||||
INFO: Application startup complete.
|
||||
```
|
||||
|
||||
Expected Status:
|
||||
|
||||
- The server process starts successfully.
|
||||
- No error logs related to HCCL or NPU initialization.
|
||||
- The container does not exit immediately.
|
||||
|
||||
## 6 Functional Verification
|
||||
|
||||
After the service is started, the model can be invoked by sending a prompt:
|
||||
|
||||
```bash
|
||||
curl http://localhost:8000/v1/chat/completions -H "Content-Type: application/json" -d '{
|
||||
"model": "moonshotai/Kimi-K2-Thinking",
|
||||
"messages": [
|
||||
{"role": "user", "content": "Who are you?"}
|
||||
],
|
||||
"temperature": 1.0
|
||||
}'
|
||||
```
|
||||
|
||||
Expected Result:
|
||||
|
||||
- The HTTP status code is `200`.
|
||||
- `choices[0].message.content` contains the generated assistant response.
|
||||
|
||||
## 7 Accuracy Evaluation
|
||||
|
||||
### Using AISBench
|
||||
|
||||
For details, please refer to [Using AISBench](../../developer_guide/evaluation/using_ais_bench.md).
|
||||
|
||||
### Using lm-eval
|
||||
|
||||
You can use [lm-eval](https://github.com/EleutherAI/lm-evaluation-harness) to evaluate the model accuracy through the OpenAI-compatible API.
|
||||
|
||||
For `lm_eval` installation, please refer to [Using lm_eval](../../developer_guide/evaluation/using_lm_eval.md).
|
||||
|
||||
Run `lm_eval` to execute the accuracy evaluation:
|
||||
|
||||
```shell
|
||||
lm_eval \
|
||||
--model local-completions \
|
||||
--model_args model=moonshotai/Kimi-K2-Thinking,base_url=http://127.0.0.1:8000/v1/completions,tokenized_requests=False,trust_remote_code=True \
|
||||
--tasks gsm8k \
|
||||
--output_path ./
|
||||
```
|
||||
|
||||
Reference configuration: `gsm8k` (5-shot), `--apply_chat_template`, `--fewshot_as_multiturn`, greedy decoding (`temperature=0.0`, `top_p=1.0`), max 2048 output tokens, batch size 1.
|
||||
|
||||
Below are reference `gsm8k` results for `Kimi-K2-Thinking` powered by `vllm-ascend:v0.20.2rc1`, evaluated on one Atlas 800 A3 node (64G × 16).
|
||||
|
||||
| task | version | filter | n-shot | metric | value | stderr |
|
||||
| --- | ---: | --- | ---: | --- | ---: | ---: |
|
||||
| `gsm8k` | 3 | `flexible-extract` | 5 | `exact_match` | 0.8992 | 0.0083 |
|
||||
| `gsm8k` | 3 | `strict-match` | 5 | `exact_match` | 0.8453 | 0.0100 |
|
||||
|
||||
## 8 Performance Evaluation
|
||||
|
||||
Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/) for more details.
|
||||
|
||||
**Test Command Example:**
|
||||
|
||||
```bash
|
||||
vllm bench serve \
|
||||
--backend openai-chat \
|
||||
--model moonshotai/Kimi-K2-Thinking \
|
||||
--endpoint /v1/chat/completions \
|
||||
--dataset-name random \
|
||||
--random-input-len 1024 \
|
||||
--random-output-len 1024 \
|
||||
--num-prompts 10 \
|
||||
--request-rate 1
|
||||
```
|
||||
|
||||
After the benchmark completes, you can get the performance result, including request throughput, output token throughput, TTFT, TPOT, and ITL.
|
||||
|
||||
The following reference results are obtained with `vllm-ascend:v0.20.2rc1` on one Atlas 800 A3 node (64G × 16), using OpenAI chat serving, random input/output lengths, 10 prompts, and `--request-rate 1`:
|
||||
|
||||
| random input len | random output len | success | duration (s) | request throughput (req/s) | output throughput (tok/s) | total throughput (tok/s) | mean TTFT (ms) | mean TPOT (ms) | mean ITL (ms) |
|
||||
| ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
|
||||
| 512 | 512 | 10 / 10 | 111.00 | 0.09 | 46.12 | 94.38 | 507.60 | 200.47 | 200.08 |
|
||||
| 1024 | 1024 | 10 / 10 | 221.52 | 0.05 | 46.23 | 93.48 | 566.39 | 208.20 | 208.00 |
|
||||
| 2048 | 2048 | 10 / 10 | 479.72 | 0.02 | 42.69 | 85.78 | 722.32 | 230.26 | 230.15 |
|
||||
|
||||
For a concurrency sweep, keep the input and output length fixed and vary `--max-concurrency`:
|
||||
|
||||
```bash
|
||||
MODEL_NAME=moonshotai/Kimi-K2-Thinking
|
||||
INPUT_LEN=1024
|
||||
OUTPUT_LEN=1024
|
||||
|
||||
for CONCURRENCY in 1 2 4 8 16 32; do
|
||||
NUM_PROMPTS=$((CONCURRENCY * 10))
|
||||
vllm bench serve \
|
||||
--backend openai-chat \
|
||||
--model "$MODEL_NAME" \
|
||||
--endpoint /v1/chat/completions \
|
||||
--dataset-name random \
|
||||
--random-input-len "$INPUT_LEN" \
|
||||
--random-output-len "$OUTPUT_LEN" \
|
||||
--num-prompts "$NUM_PROMPTS" \
|
||||
--request-rate inf \
|
||||
--max-concurrency "$CONCURRENCY"
|
||||
done
|
||||
```
|
||||
|
||||
Reference results for 1024 input tokens and 1024 output tokens are:
|
||||
|
||||
| max concurrency | prompts | success | duration (s) | request throughput (req/s) | output throughput (tok/s) | total throughput (tok/s) | mean TTFT (ms) | P99 TTFT (ms) | mean TPOT (ms) |
|
||||
| ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
|
||||
| 1 | 10 | 10 / 10 | 595.07 | 0.02 | 17.21 | 34.80 | 473.71 | 712.49 | 57.71 |
|
||||
| 2 | 20 | 20 / 20 | 623.88 | 0.03 | 32.83 | 66.35 | 708.16 | 996.59 | 60.29 |
|
||||
| 4 | 40 | 40 / 40 | 725.38 | 0.06 | 56.47 | 114.13 | 956.11 | 1137.55 | 69.97 |
|
||||
| 8 | 80 | 80 / 80 | 907.44 | 0.09 | 90.28 | 182.43 | 1361.85 | 1900.15 | 87.37 |
|
||||
| 16 | 160 | 160 / 160 | 3093.07 | 0.05 | 52.97 | 107.04 | 76766.84 | 251245.22 | 222.07 |
|
||||
|
||||
> **Note:** At concurrency levels of 16, the Mean TTFT increases significantly (76.7s), indicating severe queueing delay. For production deployment, it is recommended to limit concurrency based on your latency requirements or increase `--max-num-seqs` and `--max-num-batched-tokens` if NPU memory allows.
|
||||
|
||||
## 9 Performance Tuning
|
||||
|
||||
### 9.1 Recommended Configurations
|
||||
|
||||
> **Note:** The following configurations are validated in specific test environments and are for reference only. The optimal configuration depends on factors such as maximum input/output length, prefix cache hit rate, precision requirements, and deployment machine ratios. It is recommended to refer to Section 9.2 for tuning based on actual conditions.
|
||||
|
||||
#### Table 1: Scenario Overview
|
||||
|
||||
| Scenario | Deployment Mode | Total NPUs | Weight Version | Key Considerations |
|
||||
|----------|----------------|------------|----------------|---------------------|
|
||||
| Long Context | Single-node | 16 (A3) | bfloat16 | Keep `--max-model-len` close to the real maximum input and output length, and reduce `--max-num-seqs` first when memory pressure is high. The validated scope of this single-node baseline covers up to 2K input / 2K output in Chapter 8. |
|
||||
| Low Latency | Single-node | 16 (A3) | bfloat16 | Reduce `--max-num-seqs` and `--max-num-batched-tokens` from the validated baseline (`12` and `8192`) to reduce queueing delay. In the Chapter 8 concurrency sweep, concurrency 1-4 kept mean TTFT below 1s; validate TTFT, TPOT, and tail latency against the target latency SLO. |
|
||||
| High Throughput | Single-node | 16 (A3) | bfloat16 | Increase `--max-num-seqs` gradually and benchmark with a request rate close to the real workload. In the 1K/1K concurrency sweep in Chapter 8, concurrency 8 gave the best output throughput; validate tail latency before using higher concurrency in production. |
|
||||
|
||||
### 9.2 Tuning Guidelines
|
||||
|
||||
#### 9.2.1 General Tuning Reference
|
||||
|
||||
Please refer to the [Public Performance Tuning Documentation](../../developer_guide/performance_and_debug/optimization_and_tuning.md) for general tuning methods.
|
||||
|
||||
Please refer to the [Feature Guide](../../user_guide/support_matrix/feature_matrix.md) for detailed feature descriptions.
|
||||
|
||||
## 10 FAQ
|
||||
|
||||
> For common environment, installation, and general parameter issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html); this chapter only covers model-specific issues.
|
||||
|
||||
- **Q: API returns `{"error":"Model not found"}` or `404` when requesting with `model: "Kimi-K2-Thinking"`?**
|
||||
|
||||
A: The server registers the model under its full path `moonshotai/Kimi-K2-Thinking` by default. When the request uses the short name `Kimi-K2-Thinking` without `--served-model-name` override, the server cannot resolve the model ID. Use `"model": "moonshotai/Kimi-K2-Thinking"` in requests, or start the server with `--served-model-name Kimi-K2-Thinking` to enable the short name.
|
||||
1062
docs/source/tutorials/models/Kimi-K2.5.md
Normal file
1062
docs/source/tutorials/models/Kimi-K2.5.md
Normal file
File diff suppressed because it is too large
Load Diff
874
docs/source/tutorials/models/Kimi-K2.6.md
Normal file
874
docs/source/tutorials/models/Kimi-K2.6.md
Normal file
@@ -0,0 +1,874 @@
|
||||
# Kimi-K2.6
|
||||
|
||||
## 1 Introduction
|
||||
|
||||
Kimi K2.6 is an open-source, native multimodal agentic model built through continual pretraining on approximately 15 trillion mixed visual and text tokens atop Kimi-K2-Base. It seamlessly integrates vision and language understanding with advanced agentic capabilities, instant and thinking modes, as well as conversational and agentic paradigms.
|
||||
|
||||
This document will show the main verification steps of the model, including supported features, feature configuration, environment preparation, single-node and multi-node deployment, accuracy and performance evaluation.
|
||||
|
||||
This document is validated and written based on **vLLM-Ascend v0.20.0rc1**. The current model (Kimi-K2.6) is first supported in this version.
|
||||
|
||||
## 2 Supported Features
|
||||
|
||||
Refer to [supported features](../../user_guide/support_matrix/supported_models.md) to get the model's supported feature matrix.
|
||||
|
||||
Refer to [feature guide](../../user_guide/feature_guide/index.md) to get the feature's configuration.
|
||||
|
||||
## 3 Prerequisites
|
||||
|
||||
### 3.1 Model Weight
|
||||
|
||||
- `Kimi-K2.6-w4a8` (Quantized version for w4a8): requires 1 Atlas 800 A3 (64G × 16) node or 2 Atlas 800 A2 (64G × 8) nodes. [Download model weight](https://www.modelscope.cn/models/Eco-Tech/Kimi-K2.6-W4A8).
|
||||
- `kimi-k2.6-eagle3` (Eagle3 MTP draft model for accelerating inference of Kimi-K2.6): [Download model weight](https://huggingface.co/lightseekorg/kimi-k2.6-eagle3)
|
||||
- `Kimi-K2.5-DFlash` (a speculative decoding framework that leverages a lightweight block diffusion model for parallel drafting): [Download model weight](https://huggingface.co/z-lab/Kimi-K2.5-DFlash)
|
||||
|
||||
It is recommended to download the model weight to the shared directory of multiple nodes, such as `/root/.cache/`.
|
||||
|
||||
### 3.2 Verify Multi-node Communication (Optional)
|
||||
|
||||
If you want to deploy multi-node environment, you need to verify multi-node communication according to [verify multi-node communication environment](../../installation.md#verify-multi-node-communication).
|
||||
|
||||
## 4 Installation
|
||||
|
||||
### 4.1 Docker Image Installation
|
||||
|
||||
Select an image based on your machine type and start the docker image on your node, refer to [using docker](../../installation.md#set-up-using-docker).
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: install
|
||||
|
||||
::::{tab-item} A3 series
|
||||
:sync: A3
|
||||
|
||||
Start the docker image on each node.
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-a3
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--privileged=true \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci8 \
|
||||
--device /dev/davinci9 \
|
||||
--device /dev/davinci10 \
|
||||
--device /dev/davinci11 \
|
||||
--device /dev/davinci12 \
|
||||
--device /dev/davinci13 \
|
||||
--device /dev/davinci14 \
|
||||
--device /dev/davinci15 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} A2 series
|
||||
:sync: A2
|
||||
|
||||
Start the docker image on each node.
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--privileged=true \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
After a successful docker run, you can verify the running container service by executing the `docker ps` command.
|
||||
|
||||
### 4.2 Source Code Installation
|
||||
|
||||
If you don't want to use the docker image as above, you can also build all from source:
|
||||
|
||||
- Install `vllm-ascend` from source, refer to [installation](../../installation.md).
|
||||
|
||||
If you want to deploy multi-node environment, you need to set up environment on each node.
|
||||
|
||||
To use the tool_calls feature, please ensure that your transformers version is 4.57.6 or lower. If vllm-ascend has been upgraded to v0.21 or later, this requirement no longer applies.
|
||||
|
||||
## 5 Online Service Deployment
|
||||
|
||||
### 5.1 Single-Node Online Deployment
|
||||
|
||||
Single-node deployment completes both Prefill and Decode within the same node. The quantized model `Kimi-K2.6-w4a8` can be deployed on 1 Atlas 800 A3 (64G × 16).
|
||||
|
||||
While a single-node setup supports all input/output scenarios, consider deploying multinodes for optimal performance.
|
||||
|
||||
Startup Command:
|
||||
|
||||
```bash
|
||||
#!/bin/sh
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export VLLM_ASCEND_ENABLE_MLAPO=1
|
||||
|
||||
# [Optional] jemalloc
|
||||
# jemalloc is for better performance, if `libjemalloc.so` is installed on your machine, you can turn it on.
|
||||
export LD_PRELOAD=/usr/lib/aarch64-linux-gnu/libjemalloc.so.2:$LD_PRELOAD
|
||||
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
|
||||
sysctl -w vm.swappiness=0
|
||||
sysctl -w kernel.numa_balancing=0
|
||||
sysctl -w kernel.sched_migration_cost_ns=50000
|
||||
|
||||
export HCCL_BUFFSIZE=800
|
||||
export VLLM_ASCEND_ENABLE_FLASHCOMM1=1
|
||||
export VLLM_ASCEND_BALANCE_SCHEDULING=1
|
||||
|
||||
vllm serve Eco-Tech/Kimi-K2.6-W4A8 \
|
||||
--quantization ascend \
|
||||
--served-model-name kimi_k26 \
|
||||
--allowed-local-media-path / \
|
||||
--trust-remote-code \
|
||||
--tensor-parallel-size 4 \
|
||||
--data-parallel-size 4 \
|
||||
--no-enable-prefix-caching \
|
||||
--enable-expert-parallel \
|
||||
--port 8088 \
|
||||
--max-num-seqs 4 \
|
||||
--max-model-len 32768 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--seed 42 \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \
|
||||
--mm-processor-cache-gb 0 \
|
||||
--mm-encoder-tp-mode data \
|
||||
--speculative-config '{"method": "dflash","model": "z-lab/Kimi-K2.5-DFlash", "num_speculative_tokens": 15}'
|
||||
```
|
||||
|
||||
Key Parameter Descriptions:
|
||||
|
||||
- Setting the environment variable `VLLM_ASCEND_BALANCE_SCHEDULING=1` enables balance scheduling. This may help increase output throughput and reduce TPOT in v1 scheduler. However, TTFT may degrade in some scenarios. Furthermore, enabling this feature is not recommended in scenarios where PD is separated.
|
||||
- `--max-model-len` specifies the maximum context length - that is, the sum of input and output tokens for a single request. For performance testing with an input length of 3.5K and output length of 1.5K, a value of `16384` is sufficient, however, for precision testing, please set it at least `35000`.
|
||||
- `--no-enable-prefix-caching` indicates that prefix caching is disabled. To enable it, remove this option.
|
||||
- `--mm-encoder-tp-mode` indicates how to optimize multi-modal encoder inference using tensor parallelism (TP). If you want to test the multimodal inputs, we recommend using `data`.
|
||||
- If you use the w4a8 weight, more memory will be allocated to kvcache, and you can try increasing `--max-num-seqs` to improve system throughput.
|
||||
|
||||
Common Issues Tip: If you encounter issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html) for troubleshooting.
|
||||
|
||||
Service Verification:
|
||||
|
||||
```shell
|
||||
curl http://<node_ip>:8088/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "kimi_k26",
|
||||
"messages": [{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "The future of AI is"
|
||||
}]
|
||||
}],
|
||||
"max_tokens": 1024,
|
||||
"temperature": 1.0,
|
||||
"top_p": 0.95
|
||||
}'
|
||||
```
|
||||
|
||||
Expected Result:
|
||||
|
||||
The service returns HTTP 200 OK with a JSON response containing the `choices` field. Example output (content truncated for brevity):
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "chatcmpl-9df13fd5e539af93",
|
||||
"object": "chat.completion",
|
||||
"created": 1780971952,
|
||||
"model": "kimi_k26",
|
||||
"choices": [
|
||||
{
|
||||
"index": 0,
|
||||
"message": {
|
||||
"role": "assistant",
|
||||
"content": "The future of AI is not a destination we are passively approaching, but a design problem we are actively solving right now...",
|
||||
"reasoning": "The user is asking for my thoughts on \"The future of AI is\"...",
|
||||
"refusal": null,
|
||||
"annotations": null,
|
||||
"audio": null,
|
||||
"function_call": null
|
||||
},
|
||||
"logprobs": null,
|
||||
"finish_reason": "length",
|
||||
"stop_reason": null,
|
||||
"token_ids": null
|
||||
}
|
||||
],
|
||||
"usage": {
|
||||
"prompt_tokens": 13,
|
||||
"total_tokens": 1037,
|
||||
"completion_tokens": 1024,
|
||||
"completion_tokens_details": {
|
||||
"reasoning_tokens": 0,
|
||||
"audio_tokens": null,
|
||||
"accepted_prediction_tokens": null,
|
||||
"rejected_prediction_tokens": null
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### 5.2 Multi-Node PD Separation Deployment
|
||||
|
||||
We recommend using Mooncake for deployment: [Mooncake](../features/pd_disaggregation_mooncake_multi_node.md).
|
||||
|
||||
In the standard single-node deployment mode, Prefill (prompt processing) and Decode (token generation) tasks run on the same set of NPUs. This can lead to two issues:
|
||||
|
||||
1. **Prefill preemption interrupts Decode**: Prefill is a compute-intensive task that processes the entire input context at once, while Decode generates tokens one by one. When a new user request arrives, its Prefill phase can preempt and interrupt ongoing Decode tasks, causing jitter and higher time-per-output-token (TPOT) latency.
|
||||
2. **Inflexible resource allocation**: Prefill and Decode have fundamentally different computational characteristics — Prefill is compute-bound and memory-bandwidth-intensive, while Decode is memory-bandwidth-bound. Running them on the same hardware forces a compromise that satisfies neither optimally.
|
||||
|
||||
PD (Prefill-Decode) separation addresses these issues by running Prefill and Decode on dedicated node groups, each configured independently:
|
||||
|
||||
- **Prefill nodes** focus on high-throughput prompt processing, optimized for compute and communication (e.g., enabling FlashComm for Allreduce acceleration).
|
||||
|
||||
- **Decode nodes** focus on low-latency token generation, optimized for memory bandwidth (e.g., enabling MLAPO fusion operators).
|
||||
|
||||
This architecture is recommended for production deployments with concurrent multi-user workloads, where stable latency and high throughput are both required.
|
||||
|
||||
Take Atlas 800 A3 (64G × 16) for example, we recommend to deploy 2P1D (4 nodes) rather than 1P1D (2 nodes), because there is not enough NPU memory to serve high concurrency in 1P1D case.
|
||||
|
||||
- `Kimi-K2.6-w4a8 2P1D`: requires 4 Atlas 800 A3 (64G × 16) nodes.
|
||||
|
||||
To run the vllm-ascend `Prefill-Decode Disaggregation` service, you need to deploy a `launch_online_dp.py` script and a `run_dp_template.sh` script on each node and deploy a `proxy.sh` script on prefill master node to forward requests.
|
||||
|
||||
[launch_online_dp.py](https://github.com/vllm-project/vllm-ascend/blob/main/examples/external_online_dp/launch_online_dp.py)
|
||||
|
||||
Parameter descriptions:
|
||||
|
||||
|Parameter|Type|Required|Default|Description|
|
||||
|---------|----|--------|-------|-----------|
|
||||
|`--dp-size`|int|Yes|-|Data parallel size (total number of DP ranks across all nodes).|
|
||||
|`--tp-size`|int|No|1|Tensor parallel size within each DP rank.|
|
||||
|`--dp-size-local`|int|No|(same as `--dp-size`)|Number of DP ranks on the current node. If not set, defaults to `--dp-size`.|
|
||||
|`--dp-rank-start`|int|No|0|Starting rank offset for data parallel ranks on this node.|
|
||||
|`--dp-address`|str|Yes|-|IP address of the data parallel master node (node 0).|
|
||||
|`--dp-rpc-port`|str|No|12345|RPC port for data parallel master communication.|
|
||||
|`--vllm-start-port`|int|No|9000|Starting port for each vLLM engine instance on this node. Each DP rank's engine port = `vllm_start_port` + local rank index.|
|
||||
|
||||
1. `run_dp_template.sh` script
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: script
|
||||
|
||||
::::{tab-item} Node 0(Prefill)
|
||||
:sync: Node 0(Prefill)
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# this obtained through ifconfig
|
||||
# nic_name is the network interface name corresponding to local_ip of the current node
|
||||
nic_name="xxx"
|
||||
local_ip="141.xx.xx.1"
|
||||
|
||||
# The value of node0_ip must be consistent with the value of local_ip set in node0 (master node)
|
||||
node0_ip="xxxx"
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
|
||||
# [Optional] jemalloc
|
||||
# jemalloc is for better performance, if `libjemalloc.so` is installed on your machine, you can turn it on.
|
||||
export LD_PRELOAD=/usr/lib/aarch64-linux-gnu/libjemalloc.so.2:$LD_PRELOAD
|
||||
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
|
||||
sysctl -w vm.swappiness=0
|
||||
sysctl -w kernel.numa_balancing=0
|
||||
sysctl kernel.sched_migration_cost_ns=50000
|
||||
export VLLM_RPC_TIMEOUT=3600000
|
||||
export VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=30000
|
||||
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export LD_LIBRARY_PATH=/usr/local/Ascend/ascend-toolkit/latest/python/site-packages/mooncake:$LD_LIBRARY_PATH
|
||||
|
||||
export HCCL_BUFFSIZE=800
|
||||
export VLLM_ASCEND_ENABLE_FLASHCOMM1=1
|
||||
export ASCEND_RT_VISIBLE_DEVICES=$1
|
||||
|
||||
vllm serve Eco-Tech/Kimi-K2.6-W4A8 \
|
||||
--host 0.0.0.0 \
|
||||
--port $2 \
|
||||
--data-parallel-size $3 \
|
||||
--data-parallel-rank $4 \
|
||||
--data-parallel-address $5 \
|
||||
--data-parallel-rpc-port $6 \
|
||||
--tensor-parallel-size $7 \
|
||||
--enable-expert-parallel \
|
||||
--seed 1024 \
|
||||
--quantization ascend \
|
||||
--served-model-name kimi_k26 \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 4 \
|
||||
--max-model-len 32768 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--no-enable-prefix-caching \
|
||||
--gpu-memory-utilization 0.95 \
|
||||
--enforce-eager \
|
||||
--speculative-config '{"method": "eagle3", "model":"lightseekorg/kimi-k2.6-eagle3", "num_speculative_tokens": 1}' \
|
||||
--mm-encoder-tp-mode data \
|
||||
--kv-transfer-config \
|
||||
'{"kv_connector": "MooncakeConnectorV1",
|
||||
"kv_role": "kv_producer",
|
||||
"kv_port": "30000",
|
||||
"engine_id": "0",
|
||||
"kv_connector_extra_config": {
|
||||
"prefill": {
|
||||
"dp_size": 4,
|
||||
"tp_size": 4
|
||||
},
|
||||
"decode": {
|
||||
"dp_size": 8,
|
||||
"tp_size": 4
|
||||
}
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} Node 1(Prefill)
|
||||
:sync: Node 1(Prefill)
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# this obtained through ifconfig
|
||||
# nic_name is the network interface name corresponding to local_ip of the current node
|
||||
nic_name="xxx"
|
||||
local_ip="141.xx.xx.2"
|
||||
|
||||
# The value of node0_ip must be consistent with the value of local_ip set in node0 (master node)
|
||||
node0_ip="xxxx"
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
|
||||
# [Optional] jemalloc
|
||||
# jemalloc is for better performance, if `libjemalloc.so` is installed on your machine, you can turn it on.
|
||||
export LD_PRELOAD=/usr/lib/aarch64-linux-gnu/libjemalloc.so.2:$LD_PRELOAD
|
||||
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
|
||||
sysctl -w vm.swappiness=0
|
||||
sysctl -w kernel.numa_balancing=0
|
||||
sysctl kernel.sched_migration_cost_ns=50000
|
||||
export VLLM_RPC_TIMEOUT=3600000
|
||||
export VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=30000
|
||||
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export LD_LIBRARY_PATH=/usr/local/Ascend/ascend-toolkit/latest/python/site-packages/mooncake:$LD_LIBRARY_PATH
|
||||
|
||||
export HCCL_BUFFSIZE=800
|
||||
export VLLM_ASCEND_ENABLE_FLASHCOMM1=1
|
||||
export ASCEND_RT_VISIBLE_DEVICES=$1
|
||||
|
||||
vllm serve Eco-Tech/Kimi-K2.6-W4A8 \
|
||||
--host 0.0.0.0 \
|
||||
--port $2 \
|
||||
--data-parallel-size $3 \
|
||||
--data-parallel-rank $4 \
|
||||
--data-parallel-address $5 \
|
||||
--data-parallel-rpc-port $6 \
|
||||
--tensor-parallel-size $7 \
|
||||
--enable-expert-parallel \
|
||||
--seed 1024 \
|
||||
--quantization ascend \
|
||||
--served-model-name kimi_k26 \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 4 \
|
||||
--max-model-len 32768 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--no-enable-prefix-caching \
|
||||
--gpu-memory-utilization 0.95 \
|
||||
--enforce-eager \
|
||||
--speculative-config '{"method": "eagle3", "model":"lightseekorg/kimi-k2.6-eagle3", "num_speculative_tokens": 1}' \
|
||||
--mm-encoder-tp-mode data \
|
||||
--kv-transfer-config \
|
||||
'{"kv_connector": "MooncakeConnectorV1",
|
||||
"kv_role": "kv_producer",
|
||||
"kv_port": "30100",
|
||||
"engine_id": "1",
|
||||
"kv_connector_extra_config": {
|
||||
"prefill": {
|
||||
"dp_size": 4,
|
||||
"tp_size": 4
|
||||
},
|
||||
"decode": {
|
||||
"dp_size": 8,
|
||||
"tp_size": 4
|
||||
}
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} Node 0(Decode)
|
||||
:sync: Node 0(Decode)
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# this obtained through ifconfig
|
||||
# nic_name is the network interface name corresponding to local_ip of the current node
|
||||
nic_name="xxx"
|
||||
local_ip="141.xx.xx.3"
|
||||
|
||||
# The value of node0_ip must be consistent with the value of local_ip set in node0 (master node)
|
||||
node0_ip="xxxx"
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
|
||||
# [Optional] jemalloc
|
||||
# jemalloc is for better performance, if `libjemalloc.so` is installed on your machine, you can turn it on.
|
||||
export LD_PRELOAD=/usr/lib/aarch64-linux-gnu/libjemalloc.so.2:$LD_PRELOAD
|
||||
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
|
||||
sysctl -w vm.swappiness=0
|
||||
sysctl -w kernel.numa_balancing=0
|
||||
sysctl kernel.sched_migration_cost_ns=50000
|
||||
export VLLM_RPC_TIMEOUT=3600000
|
||||
export VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=30000
|
||||
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export LD_LIBRARY_PATH=/usr/local/Ascend/ascend-toolkit/latest/python/site-packages/mooncake:$LD_LIBRARY_PATH
|
||||
|
||||
export HCCL_BUFFSIZE=800
|
||||
export VLLM_ASCEND_ENABLE_MLAPO=1
|
||||
export ASCEND_RT_VISIBLE_DEVICES=$1
|
||||
|
||||
vllm serve Eco-Tech/Kimi-K2.6-W4A8 \
|
||||
--host 0.0.0.0 \
|
||||
--port $2 \
|
||||
--data-parallel-size $3 \
|
||||
--data-parallel-rank $4 \
|
||||
--data-parallel-address $5 \
|
||||
--data-parallel-rpc-port $6 \
|
||||
--tensor-parallel-size $7 \
|
||||
--enable-expert-parallel \
|
||||
--seed 1024 \
|
||||
--quantization ascend \
|
||||
--served-model-name kimi_k26 \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 8 \
|
||||
--max-model-len 32768 \
|
||||
--max-num-batched-tokens 32 \
|
||||
--no-enable-prefix-caching \
|
||||
--gpu-memory-utilization 0.91 \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"recompute_scheduler_enable":true,"multistream_overlap_shared_expert": false}' \
|
||||
--speculative-config '{"method": "eagle3", "model":"lightseekorg/kimi-k2.6-eagle3", "num_speculative_tokens": 3}' \
|
||||
--kv-transfer-config \
|
||||
'{"kv_connector": "MooncakeConnectorV1",
|
||||
"kv_role": "kv_consumer",
|
||||
"kv_port": "30200",
|
||||
"engine_id": "2",
|
||||
"kv_connector_extra_config": {
|
||||
"prefill": {
|
||||
"dp_size": 4,
|
||||
"tp_size": 4
|
||||
},
|
||||
"decode": {
|
||||
"dp_size": 8,
|
||||
"tp_size": 4
|
||||
}
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} Node 1(Decode)
|
||||
:sync: Node 1(Decode)
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# this obtained through ifconfig
|
||||
# nic_name is the network interface name corresponding to local_ip of the current node
|
||||
nic_name="xxx"
|
||||
local_ip="141.xx.xx.4"
|
||||
|
||||
# The value of node0_ip must be consistent with the value of local_ip set in node0 (master node)
|
||||
node0_ip="xxxx"
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
|
||||
# [Optional] jemalloc
|
||||
# jemalloc is for better performance, if `libjemalloc.so` is installed on your machine, you can turn it on.
|
||||
export LD_PRELOAD=/usr/lib/aarch64-linux-gnu/libjemalloc.so.2:$LD_PRELOAD
|
||||
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
|
||||
sysctl -w vm.swappiness=0
|
||||
sysctl -w kernel.numa_balancing=0
|
||||
sysctl kernel.sched_migration_cost_ns=50000
|
||||
export VLLM_RPC_TIMEOUT=3600000
|
||||
export VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=30000
|
||||
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export LD_LIBRARY_PATH=/usr/local/Ascend/ascend-toolkit/latest/python/site-packages/mooncake:$LD_LIBRARY_PATH
|
||||
|
||||
export HCCL_BUFFSIZE=1100
|
||||
export VLLM_ASCEND_ENABLE_MLAPO=1
|
||||
export ASCEND_RT_VISIBLE_DEVICES=$1
|
||||
|
||||
vllm serve Eco-Tech/Kimi-K2.6-W4A8 \
|
||||
--host 0.0.0.0 \
|
||||
--port $2 \
|
||||
--data-parallel-size $3 \
|
||||
--data-parallel-rank $4 \
|
||||
--data-parallel-address $5 \
|
||||
--data-parallel-rpc-port $6 \
|
||||
--tensor-parallel-size $7 \
|
||||
--enable-expert-parallel \
|
||||
--seed 1024 \
|
||||
--quantization ascend \
|
||||
--served-model-name kimi_k26 \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 8 \
|
||||
--max-model-len 32768 \
|
||||
--max-num-batched-tokens 4 \
|
||||
--no-enable-prefix-caching \
|
||||
--gpu-memory-utilization 0.91 \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"recompute_scheduler_enable":true,"multistream_overlap_shared_expert": false}' \
|
||||
--speculative-config '{"method": "eagle3", "model":"lightseekorg/kimi-k2.6-eagle3", "num_speculative_tokens": 3}' \
|
||||
--kv-transfer-config \
|
||||
'{"kv_connector": "MooncakeConnectorV1",
|
||||
"kv_role": "kv_consumer",
|
||||
"kv_port": "30300",
|
||||
"engine_id": "3",
|
||||
"kv_connector_extra_config": {
|
||||
"prefill": {
|
||||
"dp_size": 4,
|
||||
"tp_size": 4
|
||||
},
|
||||
"decode": {
|
||||
"dp_size": 8,
|
||||
"tp_size": 4
|
||||
}
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
Key Parameter Descriptions:
|
||||
|
||||
- `VLLM_ASCEND_ENABLE_FLASHCOMM1=1`: enables the communication optimization function on the prefill nodes.
|
||||
- `VLLM_ASCEND_ENABLE_MLAPO=1`: enables the fusion operator, which can significantly improve performance but consumes more NPU memory. In the Prefill-Decode (PD) separation scenario, enable MLAPO only on decode nodes.
|
||||
- `recompute_scheduler_enable: true`: enables the recomputation scheduler. When the Key-Value Cache (KV Cache) of the decode node is insufficient, requests will be sent to the prefill node to recompute the KV Cache. In the PD separation scenario, enable this configuration only on decode nodes.
|
||||
- `multistream_overlap_shared_expert: true`: When the Tensor Parallelism (TP) size is 1 or `enable_shared_expert_dp: true`, an additional stream is enabled to overlap the computation process of shared experts for improved efficiency.
|
||||
|
||||
2. Run server for each node:
|
||||
|
||||
```shell
|
||||
# p0
|
||||
python launch_online_dp.py --dp-size 4 --tp-size 4 --dp-size-local 4 --dp-rank-start 0 --dp-address 141.xx.xx.1 --dp-rpc-port 12321 --vllm-start-port 7100
|
||||
# p1
|
||||
python launch_online_dp.py --dp-size 4 --tp-size 4 --dp-size-local 4 --dp-rank-start 0 --dp-address 141.xx.xx.2 --dp-rpc-port 12321 --vllm-start-port 7100
|
||||
# d0
|
||||
python launch_online_dp.py --dp-size 8 --tp-size 4 --dp-size-local 8 --dp-rank-start 0 --dp-address 141.xx.xx.3 --dp-rpc-port 12321 --vllm-start-port 7100
|
||||
# d1
|
||||
python launch_online_dp.py --dp-size 8 --tp-size 4 --dp-size-local 8 --dp-rank-start 8 --dp-address 141.xx.xx.3 --dp-rpc-port 12321 --vllm-start-port 7100
|
||||
```
|
||||
|
||||
3. Run the `proxy.sh` script on the prefill master node
|
||||
|
||||
Run a proxy server on the same node with the prefiller service instance. You can get the proxy program in the repository's examples: [load_balance_proxy_server_example.py](https://github.com/vllm-project/vllm-ascend/blob/main/examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py)
|
||||
|
||||
```shell
|
||||
python load_balance_proxy_server_example.py \
|
||||
--port 1999 \
|
||||
--host 141.xx.xx.1 \
|
||||
--prefiller-hosts \
|
||||
141.xx.xx.1 \
|
||||
141.xx.xx.1 \
|
||||
141.xx.xx.1 \
|
||||
141.xx.xx.1 \
|
||||
141.xx.xx.2 \
|
||||
141.xx.xx.2 \
|
||||
141.xx.xx.2 \
|
||||
141.xx.xx.2 \
|
||||
--prefiller-ports \
|
||||
7100 7101 7102 7103 7100 7101 7102 7103 \
|
||||
--decoder-hosts \
|
||||
141.xx.xx.3 \
|
||||
141.xx.xx.3 \
|
||||
141.xx.xx.3 \
|
||||
141.xx.xx.3 \
|
||||
141.xx.xx.4 \
|
||||
141.xx.xx.4 \
|
||||
141.xx.xx.4 \
|
||||
141.xx.xx.4 \
|
||||
--decoder-ports \
|
||||
7100 7101 7102 7103 \
|
||||
7100 7101 7102 7103 \
|
||||
```
|
||||
|
||||
```shell
|
||||
cd vllm-ascend/examples/disaggregated_prefill_v1/
|
||||
bash proxy.sh
|
||||
```
|
||||
|
||||
Deployment Verification:
|
||||
|
||||
After the PD separation service is fully started, send a request through the proxy port on the prefill master node to verify that Prefill and Decode nodes are working correctly together:
|
||||
|
||||
```shell
|
||||
curl http://141.xx.xx.1:1999/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "kimi_k26",
|
||||
"messages": [{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "The future of AI is"
|
||||
}]
|
||||
}],
|
||||
"max_tokens": 1024,
|
||||
"temperature": 1.0,
|
||||
"top_p": 0.95
|
||||
}'
|
||||
```
|
||||
|
||||
Expected Result:
|
||||
|
||||
The proxy returns HTTP 200 OK. The JSON response contains the `choices` field with the generated text, confirming that Prefill nodes have successfully processed the prompt and Decode nodes have generated the response:
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "chatcmpl-xxxxxxxxxxxxx",
|
||||
"object": "chat.completion",
|
||||
"model": "kimi_k26",
|
||||
"choices": [
|
||||
{
|
||||
"index": 0,
|
||||
"message": {
|
||||
"role": "assistant",
|
||||
"content": "The future of AI is not a destination we are passively approaching...",
|
||||
"finish_reason": "length"
|
||||
}
|
||||
}
|
||||
],
|
||||
"usage": {
|
||||
"prompt_tokens": 13,
|
||||
"total_tokens": 1037,
|
||||
"completion_tokens": 1024
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Common Issues Tip: If you encounter issues with PD separation deployment, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html) for troubleshooting.
|
||||
|
||||
## 6 Functional Verification
|
||||
|
||||
Once your server is started, you can query the model with input prompts:
|
||||
|
||||
```shell
|
||||
curl http://<node0_ip>:<port>/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "kimi_k26",
|
||||
"messages": [{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "The future of AI is"
|
||||
}]
|
||||
}],
|
||||
"max_tokens": 1024,
|
||||
"temperature": 1.0,
|
||||
"top_p": 0.95
|
||||
}'
|
||||
```
|
||||
|
||||
Expected Result:
|
||||
|
||||
The service returns HTTP 200 OK. The JSON response contains the `choices` field with the generated text, along with usage statistics:
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "chatcmpl-9df13fd5e539af93",
|
||||
"object": "chat.completion",
|
||||
"created": 1780971952,
|
||||
"model": "kimi_k26",
|
||||
"choices": [
|
||||
{
|
||||
"index": 0,
|
||||
"message": {
|
||||
"role": "assistant",
|
||||
"content": "The future of AI is not a destination we are passively approaching, but a design problem we are actively solving right now...",
|
||||
"reasoning": "The user is asking for my thoughts on...",
|
||||
"finish_reason": "length"
|
||||
}
|
||||
}
|
||||
],
|
||||
"usage": {
|
||||
"prompt_tokens": 13,
|
||||
"total_tokens": 1037,
|
||||
"completion_tokens": 1024
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## 7 Accuracy Evaluation
|
||||
|
||||
Here is one accuracy evaluation method.
|
||||
|
||||
### Using AISBench
|
||||
|
||||
1. Refer to [Using AISBench](../../developer_guide/evaluation/using_ais_bench.md) for details.
|
||||
|
||||
2. After execution, you can get the result. Here is the result of `Kimi-K2.6-w4a8` in `vllm-ascend:v0.20.0rc1` for reference only.
|
||||
|
||||
| dataset | version | metric | mode | vllm-api-general-chat | note |
|
||||
| ----- | ----- | ----- | ----- | ----- | ----- |
|
||||
| AIME2026 | - | accuracy | gen | 90.00 | 1 Atlas 800 A3 (64G × 16) |
|
||||
| GPQA | - | accuracy | gen | 89.90 | 1 Atlas 800 A3 (64G × 16) |
|
||||
| MMMU | - | accuracy | gen | 82.67 | 1 Atlas 800 A3 (64G × 16) |
|
||||
|
||||
## 8 Performance Evaluation
|
||||
|
||||
### Using AISBench
|
||||
|
||||
Refer to [Using AISBench for performance evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation) for details.
|
||||
|
||||
### Using vLLM Benchmark
|
||||
|
||||
Run performance evaluation of `Kimi-K2.6-w4a8` as an example.
|
||||
|
||||
Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/) for more details.
|
||||
|
||||
There are three `vllm bench` subcommands:
|
||||
|
||||
- `latency`: Benchmark the latency of a single batch of requests.
|
||||
- `serve`: Benchmark the online serving throughput.
|
||||
- `throughput`: Benchmark offline inference throughput.
|
||||
|
||||
Take the `serve` as an example. Run the code as follows.
|
||||
|
||||
```shell
|
||||
vllm bench serve --model Eco-Tech/Kimi-K2.6-w4a8 --dataset-name random --random-input 1024 --num-prompts 200 --request-rate 1 --save-result --result-dir ./
|
||||
```
|
||||
|
||||
After about several minutes, you can get the performance evaluation result.
|
||||
|
||||
## 9 Performance Tuning
|
||||
|
||||
### 9.1 Recommended Configurations
|
||||
|
||||
> **Note**: The following configurations are validated in specific test environments and are for reference only. The optimal configuration depends on factors such as maximum input/output length, prefix cache hit rate, precision requirements, and deployment machine ratios. It is recommended to refer to Section 9.2 for tuning based on actual conditions.
|
||||
|
||||
#### Table 1: Scenario Overview
|
||||
|
||||
> `*Total NPUs` indicates the total number of NPUs used across all nodes. 1 node = 1 Atlas 800 A3 server (64G × 16 NPUs).
|
||||
|
||||
|Scenario|Deployment Mode|*Total NPUs|Weight Version|Key Considerations|
|
||||
|--------|---------------|-----------|--------------|------------------|
|
||||
|High Throughput<br>(16k input)|Single-Node Mixed|16 (A3)|kimi-k2.6-w4a8|Use dp2 tp8 to balance memory capacity and compute efficiency|
|
||||
|High Throughput<br>(16k input)|1P1D deployment|32 (A3)|kimi-k2.6-w4a8|dp2 tp8 on both P and D nodes; balanced latency and throughput|
|
||||
|High Throughput<br>(16k input)|2P1D deployment|64 (A3)|kimi-k2.6-w4a8|Scale from dp4 tp4 to dp8 tp4 across nodes|
|
||||
|Long Context<br>(128k, no prefix cache)|Single-Node Mixed|16 (A3)|kimi-k2.6-w4a8|dp1 tp16 to maximize TP, accommodate extreme context lengths|
|
||||
|Long Context<br>(128k, with prefix cache)|Single-Node Mixed|16 (A3)|kimi-k2.6-w4a8|dp2 tp8 to optimize memory bandwidth and improve cache utilization|
|
||||
|Multimodal<br>(1080P)|Single-Node Mixed|16 (A3)|kimi-k2.6-w4a8|dp1 tp16 for high-resolution visual inputs|
|
||||
|Multimodal<br>(1080P)|1P1D deployment|32 (A3)|kimi-k2.6-w4a8|dp2 tp8 or dp16 tp1, depending on memory and concurrency|
|
||||
|Multimodal<br>(1080P)|2P1D deployment|64 (A3)|kimi-k2.6-w4a8|dp8 tp2 to dp32 tp1, maximize throughput for heavy multimodal workloads|
|
||||
|
||||
#### Table 2: Detailed Node Configuration
|
||||
|
||||
|Scenario|Configuration|NPUs|TP|DP|Max Model Len|MTP Speculation Num|
|
||||
|--------|-------------|-----|--|--|-------------------|--------------------|
|
||||
|High Throughput / Low Latency (16k)|Server / Single Machine|16|8|2|17k|15|
|
||||
|High Throughput / Low Latency (16k)|Server-P Node|16|8|2|17k|3|
|
||||
|High Throughput / Low Latency (16k)|Server-D Node|16|8|2|17k|3|
|
||||
|Long Context (128k, no cache)|Server / Single Machine|16|16|1|130k|15|
|
||||
|Long Context (128k, with cache)|Server / Single Machine|16|8|2|130k|15|
|
||||
|Multimodal (1080P)|Server / Single Machine|16|16|1|17k|15|
|
||||
|Multimodal (1080P)|Server-P Node|16|8|2|17k|3|
|
||||
|Multimodal (1080P)|Server-D Node|16|1|16|17k|3|
|
||||
|
||||
> For complete startup commands and parameter descriptions, please refer to the deployment examples in [Chapter 5](#5-online-service-deployment).
|
||||
|
||||
**Notice:**
|
||||
`max-model-len` and `max-num-seqs` need to be set according to the actual usage scenario. For other settings, please refer to the **[Deployment](#5-online-service-deployment)** chapter.
|
||||
|
||||
### 9.2 Tuning Guidelines
|
||||
|
||||
Please refer to the [Public Performance Tuning Documentation](../../developer_guide/performance_and_debug/optimization_and_tuning.md) for tuning methods.
|
||||
|
||||
Please refer to the [Feature Guide](../../user_guide/support_matrix/feature_matrix.md) for detailed feature descriptions.
|
||||
|
||||
## 10 FAQ
|
||||
|
||||
For common environment, installation, and general parameter issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html); this chapter only covers model-specific issues.
|
||||
|
||||
- **Q: What transformer version is required for tool_calls feature?**
|
||||
|
||||
A: To use the tool_calls feature, please ensure that your transformers version is 4.57.6 or lower. If vllm-ascend has been upgraded to v0.21 or later, this requirement no longer applies.
|
||||
1063
docs/source/tutorials/models/Kimi-K3.md
Normal file
1063
docs/source/tutorials/models/Kimi-K3.md
Normal file
File diff suppressed because it is too large
Load Diff
157
docs/source/tutorials/models/LLaVA-OneVision-Qwen2-0.5B-OV.md
Normal file
157
docs/source/tutorials/models/LLaVA-OneVision-Qwen2-0.5B-OV.md
Normal file
@@ -0,0 +1,157 @@
|
||||
# LLaVA-OneVision-Qwen2-0.5B-OV
|
||||
|
||||
## Introduction
|
||||
|
||||
`llava-hf/llava-onevision-qwen2-0.5b-ov-hf` is a compact multimodal model built on top of Qwen2. It supports text-only generation together with image understanding, multi-image reasoning, and visual dialogue.
|
||||
|
||||
This document shows the main verification steps for the model on vLLM Ascend, including environment preparation, single-NPU deployment, functional verification, and the existing accuracy baseline used by the repository.
|
||||
|
||||
## Supported Features
|
||||
|
||||
Refer to [supported features](../../user_guide/support_matrix/supported_models.md) to get the model's supported feature matrix.
|
||||
|
||||
Refer to [feature guide](../../user_guide/feature_guide/index.md) to get the feature's configuration.
|
||||
|
||||
## Environment Preparation
|
||||
|
||||
### Model Weight
|
||||
|
||||
- `llava-hf/llava-onevision-qwen2-0.5b-ov-hf`: [Download model weight](https://huggingface.co/llava-hf/llava-onevision-qwen2-0.5b-ov-hf)
|
||||
|
||||
The verified single-card deployment uses one Atlas A2 NPU. It is recommended to cache model weights under `/root/.cache` in advance to reduce startup time.
|
||||
|
||||
### Installation
|
||||
|
||||
You can use the official docker image to run `LLaVA-OneVision-Qwen2-0.5B-OV` directly.
|
||||
|
||||
Select an image based on your machine type and start the docker image on your node. Refer to [using docker](../../installation.md#set-up-using-docker).
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
## Deployment
|
||||
|
||||
### Single-node Deployment
|
||||
|
||||
#### Single NPU
|
||||
|
||||
Run the following script to start the vLLM service on a single Atlas A2 NPU:
|
||||
|
||||
```bash
|
||||
export MODEL_PATH="llava-hf/llava-onevision-qwen2-0.5b-ov-hf"
|
||||
|
||||
vllm serve "${MODEL_PATH}" \
|
||||
--host 0.0.0.0 \
|
||||
--port 8000 \
|
||||
--served-model-name LLaVA-OneVision-0.5B \
|
||||
--trust-remote-code \
|
||||
--gpu-memory-utilization 0.8
|
||||
```
|
||||
|
||||
#### Multiple NPU
|
||||
|
||||
Single-NPU deployment is recommended for this 0.5B model.
|
||||
|
||||
### Prefill-Decode Disaggregation
|
||||
|
||||
Not supported yet.
|
||||
|
||||
## Functional Verification
|
||||
|
||||
If your service starts successfully, you can see logs similar to the following:
|
||||
|
||||
```bash
|
||||
INFO: Started server process [8173]
|
||||
INFO: Waiting for application startup.
|
||||
INFO: Application startup complete.
|
||||
```
|
||||
|
||||
You can first verify that the model is exposed by the OpenAI-compatible API:
|
||||
|
||||
```bash
|
||||
curl http://127.0.0.1:8000/v1/models
|
||||
```
|
||||
|
||||
### Text-only Request
|
||||
|
||||
```bash
|
||||
curl http://127.0.0.1:8000/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "LLaVA-OneVision-0.5B",
|
||||
"messages": [
|
||||
{
|
||||
"role": "user",
|
||||
"content": "Say hello in one short sentence."
|
||||
}
|
||||
],
|
||||
"max_completion_tokens": 16,
|
||||
"temperature": 0
|
||||
}'
|
||||
```
|
||||
|
||||
If the request succeeds, you can see a response similar to the following:
|
||||
|
||||
```bash
|
||||
{"choices":[{"message":{"content":"Hello! How can I assist you today?"}}]}
|
||||
```
|
||||
|
||||
### Image Understanding Request
|
||||
|
||||
```bash
|
||||
curl http://127.0.0.1:8000/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "LLaVA-OneVision-0.5B",
|
||||
"messages": [
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{
|
||||
"type": "text",
|
||||
"text": "Describe this image briefly."
|
||||
},
|
||||
{
|
||||
"type": "image_url",
|
||||
"image_url": {
|
||||
"url": "https://modelscope.oss-cn-beijing.aliyuncs.com/resource/qwen.png"
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"max_completion_tokens": 64,
|
||||
"temperature": 0
|
||||
}'
|
||||
```
|
||||
|
||||
If the request succeeds, you can see a response similar to the following:
|
||||
|
||||
```bash
|
||||
{"choices":[{"message":{"content":"The image features a logo consisting of a stylized geometric figure and the text \"TONGYI\" and \"Qwen\"..."}}]}
|
||||
```
|
||||
|
||||
## Accuracy Evaluation
|
||||
|
||||
The repository already contains an end-to-end accuracy baseline for this model in `tests/e2e/models/configs/llava-onevision-qwen2-0.5b-ov-hf.yaml`.
|
||||
|
||||
| dataset | platform | metric | value |
|
||||
|----- | ----- | ----- | ----- |
|
||||
| ceval-valid | A2 | acc,none | 0.42 |
|
||||
712
docs/source/tutorials/models/MiniMax-M2.md
Normal file
712
docs/source/tutorials/models/MiniMax-M2.md
Normal file
@@ -0,0 +1,712 @@
|
||||
# MiniMax-M2
|
||||
|
||||
## 1 Introduction
|
||||
|
||||
MiniMax-M2 is MiniMax's flagship large language model series, including **MiniMax-M2.5** and **MiniMax-M2.7**. It is reinforced for high-value scenarios such as code generation, agentic tool calling/search, and complex office workflows, with an emphasis on reasoning efficiency and end-to-end speed on challenging tasks.
|
||||
|
||||
This document will show the main verification steps for both MiniMax-M2.5 and MiniMax-M2.7, including supported features, feature configuration, environment preparation, single-node and multi-node deployment, accuracy and performance evaluation.
|
||||
|
||||
This document is written based on the latest vLLM-Ascend version. Both MiniMax-M2.5 and MiniMax-M2.7 are fully supported. To use the latest features (e.g., PD separation, EAGLE3 speculative decoding), it is recommended to use the latest version.
|
||||
|
||||
## 2 Supported Features
|
||||
|
||||
Refer to [supported features](../../user_guide/support_matrix/supported_models.md) to get the model's supported feature matrix.
|
||||
|
||||
Refer to [feature guide](../../user_guide/feature_guide/index.md) to get the feature's configuration.
|
||||
|
||||
## 3 Prerequisites
|
||||
|
||||
### 3.1 Model Weight
|
||||
|
||||
The following model weights and EAGLE3 weights are available on ModelScope. Search for the corresponding model name on [ModelScope](https://modelscope.cn) to obtain the latest weight files.
|
||||
|
||||
| Model | Description | Recommended Hardware | Source |
|
||||
|-------|-------------|---------------------|--------|
|
||||
| `MiniMax-M2.7-w8a8-QuaRot` | M2.7 W8A8 quantized version | 1× Atlas 800 A3 (64GB × 16) or 1× Atlas 800I A2 (64GB × 8) | [MiniMax-M2.7-w8a8-QuaRot](https://www.modelscope.ai/models/vllm-ascend/MiniMax-M2.7-w8a8-QuaRot) |
|
||||
| `MiniMax-M2.5-w8a8-QuaRot` | M2.5 W8A8 quantized version | 1× Atlas 800 A3 (64GB × 16) or 1× Atlas 800I A2 (64GB × 8) | [MiniMax-M2.5-w8a8-QuaRot](https://www.modelscope.cn/models/Eco-Tech/MiniMax-M2.5-w8a8-QuaRot) |
|
||||
| `MiniMax-M2.7-w8a8c8-QuaRot` | M2.7 W8A8C8 quantized version | 1× Atlas 800 A3 (64GB × 16) or 1× Atlas 800I A2 (64GB × 8) | [MiniMax-M2.7-w8a8c8-QuaRot](https://www.modelscope.ai/models/vllm-ascend/MiniMax-M2.7-w8a8c8-QuaRot) |
|
||||
| `EAGLE3` (M2.7) | M2.7 speculative decoding head model | Matches the base model node count | [MiniMax-M2.7-eagle-model](https://www.modelscope.cn/models/Eco-Tech/MiniMax-M2.7-eagle-model-short) |
|
||||
| `EAGLE3` (M2.5) | M2.5 speculative decoding head model | Matches the base model node count | [MiniMax-M2.5-eagle-model](https://www.modelscope.cn/models/vllm-ascend/MiniMax-M2.5-eagle-model-0318) |
|
||||
|
||||
It is recommended to download the model weights to a shared directory, such as `/root/.cache/`.
|
||||
|
||||
### 3.2 Verify Multi-node Communication (Optional)
|
||||
|
||||
If you need to deploy a multi-node environment, verify the multi-node communication according to [Verify Multi-node Communication Environment](../../installation.md#verify-multi-node-communication).
|
||||
|
||||
## 4 Installation
|
||||
|
||||
### 4.1 Docker Image Installation
|
||||
|
||||
You can use the official all-in-one Docker image. For the available image tags and published versions, refer to [Using Docker](../../installation.md#set-up-using-docker).
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: hardware
|
||||
|
||||
::::{tab-item} A3 series
|
||||
:sync: a3
|
||||
|
||||
**Docker Run:**
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-a3
|
||||
|
||||
docker run \
|
||||
--name vllm-ascend-env \
|
||||
--ipc host \
|
||||
--net host \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci8 \
|
||||
--device /dev/davinci9 \
|
||||
--device /dev/davinci10 \
|
||||
--device /dev/davinci11 \
|
||||
--device /dev/davinci12 \
|
||||
--device /dev/davinci13 \
|
||||
--device /dev/davinci14 \
|
||||
--device /dev/davinci15 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/Ascend/driver:/usr/local/Ascend/driver \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /usr/local/sbin:/usr/local/sbin \
|
||||
-it -d $IMAGE bash
|
||||
```
|
||||
|
||||
:::{note}
|
||||
A3 has 8 NPUs with dual-die design (16 chips total: `/dev/davinci[0-15]`).
|
||||
If you are on a shared machine, map only the chips you need (e.g., `/dev/davinci[0-7]` for NPU 0-3).
|
||||
:::
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} A2 series
|
||||
:sync: a2
|
||||
|
||||
**Docker Run:**
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
|
||||
docker run \
|
||||
--name vllm-ascend-env \
|
||||
--ipc host \
|
||||
--net host \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/Ascend/driver:/usr/local/Ascend/driver \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /usr/local/sbin:/usr/local/sbin \
|
||||
-it -d $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
:::::
|
||||
|
||||
:::{tip}
|
||||
The mounts above are the minimum required for NPU driver access. Add additional `-v` mounts (e.g., model weight paths, datasets) as needed for your environment.
|
||||
:::
|
||||
|
||||
The default workdir is `/workspace`. vLLM and vLLM-Ascend are installed as Python packages in site-packages.
|
||||
|
||||
**Installation Verification:**
|
||||
|
||||
After starting the container, run the following command to verify the installation:
|
||||
|
||||
```bash
|
||||
docker ps | grep vllm-ascend-env
|
||||
```
|
||||
|
||||
Expected result: The container is listed with status `Up`. You can also verify the vllm-ascend version inside the container:
|
||||
|
||||
```bash
|
||||
pip show vllm-ascend
|
||||
```
|
||||
|
||||
Expected result: The version information is displayed, matching the pulled image version.
|
||||
|
||||
### 4.2 Source Code Installation
|
||||
|
||||
If you prefer to build from source instead of using the Docker image, install vLLM-Ascend following the [Installation Guide](../../installation.md).
|
||||
|
||||
To verify the source installation:
|
||||
|
||||
```bash
|
||||
python -c "import vllm_ascend; print(vllm_ascend.__version__)"
|
||||
```
|
||||
|
||||
## 5 Online Service Deployment
|
||||
|
||||
:::{note}
|
||||
In this tutorial, we assume you have downloaded the model weights. Replace `/path/to/weight/` with your actual model weight path.
|
||||
:::
|
||||
|
||||
### 5.1 Single-Node Online Deployment
|
||||
|
||||
Single-node deployment completes both Prefill and Decode within the same node, suitable for development, testing, and low-to-medium throughput production scenarios.
|
||||
|
||||
**Common Issues Tip:** If you encounter OOM, HCCL port conflicts, or other startup issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html) for troubleshooting. For MiniMax-specific issues, refer to [Chapter 10 FAQ](#10-faq).
|
||||
|
||||
#### A3 (single node)
|
||||
|
||||
Below is a recommended startup configuration for short-context conditions (e.g., 3.5k input / 1.5k output) to achieve good performance.
|
||||
|
||||
Notes:
|
||||
|
||||
- If you only care about short-context low latency, you can set `--max-model-len 32768`, `--tensor-parallel-size 4`, and `--data-parallel-size 4`.
|
||||
|
||||
```{code-block} bash
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export OMP_NUM_THREADS=1
|
||||
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
|
||||
sysctl -w vm.swappiness=0
|
||||
sysctl -w kernel.numa_balancing=0
|
||||
sysctl kernel.sched_migration_cost_ns=50000
|
||||
export LD_PRELOAD=/usr/lib/aarch64-linux-gnu/libjemalloc.so.2:$LD_PRELOAD
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
|
||||
export VLLM_ASCEND_BALANCE_SCHEDULING=0
|
||||
|
||||
vllm serve /path/to/weight/MiniMax-M2.7-w8a8-QuaRot \
|
||||
--served-model-name "MiniMax-M2.7" \
|
||||
--host 0.0.0.0 \
|
||||
--port 8000 \
|
||||
--trust-remote-code \
|
||||
--quantization ascend \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
|
||||
--async-scheduling \
|
||||
--additional-config '{"enable_cpu_binding":true,
|
||||
"enable_fused_mc2":true,
|
||||
"enable_flashcomm1":true,
|
||||
"weight_nz_mode":true}' \
|
||||
--enable-expert-parallel \
|
||||
--tensor-parallel-size 4 \
|
||||
--data-parallel-size 4 \
|
||||
--max-num-seqs 48 \
|
||||
--max-model-len 40690 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--gpu-memory-utilization 0.85 \
|
||||
--speculative_config '{"enforce_eager": true, "method": "eagle3", "model": "/path/to/weight/Eagle3/", "num_speculative_tokens": 3}'
|
||||
```
|
||||
|
||||
Remarks:
|
||||
|
||||
- `minimax_m2_append_think` keeps `<think>...</think>` inside `content`.
|
||||
- If you mainly rely on the reasoning semantics of `/v1/responses`, it is recommended to use `--reasoning-parser minimax_m2` instead.
|
||||
- To achieve better performance on long-context scenarios (e.g., 128k or 64k), we recommend the following adjustments:
|
||||
|
||||
```{code-block} bash
|
||||
--tensor-parallel-size 8 \
|
||||
--data-parallel-size 1 \
|
||||
--decode-context-parallel-size 1 \
|
||||
--prefill-context-parallel-size 2 \
|
||||
--cp-kv-cache-interleave-size 128 \
|
||||
--max-num-seqs 16 \
|
||||
--max-model-len 138000 \
|
||||
--max-num-batched-tokens 65536 \
|
||||
--gpu-memory-utilization 0.85 \
|
||||
--speculative_config '{"enforce_eager": true, "method": "eagle3", "model": "/path/to/weight/Eagle3/", "num_speculative_tokens": 1}'
|
||||
```
|
||||
|
||||
> **Note**: The above parameters are validated in a specific test environment for reference only. Please adjust `--max-model-len`, `--max-num-seqs`, `--max-num-batched-tokens`, and `--gpu-memory-utilization` based on your actual input/output length, concurrency, and hardware configuration.
|
||||
|
||||
- If you need to test with `curl` and tool calling, add the following to the startup command:
|
||||
|
||||
```{code-block} bash
|
||||
--enable-auto-tool-choice \
|
||||
--tool-call-parser minimax_m2 \
|
||||
--reasoning-parser minimax_m2_append_think \
|
||||
```
|
||||
|
||||
#### A2 (single node)
|
||||
|
||||
```{code-block} bash
|
||||
export LD_PRELOAD=/usr/lib/aarch64-linux-gnu/libjemalloc.so.2:$LD_PRELOAD
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export HCCL_BUFFSIZE=512
|
||||
sysctl -w vm.swappiness=0
|
||||
sysctl -w kernel.numa_balancing=0
|
||||
sysctl kernel.sched_migration_cost_ns=50000
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export HCCL_INTRA_PCIE_ENABLE=1
|
||||
export HCCL_INTRA_ROCE_ENABLE=0
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
|
||||
vllm serve /path/to/weight/MiniMax-M2.7-w8a8-QuaRot \
|
||||
--served-model-name MiniMax-M2.7 \
|
||||
--host 0.0.0.0 \
|
||||
--port 8000 \
|
||||
--trust-remote-code \
|
||||
--tensor-parallel-size 8 \
|
||||
--quantization ascend \
|
||||
--enable-expert-parallel \
|
||||
--max-num-seqs 32 \
|
||||
--seed 1024 \
|
||||
--max-num-batched-tokens 32768 \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
|
||||
--gpu-memory-utilization 0.85 \
|
||||
--additional-config '{"enable_cpu_binding":true,
|
||||
"enable_flashcomm1":true}' \
|
||||
--model-loader-extra-config '{"enable_multithread_load":true,"num_threads":16}' \
|
||||
--speculative_config '{"method": "eagle3", "model": "/path/to/weight/Eagle3/", "num_speculative_tokens":3}'
|
||||
```
|
||||
|
||||
> **Note**: The above parameters are validated in a specific test environment for reference only. Please adjust `--max-model-len`, `--max-num-seqs`, `--max-num-batched-tokens`, and `--gpu-memory-utilization` based on your actual input/output length, concurrency, and hardware configuration.
|
||||
|
||||
- If you need to test with `curl` and tool calling, add the following to the startup command:
|
||||
|
||||
```{code-block} bash
|
||||
--enable-auto-tool-choice \
|
||||
--tool-call-parser minimax_m2 \
|
||||
--reasoning-parser minimax_m2_append_think \
|
||||
```
|
||||
|
||||
### 5.2 Multi-Node PD Separation Deployment
|
||||
|
||||
PD (Prefill-Decode) separation splits the Prefill and Decode phases across different nodes for better throughput. The following 1P1D configuration is validated for 128k input/output scenarios with `MiniMax-M2.7-W8A8`.
|
||||
|
||||
**Hardware**: 2× Atlas 800 A3 (64GB × 16), one for Prefill, one for Decode.
|
||||
|
||||
**Common Issues Tip:** For PD separation specific issues such as KV transfer timeouts or Mooncake connection errors, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html). For MiniMax-specific PD separation issues, refer to [Chapter 10 FAQ](#10-faq).
|
||||
|
||||
First, prepare `launch_online_dp.py` on each node:
|
||||
|
||||
```python
|
||||
import argparse
|
||||
import multiprocessing
|
||||
import os
|
||||
import subprocess
|
||||
import sys
|
||||
|
||||
def parse_args():
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument("--dp-size", type=int, required=True)
|
||||
parser.add_argument("--tp-size", type=int, default=1)
|
||||
parser.add_argument("--dp-size-local", type=int, default=-1)
|
||||
parser.add_argument("--dp-rank-start", type=int, default=0)
|
||||
parser.add_argument("--dp-address", type=str, required=True)
|
||||
parser.add_argument("--dp-rpc-port", type=str, default=12345)
|
||||
parser.add_argument("--vllm-start-port", type=int, default=9000)
|
||||
return parser.parse_args()
|
||||
|
||||
args = parse_args()
|
||||
dp_size, tp_size = args.dp_size, args.tp_size
|
||||
dp_size_local = args.dp_size_local if args.dp_size_local != -1 else dp_size
|
||||
|
||||
def run_command(visible_devices, dp_rank, vllm_engine_port):
|
||||
subprocess.run([
|
||||
"bash", "./run_dp_template.sh",
|
||||
visible_devices, str(vllm_engine_port),
|
||||
str(dp_size), str(dp_rank), args.dp_address,
|
||||
args.dp_rpc_port, str(tp_size),
|
||||
], check=True)
|
||||
|
||||
if __name__ == "__main__":
|
||||
for i in range(dp_size_local):
|
||||
dp_rank = args.dp_rank_start + i
|
||||
vllm_port = args.vllm_start_port + i
|
||||
visible_devices = ",".join(str(x) for x in range(i * tp_size, (i + 1) * tp_size))
|
||||
p = multiprocessing.Process(target=run_command, args=(visible_devices, dp_rank, vllm_port))
|
||||
p.start()
|
||||
p.join()
|
||||
```
|
||||
|
||||
Then prepare `run_dp_template.sh` on each node.
|
||||
|
||||
**Prefill node** (set `nic_name` and `local_ip` to your own):
|
||||
|
||||
```bash
|
||||
unset http_proxy https_proxy ftp_proxy
|
||||
|
||||
nic_name="<your_nic_name>"
|
||||
local_ip="<your_ip>"
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export OMP_NUM_THREADS=1
|
||||
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
|
||||
sysctl -w vm.swappiness=0
|
||||
sysctl -w kernel.numa_balancing=0
|
||||
sysctl kernel.sched_migration_cost_ns=50000
|
||||
export LD_PRELOAD=/usr/lib/aarch64-linux-gnu/libjemalloc.so.2:$LD_PRELOAD
|
||||
export LD_LIBRARY_PATH=/usr/local/Ascend/ascend-toolkit/latest/python/site-packages/mooncake:$LD_LIBRARY_PATH
|
||||
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export VLLM_ASCEND_ENABLE_FLASHCOMM1=1
|
||||
export VLLM_ASCEND_ENABLE_FUSED_MC2=1
|
||||
export PYTHONHASHSEED=0
|
||||
|
||||
export ASCEND_RT_VISIBLE_DEVICES=$1
|
||||
|
||||
vllm serve /path/to/weight/MiniMax-M2.7-w8a8-QuaRot \
|
||||
--host 0.0.0.0 \
|
||||
--port $2 \
|
||||
--data-parallel-size $3 \
|
||||
--data-parallel-rank $4 \
|
||||
--data-parallel-address $5 \
|
||||
--data-parallel-rpc-port $6 \
|
||||
--tensor-parallel-size $7 \
|
||||
--enable-expert-parallel \
|
||||
--served-model-name minimax \
|
||||
--max-model-len 200000 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--max-num-seqs 64 \
|
||||
--trust-remote-code \
|
||||
--gpu-memory-utilization 0.75 \
|
||||
--quantization ascend \
|
||||
--enforce-eager \
|
||||
--speculative_config '{"method": "eagle3", "model": "/path/to/weight/Eagle3/", "num_speculative_tokens": 1}' \
|
||||
--additional-config '{"enable_cpu_binding":true}' \
|
||||
--kv-transfer-config \
|
||||
'{"kv_connector": "MooncakeConnectorV1",
|
||||
"kv_role": "kv_producer",
|
||||
"kv_port": "35880",
|
||||
"engine_id": "0",
|
||||
"kv_connector_extra_config": {
|
||||
"use_ascend_direct": true,
|
||||
"prefill": {"dp_size": 2, "tp_size": 8},
|
||||
"decode": {"dp_size": 2, "tp_size": 8}
|
||||
}}'
|
||||
```
|
||||
|
||||
**Decode node** (set `nic_name` and `local_ip` to your own):
|
||||
|
||||
```bash
|
||||
unset http_proxy https_proxy ftp_proxy
|
||||
|
||||
nic_name="<your_nic_name>"
|
||||
local_ip="<your_ip>"
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
|
||||
export HCCL_BUFFSIZE=2048
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export OMP_NUM_THREADS=1
|
||||
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
|
||||
sysctl -w vm.swappiness=0
|
||||
sysctl -w kernel.numa_balancing=0
|
||||
sysctl kernel.sched_migration_cost_ns=50000
|
||||
export LD_PRELOAD=/usr/lib/aarch64-linux-gnu/libjemalloc.so.2:$LD_PRELOAD
|
||||
export LD_LIBRARY_PATH=/usr/local/Ascend/ascend-toolkit/latest/python/site-packages/mooncake:$LD_LIBRARY_PATH
|
||||
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export VLLM_ASCEND_ENABLE_FLASHCOMM1=0
|
||||
export VLLM_ASCEND_ENABLE_FUSED_MC2=1
|
||||
export PYTHONHASHSEED=0
|
||||
|
||||
export ASCEND_RT_VISIBLE_DEVICES=$1
|
||||
|
||||
vllm serve /path/to/weight/MiniMax-M2.7-w8a8-QuaRot \
|
||||
--host 0.0.0.0 \
|
||||
--port $2 \
|
||||
--data-parallel-size $3 \
|
||||
--data-parallel-rank $4 \
|
||||
--data-parallel-address $5 \
|
||||
--data-parallel-rpc-port $6 \
|
||||
--tensor-parallel-size $7 \
|
||||
--enable-expert-parallel \
|
||||
--served-model-name minimax \
|
||||
--max-model-len 200000 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--max-num-seqs 16 \
|
||||
--trust-remote-code \
|
||||
--no-enable-prefix-caching \
|
||||
--gpu-memory-utilization 0.75 \
|
||||
--quantization ascend \
|
||||
--async-scheduling \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \
|
||||
--speculative_config '{"method": "eagle3", "model": "/path/to/weight/Eagle3/", "num_speculative_tokens": 3}' \
|
||||
--additional-config '{"enable_cpu_binding":true}' \
|
||||
--kv-transfer-config \
|
||||
'{"kv_connector": "MooncakeConnectorV1",
|
||||
"kv_role": "kv_consumer",
|
||||
"kv_port": "56900",
|
||||
"engine_id": "1",
|
||||
"kv_connector_extra_config": {
|
||||
"use_ascend_direct": true,
|
||||
"prefill": {"dp_size": 2, "tp_size": 8},
|
||||
"decode": {"dp_size": 2, "tp_size": 8}
|
||||
}}'
|
||||
```
|
||||
|
||||
Once the scripts are ready, start the servers on each node.
|
||||
|
||||
**Prefill node:**
|
||||
|
||||
```bash
|
||||
python launch_online_dp.py \
|
||||
--dp-size 2 --tp-size 8 \
|
||||
--dp-size-local 2 --dp-rank-start 0 \
|
||||
--dp-address <prefill_ip> --dp-rpc-port 12321 \
|
||||
--vllm-start-port 7000
|
||||
```
|
||||
|
||||
**Decode node:**
|
||||
|
||||
```bash
|
||||
python launch_online_dp.py \
|
||||
--dp-size 2 --tp-size 8 \
|
||||
--dp-size-local 2 --dp-rank-start 0 \
|
||||
--dp-address <decode_ip> --dp-rpc-port 12321 \
|
||||
--vllm-start-port 7100
|
||||
```
|
||||
|
||||
#### Request Forwarding
|
||||
|
||||
Run the proxy on any machine that can reach both nodes. You can get the proxy script from the repository: [load_balance_proxy_server_example.py](https://github.com/vllm-project/vllm-ascend/blob/main/examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py).
|
||||
|
||||
```bash
|
||||
unset http_proxy https_proxy
|
||||
|
||||
python load_balance_proxy_server_example.py \
|
||||
--port 8009 \
|
||||
--host <prefill_ip> \
|
||||
--prefiller-hosts \
|
||||
<prefill_ip> <prefill_ip> \
|
||||
--prefiller-ports \
|
||||
7000 7001 \
|
||||
--decoder-hosts \
|
||||
<decode_ip> <decode_ip> \
|
||||
--decoder-ports \
|
||||
7100 7101
|
||||
```
|
||||
|
||||
The service is then accessible at `http://<proxy_ip>:8009`.
|
||||
|
||||
## 6 Functional Verification
|
||||
|
||||
Once your server is started, you can query the model with input prompts.
|
||||
|
||||
**Note:**
|
||||
|
||||
- `<node_ip>`: The IP address of the node where the server is running (e.g., localhost for single-node).
|
||||
- `<port>`: The port number specified in the server startup command (e.g., `8000`).
|
||||
|
||||
### Using curl
|
||||
|
||||
```bash
|
||||
curl http://<node_ip>:<port>/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "MiniMax-M2.7",
|
||||
"messages": [{"role": "user", "content": "Hello, who are you?"}],
|
||||
"stream": false,
|
||||
"temperature": 0.8,
|
||||
"max_tokens": 200
|
||||
}'
|
||||
```
|
||||
|
||||
Expected result: HTTP 200 with a JSON response containing a `choices` field with the model's reply text.
|
||||
|
||||
### Using OpenAI Python Client
|
||||
|
||||
```python
|
||||
from openai import OpenAI
|
||||
|
||||
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="na")
|
||||
|
||||
resp = client.chat.completions.create(
|
||||
model="MiniMax-M2.7",
|
||||
messages=[{"role": "user", "content": "你好,请介绍一下你自己,并展示一次工具调用的参数格式。"}],
|
||||
max_tokens=256,
|
||||
)
|
||||
print(resp.choices[0].message.content)
|
||||
```
|
||||
|
||||
Expected result: The response should contain a coherent self-introduction and tool call parameter format in the `content` field.
|
||||
|
||||
### Tool Calling Verification
|
||||
|
||||
```bash
|
||||
curl http://localhost:8000/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "MiniMax-M2.7",
|
||||
"messages": [{"role": "user", "content": "请查询上海的天气。"}],
|
||||
"tools": [{
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": "get_current_weather",
|
||||
"description": "Get weather by city",
|
||||
"parameters": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"city": {"type": "string"},
|
||||
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}
|
||||
},
|
||||
"required": ["city"]
|
||||
}
|
||||
}
|
||||
}],
|
||||
"tool_choice": "auto",
|
||||
"temperature": 0,
|
||||
"max_tokens": 512
|
||||
}'
|
||||
```
|
||||
|
||||
Expected result: HTTP 200 with a JSON response containing a `tool_calls` field with the function name and arguments.
|
||||
|
||||
## 7 Accuracy Evaluation
|
||||
|
||||
> **Note**: Post-processing parameters (e.g., `max_tokens`, `temperature`, `stop` tokens) should match those defined in the model weight's `generation_config.json`. The recommended maximum output length for GPQA-diamond and AIME2025 is 64k (65536 tokens).
|
||||
|
||||
Here are two accuracy evaluation methods.
|
||||
|
||||
### 7.1 Using AISBench
|
||||
|
||||
For details, please refer to [Using AISBench](../../developer_guide/evaluation/using_ais_bench.md).
|
||||
|
||||
### 7.2 Using Language Model Evaluation Harness
|
||||
|
||||
Using the `gsm8k` dataset as an example test dataset, run the accuracy evaluation for `MiniMax-M2.7-W8A8` in online mode.
|
||||
|
||||
1. For `lm_eval` installation, please refer to [Using lm_eval](../../developer_guide/evaluation/using_lm_eval.md).
|
||||
2. Run `lm_eval` to execute the accuracy evaluation:
|
||||
|
||||
```shell
|
||||
lm_eval \
|
||||
--model local-completions \
|
||||
--model_args model=/path/to/weight/MiniMax-M2.7-w8a8-QuaRot,base_url=http://127.0.0.1:8000/v1/completions,tokenized_requests=False,trust_remote_code=True \
|
||||
--tasks gsm8k \
|
||||
--output_path ./
|
||||
```
|
||||
|
||||
## 8 Performance Evaluation
|
||||
|
||||
### 8.1 Using AISBench
|
||||
|
||||
Refer to [Using AISBench for performance evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation) for details.
|
||||
|
||||
### 8.2 Using vLLM Benchmark
|
||||
|
||||
Run performance evaluation for `MiniMax-M2.7-W8A8` as an example.
|
||||
|
||||
Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/) for more details.
|
||||
|
||||
Take the `serve` subcommand as an example:
|
||||
|
||||
```shell
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
vllm bench serve \
|
||||
--model /path/to/weight/MiniMax-M2.7-w8a8-QuaRot \
|
||||
--dataset-name random \
|
||||
--random-input 200 \
|
||||
--num-prompts 200 \
|
||||
--request-rate 1 \
|
||||
--save-result \
|
||||
--result-dir ./
|
||||
```
|
||||
|
||||
## 9 Performance Tuning
|
||||
|
||||
> **Note**: The following configurations are validated in specific test environments and are for reference only. The optimal configuration depends on factors such as maximum input/output length, prefix cache hit rate, precision requirements, and deployment machine ratios. It is recommended to refer to Section 9.2 for tuning based on actual conditions.
|
||||
|
||||
### 9.1 Recommended Configurations
|
||||
|
||||
The following configurations are validated in internal testing and are categorized by use case.
|
||||
|
||||
| Scenario | Input/Output | Deployment | NPUs | P Config | D Config | Max Batched Tokens | Max Num Seqs (P/D) | Max Model Len | EAGLE3 | FUSED_MC2 | FlashComm1 | Async Scheduling |
|
||||
|----------|-------------|------------|------|----------|----------|-------------------|----------------|---------------|--------|-----------|------------|------------------|
|
||||
| Short Seq High Throughput | 3.5k → 1.5k | 1P2D PD separation | 24 (A3) | DP8TP2EP16 | DP32TP1EP32 | 16384 | 128 / 128 | 32k | 3 | On | On | On |
|
||||
| Short Seq Low Latency | 3.5k → 1.5k | 1P2D PD separation | 24 (A3) | DP4TP4EP16 | DP8TP4EP32 | 16384 | 128 / 128 | 32k | 3 | On | On | On |
|
||||
| Long Seq High Throughput | 128k → 1k <br> (90% cache hit) | 1P1D PD separation | 16 (A3) | DP2TP8EP16 | DP2TP8EP16 | 16384 | 64 / 16 | 200k | 3 | On | On | On |
|
||||
| Long Seq Low Latency | 128k → 1k <br> (90% cache hit) | 1P2D PD separation | 24 (A3) | DP2TP8EP16 | DP4TP8EP32 | 16384 | 64 / 16 | 200k | 3 | On | On | On |
|
||||
|
||||
> **Note**: The prefix cache hit rate for short-sequence tests is 0%; for long-sequence tests it is 90%. Adjust `max-num-seqs`, `max-model-len`, and `max-num-batched-tokens` based on your actual workload.
|
||||
|
||||
### 9.2 Tuning Guidelines
|
||||
|
||||
#### 9.2.1 General Tuning Reference
|
||||
|
||||
Please refer to the [Public Performance Tuning Documentation](../../developer_guide/performance_and_debug/optimization_and_tuning.md) for general tuning methods.
|
||||
|
||||
Please refer to the [Feature Guide](../../user_guide/support_matrix/feature_matrix.md) for detailed feature descriptions.
|
||||
|
||||
#### 9.2.2 Model-Specific Optimizations
|
||||
|
||||
##### Optimizations Enabled by Default
|
||||
|
||||
The following optimizations are enabled by default and require no additional configuration:
|
||||
|
||||
| Optimization Technique | Technical Principle | Performance Benefit |
|
||||
| ---------------------- | ------------------- | ------------------- |
|
||||
| FullGraph Optimization | Captures and replays the entire decoding graph at once using `compilation_config={"cudagraph_mode":"FULL_DECODE_ONLY"}` | Significantly reduces scheduling latency, stabilizes multi-device performance |
|
||||
| CPU Binding | Uses `--additional-config '{"enable_cpu_binding":true}'` to bind CPU cores | Reduces cross-core scheduling overhead, improving decode latency stability |
|
||||
| Multi-thread Weight Loading | Uses `--model-loader-extra-config '{"enable_multithread_load":true}'` for parallel weight loading | Reduces model loading time |
|
||||
|
||||
##### Optimizations That Require Explicit Enabling
|
||||
|
||||
| Optimization Technique | Applicable Scenarios | Enablement Method | Technical Principle | Precautions |
|
||||
| ---------------------- | -------------------- | ----------------- | ------------------- | ----------- |
|
||||
| FlashComm v1 | High-concurrency, TP scenarios | `--additional-config '{"enable_flashcomm1": true}'` | Decomposes traditional Allreduce into Reduce-Scatter and All-Gather | Threshold protection: only takes effect when the actual number of tokens exceeds the threshold |
|
||||
| Fused MC2 | TP ≥ 4 scenarios | `--additional-config '{"enable_fused_mc2": true}'` | Fuses multiple communication and computation operations | Recommended for A3; not applicable for A2 |
|
||||
| Balanced Scheduling | High DP scenarios | `export VLLM_ASCEND_BALANCE_SCHEDULING=1` | Enhances scheduling capacity between prefill and decode | Currently disabled by default (`0`). Set to `1` only when concurrency ≈ DP × max-num-seqs. Disable for long-context scenarios |
|
||||
| EAGLE3 Speculative Decoding | All scenarios | `--speculative_config '{"method": "eagle3", "model": "/path/to/Eagle3/", "num_speculative_tokens": 3}'` | Uses a draft model to predict future tokens | 1–3 tokens for long context; 3 tokens for short context |
|
||||
| jemalloc Preload | All scenarios | `export LD_PRELOAD=/usr/lib/aarch64-linux-gnu/libjemalloc.so.2` | Replaces default memory allocator to reduce fragmentation | Ensure jemalloc is installed in the container |
|
||||
|
||||
## 10 FAQ
|
||||
|
||||
For common environment, installation, and general parameter issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html). This chapter only covers MiniMax-M2 (M2.5/M2.7) model-specific issues.
|
||||
|
||||
- **Q: Does C8 quantization support EAGLE3 speculative decoding?**
|
||||
|
||||
A: Not yet. C8 quantization with EAGLE3 is currently unsupported.
|
||||
|
||||
- **Q: Which `--reasoning-parser` is recommended for tool calling tasks?**
|
||||
|
||||
A: For tool calling tasks, it is recommended to use `--reasoning-parser minimax_m2_append_think`.
|
||||
|
||||
- **Q: Why is the `reasoning` field often empty when using `minimax_m2_append_think`, and how should I choose the right `--reasoning-parser`?**
|
||||
|
||||
A: This is expected behavior. The `minimax_m2_append_think` parser retains `<think>...</think>` blocks directly inside the `content` field instead of separating them. If your downstream application relies on the standard reasoning semantics of `/v1/responses` (where the thinking process and final answer are separated), you should use `--reasoning-parser minimax_m2` to ensure the dedicated `reasoning` field is properly populated.
|
||||
|
||||
- **Q: Startup fails with HCCL port conflicts (address already bound). What should I do?**
|
||||
|
||||
A: Check whether another process is already occupying the port (e.g., `lsof -i :<port>` or `ss -tlnp | grep <port>`). If a port conflict is found, switch to a different port with `--port`, or terminate the specific process occupying that port.
|
||||
|
||||
- **Q: How to handle OOM or unstable startup?**
|
||||
|
||||
A: Refer to the upstream vLLM guide on [out-of-memory troubleshooting](https://docs.vllm.ai/en/latest/usage/troubleshooting/#out-of-memory). In short: reduce `--max-num-seqs` and `--max-num-batched-tokens` first, lower `--gpu-memory-utilization` (e.g., from 0.9 to 0.85), or decrease the number of concurrent requests.
|
||||
|
||||
- **Q: Which ports must be accessible?**
|
||||
|
||||
A: At minimum, expose the serving port (e.g., `8000`). For multi-node deployment, also ensure HCCL communication ports and DP RPC ports are accessible.
|
||||
120
docs/source/tutorials/models/Minitron-8B-Base.md
Normal file
120
docs/source/tutorials/models/Minitron-8B-Base.md
Normal file
@@ -0,0 +1,120 @@
|
||||
# Minitron-8B-Base
|
||||
|
||||
## Introduction
|
||||
|
||||
The released `Minitron-8B-Base` is a lightweight, efficient large language model developed by NVIDIA. It is designed for general-purpose text generation and reasoning tasks, and can be deployed with vLLM for online serving and evaluation on Ascend NPU hardware through `vllm-ascend`.
|
||||
|
||||
This document describes the main verification steps of the model, including supported features, environment preparation, single-node deployment, functional verification, and accuracy evaluation on the GSM8K benchmark.
|
||||
|
||||
## Environment Preparation
|
||||
|
||||
### Model Weight
|
||||
|
||||
`Minitron-8B-Base`(BF16 version): requires 1 Ascend 910B (with 1 x 64GB NPUs). [Download model weight](https://www.modelscope.cn/models/nv-community/Minitron-8B-Base)
|
||||
|
||||
It is recommended to place the model weight in a shared cache directory, such as `/root/.cache/` or a local model path like `/data/vllm-workspace/models/Minitron-8B-Base`.
|
||||
|
||||
### Installation
|
||||
|
||||
`Minitron-8B-Base` can be deployed with `vllm-ascend` in a compatible runtime environment.
|
||||
|
||||
You can use the official docker image for deployment:
|
||||
|
||||
```bash
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-v /data/vllm-workspace/models:/data/vllm-workspace/models \
|
||||
-p 8000:8000 \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
If you do not want to use the docker image, you can also build from source:
|
||||
|
||||
- Install `vllm-ascend` from source, refer to [installation](../../installation.md).
|
||||
|
||||
## Deployment
|
||||
|
||||
Start the online serving service with the following command:
|
||||
|
||||
``` bash
|
||||
vllm serve "nv-community/Minitron-8B-Base" \
|
||||
--served-model-name minitron-8b-base \
|
||||
--tensor-parallel-size 1 \
|
||||
--max-model-len 4096 \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--enforce-eager \
|
||||
--port 8000
|
||||
```
|
||||
|
||||
## Functional Verification
|
||||
|
||||
Once your server is started, you can query the model with a simple prompt:
|
||||
|
||||
```bash
|
||||
curl http://localhost:8000/v1/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "minitron-8b-base",
|
||||
"prompt": "Question: If a train travels 60 miles in 2 hours, what is its average speed in miles per hour?\nAnswer:",
|
||||
"max_tokens": 64,
|
||||
"temperature": 1.0
|
||||
}'
|
||||
```
|
||||
|
||||
A valid response indicates that the model is deployed correctly and can generate text outputs.
|
||||
|
||||
## Accuracy Evaluation
|
||||
|
||||
The GSM8K dataset was used to evaluate the reasoning capability of `Minitron-8B-Base`.
|
||||
|
||||
The current evaluation setting is:
|
||||
|
||||
- Dataset: `gsm8k`
|
||||
- Split: `test`
|
||||
- Number of samples: `1000`
|
||||
- Few-shot setting: `5-shot`
|
||||
- `apply_chat_template`: `False`
|
||||
- `fewshot_as_multiturn`: `False`
|
||||
|
||||
The current evaluation results are:
|
||||
|
||||
| Category | Dataset | Metric | Result |
|
||||
|----------|---------|--------|--------|
|
||||
| Accuracy | gsm8k / test | Total Samples | 1000 |
|
||||
| Accuracy | gsm8k / test | exact_match,strict-match | 0.5436 |
|
||||
| Accuracy | gsm8k / test | exact_match,flexible-extract | 0.5451 |
|
||||
|
||||
### Remarks on Metrics
|
||||
|
||||
- **exact_match,strict-match**: Only predictions that strictly match the expected final-answer extraction format are counted as correct.
|
||||
- **exact_match,flexible-extract**: Predictions are evaluated with a more flexible answer extraction rule, which tolerates minor formatting differences as long as the final numeric answer is correct.
|
||||
|
||||
## Performance
|
||||
|
||||
### Baseline Result
|
||||
|
||||
`Minitron-8B-Base` can be deployed through `vllm-ascend` for online inference and benchmark evaluation.
|
||||
Actual throughput and latency depend on hardware resources, prompt length, output length, concurrency, and runtime configuration.
|
||||
|
||||
### Remarks
|
||||
|
||||
This document focuses on functional verification and benchmark accuracy on GSM8K.
|
||||
Further benchmarking is recommended for:
|
||||
|
||||
- request latency
|
||||
- throughput under concurrency
|
||||
- long-context inference
|
||||
- memory utilization
|
||||
- stability under continuous serving workloads
|
||||
201
docs/source/tutorials/models/Mixtral-8x7B-Instruct-v0.1.md
Normal file
201
docs/source/tutorials/models/Mixtral-8x7B-Instruct-v0.1.md
Normal file
@@ -0,0 +1,201 @@
|
||||
# Mixtral-8x7B-Instruct-v0.1
|
||||
|
||||
## Introduction
|
||||
|
||||
Mixtral-8x7B-Instruct-v0.1 is a state-of-the-art mixture-of-experts (MoE) language model developed by Mistral AI. It features 8 expert models, each with 7B parameters, and is specifically fine-tuned for instruction following tasks.
|
||||
|
||||
Key features of Mixtral-8x7B-Instruct-v0.1 include:
|
||||
|
||||
- 8x7B parameters with sparse activation (only 2 experts activated per token)
|
||||
- Strong performance across various NLP tasks
|
||||
- Support for extended context length
|
||||
- High-quality instruction following capabilities
|
||||
|
||||
This document will show the main verification steps of the model, including supported features, feature configuration, environment preparation, single-node deployment, accuracy and performance evaluation.
|
||||
|
||||
The `Mixtral-8x7B-Instruct-v0.1` model is supported in vllm-ascend.
|
||||
|
||||
## Environment Preparation
|
||||
|
||||
### Model Weight
|
||||
|
||||
- `Mixtral-8x7B-Instruct-v0.1`(BF16 version): [Download model weight](https://huggingface.co/mistralai/Mixtral-8x7B-Instruct-v0.1)
|
||||
- Quantized versions may be available from third-party providers.
|
||||
|
||||
It is recommended to download the model weight to a local directory, such as `/data/models/`.
|
||||
|
||||
### Installation
|
||||
|
||||
You can use our official docker image to run `Mixtral-8x7B-Instruct-v0.1` directly.
|
||||
|
||||
Select an image based on your machine type and start the docker image on your node, refer to [using docker](../../installation.md#set-up-using-docker).
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# Update --device according to your device (Atlas A2: /dev/davinci[0-7] Atlas A3:/dev/davinci[0-15]).
|
||||
# Update the vllm-ascend image according to your environment.
|
||||
# Note you should download the weight to /root/.cache in advance.
|
||||
# Update the vllm-ascend image
|
||||
export IMAGE=m.daocloud.io/quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
export NAME=vllm-ascend
|
||||
|
||||
# Run the container using the defined variables
|
||||
# Note: If you are running bridge network with docker, please expose available ports for multiple nodes communication in advance.
|
||||
docker run --rm \
|
||||
--name $NAME \
|
||||
--net=host \
|
||||
--shm-size=1g \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
## Deployment
|
||||
|
||||
### Single-node Deployment
|
||||
|
||||
- `Mixtral-8x7B-Instruct-v0.1` can be deployed on 1 Atlas 800 A3 (64GB × 16) or 1 Atlas 800 A2 (64GB × 8).
|
||||
|
||||
Run the following script to execute online inference.
|
||||
|
||||
``` bash
|
||||
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=10
|
||||
export VLLM_USE_V1=1
|
||||
export HCCL_BUFFSIZE=200
|
||||
export VLLM_ASCEND_ENABLE_MLAPO=1
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export VLLM_ASCEND_ENABLE_FLASHCOMM1=1
|
||||
```
|
||||
|
||||
``` bash
|
||||
|
||||
vllm serve "mistralai/Mixtral-8x7B-Instruct-v0.1" \
|
||||
--tensor-parallel-size 4 \
|
||||
--max-model-len 4096 \
|
||||
--dtype bfloat16 \
|
||||
--trust-remote-code \
|
||||
--enforce-eager \
|
||||
--block-size 128 \
|
||||
--gpu-memory-utilization 0.7
|
||||
```
|
||||
|
||||
**Notice:**
|
||||
The parameters are explained as follows:
|
||||
|
||||
- Setting the environment variable `VLLM_ASCEND_BALANCE_SCHEDULING=1` enables balance scheduling. This may help increase output throughput and reduce TPOT in v1 scheduler. However, TTFT may degrade in some scenarios.
|
||||
- `--max-model-len` specifies the maximum context length - that is, the sum of input and output tokens for a single request. For testing purposes, a value of `4096` is used here.
|
||||
- `--dtype float16` specifies the data type for model weights and computations.
|
||||
- `--trust-remote-code` allows loading models with custom code.
|
||||
- `--enforce-eager` forces the use of eager execution mode instead of graph compilation, which can be more stable for some models.
|
||||
- `--block-size` specifies the block size for KV cache management, with a value of `128` used here.
|
||||
- `--gpu-memory-utilization` sets the proportion of NPU memory to use for the model, with a value of `0.7` used here to reduce memory usage.
|
||||
|
||||
## Functional Verification
|
||||
|
||||
Once your server is started, you can query the model with input prompts. Mixtral-8x7B-Instruct-v0.1 uses a specific prompt format with [INST] and [/INST] tags:
|
||||
|
||||
```shell
|
||||
curl http://localhost:8000/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "mistralai/Mixtral-8x7B-Instruct-v0.1",
|
||||
"messages": [
|
||||
{"role": "user", "content": "你好,介绍一下你自己"}
|
||||
],
|
||||
"max_tokens": 100,
|
||||
"temperature": 0.7
|
||||
}'
|
||||
```
|
||||
|
||||
For instruction following tasks, you can use prompts like:
|
||||
|
||||
```shell
|
||||
curl http://localhost:8000/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "mistralai/Mixtral-8x7B-Instruct-v0.1",
|
||||
"messages": [
|
||||
{"role": "user", "content": "扮演一位资深架构师,评价一下在昇腾 Atlas A2 上部署 vLLM 的优势。"}
|
||||
],
|
||||
"max_tokens": 100,
|
||||
"temperature": 0.7
|
||||
}'
|
||||
```
|
||||
|
||||
For MoE-related questions:
|
||||
|
||||
```shell
|
||||
curl http://localhost:8000/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "mistralai/Mixtral-8x7B-Instruct-v0.1",
|
||||
"messages": [
|
||||
{"role": "user", "content": "简单解释一下为什么 Mixtral 模型被称为\"混合专家模型\"(MoE)?"}
|
||||
],
|
||||
"max_tokens": 100,
|
||||
"temperature": 0.7
|
||||
}'
|
||||
```
|
||||
|
||||
## Accuracy Evaluation
|
||||
|
||||
### Using AISBench
|
||||
|
||||
1. Refer to [Using AISBench](../../developer_guide/evaluation/using_ais_bench.md) for details.
|
||||
|
||||
2. After execution, you can get the result. For reference, Mixtral-8x7B-Instruct-v0.1 typically performs well on various benchmarks including reasoning, comprehension, and instruction following tasks.
|
||||
|
||||
## Performance Evaluation
|
||||
|
||||
### Using AISBench
|
||||
|
||||
Refer to [Using AISBench for performance evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation) for details.
|
||||
|
||||
### Using vLLM Benchmark
|
||||
|
||||
Run performance evaluation of `Mixtral-8x7B-Instruct-v0.1` as an example.
|
||||
|
||||
Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/) for more details.
|
||||
|
||||
There are three `vllm bench` subcommands:
|
||||
|
||||
- `latency`: Benchmark the latency of a single batch of requests.
|
||||
- `serve`: Benchmark the online serving throughput.
|
||||
- `throughput`: Benchmark offline inference throughput.
|
||||
|
||||
Take the `serve` as an example. First, start the server:
|
||||
|
||||
```shell
|
||||
python -m vllm.entrypoints.openai.api_server \
|
||||
--model mistralai/Mixtral-8x7B-Instruct-v0.1 \
|
||||
--host 0.0.0.0 \
|
||||
--port 8000 \
|
||||
--tensor-parallel-size 4 \
|
||||
--max-model-len 512 \
|
||||
--dtype float16 \
|
||||
--trust-remote-code \
|
||||
--enforce-eager \
|
||||
--block-size 128 \
|
||||
--gpu-memory-utilization 0.7
|
||||
```
|
||||
|
||||
## Conclusion
|
||||
|
||||
Mixtral-8x7B-Instruct-v0.1 is a powerful MoE model that offers excellent performance for instruction following tasks. With proper deployment on Ascend hardware using vllm-ascend, you can achieve high throughput and low latency for your AI applications.
|
||||
|
||||
For more details about model capabilities and best practices, refer to the official Mixtral documentation and vllm-ascend user guide.
|
||||
417
docs/source/tutorials/models/PaddleOCR-VL.md
Normal file
417
docs/source/tutorials/models/PaddleOCR-VL.md
Normal file
@@ -0,0 +1,417 @@
|
||||
# PaddleOCR-VL
|
||||
|
||||
## 1 Introduction
|
||||
|
||||
PaddleOCR-VL is a SOTA and resource-efficient model tailored for document parsing. Its core component is PaddleOCR-VL-0.9B, a compact yet powerful vision-language model (VLM) that integrates a NaViT-style dynamic resolution visual encoder with the ERNIE-4.5-0.3B language model to enable accurate element recognition.
|
||||
|
||||
This document provides a detailed workflow for the complete deployment and verification of the model, including supported features, environment preparation, single-node deployment, and functional verification. It is designed to help users quickly complete model deployment and verification.
|
||||
|
||||
This document is validated and written based on **vLLM-Ascend v0.21.0rc1**. The current model (PaddleOCR-VL) is supported in this version. It is recommended to use this version or another updated official version for deployment.
|
||||
|
||||
## 2 Supported Features
|
||||
|
||||
Refer to [Supported Features List](../../user_guide/support_matrix/supported_models.md) to get the model's supported feature matrix.
|
||||
|
||||
Refer to [Feature Guide](../../user_guide/feature_guide/index.md) to get the feature's configuration.
|
||||
|
||||
## 3 Prerequisites
|
||||
|
||||
### 3.1 Model Weight
|
||||
|
||||
- `PaddleOCR-VL-0.9B`: [PaddleOCR-VL-0.9B](https://www.modelscope.cn/models/PaddlePaddle/PaddleOCR-VL)
|
||||
|
||||
It is recommended to download the model weights to the cache directory and set `VLLM_USE_MODELSCOPE=True` to load the model automatically. If you have downloaded the weights to a local directory, update the `MODEL_PATH` variable in the deployment script accordingly.
|
||||
|
||||
## 4 Installation
|
||||
|
||||
### 4.1 Docker Image Installation
|
||||
|
||||
You can use our official docker image to run `PaddleOCR-VL` directly.
|
||||
|
||||
Select an image based on your machine type and start the docker image on your node, refer to [using docker](../../installation.md#set-up-using-docker).
|
||||
|
||||
For Atlas 300I DUO, use `vllm-ascend:nightly-releases-v0.23.0-310p` (or a later `-310p` image).
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: install
|
||||
|
||||
::::{tab-item} A2 series
|
||||
:sync: A2
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
:sync: atlas300
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-310p
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
After a successful docker run, you can verify the running container service by executing the `docker ps` command.
|
||||
|
||||
### 4.2 Source Code Installation
|
||||
|
||||
If you don't want to use the docker image as above, you can also build all from source:
|
||||
|
||||
- Install `vllm-ascend` from source, refer to [installation](../../installation.md).
|
||||
|
||||
## 5 Online Service Deployment
|
||||
|
||||
### 5.1 Single-Node Online Deployment
|
||||
|
||||
PaddleOCR-VL supports single-node single-card deployment on the A2 series and Atlas 300I DUO platform. Single-node deployment completes both Prefill and Decode within the same node.
|
||||
|
||||
Follow these steps to start the inference service:
|
||||
|
||||
1. Prepare model weights: Ensure the model weights are accessible. With `VLLM_USE_MODELSCOPE=True`, the model will be loaded automatically from ModelScope.
|
||||
2. Set the `MODEL_PATH` environment variable to point to your model directory.
|
||||
3. Create and execute the deployment script (save as `deploy.sh`).
|
||||
|
||||
Startup Command:
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: install
|
||||
|
||||
::::{tab-item} A2 series
|
||||
:sync: A2
|
||||
|
||||
```bash
|
||||
#!/bin/sh
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
export MODEL_PATH="PaddlePaddle/PaddleOCR-VL"
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export CPU_AFFINITY_CONF=1
|
||||
export PYTORCH_NPU_ALLOC_CONF="expandable_segments:True"
|
||||
|
||||
vllm serve ${MODEL_PATH} \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--served-model-name PaddleOCR-VL-0.9B \
|
||||
--trust-remote-code \
|
||||
--no-enable-prefix-caching \
|
||||
--mm-processor-cache-gb 0 \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \
|
||||
--additional_config '{"enable_cpu_binding":true}' \
|
||||
--port 8000
|
||||
```
|
||||
|
||||
Key Parameter Descriptions:
|
||||
|
||||
- `--max-num-batched-tokens` specifies the maximum number of tokens batched in a single forward pass. Adjust this parameter for throughput optimization.
|
||||
- `--no-enable-prefix-caching` indicates that prefix caching is disabled. To enable it, remove this option.
|
||||
- `--mm-processor-cache-gb` sets the size of the multimodal processor cache (in GB). A value of `0` disables caching.
|
||||
- `--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}'` enables full decode graph compilation for improved performance.
|
||||
- `--additional_config '{"enable_cpu_binding":true}'` enables CPU binding to improve performance.
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
:sync: atlas300
|
||||
|
||||
```bash
|
||||
#!/bin/sh
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
export MODEL_PATH="PaddlePaddle/PaddleOCR-VL"
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export PYTORCH_NPU_ALLOC_CONF="expandable_segments:True"
|
||||
|
||||
vllm serve ${MODEL_PATH} \
|
||||
--max_model_len 16384 \
|
||||
--served-model-name PaddleOCR-VL-0.9B \
|
||||
--trust-remote-code \
|
||||
--no-enable-prefix-caching \
|
||||
--mm-processor-cache-gb 0 \
|
||||
--dtype float16 \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \
|
||||
--additional_config '{"enable_cpu_binding":true}' \
|
||||
--port 8000
|
||||
```
|
||||
|
||||
:::{note}
|
||||
|
||||
On Atlas 300I DUO:
|
||||
|
||||
- Only `float16` dtype is supported.
|
||||
- Graph compilation (`--compilation-config`) requires **CANN version >= 9.0.0**. If your CANN version is lower, please revert to eager mode by replacing the `--compilation-config` argument with `--enforce-eager`.
|
||||
:::
|
||||
|
||||
Key Parameter Descriptions:
|
||||
|
||||
- `--max_model_len` specifies the maximum context length — that is, the sum of input and output tokens for a single request.
|
||||
- `--no-enable-prefix-caching` indicates that prefix caching is disabled. To enable it, remove this option.
|
||||
- `--mm-processor-cache-gb` sets the size of the multimodal processor cache (in GB). A value of `0` disables caching.
|
||||
- `--dtype float16` specifies the model dtype. On Atlas 300I DUO, only `float16` is supported.
|
||||
- `--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}'` enables full decode graph compilation for improved performance. On Atlas 300I DUO, `fuse_norm_quant` in graph compilation is disabled by default in `--additional_config`.
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
Common Issues Tip: If you encounter startup issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html) for troubleshooting.
|
||||
|
||||
### 5.2 Multi-Node PD Separation Deployment
|
||||
|
||||
Not supported yet.
|
||||
|
||||
### 5.3 Special Deployment Modes
|
||||
|
||||
#### 5.3.1 Offline Inference with vLLM and PP-DocLayoutV2
|
||||
|
||||
In the above example, we demonstrated how to use vLLM to infer the PaddleOCR-VL-0.9B model. Typically, we also need to integrate the PP-DocLayoutV2 model to fully unleash the capabilities of the PaddleOCR-VL model, making it more consistent with the examples provided by the official PaddlePaddle documentation.
|
||||
|
||||
:::{note}
|
||||
Use separate virtual environments for VLLM and PP-DocLayoutV2 to prevent dependency conflicts.
|
||||
:::
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: install
|
||||
|
||||
::::{tab-item} A2 series
|
||||
:sync: A2
|
||||
|
||||
The A2 series device supports inference using the PaddlePaddle framework.
|
||||
|
||||
1. Pull the PaddlePaddle-compatible CANN image
|
||||
|
||||
```bash
|
||||
docker pull ccr-2vdh3abv-pub.cnc.bj.baidubce.com/device/paddle-npu:cann800-ubuntu20-npu-910b-base-aarch64-gcc84
|
||||
```
|
||||
|
||||
Start the container using the following command:
|
||||
|
||||
```bash
|
||||
docker run -it --name paddle-npu-dev -v $(pwd):/work \
|
||||
--privileged --network=host --shm-size=128G -w=/work \
|
||||
-v /usr/local/Ascend/driver:/usr/local/Ascend/driver \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-e ASCEND_RT_VISIBLE_DEVICES="0,1,2,3,4,5,6,7" \
|
||||
ccr-2vdh3abv-pub.cnc.bj.baidubce.com/device/paddle-npu:cann800-ubuntu20-npu-910b-base-$(uname -m)-gcc84 /bin/bash
|
||||
```
|
||||
|
||||
2. Install [PaddlePaddle](https://www.paddlepaddle.org.cn/install/quick?docurl=undefined) and [PaddleOCR](https://github.com/PaddlePaddle/PaddleOCR)
|
||||
|
||||
```bash
|
||||
python -m pip install paddlepaddle==3.2.0
|
||||
wget https://paddle-whl.bj.bcebos.com/stable/npu/paddle-custom-npu/paddle_custom_npu-3.2.0-cp310-cp310-linux_aarch64.whl
|
||||
pip install paddle_custom_npu-3.2.0-cp310-cp310-linux_aarch64.whl
|
||||
python -m pip install -U "paddleocr[doc-parser]"
|
||||
pip install safetensors
|
||||
```
|
||||
|
||||
:::{note}
|
||||
The OpenCV component may be missing:
|
||||
|
||||
```bash
|
||||
apt-get update
|
||||
apt-get install -y libgl1 libglib2.0-0
|
||||
```
|
||||
|
||||
CANN-8.0.0 does not support some versions of NumPy and OpenCV. It is recommended to install the specified versions.
|
||||
|
||||
```bash
|
||||
python -m pip install numpy==1.26.4
|
||||
python -m pip install opencv-python==3.4.18.65
|
||||
```
|
||||
|
||||
:::
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
:sync: atlas300
|
||||
|
||||
The Atlas 300I DUO supports only the OM model inference. For details about the process, see the guide provided in [ModelZoo](https://gitcode.com/Ascend/ModelZoo-PyTorch/tree/master/ACL_PyTorch/built-in/ocr/PP-DocLayoutV2).
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
#### 5.3.2 Using vLLM as the backend, combined with PP-DocLayoutV2 for offline inference
|
||||
|
||||
```python
|
||||
from paddleocr import PaddleOCRVL
|
||||
|
||||
doclayout_model_path = "/path/to/your/PP-DocLayoutV2/"
|
||||
|
||||
pipeline = PaddleOCRVL(vl_rec_backend="vllm-server",
|
||||
vl_rec_server_url="http://localhost:8000/v1",
|
||||
layout_detection_model_name="PP-DocLayoutV2",
|
||||
layout_detection_model_dir=doclayout_model_path,
|
||||
device="npu")
|
||||
|
||||
output = pipeline.predict("https://paddle-model-ecology.bj.bcebos.com/paddlex/imgs/demo_image/paddleocr_vl_demo.png")
|
||||
|
||||
for i, res in enumerate(output):
|
||||
res.save_to_json(save_path=f"output_{i}.json")
|
||||
res.save_to_markdown(save_path=f"output_{i}.md")
|
||||
```
|
||||
|
||||
## 6 Functional Verification
|
||||
|
||||
If your service starts successfully, you can see the info shown below:
|
||||
|
||||
```bash
|
||||
INFO: Started server process [87471]
|
||||
INFO: Waiting for application startup.
|
||||
INFO: Application startup complete.
|
||||
```
|
||||
|
||||
Once your server is started, you can use the OpenAI API client to make queries.
|
||||
|
||||
```python
|
||||
from openai import OpenAI
|
||||
|
||||
client = OpenAI(
|
||||
api_key="EMPTY",
|
||||
base_url="http://localhost:8000/v1",
|
||||
timeout=3600
|
||||
)
|
||||
|
||||
# Task-specific base prompts
|
||||
TASKS = {
|
||||
"ocr": "OCR:",
|
||||
"table": "Table Recognition:",
|
||||
"formula": "Formula Recognition:",
|
||||
"chart": "Chart Recognition:"
|
||||
}
|
||||
|
||||
messages = [
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{
|
||||
"type": "image_url",
|
||||
"image_url": {
|
||||
"url": "https://ofasys-multimodal-wlcb-3-toshanghai.oss-accelerate.aliyuncs.com/wpf272043/keepme/image/receipt.png"
|
||||
}
|
||||
},
|
||||
{
|
||||
"type": "text",
|
||||
"text": TASKS["ocr"]
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
|
||||
response = client.chat.completions.create(
|
||||
model="PaddleOCR-VL-0.9B",
|
||||
messages=messages,
|
||||
temperature=0.0,
|
||||
)
|
||||
print(f"Generated text: {response.choices[0].message.content}")
|
||||
```
|
||||
|
||||
Expected Result:
|
||||
|
||||
If you query the server successfully, you can see the info shown below (client):
|
||||
|
||||
```bash
|
||||
Generated text: CINNAMON SUGAR
|
||||
1 x 17,000
|
||||
17,000
|
||||
SUB TOTAL
|
||||
17,000
|
||||
GRAND TOTAL
|
||||
17,000
|
||||
CASH IDR
|
||||
20,000
|
||||
CHANGE DUE
|
||||
3,000
|
||||
```
|
||||
|
||||
## 7 Accuracy Evaluation
|
||||
|
||||
For the accuracy evaluation of PaddleOCR-VL, please refer to the official [ModelZoo](https://gitcode.com/Ascend/ModelZoo-PyTorch/tree/master/ACL_PyTorch/built-in/ocr/PP-DocLayoutV2) for the evaluation process and results.
|
||||
|
||||
## 8 Performance Evaluation
|
||||
|
||||
For the performance evaluation of PaddleOCR-VL, please refer to the official [ModelZoo](https://gitcode.com/Ascend/ModelZoo-PyTorch/tree/master/ACL_PyTorch/built-in/ocr/PP-DocLayoutV2) for the benchmark methodology and results.
|
||||
|
||||
## 9 Performance Tuning
|
||||
|
||||
### 9.1 Recommended Configurations
|
||||
|
||||
> **Note**: The following configurations are validated in specific test environments and are for reference only. The optimal configuration depends on factors such as maximum input/output length, precision requirements, and actual hardware specifications. It is recommended to refer to Section 9.2 for tuning based on actual conditions.
|
||||
|
||||
PaddleOCR-VL is a lightweight model that runs on a single NPU. The key tuning parameters differ between hardware platforms.
|
||||
|
||||
#### Table 1: Scenario Overview
|
||||
|
||||
| Scenario | Hardware | *Total NPUs | Weight Version | Key Considerations |
|
||||
|----------|----------|------------|---------------|-------------------|
|
||||
| High Throughput | A2 series | 1 | PaddleOCR-VL-0.9B | - |
|
||||
| High Throughput | Atlas 300I DUO | 1 | PaddleOCR-VL-0.9B | Graph compilation requires **CANN >= 9.0.0** |
|
||||
|
||||
> `*Total NPUs` indicates the total number of NPUs used across all nodes.
|
||||
|
||||
#### Table 2: Detailed Node Configuration
|
||||
|
||||
| Scenario | Configuration | NPUs | TP | DP | Max Model Len | Max Num Batched Tokens | Graph Compilation | dtype |
|
||||
|----------|-------------|------|----|----|---------------|------------------------|--------------------|-------|
|
||||
| High Throughput | A2 series / Single Machine | 1 | — | — | — | — | FULL_DECODE_ONLY | bfloat16 (default) |
|
||||
| High Throughput | Atlas 300I DUO / Single Machine | 1 | — | — | — | — | FULL_DECODE_ONLY; otherwise enforce-eager | float16 |
|
||||
|
||||
> For complete startup commands and parameter descriptions, please refer to the deployment examples in [Section 5.1](#51-single-node-online-deployment).
|
||||
|
||||
### 9.2 Tuning Guidelines
|
||||
|
||||
#### 9.2.1 General Tuning Reference
|
||||
|
||||
For performance tuning, please refer to the [Public Performance Tuning Documentation](../../developer_guide/performance_and_debug/optimization_and_tuning.md) for general tuning methods, including OS optimization (jemalloc, tcmalloc), `torch_npu` optimization (memory and scheduling), and CANN optimization.
|
||||
|
||||
Please refer to the [Feature Guide](../../user_guide/support_matrix/feature_matrix.md) for detailed feature descriptions.
|
||||
|
||||
## 10 FAQ
|
||||
|
||||
For common environment, installation, and general parameter issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html); this chapter only covers model-specific issues.
|
||||
|
||||
- **Q: What are the deployment requirements for Atlas 300I DUO?**
|
||||
|
||||
A: On Atlas 300I DUO, only `float16` dtype is supported. Graph compilation (`--compilation-config`) requires **CANN version >= 9.0.0**; if your CANN version is lower, use `--enforce-eager` instead.
|
||||
|
||||
- **Q: What should I do if I encounter dependency conflicts during installation on Atlas 300I DUO?**
|
||||
|
||||
A: Uninstall `triton` and `triton-ascend` before starting the service:
|
||||
|
||||
```bash
|
||||
pip uninstall -y triton triton-ascend
|
||||
```
|
||||
503
docs/source/tutorials/models/Qwen-VL-Dense.md
Normal file
503
docs/source/tutorials/models/Qwen-VL-Dense.md
Normal file
@@ -0,0 +1,503 @@
|
||||
# Qwen-VL-Dense(Qwen3-VL-8B/32B)
|
||||
|
||||
## 1 Introduction
|
||||
|
||||
The Qwen-VL(Vision-Language)series from Alibaba Cloud comprises a family of powerful Large Vision-Language Models (LVLMs) designed for comprehensive multimodal understanding. They accept images, text, and bounding boxes as input, and output text and detection boxes, enabling advanced functions like image detection, multi-modal dialogue, and multi-image reasoning.
|
||||
|
||||
This document will show the main verification steps of the model, including supported features, feature configuration, environment preparation, NPU deployment, accuracy and performance evaluation.
|
||||
|
||||
This tutorial uses the vLLM-Ascend `v0.11.0rc3-a3` version for demonstration, showcasing the `Qwen3-VL-8B-Instruct` model as an example for single NPU and multi-NPU deployment.
|
||||
|
||||
:::{note}
|
||||
For **Atlas inference products**, Qwen3-VL Dense requires vLLM-Ascend `v0.18.0` or later(for Ascend950DT, the model is supported from `vllm-ascend:v0.23.0rc1`). Do not use the demonstration version above on this hardware.
|
||||
:::
|
||||
|
||||
## 2 Supported Features
|
||||
|
||||
Refer to [Supported Features List](../../user_guide/support_matrix/supported_models.md) to get the model's supported feature matrix.
|
||||
|
||||
Refer to [Feature Guide](../../user_guide/feature_guide/index.md) to get the feature's configuration.
|
||||
|
||||
## 3 Prerequisites
|
||||
|
||||
### 3.1 Model Weight
|
||||
|
||||
Requires 1 card on Atlas 800I A2 (64G × 8), Atlas 800 A3 (64G × 16), or Atlas 300I DUO:
|
||||
|
||||
- `Qwen3-VL-8B-Instruct`: [Download model weight](https://modelscope.cn/models/Qwen/Qwen3-VL-8B-Instruct)
|
||||
|
||||
Requires 1 card on Ascend950DT series (96G × 8) node.
|
||||
|
||||
- `Qwen3-VL-8B-Instruct-w8a8`(Quantized version): [Download model weight](https://modelscope.cn/models/Eco-Tech/Qwen3-VL-8B-Instruct-w8a8-mxfp8)
|
||||
|
||||
Requires 2 cards on Atlas 800I A2 (64G × 8), Atlas 800 A3 (64G × 16), or Atlas inference products:
|
||||
|
||||
- `Qwen3-VL-32B-Instruct`: [Download model weight](https://www.modelscope.cn/models/Qwen/Qwen3-VL-32B-Instruct)
|
||||
|
||||
Requires 1 card on Ascend950DT series (96G × 8) node.
|
||||
|
||||
- `Qwen3-VL-32B-Instruct-w8a8`(Quantized version): [Download model weight](https://modelscope.cn/models/Eco-Tech/Qwen3-VL-32B-Instruct-w8a8-mxfp8)
|
||||
|
||||
It is recommended to download the model weight to the shared directory of multiple nodes, such as `/root/.cache/`.
|
||||
|
||||
## 4 Installation
|
||||
|
||||
### 4.1 Docker Image Installation
|
||||
|
||||
Select an image based on your machine type and start the docker image on your node, refer to [using docker](../../installation.md#set-up-using-docker).
|
||||
|
||||
For Atlas 300I DUO, use `vllm-ascend:nightly-releases-v0.23.0-310p` (or a later `-310p` image).
|
||||
|
||||
::::::{tab-set}
|
||||
:sync-group: hardware
|
||||
|
||||
::::{tab-item} Ascend950DT series
|
||||
:sync: 950dt
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
Start the docker image on your each node.
|
||||
|
||||
```shell
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-#TODO
|
||||
export NAME=vllm-ascend
|
||||
|
||||
docker run --rm \
|
||||
--name $NAME \
|
||||
--net=host \
|
||||
--shm-size=1g \
|
||||
--privileged=true \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/hisi_hdc \
|
||||
--device /dev/ummu \
|
||||
--device /dev/uburma \
|
||||
-v /usr/local/Ascend/driver:/usr/local/Ascend/driver \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /etc/hccl_rootinfo.json:/etc/hccl_rootinfo.json \
|
||||
-v /etc/hixlep/:/etc/hixlep/ \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-v /usr/local/sbin:/usr/local/sbin \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/sbin/npu-smi:/usr/local/sbin/npu-smi \
|
||||
-v /usr/bin/urma_admin:/usr/bin/urma_admin \
|
||||
-v /lib/route.conf:/lib/route.conf \
|
||||
-v /usr/lib64:/usr/lib64 \
|
||||
-itd $IMAGE bash
|
||||
```
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} A2 / A3 series
|
||||
:sync: a2a3
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
# Update the vllm-ascend image
|
||||
# A2: quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
# A3: quay.io/ascend/vllm-ascend:|vllm_ascend_version|-a3
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-p 8000:8000 \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
:sync: atlas
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
# Use the vllm-ascend image
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-310p
|
||||
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=10g \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-p 8000:8000 \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
::::::
|
||||
|
||||
**Installation Verification:**
|
||||
|
||||
After starting the container, run the following command to verify the installation:
|
||||
|
||||
```bash
|
||||
docker ps | grep vllm-ascend
|
||||
```
|
||||
|
||||
Expected result: The container is listed with status `Up`. You can also verify the vllm-ascend version inside the container:
|
||||
|
||||
```bash
|
||||
pip show vllm-ascend
|
||||
```
|
||||
|
||||
Expected result: The version information is displayed, matching the pulled image version.
|
||||
|
||||
### 4.2 Source Code Installation
|
||||
|
||||
If you prefer not to use the Docker image, you can build from source. Install vLLM from source first:
|
||||
|
||||
1. Clone and install vLLM:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/vllm-project/vllm.git
|
||||
cd vllm
|
||||
pip install -e .
|
||||
```
|
||||
|
||||
2. Clone and install the vLLM-Ascend repository:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/vllm-project/vllm-ascend.git
|
||||
cd vllm-ascend
|
||||
pip install -e .
|
||||
```
|
||||
|
||||
:::{note}
|
||||
Atlas 300I DUO does not support `triton` or `triton-ascend`. Source installation may pull them in automatically; uninstall them manually before running:
|
||||
|
||||
```bash
|
||||
pip uninstall -y triton-ascend triton
|
||||
```
|
||||
|
||||
:::
|
||||
|
||||
**Installation Verification:**
|
||||
|
||||
```bash
|
||||
pip show vllm vllm-ascend
|
||||
```
|
||||
|
||||
Expected result: The version information for both packages is displayed, confirming a successful installation.
|
||||
|
||||
:::{note}
|
||||
If deploying a multi-node environment, set up the environment on each node.
|
||||
:::
|
||||
|
||||
For more details, please refer to the [Installation Guide](../../installation.md).
|
||||
|
||||
## 5 Online Service Deployment
|
||||
|
||||
### 5.1 Single-Node Online Deployment
|
||||
|
||||
Run docker container to start the vLLM server on single-NPU:
|
||||
|
||||
::::::{tab-set}
|
||||
:sync-group: hardware
|
||||
|
||||
::::{tab-item} Ascend950DT series
|
||||
:sync: 950dt
|
||||
|
||||
```bash
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export ASCEND_RT_VISIBLE_DEVICES=$1
|
||||
|
||||
vllm serve Qwen/Qwen3-VL-8B-Instruct \
|
||||
--host 0.0.0.0 \
|
||||
--port $2 \
|
||||
--quantization ascend \
|
||||
--served-model-name qwen3vl \
|
||||
--no-enable-prefix-caching \
|
||||
--data-parallel-size $3 \
|
||||
--tensor-parallel-size $4 \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 128 \
|
||||
--max-model-len 32768 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--gpu-memory-utilization 0.91 \
|
||||
--async-scheduling \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY", "cudagraph_capture_sizes": [1,2,4,8,16,32]}' \
|
||||
--mm-processor-cache-gb 0
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} A2 / A3 series
|
||||
:sync: a2a3
|
||||
|
||||
```bash
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export ASCEND_RT_VISIBLE_DEVICES=$1
|
||||
|
||||
vllm serve Qwen/Qwen3-VL-8B-Instruct \
|
||||
--host 0.0.0.0 \
|
||||
--port $2 \
|
||||
--dtype bfloat16 \
|
||||
--served-model-name qwen3vl \
|
||||
--no-enable-prefix-caching \
|
||||
--data-parallel-size $3 \
|
||||
--tensor-parallel-size $4 \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 128 \
|
||||
--max-model-len 32768 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--gpu-memory-utilization 0.91 \
|
||||
--async-scheduling \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY", "cudagraph_capture_sizes": [1,2,4,8,16,32]}' \
|
||||
--mm-processor-cache-gb 0
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
:sync: atlas
|
||||
|
||||
```bash
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export ASCEND_RT_VISIBLE_DEVICES=$1
|
||||
|
||||
vllm serve Qwen/Qwen3-VL-8B-Instruct \
|
||||
--dtype float16 \
|
||||
--max_model_len 16384 \
|
||||
--host 0.0.0.0 \
|
||||
--port $2 \
|
||||
--dtype bfloat16 \
|
||||
--served-model-name qwen3vl \
|
||||
--no-enable-prefix-caching \
|
||||
--data-parallel-size $3 \
|
||||
--tensor-parallel-size $4 \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 128 \
|
||||
--max-model-len 32768 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--gpu-memory-utilization 0.91 \
|
||||
--async-scheduling \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY", "cudagraph_capture_sizes": [1,2,4,8,16,32]}' \
|
||||
--additional-config '{"ascend_compilation_config": {"enable_npugraph_ex":false}}' \
|
||||
--mm-processor-cache-gb 0
|
||||
```
|
||||
|
||||
:::{note}
|
||||
On Atlas 300I DUO:
|
||||
|
||||
- Only `float16` dtype is supported.
|
||||
- Graph compilation (`--compilation-config`) requires **CANN version >= 9.0.0**. If your CANN version is lower, replace `--compilation-config` with `--enforce-eager`.
|
||||
- `--additional-config` with `"ascend_compilation_config": {"enable_npugraph_ex": false}` is required because `enable_npugraph_ex` is not supported on Atlas 300I DUO.
|
||||
:::
|
||||
|
||||
::::
|
||||
::::::
|
||||
|
||||
Key Parameter Descriptions:
|
||||
|
||||
- Add the `--max_model_len` option to avoid the ValueError that occurs when the Qwen3-VL-8B-Instruct model's max-seq-len (256000) exceeds the maximum number of tokens that can be stored in KV cache. This will differ with different NPU series based on the on-chip memory size. Please modify the value according to a suitable value for your NPU series.
|
||||
|
||||
If your service starts successfully, you can see the info shown below:
|
||||
|
||||
```bash
|
||||
INFO: Started server process [2736]
|
||||
INFO: Waiting for application startup.
|
||||
INFO: Application startup complete.
|
||||
```
|
||||
|
||||
## 6 Functional Verification
|
||||
|
||||
Once your server is started, you can query the model with input prompts:
|
||||
|
||||
```bash
|
||||
curl http://localhost:8000/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "Qwen/Qwen3-VL-8B-Instruct",
|
||||
"messages": [
|
||||
{"role": "system", "content": "You are a helpful assistant."},
|
||||
{"role": "user", "content": [
|
||||
{"type": "image_url", "image_url": {"url": "https://modelscope.oss-cn-beijing.aliyuncs.com/resource/qwen.png"}},
|
||||
{"type": "text", "text": "What is the text in the illustration?"}
|
||||
]}
|
||||
]
|
||||
}'
|
||||
```
|
||||
|
||||
Expected Result:
|
||||
|
||||
The service returns HTTP 200 OK.
|
||||
|
||||
```bash
|
||||
{"id":"chatcmpl-d3270d4a16cb4b98936f71ee3016451f","object":"chat.completion","created":1764924127,"model":"Qwen/Qwen3-VL-8B-Instruct","choices":[{"index":0,"message":{"role":"assistant","content":"The text in the illustration is: **TONGYI Qwen**","refusal":null,"annotations":null,"audio":null,"function_call":null,"tool_calls":[],"reasoning_content":null},"logprobs":null,"finish_reason":"stop","stop_reason":null,"token_ids":null}],"service_tier":null,"system_fingerprint":null,"usage":{"prompt_tokens":107,"total_tokens":123,"completion_tokens":16,"prompt_tokens_details":null},"prompt_logprobs":null,"prompt_token_ids":null,"kv_transfer_params":null}
|
||||
```
|
||||
|
||||
## 7 Accuracy Evaluation
|
||||
|
||||
The accuracy of some models is already within our CI monitoring scope, including:
|
||||
|
||||
- `Qwen3-VL-8B-Instruct`
|
||||
|
||||
::::::{tab-set}
|
||||
:sync-group: hardware
|
||||
|
||||
::::{tab-item} A2 / A3 series
|
||||
:sync: a2a3
|
||||
|
||||
**Using Language Model Evaluation Harness**
|
||||
|
||||
As an example, take the `mmmu_val` dataset as a test dataset, and run accuracy evaluation of `Qwen3-VL-8B-Instruct` in offline mode.
|
||||
|
||||
1. Refer to [Using lm_eval](../../developer_guide/evaluation/using_lm_eval.md) for more details on `lm_eval` installation.
|
||||
|
||||
```shell
|
||||
pip install lm_eval
|
||||
```
|
||||
|
||||
2. Run `lm_eval` to execute the accuracy evaluation.
|
||||
|
||||
```shell
|
||||
lm_eval \
|
||||
--model vllm-vlm \
|
||||
--model_args pretrained=Qwen/Qwen3-VL-8B-Instruct,max_model_len=8192,gpu_memory_utilization=0.7 \
|
||||
--tasks mmmu_val \
|
||||
--batch_size 32 \
|
||||
--apply_chat_template \
|
||||
--trust_remote_code \
|
||||
--output_path ./results
|
||||
```
|
||||
|
||||
3. After execution, you can get the result, here is the result of `Qwen3-VL-8B-Instruct` in `vllm-ascend:0.11.0rc3` for reference only.
|
||||
|
||||
| Tasks | Value | Stderr |
|
||||
| -------- | ------ | ------ |
|
||||
| mmmu_val | 0.5389 | 0.0159 |
|
||||
|
||||
::::
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
:sync: atlas
|
||||
|
||||
**Using AISBench**
|
||||
|
||||
Take the `text_vqa` dataset as an example, and run accuracy evaluation of `Qwen3-VL-8B-Instruct`.
|
||||
|
||||
1. Refer to [Using AISBench](../../developer_guide/evaluation/using_ais_bench.md) for installation, dataset download, and configuration details.
|
||||
|
||||
2. Run `ais_bench` to execute the accuracy evaluation.
|
||||
|
||||
```shell
|
||||
ais_bench --models vllm_api_general_chat --datasets textvqa_gen_base64 --mode all --debug
|
||||
```
|
||||
|
||||
3. After execution, you can get the result, here is the result of `Qwen3-VL-8B-Instruct` in `vllm-ascend:0.23.0rc1` for reference only.
|
||||
|
||||
| dataset | metric | mode | vllm-api-general-chat |
|
||||
| -------- | -------- | ---- | --------------------- |
|
||||
| text_vqa | accuracy | gen | 80.57 |
|
||||
|
||||
::::
|
||||
::::::
|
||||
|
||||
## 8 Performance Evaluation
|
||||
|
||||
### Using vLLM Benchmark
|
||||
|
||||
Refer to [vLLM Benchmark](https://docs.vllm.ai/en/latest/benchmarking/) for more details.
|
||||
|
||||
There are three `vllm bench` subcommands:
|
||||
|
||||
- `latency`: Benchmark the latency of a single batch of requests.
|
||||
- `serve`: Benchmark the online serving throughput.
|
||||
- `throughput`: Benchmark offline inference throughput.
|
||||
|
||||
The performance evaluation must be conducted in an online mode. Take the `serve` as an example. Run the code as follows.
|
||||
|
||||
```shell
|
||||
vllm bench serve --model Qwen/Qwen3-VL-8B-Instruct --dataset-name random --random-input 200 --num-prompts 200 --request-rate 1 --save-result --result-dir ./
|
||||
```
|
||||
|
||||
After several minutes, you can get the performance evaluation result.
|
||||
|
||||
## 9 Performance Tuning
|
||||
|
||||
### 9.1 Recommended Configurations
|
||||
|
||||
> **Note**: The following configurations are validated in specific test environments and are for reference only. The optimal configuration depends on factors such as maximum input/output length, prefix cache hit rate, precision requirements, and deployment machine ratios. It is recommended to refer to Section 9.2 for tuning based on actual conditions.
|
||||
|
||||
#### Table 1: Scenario Overview
|
||||
|
||||
|Scenario|Deployment Mode|*Total NPUs|Weight Version|Key Considerations|
|
||||
|--------|---------------|-----------|--------------|------------------|
|
||||
|High Throughput<br>(16k context)|Single-Node Mixed|1 (A3)|Qwen3-VL-8B-Instruct|Use tp2 for high-resolution text inputs|
|
||||
|Long Context<br>(128k, no prefix cache)|Single-Node Mixed|1 (A3)|Qwen3-VL-8B-Instruct|tp2 for high-resolution text inputs|
|
||||
|Long Context<br>(128k, with prefix cache)|Single-Node Mixed|1 (A3)|Qwen3-VL-8B-Instruct|tp2 for high-resolution text inputs|
|
||||
|Multimodal<br>(1080p)|Single-Node Mixed|1 (A3)|Qwen3-VL-8B-Instruct|tp2 for high-resolution visual inputs|
|
||||
|
||||
> `*Total NPUs` indicates the total number of NPUs used across all nodes. 1 node = 1 Atlas 800 A3 server (64G × 16 NPUs).
|
||||
|
||||
#### Table 2: Detailed Node Configuration
|
||||
|
||||
|Scenario|Configuration|NPUs|TP|DP|Max Model Len|MTP Speculation Num|Weight Version|
|
||||
|--------|-------------|-----|--|--|-------------------|--------------------|---|
|
||||
|High Throughput / Low Latency (16k)|Server / Single Machine|1|1|1|~16k|3|Qwen3-VL-8B-Instruct|
|
||||
|Long Context (128k, no cache)|Server / Single Machine|1|1|1|128k|3|Qwen3-VL-8B-Instruct|
|
||||
|Long Context (128k, with cache)|Server / Single Machine|1|1|1|128k|3|Qwen3-VL-8B-Instruct|
|
||||
|Multimodal (1080p)|Server / Single Machine|1|1|1|~16k|3|Qwen3-VL-8B-Instruct|
|
||||
|
||||
> For complete startup commands and parameter descriptions, please refer to the deployment examples in [Chapter 5](#5-online-service-deployment).
|
||||
|
||||
**Notice:**
|
||||
`max-model-len` and `max-num-seqs` need to be set according to the actual usage scenario. For other settings, please refer to the **[Deployment](#5-online-service-deployment)** chapter.
|
||||
|
||||
### 9.2 Tuning Guidelines
|
||||
|
||||
#### 9.2.1 General Tuning Reference
|
||||
|
||||
Please refer to the [Public Performance Tuning Documentation](../../developer_guide/performance_and_debug/optimization_and_tuning.md) for tuning methods.
|
||||
|
||||
Please refer to the [Feature Matrix](../../user_guide/support_matrix/feature_matrix.md) for detailed feature descriptions.
|
||||
|
||||
## 10 FAQ
|
||||
|
||||
For common environment, installation, and general parameter issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html).
|
||||
143
docs/source/tutorials/models/Qwen2.5-Math-RM-72B.md
Normal file
143
docs/source/tutorials/models/Qwen2.5-Math-RM-72B.md
Normal file
@@ -0,0 +1,143 @@
|
||||
# Qwen2.5-Math-RM-72B
|
||||
|
||||
## Introduction
|
||||
|
||||
Qwen2.5-Math-RM-72B is a 72-billion parameter reward model designed for mathematical reasoning and evaluation. It is part of Alibaba Cloud's Qwen 2.5 series, specifically optimized for scoring and ranking mathematical problem solutions. The model supports a maximum context window of 128k tokens and delivers enhanced capabilities in mathematical computation, step-by-step reasoning evaluation, and solution quality assessment.
|
||||
|
||||
This document provides a detailed workflow for the complete deployment and verification of the model, including supported features, environment preparation, single-node deployment, functional verification, and performance evaluation.
|
||||
|
||||
The `Qwen2.5-Math-RM-72B` model is supported since `vllm-ascend:v0.9.0`.
|
||||
|
||||
## Supported Features
|
||||
|
||||
Refer to [supported features](../../user_guide/support_matrix/supported_models.md) to get the model's supported feature matrix.
|
||||
|
||||
Refer to [feature guide](../../user_guide/feature_guide/index.md) to get the feature's configuration.
|
||||
|
||||
## Environment Preparation
|
||||
|
||||
### Model Weight
|
||||
|
||||
- `Qwen2.5-Math-RM-72B` (BF16 version):
|
||||
- With CPU offloading: requires at least 1 Atlas 910B4 (32GB × 1) card or higher
|
||||
- Without CPU offloading: requires at least 4 Atlas 910B4 (32GB × 4) cards or higher
|
||||
[Download model weight](https://www.modelscope.cn/models/Qwen/Qwen2.5-Math-RM-72B)
|
||||
|
||||
It is recommended to download the model weights to a local directory (e.g., `./Qwen2.5-Math-RM-72B/`) for quick access during deployment.
|
||||
|
||||
### Installation
|
||||
|
||||
You can use our official docker image to run `Qwen2.5-Math-RM-72B` directly.
|
||||
|
||||
These versions support multi-NPU deployment, allowing the model to utilize all available NPU devices (e.g., 4 NPUs) for improved performance.
|
||||
|
||||
Select an image based on your machine type and start the docker image on your node, refer to [using docker](../../installation.md#set-up-using-docker).
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
docker run --rm \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
## Deployment
|
||||
|
||||
### Single-node Deployment
|
||||
|
||||
Qwen2.5-Math-RM-72B supports single-node single-card deployment on the 910B4 platform. Follow these steps to start the inference service:
|
||||
|
||||
1. Prepare model weights: Ensure the downloaded model weights are stored in the `./Qwen2.5-Math-RM-72B/` directory.
|
||||
2. Create and execute the deployment script (save as `deploy.sh`):
|
||||
|
||||
```shell
|
||||
#!/bin/sh
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0
|
||||
export MODEL_PATH="Qwen/Qwen2.5-Math-RM-72B"
|
||||
|
||||
vllm serve ${MODEL_PATH} \
|
||||
--host 0.0.0.0 \
|
||||
--port 8000 \
|
||||
--served-model-name qwen2.5-math-rm-72b \
|
||||
--trust-remote-code \
|
||||
--max-model-len 32768 \
|
||||
--task reward
|
||||
```
|
||||
|
||||
:::{note}
|
||||
The `--task reward` parameter is required to run the model in reward model mode for scoring mathematical solutions.
|
||||
:::
|
||||
|
||||
## Functional Verification
|
||||
|
||||
After starting the service, verify functionality using a `curl` request:
|
||||
|
||||
```shell
|
||||
curl http://localhost:8000/v1/reward \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "qwen2.5-math-rm-72b",
|
||||
"messages": [
|
||||
{"role": "system", "content": "You are a helpful math assistant."},
|
||||
{"role": "user", "content": "What is 2+2?"},
|
||||
{"role": "assistant", "content": "2+2 equals 4."}
|
||||
]
|
||||
}'
|
||||
```
|
||||
|
||||
A valid response (e.g., `{"reward_score": 1.69}`) indicates successful deployment.
|
||||
|
||||
### Batch Reward Scoring
|
||||
|
||||
You can also score multiple responses for comparison:
|
||||
|
||||
```shell
|
||||
curl http://localhost:8000/v1/reward/batch \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "qwen2.5-math-rm-72b",
|
||||
"conversations": [
|
||||
[
|
||||
{"role": "system", "content": "You are a helpful math assistant."},
|
||||
{"role": "user", "content": "What is 2+2?"},
|
||||
{"role": "assistant", "content": "2+2 equals 4."}
|
||||
],
|
||||
[
|
||||
{"role": "system", "content": "You are a helpful math assistant."},
|
||||
{"role": "user", "content": "What is 2+2?"},
|
||||
{"role": "assistant", "content": "2+2 equals 5."}
|
||||
]
|
||||
],
|
||||
"batch_rewards": [
|
||||
{
|
||||
"index": 0,
|
||||
"score": 9.85,
|
||||
"reasoning": "The answer is mathematically correct and concise."
|
||||
},
|
||||
{
|
||||
"index": 1,
|
||||
"score": 1.20,
|
||||
"reasoning": "The answer contains a factual mathematical error (2+2 is not 5)."
|
||||
}
|
||||
]
|
||||
}'
|
||||
```
|
||||
|
||||
## References
|
||||
|
||||
- [Qwen2.5-Math Technical Report](https://arxiv.org/abs/2409.12122)
|
||||
- [HuggingFace Model Card](https://huggingface.co/Qwen/Qwen2.5-Math-RM-72B)
|
||||
- [vLLM Documentation](https://docs.vllm.ai/)
|
||||
926
docs/source/tutorials/models/Qwen3-235B-A22B.md
Normal file
926
docs/source/tutorials/models/Qwen3-235B-A22B.md
Normal file
@@ -0,0 +1,926 @@
|
||||
# Qwen3-235B-A22B
|
||||
|
||||
## 1 Introduction
|
||||
|
||||
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support. Qwen3-235B-A22B is the largest MoE variant, featuring 235B total parameters with 22B activated per token.
|
||||
|
||||
This document will demonstrate the main validation steps for Qwen3-235B-A22B in the vLLM-Ascend environment, including supported features, environment preparation, single-node and multi-node deployment, accuracy and performance evaluation.
|
||||
|
||||
The Qwen3-235B-A22B model is first supported in **v0.8.4rc2**. This document is validated and written based on **vLLM-Ascend v0.21.0**. All **v0.21.0 and later versions** can run stably. To use the latest features, it is recommended to use the latest release candidate or official version.
|
||||
|
||||
## 2 Supported Features
|
||||
|
||||
Please refer to the [Supported Features List](../../user_guide/support_matrix/supported_models.md) for the model support matrix.
|
||||
|
||||
Please refer to the [Feature Guide](../../user_guide/feature_guide/index.md) for feature configuration information.
|
||||
|
||||
## 3 Prerequisites
|
||||
|
||||
### 3.1 Model Weight
|
||||
|
||||
The following model variants are available. It is recommended to download the model weight to a shared directory accessible to all nodes.
|
||||
|
||||
**BF16 Version:**
|
||||
|
||||
| Model | Hardware Requirement | Download |
|
||||
|-------|---------------------|----------|
|
||||
| Qwen3-235B-A22B (BF16) | 1 Atlas 800I A3 (64GB × 16), 1 Atlas 800I A2 (64GB × 8), 2 Atlas 800I A2 (32GB × 8)| [Download](https://www.modelscope.cn/models/Qwen/Qwen3-235B-A22B) |
|
||||
|
||||
**Quantized Version (Pre-converted):**
|
||||
|
||||
| Model | Quantization | Hardware Requirement | Download |
|
||||
|-------|-------------|---------------------|----------|
|
||||
| Qwen3-235B-A22B-W8A8 | W8A8 | 1 Atlas 800I A3 (64GB × 16), 1 Atlas 800I A2 (64GB × 8), 2 Atlas 800I A2 (32GB × 8)| [Download](https://modelers.cn/models/Modelers_Park/Qwen3-235B-A22B-w8a8) |
|
||||
|
||||
These are the recommended numbers of cards, which can be adjusted according to the actual situation.
|
||||
|
||||
### 3.2 Model Quantization
|
||||
|
||||
**Install msmodelslim:**
|
||||
|
||||
```shell
|
||||
# 1. Clone the msmodelslim repository.
|
||||
git clone https://gitcode.com/Ascend/msmodelslim.git
|
||||
|
||||
# 2. Enter the msmodelslim directory and run the installation script.
|
||||
cd msmodelslim
|
||||
bash install.sh
|
||||
|
||||
# The following message indicates that msmodelslim has been installed successfully.
|
||||
Successfully installed msmodelslim-{version}
|
||||
```
|
||||
|
||||
**Run quantization:**
|
||||
|
||||
```shell
|
||||
cd example/Qwen3-MOE
|
||||
# Run the following command to quantize the model.
|
||||
python3 quant_qwen_moe_w8a8.py --model_path /path/to/your/Qwen3-235B-A22B \
|
||||
--save_path /path/to/your/Qwen3-235B-A22B-W8A8-rot \
|
||||
--anti_dataset ../common/qwen3-moe_anti_prompt_50.json \
|
||||
--calib_dataset ../common/qwen3-moe_calib_prompt_50.json \
|
||||
--trust_remote_code True \
|
||||
--rot
|
||||
```
|
||||
|
||||
### 3.3 Verify Multi-node Communication
|
||||
|
||||
If you need to deploy a multi-node environment, verify the multi-node communication according to [Verify Multi-node Communication Environment](../../installation.md#verify-multi-node-communication).
|
||||
|
||||
## 4 Installation
|
||||
|
||||
### 4.1 Docker Image Installation
|
||||
|
||||
You can use the official all-in-one Docker image for Qwen3 MoE models.
|
||||
|
||||
**Docker Pull:**
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
docker pull quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
```
|
||||
|
||||
**Docker Run:**
|
||||
|
||||
Start the docker image on each node.
|
||||
|
||||
:::::{tab-set}
|
||||
::::{tab-item} A3 series
|
||||
:sync: A3
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-a3
|
||||
|
||||
docker run --rm \
|
||||
--name vllm-ascend-env \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci8 \
|
||||
--device /dev/davinci9 \
|
||||
--device /dev/davinci10 \
|
||||
--device /dev/davinci11 \
|
||||
--device /dev/davinci12 \
|
||||
--device /dev/davinci13 \
|
||||
--device /dev/davinci14 \
|
||||
--device /dev/davinci15 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
:::{note}
|
||||
A3 has 8 NPUs with dual-die design (16 chips total: `/dev/davinci[0-15]`).
|
||||
If you are on a shared machine, map only the chips you need (e.g., `/dev/davinci[0-7]` for NPU 0-3).
|
||||
:::
|
||||
|
||||
::::
|
||||
::::{tab-item} A2 series
|
||||
:sync: A2
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
|
||||
docker run --rm \
|
||||
--name vllm-ascend-env \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
The default workdir is `/workspace`. vLLM and vLLM-Ascend are installed as Python packages in site-packages.
|
||||
|
||||
Installation Verification:
|
||||
After starting the container, run the following command to verify the installation:
|
||||
|
||||
```bash
|
||||
docker ps | grep vllm-ascend-env
|
||||
```
|
||||
|
||||
Expected result: The container is listed with status Up. You can also verify the vllm-ascend version inside the container:
|
||||
|
||||
```bash
|
||||
pip show vllm-ascend
|
||||
```
|
||||
|
||||
Expected result: The version information is displayed, matching the pulled image version.
|
||||
|
||||
### 4.2 Source Code Installation
|
||||
|
||||
If you prefer to build from source instead of using the Docker image, install vLLM-Ascend following the [Installation Guide](../../installation.md).
|
||||
|
||||
To verify the source installation:
|
||||
|
||||
```bash
|
||||
pip show vllm-ascend
|
||||
```
|
||||
|
||||
Expected result: The version information is displayed, confirming a successful installation.
|
||||
|
||||
:::{note}
|
||||
If deploying a multi-node environment, set up the environment on each node.
|
||||
:::
|
||||
|
||||
## 5 Online Service Deployment
|
||||
|
||||
### 5.1 Single-Node Online Deployment
|
||||
|
||||
Single-node deployment completes both Prefill and Decode within the same node, suitable for development, testing, and small-to-medium scale inference scenarios.
|
||||
|
||||
**Start the server:**
|
||||
> The following command is an example configuration. Adjust the parameters based on your actual scenario.
|
||||
|
||||
Atlas 800I A2/A3:
|
||||
|
||||
```shell
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3,4,5,6,7
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export HCCL_BUFFSIZE=512
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
|
||||
vllm serve your_model_path \
|
||||
--host <host_ip> \
|
||||
--port <port> \
|
||||
--tensor-parallel-size 8 \
|
||||
--data-parallel-size 1 \
|
||||
--seed 1024 \
|
||||
--quantization ascend \
|
||||
--served-model-name qwen3 \
|
||||
--max-num-seqs 32 \
|
||||
--max-model-len 131072 \
|
||||
--max-num-batched-tokens 8096 \
|
||||
--enable-expert-parallel \
|
||||
--trust-remote-code \
|
||||
--gpu-memory-utilization 0.95 \
|
||||
--hf-overrides '{"rope_parameters": {"rope_type":"yarn","rope_theta":1000000,"factor":4,"original_max_position_embeddings":32768}}' \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"enable_flashcomm1": true}' \
|
||||
--async-scheduling
|
||||
```
|
||||
|
||||
:::{note}
|
||||
|
||||
- [vLLM Serving Arguments documentation](https://docs.vllm.ai/en/latest/cli/serve/#arguments) — Additional parameter details for vLLM serve commands.
|
||||
- [Environment Variables](../../user_guide/configuration/env_vars.md) — Ascend-specific environment variables (`HCCL_*`, etc.).
|
||||
:::
|
||||
|
||||
**Service Verification:**
|
||||
|
||||
If the service starts successfully, the following startup log will be displayed:
|
||||
|
||||
```text
|
||||
(APIServer pid=<pid>) INFO: Started server process [<pid>]
|
||||
(APIServer pid=<pid>) INFO: Waiting for application startup.
|
||||
(APIServer pid=<pid>) INFO: Application startup complete.
|
||||
```
|
||||
|
||||
### 5.2 Multi-Node PD Separation Deployment
|
||||
|
||||
PD (Prefill-Decode) separation splits the Prefill and Decode phases across different nodes for better throughput. The following example shows the parameter configuration for a three-node A3 PD disaggregation scenario (one Prefill node + two Decode nodes):
|
||||
|
||||
For the detailed deployment guide, please refer to [Prefill-Decode Disaggregation Mooncake Verification](../features/pd_disaggregation_mooncake_multi_node.md).
|
||||
|
||||
**Hardware**: 3 × Atlas 800 A3 (64GB × 16), one for Prefill, two for Decode.
|
||||
|
||||
First, prepare `launch_online_dp.py` on each node:
|
||||
|
||||
```python
|
||||
import argparse
|
||||
import multiprocessing
|
||||
import os
|
||||
import subprocess
|
||||
import sys
|
||||
|
||||
def parse_args():
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument("--dp-size", type=int, required=True, help="Data parallel size.")
|
||||
parser.add_argument("--tp-size", type=int, default=1, help="Tensor parallel size.")
|
||||
parser.add_argument("--dp-size-local", type=int, default=-1, help="Local data parallel size.")
|
||||
parser.add_argument("--dp-rank-start", type=int, default=0, help="Starting rank for data parallel.")
|
||||
parser.add_argument("--dp-address", type=str, required=True, help="IP address for data parallel master node.")
|
||||
parser.add_argument("--dp-rpc-port", type=str, default=12345, help="Port for data parallel master node.")
|
||||
parser.add_argument("--vllm-start-port", type=int, default=9000, help="Starting port for the engine.")
|
||||
return parser.parse_args()
|
||||
|
||||
args = parse_args()
|
||||
dp_size = args.dp_size
|
||||
tp_size = args.tp_size
|
||||
dp_size_local = args.dp_size_local
|
||||
if dp_size_local == -1:
|
||||
dp_size_local = dp_size
|
||||
dp_rank_start = args.dp_rank_start
|
||||
dp_address = args.dp_address
|
||||
dp_rpc_port = args.dp_rpc_port
|
||||
vllm_start_port = args.vllm_start_port
|
||||
|
||||
def run_command(visible_devices, dp_rank, vllm_engine_port):
|
||||
command = [
|
||||
"bash",
|
||||
"./run_dp_template.sh",
|
||||
visible_devices,
|
||||
str(vllm_engine_port),
|
||||
str(dp_size),
|
||||
str(dp_rank),
|
||||
dp_address,
|
||||
dp_rpc_port,
|
||||
str(tp_size),
|
||||
]
|
||||
subprocess.run(command, check=True)
|
||||
|
||||
if __name__ == "__main__":
|
||||
template_path = "./run_dp_template.sh"
|
||||
if not os.path.exists(template_path):
|
||||
print(f"Template file {template_path} does not exist.")
|
||||
sys.exit(1)
|
||||
|
||||
processes = []
|
||||
num_cards = dp_size_local * tp_size
|
||||
for i in range(dp_size_local):
|
||||
dp_rank = dp_rank_start + i
|
||||
vllm_engine_port = vllm_start_port + i
|
||||
visible_devices = ",".join(str(x) for x in range(i * tp_size, (i + 1) * tp_size))
|
||||
process = multiprocessing.Process(target=run_command, args=(visible_devices, dp_rank, vllm_engine_port))
|
||||
processes.append(process)
|
||||
process.start()
|
||||
|
||||
for process in processes:
|
||||
process.join()
|
||||
```
|
||||
|
||||
Then prepare `run_dp_template.sh` on each node.
|
||||
|
||||
**Prefill node** (set `nic_name` and `local_ip` to your own):
|
||||
|
||||
```bash
|
||||
nic_name="<your_nic_name>"
|
||||
local_ip="<your_ip>"
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
|
||||
export HCCL_BUFFSIZE=512
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
|
||||
export OMP_NUM_THREADS=1
|
||||
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
|
||||
sysctl -w vm.swappiness=0
|
||||
sysctl -w kernel.numa_balancing=0
|
||||
sysctl kernel.sched_migration_cost_ns=50000
|
||||
export LD_PRELOAD=/usr/lib/aarch64-linux-gnu/libjemalloc.so.2:$LD_PRELOAD
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export LD_LIBRARY_PATH=/usr/local/Ascend/ascend-toolkit/latest/python/site-packages/mooncake:$LD_LIBRARY_PATH
|
||||
|
||||
export ASCEND_RT_VISIBLE_DEVICES=$1
|
||||
|
||||
vllm serve "/data/weights/Qwen3-235B-A22B-w8a8-rot" \
|
||||
--host 0.0.0.0 \
|
||||
--port $2 \
|
||||
--data-parallel-size $3 \
|
||||
--data-parallel-rank $4 \
|
||||
--data-parallel-address $5 \
|
||||
--data-parallel-rpc-port $6 \
|
||||
--tensor-parallel-size $7 \
|
||||
--enable-expert-parallel \
|
||||
--served-model-name qwen3_235b \
|
||||
--max-model-len 40960 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--max-num-seqs 24 \
|
||||
--trust-remote-code \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--quantization ascend \
|
||||
--no-enable-prefix-caching \
|
||||
--enforce-eager \
|
||||
--additional-config '{"enable_flashcomm1": true, "enable_fused_mc2": 1}' \
|
||||
--kv-transfer-config \
|
||||
'{"kv_connector": "MooncakeConnectorV1",
|
||||
"kv_role": "kv_producer",
|
||||
"kv_port": "30000",
|
||||
"engine_id": "0",
|
||||
"kv_connector_extra_config": {
|
||||
"use_ascend_direct": true,
|
||||
"prefill": {
|
||||
"dp_size": 2,
|
||||
"tp_size": 8
|
||||
},
|
||||
"decode": {
|
||||
"dp_size": 8,
|
||||
"tp_size": 4
|
||||
}
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
**Decode node 0** (set `nic_name` and `local_ip` to your own):
|
||||
|
||||
```bash
|
||||
nic_name="<your_nic_name>"
|
||||
local_ip="<your_ip>"
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
|
||||
export OMP_NUM_THREADS=1
|
||||
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
|
||||
sysctl -w vm.swappiness=0
|
||||
sysctl -w kernel.numa_balancing=0
|
||||
sysctl kernel.sched_migration_cost_ns=50000
|
||||
export LD_PRELOAD=/usr/lib/aarch64-linux-gnu/libjemalloc.so.2:$LD_PRELOAD
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export LD_LIBRARY_PATH=/usr/local/Ascend/ascend-toolkit/latest/python/site-packages/mooncake:$LD_LIBRARY_PATH
|
||||
|
||||
export VLLM_TORCH_PROFILER_WITH_STACK=0
|
||||
export ASCEND_RT_VISIBLE_DEVICES=$1
|
||||
|
||||
vllm serve "/data/weights/Qwen3-235B-A22B-w8a8-rot" \
|
||||
--host 0.0.0.0 \
|
||||
--port $2 \
|
||||
--data-parallel-size $3 \
|
||||
--data-parallel-rank $4 \
|
||||
--data-parallel-address $5 \
|
||||
--data-parallel-rpc-port $6 \
|
||||
--tensor-parallel-size $7 \
|
||||
--enable-expert-parallel \
|
||||
--served-model-name qwen3_235b \
|
||||
--max-model-len 40960 \
|
||||
--max-num-batched-tokens 512 \
|
||||
--max-num-seqs 128 \
|
||||
--trust-remote-code \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--quantization ascend \
|
||||
--no-enable-prefix-caching \
|
||||
--async-scheduling \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"enable_flashcomm1": true, "enable_fused_mc2": 2}' \
|
||||
--kv-transfer-config \
|
||||
'{"kv_connector": "MooncakeConnectorV1",
|
||||
"kv_role": "kv_consumer",
|
||||
"kv_port": "30100",
|
||||
"engine_id": "1",
|
||||
"kv_connector_extra_config": {
|
||||
"use_ascend_direct": true,
|
||||
"prefill": {
|
||||
"dp_size": 2,
|
||||
"tp_size": 8
|
||||
},
|
||||
"decode": {
|
||||
"dp_size": 8,
|
||||
"tp_size": 4
|
||||
}
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
**Decode node 1** (set `nic_name` and `local_ip` to your own):
|
||||
|
||||
```bash
|
||||
nic_name="<your_nic_name>"
|
||||
local_ip="<your_ip>"
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
|
||||
export OMP_NUM_THREADS=1
|
||||
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
|
||||
sysctl -w vm.swappiness=0
|
||||
sysctl -w kernel.numa_balancing=0
|
||||
sysctl kernel.sched_migration_cost_ns=50000
|
||||
export LD_PRELOAD=/usr/lib/aarch64-linux-gnu/libjemalloc.so.2:$LD_PRELOAD
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export LD_LIBRARY_PATH=/usr/local/Ascend/ascend-toolkit/latest/python/site-packages/mooncake:$LD_LIBRARY_PATH
|
||||
|
||||
export VLLM_TORCH_PROFILER_WITH_STACK=0
|
||||
export ASCEND_RT_VISIBLE_DEVICES=$1
|
||||
|
||||
vllm serve "/data/weights/Qwen3-235B-A22B-w8a8-rot" \
|
||||
--host 0.0.0.0 \
|
||||
--port $2 \
|
||||
--data-parallel-size $3 \
|
||||
--data-parallel-rank $4 \
|
||||
--data-parallel-address $5 \
|
||||
--data-parallel-rpc-port $6 \
|
||||
--tensor-parallel-size $7 \
|
||||
--enable-expert-parallel \
|
||||
--served-model-name qwen3_235b \
|
||||
--max-model-len 40960 \
|
||||
--max-num-batched-tokens 512 \
|
||||
--max-num-seqs 128 \
|
||||
--trust-remote-code \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--quantization ascend \
|
||||
--no-enable-prefix-caching \
|
||||
--async-scheduling \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"enable_flashcomm1": true, "enable_fused_mc2": 2}' \
|
||||
--kv-transfer-config \
|
||||
'{"kv_connector": "MooncakeConnectorV1",
|
||||
"kv_role": "kv_consumer",
|
||||
"kv_port": "30100",
|
||||
"engine_id": "1",
|
||||
"kv_connector_extra_config": {
|
||||
"use_ascend_direct": true,
|
||||
"prefill": {
|
||||
"dp_size": 2,
|
||||
"tp_size": 8
|
||||
},
|
||||
"decode": {
|
||||
"dp_size": 8,
|
||||
"tp_size": 4
|
||||
}
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
Once the scripts are ready, start the servers on each node.
|
||||
|
||||
**Prefill node:**
|
||||
|
||||
```bash
|
||||
python launch_online_dp.py \
|
||||
--dp-size 2 --tp-size 8 \
|
||||
--dp-size-local 2 --dp-rank-start 0 \
|
||||
--dp-address <prefill_ip> --dp-rpc-port 54951 \
|
||||
--vllm-start-port 9123
|
||||
```
|
||||
|
||||
**Decode node 0:**
|
||||
|
||||
```bash
|
||||
python launch_online_dp.py \
|
||||
--dp-size 8 --tp-size 4 \
|
||||
--dp-size-local 4 --dp-rank-start 0 \
|
||||
--dp-address <decode_ip> --dp-rpc-port 54951 \
|
||||
--vllm-start-port 9123
|
||||
```
|
||||
|
||||
**Decode node 1:**
|
||||
|
||||
```bash
|
||||
python launch_online_dp.py \
|
||||
--dp-size 8 --tp-size 4 \
|
||||
--dp-size-local 4 --dp-rank-start 4 \
|
||||
--dp-address <decode_ip> --dp-rpc-port 54951 \
|
||||
--vllm-start-port 9123
|
||||
```
|
||||
|
||||
**Request Forwarding:**
|
||||
|
||||
Run the proxy on any machine that can reach both nodes. You can get the proxy script from the repository: [load_balance_proxy_server_example.py](https://github.com/vllm-project/vllm-ascend/blob/main/examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py).
|
||||
|
||||
```bash
|
||||
unset http_proxy https_proxy
|
||||
|
||||
python load_balance_proxy_server_example.py \
|
||||
--port 38085 \
|
||||
--host <prefill_ip> \
|
||||
--prefiller-hosts \
|
||||
<prefill_ip> <prefill_ip> \
|
||||
--prefiller-ports \
|
||||
9123 9124 \
|
||||
--decoder-hosts \
|
||||
<decode0_ip> <decode0_ip> <decode0_ip> <decode0_ip> \
|
||||
<decode1_ip> <decode1_ip> <decode1_ip> <decode1_ip> \
|
||||
--decoder-ports \
|
||||
9123 9124 9125 9126 \
|
||||
9123 9124 9125 9126 \
|
||||
```
|
||||
|
||||
:::{note}
|
||||
|
||||
- [vLLM Serving Arguments documentation](https://docs.vllm.com.cn/en/latest/cli/serve/?h=block+size#arguments) — Additional parameter details for vLLM serve commands.
|
||||
- [Environment Variables](../../user_guide/configuration/env_vars.md) — Ascend-specific environment variables (`HCCL_*`, etc.).
|
||||
:::
|
||||
|
||||
**Service Verification:**
|
||||
|
||||
If the service starts successfully, the following startup log will be displayed:
|
||||
|
||||
```text
|
||||
(APIServer pid=<pid>) INFO: Started server process [<pid>]
|
||||
(APIServer pid=<pid>) INFO: Waiting for application startup.
|
||||
(APIServer pid=<pid>) INFO: Application startup complete.
|
||||
```
|
||||
|
||||
## 6 Functional Verification
|
||||
|
||||
After the service is started, the model can be invoked by sending a prompt:
|
||||
|
||||
```shell
|
||||
curl http://<node0_ip>:<port>/v1/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "qwen3",
|
||||
"prompt": "The future of AI is",
|
||||
"max_completion_tokens": 50,
|
||||
"temperature": 0
|
||||
}'
|
||||
```
|
||||
|
||||
Expected result: HTTP 200 with a JSON response containing the `choices` field with generated text.
|
||||
|
||||
## 7 Accuracy Evaluation
|
||||
|
||||
### Using AISBench
|
||||
|
||||
For setup details, including installation, dataset download, and configuration, please refer to [Using AISBench](../../developer_guide/evaluation/using_ais_bench.md).
|
||||
|
||||
The following is an example configuration for the accuracy evaluation config file:
|
||||
|
||||
**Accuracy Evaluation Config File:**
|
||||
|
||||
```bash
|
||||
# Example configuration: benchmarks/ais_bench/benchmark/configs/models/vllm_api/vllm_api_general_chat.py
|
||||
from ais_bench.benchmark.models import VLLMCustomAPIChat
|
||||
from ais_bench.benchmark.utils.model_postprocessors import extract_non_reasoning_content
|
||||
|
||||
models = [
|
||||
dict(
|
||||
attr="service",
|
||||
type=VLLMCustomAPIChat,
|
||||
abbr='vllm-api-general-chat',
|
||||
path="your_model_path",
|
||||
model="qwen3",
|
||||
request_rate = 0,
|
||||
retry = 2,
|
||||
host_ip = "127.0.0.1",
|
||||
host_port = 2001,
|
||||
max_out_len = 32768,
|
||||
batch_size = 32,
|
||||
trust_remote_code=False,
|
||||
generation_kwargs = dict(
|
||||
temperature = 0.6,
|
||||
top_k = 20,
|
||||
top_p = 0.95,
|
||||
),
|
||||
pred_postprocessor=dict(type=extract_non_reasoning_content)
|
||||
)
|
||||
]
|
||||
```
|
||||
|
||||
**Run the accuracy evaluation using the aime2024 dataset as an example:**
|
||||
|
||||
```bash
|
||||
ais_bench --models vllm_api_general_chat --datasets aime2024_gen_0_shot_chat_prompt --debug
|
||||
```
|
||||
|
||||
> The --models parameter value corresponds to the abbr field in the configuration file above. Adjust max_out_len, batch_size, and dataset tasks based on your scenario.
|
||||
|
||||
## 8 Performance Evaluation
|
||||
|
||||
### Using AISBench
|
||||
|
||||
For setup details, including installation, dataset download, and configuration, please refer to [Using AISBench](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation) for details.
|
||||
|
||||
The following is an example configuration for the performance evaluation config file:
|
||||
|
||||
```bash
|
||||
# Example configuration: benchmarks/ais_bench/benchmark/configs/models/vllm_api/vllm_api_stream_chat.py
|
||||
from ais_bench.benchmark.models import VLLMCustomAPIChat
|
||||
from ais_bench.benchmark.utils.postprocess.model_postprocessors import extract_non_reasoning_content
|
||||
|
||||
models = [
|
||||
dict(
|
||||
attr="service",
|
||||
type=VLLMCustomAPIChat,
|
||||
abbr="vllm-api-stream-chat",
|
||||
path="your_model_path",
|
||||
model="qwen",
|
||||
stream=True,
|
||||
request_rate=0,
|
||||
use_timestamp=False,
|
||||
retry=2,
|
||||
host_ip="localhost",
|
||||
host_port=20002,
|
||||
max_out_len=1500,
|
||||
batch_size=140,
|
||||
trust_remote_code=False,
|
||||
generation_kwargs=dict(
|
||||
temperature=0,
|
||||
ignore_eos = True
|
||||
),
|
||||
)
|
||||
]
|
||||
```
|
||||
|
||||
**Run the performance evaluation using the GSM8K dataset as an example:**
|
||||
|
||||
```bash
|
||||
ais_bench --models vllm_api_stream_chat --datasets gsm8k_gen_0_shot_cot_str_perf --debug --summarizer default_perf --mode perf --num-prompts 560
|
||||
```
|
||||
|
||||
### Using vLLM Benchmark
|
||||
|
||||
Refer to [vLLM benchmark](https://docs.vllm.ai/en/latest/benchmarking/) for more details.
|
||||
|
||||
There are three `vllm bench` subcommands:
|
||||
|
||||
- `latency`: Benchmark the latency of a single batch of requests.
|
||||
- `serve`: Benchmark the online serving throughput.
|
||||
- `throughput`: Benchmark offline inference throughput.
|
||||
|
||||
Take `serve` as an example:
|
||||
|
||||
```shell
|
||||
vllm bench serve \
|
||||
--model your_model_path \
|
||||
--dataset-name random \
|
||||
--random-input 200 \
|
||||
--num-prompts 200 \
|
||||
--request-rate 1 \
|
||||
--save-result \
|
||||
--result-dir ./
|
||||
```
|
||||
|
||||
After several minutes, you will get the performance evaluation result.
|
||||
|
||||
## 9 Performance Tuning
|
||||
|
||||
### 9.1 Recommended Configurations
|
||||
|
||||
> **Note**: The following configurations are validated in specific test environments and are for reference only. The optimal configuration depends on factors such as maximum input/output length, prefix cache hit rate, precision requirements, and deployment machine ratios. It is recommended to refer to Section 9.2 for tuning based on actual conditions.
|
||||
|
||||
#### Table 1: Scenario Overview
|
||||
|
||||
| Scenario | Deployment Mode | *Total NPUs | Weight Version | Key Considerations |
|
||||
|----------|----------------|-------------|----------------|---------------------|
|
||||
| High Throughput | Single-Node (TP4, DP4) | 16 (A3) | W8A8 | DP and TP distribute MoE experts across 16 NPUs for maximum throughput |
|
||||
| High Throughput | PD Disaggregation (3 nodes) | 48 (3×A3) | W8A8 | 3-node PD separation balances prefill and decode resources for high throughput |
|
||||
| Low Latency | Single-Node (TP16) | 16 (A3) | W8A8 | 16-NPU TP minimizes per-token latency with speculative decoding |
|
||||
| Long Context | Single-Node (TP8, CP2) | 16 (A3) | W8A8 | 16-NPU TP with Context Parallelism extends context to 135k tokens |
|
||||
|
||||
> `*Total NPUs` indicates the total number of NPUs used across all nodes.
|
||||
|
||||
#### Table 2: Detailed Node Configuration
|
||||
|
||||
| Scenario | Configuration | NPUs | TP | DP | MTP Speculation Num | FUSED_MC2 | EP Switch | Async Scheduling |
|
||||
|----------|---------------|-------|----|-------------|--------------------|-----------|-----------|--------------|
|
||||
| High Throughput | Single-Node | 16 | 4 | 4 | none | On | On | On |
|
||||
| Low Latency | Single-Node | 16 | 16 | 1 | 3 | Off | On | On |
|
||||
| Long Context | Single-Node | 16 | 8 | 1 | none | On | On | Off |
|
||||
|
||||
> For additional parameter details, please refer to the deployment examples in [Section 5.1](#51-single-node-online-deployment)
|
||||
|
||||
<u>Single-node PD Hybrid — High Throughput:</u>
|
||||
|
||||
Single-node PD hybrid deployment optimized for maximum throughput on Atlas 800I A3 (64GB × 16):
|
||||
|
||||
```bash
|
||||
export HCCL_IF_IP=<node_ip>
|
||||
export GLOO_SOCKET_IFNAME=<ifname>
|
||||
export TP_SOCKET_IFNAME=<ifname>
|
||||
export HCCL_SOCKET_IFNAME=<ifname>
|
||||
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
|
||||
export OMP_NUM_THREADS=1
|
||||
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
|
||||
sysctl -w vm.swappiness=0
|
||||
sysctl -w kernel.numa_balancing=0
|
||||
sysctl kernel.sched_migration_cost_ns=50000
|
||||
export LD_PRELOAD=/usr/lib/aarch64-linux-gnu/libjemalloc.so.2:$LD_PRELOAD
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
|
||||
vllm serve your_model_path \
|
||||
--served-model-name qwen3 \
|
||||
--host <host_ip> \
|
||||
--port <port> \
|
||||
--async-scheduling \
|
||||
--tensor-parallel-size 4 \
|
||||
--data-parallel-size 4 \
|
||||
--data-parallel-size-local 4 \
|
||||
--data-parallel-start-rank 0 \
|
||||
--data-parallel-address <node_ip> \
|
||||
--data-parallel-rpc-port <rpc_port> \
|
||||
--enable-expert-parallel \
|
||||
--max-num-seqs 128 \
|
||||
--max-model-len 32768 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--trust-remote-code \
|
||||
--quantization ascend \
|
||||
--no-enable-prefix-caching \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"enable_cpu_binding":true, "enable_flashcomm1": true, "enable_fused_mc2": 1}'
|
||||
```
|
||||
|
||||
<u>Single-node PD Hybrid — Low Latency:</u>
|
||||
|
||||
Single-node PD hybrid deployment optimized for low latency with speculative decoding (Eagle3):
|
||||
|
||||
```bash
|
||||
export HCCL_IF_IP=<node_ip>
|
||||
export GLOO_SOCKET_IFNAME=<ifname>
|
||||
export TP_SOCKET_IFNAME=<ifname>
|
||||
export HCCL_SOCKET_IFNAME=<ifname>
|
||||
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
|
||||
export OMP_NUM_THREADS=1
|
||||
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
|
||||
sysctl -w vm.swappiness=0
|
||||
sysctl -w kernel.numa_balancing=0
|
||||
sysctl kernel.sched_migration_cost_ns=50000
|
||||
export LD_PRELOAD=/usr/lib/aarch64-linux-gnu/libjemalloc.so.2:$LD_PRELOAD
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
|
||||
vllm serve your_model_path \
|
||||
--served-model-name qwen3 \
|
||||
--host <host_ip> \
|
||||
--port <port> \
|
||||
--async-scheduling \
|
||||
--tensor-parallel-size 16 \
|
||||
--data-parallel-size 1 \
|
||||
--data-parallel-size-local 1 \
|
||||
--data-parallel-start-rank 0 \
|
||||
--data-parallel-address <node_ip> \
|
||||
--data-parallel-rpc-port <rpc_port> \
|
||||
--enable-expert-parallel \
|
||||
--max-num-seqs 128 \
|
||||
--max-model-len 32768 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--trust-remote-code \
|
||||
--quantization ascend \
|
||||
--no-enable-prefix-caching \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
|
||||
--speculative-config '{"method": "eagle3", "model":"your_eagle3_model_path", "num_speculative_tokens": 3}' \
|
||||
--additional-config '{"enable_cpu_binding":true, "enable_flashcomm1": true}'
|
||||
```
|
||||
|
||||
<u>Single-node PD Hybrid — Long Context:</u>
|
||||
|
||||
Single-node PD hybrid deployment optimized for long context with Context Parallelism and yarn rope-scaling:
|
||||
|
||||
```bash
|
||||
export HCCL_IF_IP=<node_ip>
|
||||
export GLOO_SOCKET_IFNAME=<ifname>
|
||||
export TP_SOCKET_IFNAME=<ifname>
|
||||
export HCCL_SOCKET_IFNAME=<ifname>
|
||||
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
|
||||
export OMP_NUM_THREADS=1
|
||||
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
|
||||
sysctl -w vm.swappiness=0
|
||||
sysctl -w kernel.numa_balancing=0
|
||||
sysctl kernel.sched_migration_cost_ns=50000
|
||||
export LD_PRELOAD=/usr/lib/aarch64-linux-gnu/libjemalloc.so.2:$LD_PRELOAD
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
|
||||
vllm serve your_model_path \
|
||||
--served-model-name qwen3 \
|
||||
--host <host_ip> \
|
||||
--port <port> \
|
||||
--tensor-parallel-size 8 \
|
||||
--data-parallel-size 1 \
|
||||
--decode-context-parallel-size 2 \
|
||||
--prefill-context-parallel-size 2 \
|
||||
--enable-expert-parallel \
|
||||
--cp-kv-cache-interleave-size 128 \
|
||||
--max-num-seqs 32 \
|
||||
--max-model-len 135000 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--gpu-memory-utilization 0.85 \
|
||||
--trust-remote-code \
|
||||
--quantization ascend \
|
||||
--no-enable-prefix-caching \
|
||||
--hf-overrides '{"rope_parameters": {"rope_type":"yarn","rope_theta":1000000,"factor":4,"original_max_position_embeddings":131072}}' \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"enable_cpu_binding":true, "enable_flashcomm1": true, "enable_fused_mc2": 1}'
|
||||
```
|
||||
|
||||
### 9.2 Tuning Guidelines
|
||||
|
||||
#### 9.2.1 General Tuning Reference
|
||||
|
||||
Please refer to the [Public Performance Tuning Documentation](../../developer_guide/performance_and_debug/optimization_and_tuning.md) for tuning methods.
|
||||
Please refer to the [Feature Guide](../../user_guide/support_matrix/feature_matrix.md) for detailed feature descriptions.
|
||||
|
||||
## 10 FAQ
|
||||
|
||||
For common environment, installation, and general parameter issues, please refer to the [vLLM-Ascend FAQs](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html). This section only covers issues specific to Qwen3-235B-A22B.
|
||||
|
||||
### Q: What hardware is required for Qwen3-235B-A22B?
|
||||
|
||||
For BF16: 1 Atlas 800I A3 (64GB × 16) node, 1 Atlas 800I A2 (64GB × 8) node, or 2 Atlas 800I A2 (32GB × 8) nodes. For W8A8 quantized version, the hardware requirements are similar.
|
||||
|
||||
### Q: How do I enable long context beyond 40k?
|
||||
|
||||
Use yarn rope-scaling. For vLLM >= v0.12.0: `--hf-overrides '{"rope_parameters": {"rope_type":"yarn","rope_theta":1000000,"factor":4,"original_max_position_embeddings":32768}}'`. For older versions, use `--rope_scaling`. Model variants like Qwen3-235B-A22B-Instruct-2507 natively support long contexts and don't need this parameter.
|
||||
|
||||
### Q: When should I use PD disaggregation vs single-node deployment?
|
||||
|
||||
Single-node deployment is simpler and recommended when the model fits within a single node. PD disaggregation separates Prefill and Decode across nodes, enabling higher throughput for large-scale serving. For Qwen3-235B-A22B, three A3 nodes with PD disaggregation can achieve ~3× the throughput of single-node deployment.
|
||||
|
||||
### Q: What is the difference between `enable_fused_mc2=1` and `=2`?
|
||||
|
||||
Value `1` enables the base MoE fused operator, suitable for typical EP configurations. Value `2` enables an alternative fusion strategy optimized for large-scale EP (e.g., EP32 in PD disaggregation scenarios). Both are experimental and currently only support W8A8 quantization on Atlas A3 servers.
|
||||
|
||||
### Q: When should I use Expert Parallelism?
|
||||
|
||||
Expert Parallelism (EP) should always be enabled for Qwen3-235B-A22B (an MoE model) via `--enable-expert-parallel`. It distributes FFN experts across NPUs to reduce per-device computation. EP works alongside TP, where MoE layers use EP and non-MoE layers use TP.
|
||||
|
||||
### Q: How do I choose between Context Parallelism and PD Disaggregation?
|
||||
|
||||
Context Parallelism (CP) splits the KV cache of a single request across multiple NPUs, suitable for long context scenarios on a single node. PD Disaggregation separates Prefill and Decode across nodes, suitable for high-throughput serving with many concurrent requests.
|
||||
652
docs/source/tutorials/models/Qwen3-30B-A3B.md
Normal file
652
docs/source/tutorials/models/Qwen3-30B-A3B.md
Normal file
@@ -0,0 +1,652 @@
|
||||
# Qwen3-30B-A3B
|
||||
|
||||
## 1 Introduction
|
||||
|
||||
Qwen3-30B-A3B is a Mixture-of-Experts (MoE) model in the Qwen3 series, featuring 30.5B total parameters with 3.3B activated per token. The sparse MoE architecture enables efficient training and inference, delivering strong performance across reasoning, instruction-following, and agent capabilities while maintaining lower computational cost compared to dense models of similar capability.
|
||||
|
||||
This document will demonstrate the main validation steps for Qwen3-30B-A3B in the vLLM-Ascend environment, including supported features, environment preparation, single-node deployment, as well as accuracy and performance evaluation.
|
||||
|
||||
The Qwen3-30B-A3B model is first supported in v0.8.4rc2. This document is validated and written based on **vLLM-Ascend v0.22.1rc**. All **v0.22.1rc and later versions** can run stably. To use the latest features, it is recommended to use the latest release candidate or official version.
|
||||
|
||||
## 2 Supported Features
|
||||
|
||||
Please refer to the [Supported Features List](../../user_guide/support_matrix/supported_models.md) for the model support matrix.
|
||||
|
||||
Please refer to the [Feature Guide](../../user_guide/feature_guide/index.md) for feature configuration information.
|
||||
|
||||
## 3 Prerequisites
|
||||
|
||||
### 3.1 Model Weight
|
||||
|
||||
The following model variants are available. It is recommended to download the model weight to a shared directory accessible to all nodes.
|
||||
|
||||
| Model | Hardware Requirement | Download |
|
||||
| -------------------- | ------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------ |
|
||||
| Qwen3-30B-A3B (BF16) | Atlas 800I A3 (64GB, 1\~2 cards)<br>Atlas 800I A2 (64GB, 2\~4 cards) | [Download](https://www.modelscope.cn/models/Qwen/Qwen3-30B-A3B) |
|
||||
| Qwen3-30B-A3B-W8A8 | Atlas 800I A3 (64GB, 1\~2 cards)<br>Atlas 800I A2 (64GB, 2\~4 cards) | [Download](https://www.modelscope.cn/models/Eco-Tech/Qwen3-30B-A3B-w8a8) |
|
||||
| Eagle3 Draft Model | NA | [Download](https://www.modelscope.cn/models/Eco-Tech/Qwen3-30B-A3B-w8a8-QuaRot-310) |
|
||||
|
||||
**Quantized Versions for Atlas 300I DUO:**
|
||||
|
||||
| Model | Quantization | Hardware Requirement | Download |
|
||||
|-------|-------------|---------------------|----------|
|
||||
| Qwen3-30B-A3B-w8a8-QuaRot-310 |W8A8 | Atlas 300I DUO (48GB,2 cards) | [Download](https://www.modelscope.cn/models/Eco-Tech/Qwen3-30B-A3B-w8a8-QuaRot-310) |
|
||||
|
||||
These are the recommended numbers of cards, which can be adjusted according to the actual situation.
|
||||
|
||||
If the W8A8 quantized weights are not available for direct download, you can obtain them by quantizing the BF16 model using **msmodelslim**. Refer to the [Quantization Guide](../../user_guide/feature_guide/quantization.md) for details. All model paths in this document should be adjusted to your actual local paths.
|
||||
|
||||
:::{note}
|
||||
|
||||
Qwen3-30B-A3B-W8A8 adopts a hybrid quantization strategy (ordered by model structure):
|
||||
|
||||
- **Embedding layer**: BF16 (no quantization)
|
||||
- **Q/K normalization** (q_norm, k_norm): BF16
|
||||
- **Attention projections** (q/k/v/o_proj): Static W8A8 with pre-computed per-tensor scales
|
||||
- **MoE routing gate** (mlp.gate): BF16
|
||||
- **MoE expert projections** (gate/up/down_proj): Dynamic W8A8 where input scales are computed on-the-fly during inference
|
||||
:::
|
||||
|
||||
## 4 Installation
|
||||
|
||||
### 4.1 Docker Image Installation
|
||||
|
||||
You can use the official all-in-one Docker image for Qwen3 MoE models.
|
||||
|
||||
For Atlas 300I DUO, use `vllm-ascend:nightly-releases-v0.23.0-310p` (or a later `-310p` image).
|
||||
|
||||
:::::{tab-set}
|
||||
|
||||
::::{tab-item} Atlas A3 inference products
|
||||
:sync: A3
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-a3
|
||||
|
||||
docker run \
|
||||
--name vllm-ascend-env \
|
||||
--ipc host \
|
||||
--net host \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci8 \
|
||||
--device /dev/davinci9 \
|
||||
--device /dev/davinci10 \
|
||||
--device /dev/davinci11 \
|
||||
--device /dev/davinci12 \
|
||||
--device /dev/davinci13 \
|
||||
--device /dev/davinci14 \
|
||||
--device /dev/davinci15 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/Ascend/driver:/usr/local/Ascend/driver \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /usr/local/sbin:/usr/local/sbin \
|
||||
-it -d $IMAGE bash
|
||||
```
|
||||
|
||||
:::{note}
|
||||
|
||||
A3 has 8 NPUs with dual-die design (16 chips total: `/dev/davinci[0-15]`).
|
||||
If you are on a shared machine, map only the chips you need (e.g., `/dev/davinci[0-7]` for NPU 0-3).
|
||||
:::
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Atlas A2 inference products
|
||||
:sync: A2
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
|
||||
docker run \
|
||||
--name vllm-ascend-env \
|
||||
--ipc host \
|
||||
--net host \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/Ascend/driver:/usr/local/Ascend/driver \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /usr/local/sbin:/usr/local/sbin \
|
||||
-it -d $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-310p
|
||||
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
The default workdir is `/workspace`. vLLM and vLLM-Ascend are installed as Python packages in site-packages.
|
||||
|
||||
**Installation Verification:**
|
||||
|
||||
After starting the container, run the following command to verify the installation:
|
||||
|
||||
```bash
|
||||
docker ps | grep vllm-ascend-env
|
||||
```
|
||||
|
||||
Expected result: The container is listed with status `Up`. You can also verify the vllm-ascend version inside the container:
|
||||
|
||||
```bash
|
||||
pip show vllm-ascend
|
||||
```
|
||||
|
||||
Expected result: The version information is displayed, matching the pulled image version.
|
||||
|
||||
### 4.2 Source Code Installation
|
||||
|
||||
If you prefer not to use the Docker image, you can build from source. Install vLLM from source first:
|
||||
|
||||
1. Clone and install vLLM:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/vllm-project/vllm.git
|
||||
cd vllm
|
||||
pip install -e .
|
||||
```
|
||||
|
||||
2. Clone and install the vLLM-Ascend repository:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/vllm-project/vllm-ascend.git
|
||||
cd vllm-ascend
|
||||
pip install -e .
|
||||
```
|
||||
|
||||
:::{note}
|
||||
|
||||
For Atlas 300I DUO, source installation may pull in `triton` and `triton-ascend`. Uninstall them before running vLLM-Ascend on Atlas 300I DUO:
|
||||
|
||||
```bash
|
||||
pip uninstall -y triton-ascend triton
|
||||
```
|
||||
|
||||
:::
|
||||
**Installation Verification:**
|
||||
|
||||
```bash
|
||||
pip show vllm vllm-ascend
|
||||
```
|
||||
|
||||
Expected result: The version information for both packages is displayed, confirming a successful installation.
|
||||
|
||||
:::{note}
|
||||
|
||||
If deploying a multi-node environment, set up the environment on each node.
|
||||
:::
|
||||
For more details, please refer to the [Installation Guide](../../installation.md).
|
||||
|
||||
## 5 Online Service Deployment
|
||||
|
||||
### 5.1 Single-Node Online Deployment
|
||||
|
||||
Single-node deployment completes both Prefill and Decode within the same node, suitable for development, testing, and small-to-medium scale inference scenarios. For the Qwen3-30B-A3B MoE model, Expert Parallelism (EP) is required to distribute experts across NPUs.
|
||||
|
||||
> The following command is an example configuration. Adjust the parameters based on your actual scenario.
|
||||
|
||||
:::::{tab-set}
|
||||
::::{tab-item} Atlas A2 inference products / Atlas A3 inference products
|
||||
|
||||
```bash
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3
|
||||
export HCCL_OP_EXPANSION_MODE="AIV" # not needed on A2
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
|
||||
vllm serve your_model_path \
|
||||
--served-model-name qwen3 \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 100 \
|
||||
--max-model-len 40960 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--tensor-parallel-size 4 \
|
||||
--enable-expert-parallel \
|
||||
--quantization ascend \
|
||||
--distributed_executor_backend "mp" \
|
||||
--no-enable-prefix-caching \
|
||||
--async-scheduling \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"enable_flashcomm1": true, "weight_nz_mode": 2}' \
|
||||
--gpu-memory-utilization 0.95 \
|
||||
--port 8000 \
|
||||
--speculative-config '{"method": "eagle3", "model": "your_eagle3_model_path", "num_speculative_tokens": 3}'
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
|
||||
```bash
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1
|
||||
|
||||
vllm serve your_model_path \
|
||||
--host 127.0.0.1 \
|
||||
--port 8000 \
|
||||
--tensor-parallel-size 2 \
|
||||
--max-num-seqs 32 \
|
||||
--served_model_name qwen3 \
|
||||
--dtype float16 \
|
||||
--quantization ascend \
|
||||
--max-model-len 16384 \
|
||||
--additional-config '{"ascend_compilation_config": {"fuse_norm_quant": false,"enable_npu_graph_ex":false}}' \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY", "cudagraph_capture_sizes": [1,32]}' \
|
||||
--no-enable-prefix-caching
|
||||
```
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
Key parameters:
|
||||
|
||||
- `--tensor-parallel-size 2` maps the model across two Atlas inference devices. Adjust it together with `ASCEND_RT_VISIBLE_DEVICES` according to the available devices and memory.
|
||||
- `--dtype float16` is used for Atlas 300I DUO to match the Atlas inference execution path.
|
||||
- `--max-model-len 16384` is intentionally conservative. On Atlas 300I DUO, large context lengths allocate large attention masks, so do not rely on automatic max-model-len detection.
|
||||
- `--max-num-seqs 16` limits concurrent active requests to reduce KV cache and graph capture pressure on Atlas 300I DUO.
|
||||
- `--gpu-memory-utilization` controls KV cache capacity. Reduce it if startup or runtime requests report OOM.
|
||||
- `--additional-config '{"ascend_compilation_config": {"fuse_norm_quant": false}}'` disables norm-quant fusion for the Atlas 300I DUO serving path.
|
||||
- `--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY", "cudagraph_capture_sizes": [1,32]}'` enables decode ACLGraph replay and explicitly limits capture sizes for Atlas 300I DUO.
|
||||
- `--no-enable-prefix-caching` is the default recommendation for this Atlas 300I DUO example to reduce memory pressure.
|
||||
- `--quantization ascend` enables Ascend quantization for the W8A8 model. Remove this option when deploying the BF16 model.
|
||||
|
||||
:::{note}
|
||||
|
||||
- `ASCEND_RT_VISIBLE_DEVICES`: must be set to the NPU chip IDs allocated to your environment (e.g., `0,1,2,3` for 4 chips).
|
||||
- `--port`: adjust to avoid conflicts with other services running on the same machine.
|
||||
- `--no-enable-prefix-caching`: disabled by default as prefix caching effectiveness for this model on Ascend NPUs has not been fully characterized. You can try enabling it to evaluate the cache hit rate for your workload.
|
||||
- `--quantization ascend`: required for W8A8 quantized models. Remove this parameter when using BF16 weights.
|
||||
:::
|
||||
|
||||
:::{tip}
|
||||
|
||||
For parameter details, refer to:
|
||||
|
||||
- [vLLM CLI documentation](https://docs.vllm.ai/en/stable/cli/) — standard serve parameters (`--host`, `--port`, `--max-model-len`, etc.)
|
||||
- [Environment Variables](../../user_guide/configuration/env_vars.md) — Ascend-specific environment variables (`HCCL_*`, etc.)
|
||||
- [Additional Configuration](../../user_guide/configuration/additional_config.md) — `--additional-config` format and options
|
||||
:::
|
||||
|
||||
**Service Verification:**
|
||||
|
||||
After the service is started, verify it is running by sending a prompt. Refer to [Section 6](#6-functional-verification) for a usage example.
|
||||
|
||||
## 6 Functional Verification
|
||||
|
||||
After the service is started, the model can be invoked by sending a prompt.
|
||||
|
||||
**Chat Completions API:**
|
||||
|
||||
```bash
|
||||
curl http://localhost:8000/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "qwen3",
|
||||
"messages": [
|
||||
{"role": "user", "content": "Give me a short introduction to large language models."}
|
||||
],
|
||||
"temperature": 0.6,
|
||||
"top_p": 0.95,
|
||||
"top_k": 20,
|
||||
"max_completion_tokens": 4096
|
||||
}'
|
||||
```
|
||||
|
||||
:::{note}
|
||||
Adjust the following fields based on your deployment:
|
||||
|
||||
- **URL** (`http://localhost:8000`): Replace `localhost` and `8000` with your server IP and the `--port` value from the `vllm serve` command.
|
||||
- **`model`**: Must match the `--served-model-name` value from the `vllm serve` command (e.g., `qwen3`).
|
||||
:::
|
||||
Expected result: HTTP 200 with a JSON response containing the `choices` field with generated text.
|
||||
|
||||
## 7 Accuracy Evaluation
|
||||
|
||||
### Using AISBench
|
||||
|
||||
For setup details, including installation, dataset download, and configuration, please refer to [Using AISBench](../../developer_guide/evaluation/using_ais_bench.md).
|
||||
|
||||
The following is an example configuration for the accuracy evaluation config file, demonstrated using the GSM8K dataset:
|
||||
|
||||
```python
|
||||
# Example configuration: benchmarks/ais_bench/benchmark/configs/models/vllm_api/vllm_api_general_chat.py
|
||||
from ais_bench.benchmark.models import VLLMCustomAPIChat
|
||||
|
||||
models = [
|
||||
dict(
|
||||
attr="service",
|
||||
type=VLLMCustomAPIChat,
|
||||
abbr='vllm-api-general-chat',
|
||||
path="your_model_path",
|
||||
model="qwen3",
|
||||
request_rate=0,
|
||||
retry=2,
|
||||
host_ip="localhost",
|
||||
host_port=8000,
|
||||
max_out_len=32768,
|
||||
batch_size=32,
|
||||
trust_remote_code=True,
|
||||
generation_kwargs=dict(
|
||||
temperature=0.6,
|
||||
top_k=20,
|
||||
top_p=0.95,
|
||||
),
|
||||
)
|
||||
]
|
||||
```
|
||||
|
||||
Run the accuracy evaluation using the `gsm8k` dataset as an example:
|
||||
|
||||
```shell
|
||||
ais_bench --models vllm_api_general_chat --datasets gsm8k_gen_4_shot_cot_str --mode all --dump-eval-details --debug
|
||||
```
|
||||
|
||||
The following table lists the `--datasets` parameter for each evaluation dataset:
|
||||
|
||||
| Dataset | `--datasets` Parameter |
|
||||
| ------------- | ------------------------------------ |
|
||||
| GSM8K | `gsm8k_gen_4_shot_cot_str` |
|
||||
| GPQA-Diamond | `gpqa_gen_0_shot_cot_chat_prompt` |
|
||||
| AIME 2024 | `aime2024_gen_0_shot_str` |
|
||||
| LiveCodeBench | `livecodebench_0_shot_chat_v4_v5_v6` |
|
||||
|
||||
> The `--models` parameter value corresponds to the configuration file name (e.g., `vllm_api_general_chat` for `vllm_api_general_chat.py`). Adjust `max_out_len`, `batch_size`, and dataset tasks based on your scenario.
|
||||
|
||||
For dataset preparation, please refer to the [AISBench Datasets Guide](https://github.com/AISBench/benchmark/blob/master/docs/source_zh_cn/get_started/datasets.md).
|
||||
|
||||
:::{note}
|
||||
|
||||
vLLM-Ascend also supports the following evaluation tools:
|
||||
|
||||
- [lm_eval](../../developer_guide/evaluation/using_lm_eval.md)
|
||||
- [OpenCompass](../../developer_guide/evaluation/using_opencompass.md)
|
||||
- [EvalScope](../../developer_guide/evaluation/using_evalscope.md)
|
||||
:::
|
||||
**Accuracy Results (Atlas 800I A3, vLLM-Ascend v0.22.1rc, W8A8):**
|
||||
|
||||
| Dataset | Metric | Score |
|
||||
| ------------- | --------------------- | ------ |
|
||||
| GSM8K | accuracy (4-shot CoT) | 92.87% |
|
||||
| GPQA-Diamond | accuracy (0-shot CoT) | 60.10% |
|
||||
| LiveCodeBench | pass@1 (0-shot) | 60.05% |
|
||||
| AIME 2024 | accuracy (0-shot) | 76.67% |
|
||||
|
||||
## 8 Performance Evaluation
|
||||
|
||||
### Using AISBench
|
||||
|
||||
For setup details, please refer to [Using AISBench for performance evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation).
|
||||
|
||||
First, configure the model for streaming performance testing (`ais_bench/benchmark/configs/models/vllm_api/vllm_api_stream_chat.py`):
|
||||
|
||||
```python
|
||||
from ais_bench.benchmark.models import VLLMCustomAPIChat
|
||||
from ais_bench.benchmark.utils.postprocess.model_postprocessors import extract_non_reasoning_content
|
||||
|
||||
models = [
|
||||
dict(
|
||||
attr="service",
|
||||
type=VLLMCustomAPIChat,
|
||||
abbr='vllm-api-stream-chat',
|
||||
path="your_model_path",
|
||||
model="qwen3",
|
||||
stream=True,
|
||||
request_rate=0,
|
||||
retry=2,
|
||||
host_ip="localhost",
|
||||
host_port=8000,
|
||||
max_out_len=1500,
|
||||
batch_size=32,
|
||||
trust_remote_code=True,
|
||||
generation_kwargs=dict(
|
||||
temperature=0.01,
|
||||
ignore_eos=True,
|
||||
),
|
||||
pred_postprocessor=dict(type=extract_non_reasoning_content),
|
||||
)
|
||||
]
|
||||
```
|
||||
|
||||
> Key differences from the accuracy config: `stream=True`, `ignore_eos=True` (ensures output reaches `max_out_len` for consistent TPOT measurement), and `batch_size` controls concurrency.
|
||||
|
||||
Then, configure the synthetic dataset distribution (`ais_bench/datasets/synthetic/synthetic_config.py`). Adjust the configuration based on your actual scenario. Note that random synthetic data is not suitable for benchmarking scenarios where prefix caching is enabled, as random inputs produce zero cache hit rate.
|
||||
|
||||
```python
|
||||
synthetic_config = {
|
||||
"Type": "string",
|
||||
"RequestCount": 200,
|
||||
"StringConfig": {
|
||||
"Input": {
|
||||
"Method": "uniform",
|
||||
"Params": {"MinValue": 3500, "MaxValue": 3500}
|
||||
},
|
||||
"Output": {
|
||||
"Method": "uniform",
|
||||
"Params": {"MinValue": 1500, "MaxValue": 1500}
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Then run the performance evaluation:
|
||||
|
||||
```shell
|
||||
ais_bench --models vllm_api_stream_chat --datasets synthetic_gen --mode perf --debug
|
||||
```
|
||||
|
||||
> The `--models` value should match the `abbr` in your model config file. Use `--num-prompts` to limit the number of test requests.
|
||||
|
||||
### Using vLLM Benchmark
|
||||
|
||||
Refer to [vLLM benchmark](https://docs.vllm.ai/en/latest/benchmarking/) for more details.
|
||||
|
||||
Take the `serve` subcommand as an example. The `--random-output-len` parameter controls the number of output tokens per request; adjust it based on your target scenario (e.g., 2048 for short outputs, 32768 for long outputs).
|
||||
|
||||
```shell
|
||||
vllm bench serve \
|
||||
--model your_model_path \
|
||||
--served-model-name qwen3 \
|
||||
--port 8000 \
|
||||
--dataset-name random \
|
||||
--random-input 200 \
|
||||
--random-output-len 2048 \
|
||||
--num-prompts 200 \
|
||||
--request-rate 1 \
|
||||
--save-result \
|
||||
--result-dir ./
|
||||
```
|
||||
|
||||
## 9 Performance Tuning
|
||||
|
||||
### 9.1 Recommended Configurations
|
||||
|
||||
> **Note**: The following configurations are validated in specific test environments and are for reference only. The optimal configuration depends on factors such as maximum input/output length, prefix cache hit rate, precision requirements, and deployment machine ratios. It is recommended to refer to Section 9.2 for tuning based on actual conditions.
|
||||
|
||||
#### Table 1: Scenario Overview
|
||||
|
||||
| Scenario | Deployment Mode | *Total NPUs | Weight Version | Key Considerations |
|
||||
| --------------- | ----------------- | ---------------- | -------------- | ---------------------------------------------------------------- |
|
||||
| High Throughput | Single-Node (TP1) | 1 (A3)<br>2 (A2) | W8A8 | Single-card deployment maximizes concurrent request processing |
|
||||
| Low Latency | Single-Node (TP4) | 2 (A3)<br>4 (A2) | W8A8 | Multi-card TP reduces per-token latency with expert parallelism |
|
||||
| Long Context | Single-Node (TP4) | 2 (A3)<br>4 (A2) | W8A8 | Reduces concurrent sequences to accommodate longer max-model-len |
|
||||
|
||||
> `*Total NPUs` indicates the total number of NPUs used across all nodes. On Atlas 800I A3, each NPU contains two dies (chips), so TP4 requires 4 chips = 2 NPUs.
|
||||
|
||||
#### Table 2: Detailed Node Configuration
|
||||
|
||||
| Scenario | NPUs | TP | max-model-len | max-num-seqs | FUSED_MC2 | EP | hf-overrides |
|
||||
| --------------- | ------ | --- | ------------- | ------------ | --------- | --- | ------------ |
|
||||
| High Throughput | 1 (A3) | 1 | 37364 | 100 | Off | Off | - |
|
||||
| Low Latency | 2 (A3) | 4 | 37364 | 100 | Off | On | - |
|
||||
| Long Context | 2 (A3) | 4 | 131072 | 14 | Off | On | YaRN |
|
||||
|
||||
> For detailed parameter descriptions, please refer to the deployment examples in Section 5.
|
||||
|
||||
**Low Latency Configuration:**
|
||||
|
||||
```shell
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
|
||||
vllm serve your_model_path \
|
||||
--served-model-name qwen3 \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 100 \
|
||||
--max-model-len 37364 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--tensor-parallel-size 4 \
|
||||
--enable-expert-parallel \
|
||||
--distributed_executor_backend "mp" \
|
||||
--no-enable-prefix-caching \
|
||||
--async-scheduling \
|
||||
--quantization ascend \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"enable_flashcomm1": true, "weight_nz_mode": 2}' \
|
||||
--gpu-memory-utilization 0.95 \
|
||||
--port 8000 \
|
||||
--speculative-config '{"method": "eagle3","model": "your_eagle3_model_path", "num_speculative_tokens": 3}'
|
||||
```
|
||||
|
||||
:::{tip}
|
||||
|
||||
Example AISBench settings for this configuration:
|
||||
|
||||
- `request_rate`: 0
|
||||
- `batch_size`: 32
|
||||
- Input/Output length: 2048/2048 or 3500/1500
|
||||
:::
|
||||
**High Throughput Configuration:**
|
||||
|
||||
```shell
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
|
||||
vllm serve your_model_path \
|
||||
--served-model-name qwen3 \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 100 \
|
||||
--max-model-len 37364 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--tensor-parallel-size 1 \
|
||||
--distributed_executor_backend "mp" \
|
||||
--no-enable-prefix-caching \
|
||||
--async-scheduling \
|
||||
--quantization ascend \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"weight_nz_mode": 2}' \
|
||||
--gpu-memory-utilization 0.95 \
|
||||
--port 8000 \
|
||||
--speculative-config '{"method": "eagle3","model": "your_eagle3_model_path", "num_speculative_tokens": 3}'
|
||||
```
|
||||
|
||||
:::{tip}
|
||||
|
||||
Example AISBench settings for this configuration:
|
||||
|
||||
- `request_rate`: 0
|
||||
- `batch_size`: 32
|
||||
- Input/Output length: 2048/2048 or 3500/1500
|
||||
:::
|
||||
**Long Context Configuration:**
|
||||
|
||||
```shell
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
|
||||
vllm serve your_model_path \
|
||||
--served-model-name qwen3 \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 14 \
|
||||
--max-model-len 131072 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--tensor-parallel-size 4 \
|
||||
--enable-expert-parallel \
|
||||
--distributed_executor_backend "mp" \
|
||||
--no-enable-prefix-caching \
|
||||
--async-scheduling \
|
||||
--quantization ascend \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"enable_flashcomm1": true, "weight_nz_mode": 2}' \
|
||||
--gpu-memory-utilization 0.95 \
|
||||
--port 8000 \
|
||||
--speculative-config '{"method": "eagle3","model": "your_eagle3_model_path", "num_speculative_tokens": 3}' \
|
||||
--hf-overrides '{"rope_parameters": {"rope_type":"yarn","factor":4,"original_max_position_embeddings":32768}}'
|
||||
```
|
||||
|
||||
:::{tip}
|
||||
|
||||
Example AISBench settings for this configuration:
|
||||
|
||||
- `request_rate`: 0
|
||||
- `batch_size`: 32
|
||||
- Input/Output length: 65536/1024 or 131072/1024
|
||||
:::
|
||||
|
||||
### 9.2 Tuning Guidelines
|
||||
|
||||
Please refer to the [Public Performance Tuning Documentation](../../developer_guide/performance_and_debug/optimization_and_tuning.md) for tuning methods.
|
||||
Please refer to the [Feature Guide](../../user_guide/support_matrix/feature_matrix.md) for detailed feature descriptions.
|
||||
|
||||
## 10 FAQ
|
||||
|
||||
For common environment, installation, and general parameter issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html). This chapter only covers model-specific issues.
|
||||
229
docs/source/tutorials/models/Qwen3-ASR-1.7B.md
Normal file
229
docs/source/tutorials/models/Qwen3-ASR-1.7B.md
Normal file
@@ -0,0 +1,229 @@
|
||||
# Qwen3-ASR-1.7B
|
||||
|
||||
## 1 Introduction
|
||||
|
||||
Qwen3-ASR-1.7B is a 1.7B-parameter automatic speech recognition (ASR) model from the Qwen team. It supports Chinese and English speech, Chinese dialects, multilingual speech, and singing voice transcription, and provides long-audio and streaming inference capabilities.
|
||||
|
||||
This document describes the supported features, environment preparation, single-node deployment, functional verification, and evaluation workflow for Qwen3-ASR-1.7B on Ascend NPUs.
|
||||
|
||||
Qwen3-ASR-1.7B was introduced with upstream vLLM v0.19.0. Use a vLLM-Ascend image that matches your vLLM version, and refer to the support matrix for the current release status.
|
||||
|
||||
## 2 Supported Features
|
||||
|
||||
Please refer to the [Supported Features List](../../user_guide/support_matrix/supported_models.md) for the model support matrix.
|
||||
|
||||
Please refer to the [Feature Guide](../../user_guide/feature_guide/index.md) for feature configuration information.
|
||||
|
||||
## 3 Prerequisites
|
||||
|
||||
### 3.1 Model Weight
|
||||
|
||||
The BF16 model can be deployed with one Ascend 910B 64 GB NPU or one Ascend Atlas 300I DUO 48 GB NPU. Download the model weights from [ModelScope](https://www.modelscope.cn/models/Qwen/Qwen3-ASR-1.7B).
|
||||
|
||||
Download the weights to a directory that is accessible from the deployment environment. For multi-node deployments, use a shared directory; for example, `/root/.cache/`.
|
||||
|
||||
## 4 Installation
|
||||
|
||||
### 4.1 Docker Image Installation
|
||||
|
||||
Use the vLLM-Ascend Docker image that corresponds to your hardware. Replace the model-weight mount with the path used in your environment.
|
||||
|
||||
For Atlas 300I DUO, use `vllm-ascend:nightly-releases-v0.23.0-310p` (or a later `-310p` image).
|
||||
|
||||
:::::{tab-set}
|
||||
|
||||
::::{tab-item} Atlas A2 inference products
|
||||
:sync: A2
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--net host \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64:/usr/local/Ascend/driver/lib64 \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it -d $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-310p
|
||||
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=10g \
|
||||
--net host \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64:/usr/local/Ascend/driver/lib64 \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it -d $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
Verify that the container is running and that the installed package version matches the image tag:
|
||||
|
||||
```bash
|
||||
docker ps --filter name=vllm-ascend
|
||||
pip show vllm vllm-ascend
|
||||
```
|
||||
|
||||
Expected result: `docker ps` lists the container with status `Up`, and `pip show` displays version information for both packages.
|
||||
|
||||
### 4.2 Source Code Installation
|
||||
|
||||
f you prefer to build from source instead of using the Docker image, install vLLM-Ascend following the [Installation Guide](../../installation.md).
|
||||
|
||||
:::{note}
|
||||
|
||||
For Atlas 300I DUO, source installation may pull in `triton` and `triton-ascend`. Uninstall them before running vLLM-Ascend on Atlas 300I DUO:
|
||||
|
||||
```bash
|
||||
pip uninstall -y triton-ascend triton
|
||||
```
|
||||
|
||||
To verify the source installation:
|
||||
|
||||
```bash
|
||||
pip show vllm-ascend
|
||||
```
|
||||
|
||||
:::
|
||||
|
||||
## 5 Online Service Deployment
|
||||
|
||||
### 5.1 Single-Node Online Deployment
|
||||
|
||||
Single-node deployment runs both audio prefill and decoding on one NPU, making it suitable for development, testing, and small-scale ASR services. Replace `your_model_path` with the local model directory, or use `Qwen/Qwen3-ASR-1.7B` to download the model through the configured model hub.
|
||||
|
||||
:::::{tab-set}
|
||||
::::{tab-item} Atlas A2 inference products
|
||||
|
||||
```shell
|
||||
vllm serve your_model_path \
|
||||
--served-model-name qwen3-asr \
|
||||
--tensor-parallel-size 1 \
|
||||
--max-model-len 4096 \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--enforce-eager \
|
||||
--port 8000
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
vllm serve your_model_path \
|
||||
--served-model-name qwen3-asr \
|
||||
--tensor-parallel-size 1 \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--dtype float16 \
|
||||
--max-model-len 4096 \
|
||||
--additional-config '{"ascend_compilation_config": {"fuse_norm_quant": false,"enable_npu_graph_ex":false}}' \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY", "cudagraph_capture_sizes": [1,4]}' \
|
||||
--port 8000
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
:::::
|
||||
|
||||
Key parameters Descriptions:
|
||||
|
||||
- `--tensor-parallel-size 1` uses one NPU. Increase it only after confirming that the hardware and deployment topology support the chosen parallel configuration.
|
||||
- `--max-model-len 4096` limits the maximum sequence length. On Atlas 300I DUO, always specify a conservative value explicitly; automatic detection can allocate an oversized attention mask and cause an out-of-memory error.
|
||||
- `--gpu-memory-utilization 0.9` sets the fraction of device memory available to the vLLM executor. Lower this value if other workloads share the NPU.
|
||||
- `--enforce-eager` disables graph execution. It is used in the Atlas 300I A2 2UP example for compatibility.
|
||||
|
||||
When the service starts successfully, the log contains `Application startup complete`. If startup fails, see the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html).
|
||||
|
||||
## 6 Functional Verification
|
||||
|
||||
After the service is started, the model can be invoked by sending a prompt.
|
||||
|
||||
**Chat Completions API:**
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
curl http://localhost:8000/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "qwen3-asr",
|
||||
"messages": [
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{
|
||||
"type": "audio_url",
|
||||
"audio_url": {
|
||||
"url": "https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen3-ASR-Repo/asr_en.wav"
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}'
|
||||
```
|
||||
|
||||
Replace `localhost`, `8000`, and `qwen3-asr` with the address, port, and `--served-model-name` used by your deployment. Expected result: HTTP 200 and a JSON response containing the transcription in the `choices` field.
|
||||
|
||||
## 7 Accuracy Evaluation
|
||||
|
||||
Evaluate transcription quality with Word Error Rate (WER) for word-level recognition and Character Error Rate (CER) for character-level recognition.
|
||||
|
||||
## 8 Performance Evaluation
|
||||
|
||||
Measure ASR serving performance with audio samples that represent the production workload. Record at least the audio duration, request concurrency, end-to-end latency, real-time factor, and throughput. This ensures that audio preprocessing, request construction, API communication, inference, and response parsing are included in the result.
|
||||
|
||||
Actual performance varies with hardware, audio duration, concurrency, and deployment configuration. Evaluate short audio, long audio, and concurrent requests separately before selecting a production configuration.
|
||||
|
||||
## 9 Performance Tuning
|
||||
|
||||
The following settings are starting points rather than globally optimal configurations. Tune them according to audio duration, concurrency, latency requirements, and available NPU memory.
|
||||
|
||||
| Scenario | Recommended Starting Point | Key Considerations |
|
||||
| --- | --- | --- |
|
||||
| Low latency | `--tensor-parallel-size 1`, `--max-model-len 4096` | Use short audio inputs and avoid sharing the NPU with other workloads. |
|
||||
| High throughput | Increase request concurrency after establishing the latency baseline | Monitor NPU memory and end-to-end latency; do not use synthetic text-only requests as a proxy for ASR traffic. |
|
||||
| Long audio | Increase `--max-model-len` only as required | On Atlas 300I DUO, keep the value conservative because attention-mask memory grows with the configured maximum length. |
|
||||
|
||||
For general parameter tuning, refer to the [Performance Tuning Guide](../../developer_guide/performance_and_debug/optimization_and_tuning.md).
|
||||
|
||||
## 10 FAQ
|
||||
|
||||
For common environment, installation, and general parameter issues, see the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html). This section covers model- and hardware-specific guidance.
|
||||
|
||||
### Atlas 300I DUO runs out of memory during startup
|
||||
|
||||
**Symptom:** The server fails with an out-of-memory error while initializing attention.
|
||||
|
||||
**Cause:** On Atlas 300I DUO, an automatically detected large context length can create a full causal attention mask whose memory consumption grows quadratically with `max_model_len`.
|
||||
|
||||
**Solution:** Always set `--max-model-len` explicitly to a conservative value, such as `4096`, and increase it only after verifying available NPU memory.
|
||||
594
docs/source/tutorials/models/Qwen3-Coder-30B-A3B.md
Normal file
594
docs/source/tutorials/models/Qwen3-Coder-30B-A3B.md
Normal file
@@ -0,0 +1,594 @@
|
||||
# Qwen3-Coder-30B-A3B
|
||||
|
||||
## 1 Introduction
|
||||
|
||||
Qwen3-Coder-30B-A3B is a Mixture-of-Experts (MoE) model in the Qwen3 Coder series, sharing the same architecture as Qwen3-30B-A3B with 30.5B total parameters and 3.3B activated per token. Built upon the Qwen3 base architecture, it delivers significant optimizations in agentic coding, extended context support of up to 1M tokens, and versatile function calling capabilities.
|
||||
|
||||
This document will demonstrate the main validation steps for Qwen3-Coder-30B-A3B in the vLLM-Ascend environment, including supported features, environment preparation, single-node deployment, as well as accuracy and performance evaluation.
|
||||
|
||||
The Qwen3-Coder-30B-A3B model is first supported in **v0.10.0rc1**. This document is validated and written based on **vLLM-Ascend v0.22.1rc**. All **v0.22.1rc and later versions** can run stably. To use the latest features, it is recommended to use the latest release candidate or official version.
|
||||
|
||||
## 2 Supported Features
|
||||
|
||||
Please refer to the [Supported Features List](../../user_guide/support_matrix/supported_models.md) for the model support matrix.
|
||||
|
||||
Please refer to the [Feature Guide](../../user_guide/feature_guide/index.md) for feature configuration information.
|
||||
|
||||
## 3 Prerequisites
|
||||
|
||||
### 3.1 Model Weight
|
||||
|
||||
The following model variants are available. It is recommended to download the model weight to a shared directory accessible to all nodes.
|
||||
|
||||
| Model | Hardware Requirement | Download |
|
||||
| ----------------------------------- | ------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------- |
|
||||
| Qwen3-Coder-30B-A3B-Instruct (BF16) | Atlas 800I A3 (64G, 1\~2 cards)<br>Atlas 800I A2 (64G, 2\~4 cards) | [Download](https://www.modelscope.cn/models/Qwen/Qwen3-Coder-30B-A3B-Instruct) |
|
||||
| Qwen3-Coder-30B-A3B-Instruct-W8A8 | Atlas 800I A3 (64G, 1\~2 cards)<br>Atlas 800I A2 (64G, 2\~4 cards) | [Download](https://www.modelscope.cn/models/Eco-Tech/Qwen3-Coder-30B-A3B-Instruct-w8a8) |
|
||||
| Eagle3 Draft Model | NA | [Download](https://huggingface.co/AngelSlim/Qwen3-a3B_eagle3) |
|
||||
|
||||
These are the recommended numbers of cards, which can be adjusted according to the actual situation.
|
||||
|
||||
If the W8A8 quantized weights are not available for direct download, you can obtain them by quantizing the BF16 model using **msmodelslim**. Refer to the [Quantization Guide](../../user_guide/feature_guide/quantization.md) for details. All model paths in this document should be adjusted to your actual local paths.
|
||||
|
||||
:::{note}
|
||||
Qwen3-Coder-30B-A3B-W8A8 adopts a hybrid quantization strategy (ordered by model structure):
|
||||
|
||||
- **Embedding layer**: BF16 (no quantization)
|
||||
- **Q/K normalization** (q_norm, k_norm): BF16 (weights and biases)
|
||||
- **Attention projections** (q/k/v/o_proj): Static W8A8 with pre-computed per-tensor scales; biases kept in BF16
|
||||
- **MoE routing gate** (mlp.gate): BF16
|
||||
- **MoE expert projections** (gate/up/down_proj): Dynamic W8A8 where input scales are computed on-the-fly during inference
|
||||
:::
|
||||
|
||||
## 4 Installation
|
||||
|
||||
### 4.1 Docker Image Installation
|
||||
|
||||
You can use the official all-in-one Docker image for Qwen3 MoE models.
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: hardware
|
||||
|
||||
::::{tab-item} A3 series
|
||||
:sync: a3
|
||||
|
||||
**Docker Run:**
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-a3
|
||||
|
||||
docker run \
|
||||
--name vllm-ascend-env \
|
||||
--ipc host \
|
||||
--net host \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci8 \
|
||||
--device /dev/davinci9 \
|
||||
--device /dev/davinci10 \
|
||||
--device /dev/davinci11 \
|
||||
--device /dev/davinci12 \
|
||||
--device /dev/davinci13 \
|
||||
--device /dev/davinci14 \
|
||||
--device /dev/davinci15 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/Ascend/driver:/usr/local/Ascend/driver \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /usr/local/sbin:/usr/local/sbin \
|
||||
-it -d $IMAGE bash
|
||||
```
|
||||
|
||||
:::{note}
|
||||
A3 has 8 NPUs with dual-die design (16 chips total: `/dev/davinci[0-15]`).
|
||||
If you are on a shared machine, map only the chips you need (e.g., `/dev/davinci[0-7]` for NPU 0-3).
|
||||
:::
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} A2 series
|
||||
:sync: a2
|
||||
|
||||
**Docker Run:**
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
|
||||
docker run \
|
||||
--name vllm-ascend-env \
|
||||
--ipc host \
|
||||
--net host \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/Ascend/driver:/usr/local/Ascend/driver \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /usr/local/sbin:/usr/local/sbin \
|
||||
-it -d $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
:::::
|
||||
|
||||
:::{tip}
|
||||
The mounts above are the minimum required for NPU driver access. Add additional `-v` mounts (e.g., model weight paths, datasets) as needed for your environment.
|
||||
:::
|
||||
|
||||
The default workdir is `/workspace`. vLLM and vLLM-Ascend are installed as Python packages in site-packages.
|
||||
|
||||
**Installation Verification:**
|
||||
|
||||
After starting the container, run the following command to verify the installation:
|
||||
|
||||
```bash
|
||||
docker ps | grep vllm-ascend-env
|
||||
```
|
||||
|
||||
Expected result: The container is listed with status `Up`. You can also verify the vllm-ascend version inside the container:
|
||||
|
||||
```bash
|
||||
pip show vllm-ascend
|
||||
```
|
||||
|
||||
Expected result: The version information is displayed, matching the pulled image version.
|
||||
|
||||
### 4.2 Source Code Installation
|
||||
|
||||
If you prefer not to use the Docker image, you can build from source. Install vLLM from source first:
|
||||
|
||||
1. Clone and install vLLM:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/vllm-project/vllm.git
|
||||
cd vllm
|
||||
pip install -e .
|
||||
```
|
||||
|
||||
2. Clone and install the vLLM-Ascend repository:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/vllm-project/vllm-ascend.git
|
||||
cd vllm-ascend
|
||||
pip install -e .
|
||||
```
|
||||
|
||||
**Installation Verification:**
|
||||
|
||||
```bash
|
||||
pip show vllm vllm-ascend
|
||||
```
|
||||
|
||||
Expected result: The version information for both packages is displayed, confirming a successful installation.
|
||||
|
||||
:::{note}
|
||||
If deploying a multi-node environment, set up the environment on each node.
|
||||
:::
|
||||
|
||||
For more details, please refer to the [Installation Guide](../../installation.md).
|
||||
|
||||
## 5 Online Service Deployment
|
||||
|
||||
### 5.1 Single-Node Online Deployment
|
||||
|
||||
Single-node deployment completes both Prefill and Decode within the same node, suitable for development, testing, and small-to-medium scale inference scenarios. For the Qwen3-Coder-30B-A3B MoE model, Expert Parallelism (EP) is required to distribute experts across NPUs.
|
||||
|
||||
> The following command is an example configuration. Adjust the parameters based on your actual scenario.
|
||||
|
||||
**Atlas 800I A2/A3:**
|
||||
|
||||
```bash
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3
|
||||
export HCCL_OP_EXPANSION_MODE="AIV" # not needed on A2
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
|
||||
vllm serve your_model_path \
|
||||
--served-model-name qwen3-coder \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 100 \
|
||||
--max-model-len 40960 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--tensor-parallel-size 4 \
|
||||
--enable-expert-parallel \
|
||||
--quantization ascend \
|
||||
--distributed_executor_backend "mp" \
|
||||
--no-enable-prefix-caching \
|
||||
--async-scheduling \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"enable_flashcomm1": true, "weight_nz_mode": 2}' \
|
||||
--gpu-memory-utilization 0.95 \
|
||||
--port 8000 \
|
||||
--speculative-config '{"method": "eagle3", "model": "your_eagle3_model_path", "num_speculative_tokens": 3}'
|
||||
```
|
||||
|
||||
:::{note}
|
||||
|
||||
- `ASCEND_RT_VISIBLE_DEVICES`: must be set to the NPU chip IDs allocated to your environment (e.g., `0,1,2,3` for 4 chips).
|
||||
- `--port`: adjust to avoid conflicts with other services running on the same machine.
|
||||
- `--no-enable-prefix-caching`: disabled by default as prefix caching effectiveness for this model on Ascend NPUs has not been fully characterized. You can try enabling it to evaluate the cache hit rate for your workload.
|
||||
- `--quantization ascend`: required for W8A8 quantized models. Remove this parameter when using BF16 weights.
|
||||
|
||||
:::
|
||||
|
||||
:::{tip}
|
||||
For parameter details, refer to:
|
||||
|
||||
- [vLLM CLI documentation](https://docs.vllm.ai/en/stable/cli/) — standard serve parameters (`--host`, `--port`, `--max-model-len`, etc.)
|
||||
- [Environment Variables](../../user_guide/configuration/env_vars.md) — Ascend-specific environment variables (`HCCL_*`, etc.)
|
||||
- [Additional Configuration](../../user_guide/configuration/additional_config.md) — `--additional-config` format and options
|
||||
:::
|
||||
|
||||
**Service Verification:**
|
||||
|
||||
After the service is started, verify it is running by sending a prompt. Refer to [Section 6](#6-functional-verification) for a usage example.
|
||||
|
||||
## 6 Functional Verification
|
||||
|
||||
After the service is started, the model can be invoked by sending a prompt.
|
||||
|
||||
**Chat Completions API:**
|
||||
|
||||
```shell
|
||||
curl http://localhost:8000/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "qwen3-coder",
|
||||
"messages": [
|
||||
{"role": "user", "content": "Give me a short introduction to large language models."}
|
||||
],
|
||||
"temperature": 0.6,
|
||||
"top_p": 0.95,
|
||||
"top_k": 20,
|
||||
"max_completion_tokens": 4096
|
||||
}'
|
||||
```
|
||||
|
||||
:::{note}
|
||||
Adjust the following fields based on your deployment:
|
||||
|
||||
- **URL** (`http://localhost:8000`): Replace `localhost` and `8000` with your server IP and the `--port` value from the `vllm serve` command.
|
||||
- **`model`**: Must match the `--served-model-name` value from the `vllm serve` command (e.g., `qwen3-coder`).
|
||||
:::
|
||||
|
||||
Expected result: HTTP 200 with a JSON response containing the `choices` field with generated text.
|
||||
|
||||
## 7 Accuracy Evaluation
|
||||
|
||||
### Using AISBench
|
||||
|
||||
For setup details, including installation, dataset download, and configuration, please refer to [Using AISBench](../../developer_guide/evaluation/using_ais_bench.md).
|
||||
|
||||
The following is an example configuration for the accuracy evaluation config file, demonstrated using the GSM8K dataset:
|
||||
|
||||
```python
|
||||
# Example configuration: benchmarks/ais_bench/benchmark/configs/models/vllm_api/vllm_api_general_chat.py
|
||||
from ais_bench.benchmark.models import VLLMCustomAPIChat
|
||||
|
||||
models = [
|
||||
dict(
|
||||
attr="service",
|
||||
type=VLLMCustomAPIChat,
|
||||
abbr='vllm-api-general-chat',
|
||||
path="your_model_path",
|
||||
model="qwen3-coder",
|
||||
request_rate=0,
|
||||
retry=2,
|
||||
host_ip="localhost",
|
||||
host_port=8000,
|
||||
max_out_len=32768,
|
||||
batch_size=32,
|
||||
trust_remote_code=True,
|
||||
generation_kwargs=dict(
|
||||
temperature=0.6,
|
||||
top_k=20,
|
||||
top_p=0.95,
|
||||
),
|
||||
)
|
||||
]
|
||||
```
|
||||
|
||||
Run the accuracy evaluation using the `gsm8k` dataset as an example:
|
||||
|
||||
```shell
|
||||
ais_bench --models vllm_api_general_chat --datasets gsm8k_gen_4_shot_cot_str --mode all --dump-eval-details --debug
|
||||
```
|
||||
|
||||
The following table lists the `--datasets` parameter for each evaluation dataset:
|
||||
|
||||
| Dataset | `--datasets` Parameter |
|
||||
| ------------- | ------------------------------------ |
|
||||
| GSM8K | `gsm8k_gen_4_shot_cot_str` |
|
||||
| GPQA-Diamond | `gpqa_gen_0_shot_cot_chat_prompt` |
|
||||
| AIME 2024 | `aime2024_gen_0_shot_str` |
|
||||
| LiveCodeBench | `livecodebench_0_shot_chat_v4_v5_v6` |
|
||||
|
||||
> The `--models` parameter value corresponds to the configuration file name (e.g., `vllm_api_general_chat` for `vllm_api_general_chat.py`). Adjust `max_out_len`, `batch_size`, and dataset tasks based on your scenario.
|
||||
|
||||
For dataset preparation, please refer to the [AISBench Datasets Guide](https://github.com/AISBench/benchmark/blob/master/docs/source_zh_cn/get_started/datasets.md).
|
||||
|
||||
:::{note}
|
||||
vLLM-Ascend also supports the following evaluation tools:
|
||||
|
||||
- [lm_eval](../../developer_guide/evaluation/using_lm_eval.md)
|
||||
- [OpenCompass](../../developer_guide/evaluation/using_opencompass.md)
|
||||
- [EvalScope](../../developer_guide/evaluation/using_evalscope.md)
|
||||
:::
|
||||
|
||||
**Accuracy Results (Atlas 800I A3, vLLM-Ascend v0.22.1rc, W8A8):**
|
||||
|
||||
| Dataset | Metric | Score |
|
||||
| ------------- | --------------------- | ------ |
|
||||
| GSM8K | accuracy (4-shot CoT) | 90.14% |
|
||||
| GPQA-Diamond | accuracy (0-shot CoT) | 53.54% |
|
||||
| LiveCodeBench | pass@1 (0-shot) | 38.60% |
|
||||
| AIME 2024 | accuracy (0-shot) | 33.33% |
|
||||
|
||||
## 8 Performance Evaluation
|
||||
|
||||
### Using AISBench
|
||||
|
||||
For setup details, please refer to [Using AISBench for performance evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation).
|
||||
|
||||
First, configure the model for streaming performance testing (`ais_bench/benchmark/configs/models/vllm_api/vllm_api_stream_chat.py`):
|
||||
|
||||
```python
|
||||
from ais_bench.benchmark.models import VLLMCustomAPIChat
|
||||
from ais_bench.benchmark.utils.postprocess.model_postprocessors import extract_non_reasoning_content
|
||||
|
||||
models = [
|
||||
dict(
|
||||
attr="service",
|
||||
type=VLLMCustomAPIChat,
|
||||
abbr='vllm-api-stream-chat',
|
||||
path="your_model_path",
|
||||
model="qwen3-coder",
|
||||
stream=True,
|
||||
request_rate=0,
|
||||
retry=2,
|
||||
host_ip="localhost",
|
||||
host_port=8000,
|
||||
max_out_len=1500,
|
||||
batch_size=32,
|
||||
trust_remote_code=True,
|
||||
generation_kwargs=dict(
|
||||
temperature=0.01,
|
||||
ignore_eos=True,
|
||||
),
|
||||
pred_postprocessor=dict(type=extract_non_reasoning_content),
|
||||
)
|
||||
]
|
||||
```
|
||||
|
||||
> Key differences from the accuracy config: `stream=True`, `ignore_eos=True` (ensures output reaches `max_out_len` for consistent TPOT measurement), and `batch_size` controls concurrency.
|
||||
|
||||
Then, configure the synthetic dataset distribution (`ais_bench/datasets/synthetic/synthetic_config.py`). Adjust the configuration based on your actual scenario. Note that random synthetic data is not suitable for benchmarking scenarios where prefix caching is enabled, as random inputs produce zero cache hit rate.
|
||||
|
||||
```python
|
||||
synthetic_config = {
|
||||
"Type": "string",
|
||||
"RequestCount": 200,
|
||||
"StringConfig": {
|
||||
"Input": {
|
||||
"Method": "uniform",
|
||||
"Params": {"MinValue": 3500, "MaxValue": 3500}
|
||||
},
|
||||
"Output": {
|
||||
"Method": "uniform",
|
||||
"Params": {"MinValue": 1500, "MaxValue": 1500}
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Then run the performance evaluation:
|
||||
|
||||
```shell
|
||||
ais_bench --models vllm_api_stream_chat --datasets synthetic_gen --mode perf --debug
|
||||
```
|
||||
|
||||
> The `--models` value should match the `abbr` in your model config file. Use `--num-prompts` to limit the number of test requests.
|
||||
|
||||
### Using vLLM Benchmark
|
||||
|
||||
Refer to [vLLM benchmark](https://docs.vllm.ai/en/latest/benchmarking/) for more details.
|
||||
|
||||
Take the `serve` subcommand as an example. The `--random-output-len` parameter controls the number of output tokens per request; adjust it based on your target scenario (e.g., 2048 for short outputs, 32768 for long outputs).
|
||||
|
||||
```shell
|
||||
vllm bench serve \
|
||||
--model your_model_path \
|
||||
--served-model-name qwen3-coder \
|
||||
--port 8000 \
|
||||
--dataset-name random \
|
||||
--random-input 200 \
|
||||
--random-output-len 2048 \
|
||||
--num-prompts 200 \
|
||||
--request-rate 1 \
|
||||
--save-result \
|
||||
--result-dir ./
|
||||
```
|
||||
|
||||
## 9 Performance Tuning
|
||||
|
||||
### 9.1 Recommended Configurations
|
||||
|
||||
> **Note**: The following configurations are validated in specific test environments and are for reference only. The optimal configuration depends on factors such as maximum input/output length, prefix cache hit rate, precision requirements, and deployment machine ratios. It is recommended to refer to Section 9.2 for tuning based on actual conditions.
|
||||
|
||||
#### Table 1: Scenario Overview
|
||||
|
||||
| Scenario | Deployment Mode | *Total NPUs | Weight Version | Key Considerations |
|
||||
| --------------- | ----------------- | ---------------- | -------------- | ---------------------------------------------------------------- |
|
||||
| High Throughput | Single-Node (TP1) | 1 (A3)<br>2 (A2) | W8A8 | Single-card deployment maximizes concurrent request processing |
|
||||
| Low Latency | Single-Node (TP4) | 2 (A3)<br>4 (A2) | W8A8 | Multi-card TP reduces per-token latency with expert parallelism |
|
||||
| Long Context | Single-Node (TP4) | 2 (A3)<br>4 (A2) | W8A8 | Reduces concurrent sequences to accommodate longer max-model-len |
|
||||
|
||||
> `*Total NPUs` indicates the total number of NPUs used across all nodes. On Atlas 800I A3, each NPU contains two dies (chips), so TP4 requires 4 chips = 2 NPUs.
|
||||
|
||||
#### Table 2: Detailed Node Configuration
|
||||
|
||||
| Scenario | NPUs | TP | max-model-len | max-num-seqs | FUSED_MC2 | EP | hf-overrides |
|
||||
| --------------- | ------ | --- | ------------- | ------------ | --------- | --- | ------------ |
|
||||
| High Throughput | 1 (A3) | 1 | 37364 | 100 | Off | Off | - |
|
||||
| Low Latency | 2 (A3) | 4 | 37364 | 100 | Off | On | - |
|
||||
| Long Context | 2 (A3) | 4 | 131072 | 14 | Off | On | - |
|
||||
|
||||
> For detailed parameter descriptions, please refer to the deployment examples in Section 5.
|
||||
|
||||
**Low Latency Configuration:**
|
||||
|
||||
```shell
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
|
||||
vllm serve your_model_path \
|
||||
--served-model-name qwen3-coder \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 100 \
|
||||
--max-model-len 37364 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--tensor-parallel-size 4 \
|
||||
--enable-expert-parallel \
|
||||
--distributed_executor_backend "mp" \
|
||||
--no-enable-prefix-caching \
|
||||
--async-scheduling \
|
||||
--quantization ascend \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"enable_flashcomm1": true, "weight_nz_mode": 2}' \
|
||||
--gpu-memory-utilization 0.95 \
|
||||
--port 8000 \
|
||||
--speculative-config '{"method": "eagle3","model": "your_eagle3_model_path", "num_speculative_tokens": 3}'
|
||||
```
|
||||
|
||||
:::{tip}
|
||||
Example AISBench settings for this configuration:
|
||||
|
||||
- `request_rate`: 0
|
||||
- `batch_size`: 32
|
||||
- Input/Output length: 2048/2048 or 3500/1500
|
||||
:::
|
||||
|
||||
**High Throughput Configuration:**
|
||||
|
||||
```shell
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
|
||||
vllm serve your_model_path \
|
||||
--served-model-name qwen3-coder \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 100 \
|
||||
--max-model-len 37364 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--tensor-parallel-size 1 \
|
||||
--distributed_executor_backend "mp" \
|
||||
--no-enable-prefix-caching \
|
||||
--async-scheduling \
|
||||
--quantization ascend \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"weight_nz_mode": 2}' \
|
||||
--gpu-memory-utilization 0.95 \
|
||||
--port 8000 \
|
||||
--speculative-config '{"method": "eagle3","model": "your_eagle3_model_path", "num_speculative_tokens": 3}'
|
||||
```
|
||||
|
||||
:::{tip}
|
||||
Example AISBench settings for this configuration:
|
||||
|
||||
- `request_rate`: 0
|
||||
- `batch_size`: 32
|
||||
- Input/Output length: 2048/2048 or 3500/1500
|
||||
:::
|
||||
|
||||
**Long Context Configuration:**
|
||||
|
||||
```shell
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
|
||||
vllm serve your_model_path \
|
||||
--served-model-name qwen3-coder \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 14 \
|
||||
--max-model-len 131072 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--tensor-parallel-size 4 \
|
||||
--enable-expert-parallel \
|
||||
--distributed_executor_backend "mp" \
|
||||
--no-enable-prefix-caching \
|
||||
--async-scheduling \
|
||||
--quantization ascend \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"enable_flashcomm1": true, "weight_nz_mode": 2}' \
|
||||
--gpu-memory-utilization 0.95 \
|
||||
--port 8000 \
|
||||
--speculative-config '{"method": "eagle3","model": "your_eagle3_model_path", "num_speculative_tokens": 3}'
|
||||
```
|
||||
|
||||
:::{tip}
|
||||
Example AISBench settings for this configuration:
|
||||
|
||||
- `request_rate`: 0
|
||||
- `batch_size`: 32
|
||||
- Input/Output length: 65536/1024 or 131072/1024
|
||||
:::
|
||||
|
||||
### 9.2 Tuning Guidelines
|
||||
|
||||
Please refer to the [Public Performance Tuning Documentation](../../developer_guide/performance_and_debug/optimization_and_tuning.md) for tuning methods.
|
||||
Please refer to the [Feature Guide](../../user_guide/support_matrix/feature_matrix.md) for detailed feature descriptions.
|
||||
|
||||
## 10 FAQ
|
||||
|
||||
For common environment, installation, and general parameter issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html). This chapter only covers model-specific issues.
|
||||
|
||||
### Q: How do I enable long context (beyond 256K)?
|
||||
|
||||
Qwen3-Coder-30B-A3B natively supports 256K token context length. For contexts beyond 256K, YaRN rope scaling is required to extend up to 1M. Enable YaRN via `--hf-overrides`:
|
||||
|
||||
```bash
|
||||
--hf-overrides '{"rope_parameters": {"rope_type":"yarn","factor":4,"original_max_position_embeddings":262144}}'
|
||||
```
|
||||
|
||||
For contexts within the native 256K range, no additional configuration is needed. Just set `--max-model-len` to your desired length.
|
||||
|
||||
### Q: What makes Qwen3-Coder different from Qwen3-30B-A3B?
|
||||
|
||||
Qwen3-Coder-30B-A3B shares the same MoE architecture (30.5B/3.3B) as the base Qwen3-30B-A3B but is specifically fine-tuned for coding tasks, with optimizations for agentic coding, function calling, and extended context support up to 1M tokens.
|
||||
651
docs/source/tutorials/models/Qwen3-Dense.md
Normal file
651
docs/source/tutorials/models/Qwen3-Dense.md
Normal file
@@ -0,0 +1,651 @@
|
||||
# Qwen3-Dense (Qwen3-0.6B/1.7B/4B/8B/14B/32B, W8A8, W4A8, W4A4, W8A8SC-310)
|
||||
|
||||
## 1 Introduction
|
||||
|
||||
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support. The Dense variants covered in this document include Qwen3-0.6B, 1.7B, 4B, 8B, 14B, and 32B, along with their quantized versions (W8A8, W4A8, W4A4, and W8A8SC-310) optimized for Ascend NPU deployment.
|
||||
|
||||
This document will demonstrate the main validation steps for Qwen3 Dense models in the vLLM-Ascend environment, including supported features, environment preparation, model quantization, single-node and multi-node deployment, as well as accuracy and performance evaluation. By tailoring service-level configurations to fit different use cases, you can ensure optimal performance across various scenarios.
|
||||
|
||||
The Qwen3 Dense models are first supported in v0.8.4rc2. W8A8 quantization was first supported in v0.8.4rc2, W4A8 quantization is supported since v0.9.1rc2, and W4A4 is supported since v0.11.0rc1. Atlas 300I DUO uses the W8A8SC-310 quantized weights listed in this tutorial. This document is validated and written based on **vLLM-Ascend v0.21.0**. All **v0.21.0 and later versions** can run stably. To use the latest features, it is recommended to use the latest release candidate or official version.
|
||||
|
||||
## 2 Supported Features
|
||||
|
||||
Please refer to the [Supported Features List](../../user_guide/support_matrix/supported_models.md) for the model support matrix.
|
||||
|
||||
Please refer to the [Feature Guide](../../user_guide/feature_guide/index.md) for feature configuration information.
|
||||
|
||||
## 3 Prerequisites
|
||||
|
||||
### 3.1 Model Weight
|
||||
|
||||
The following model variants are available. It is recommended to download the model weight to a shared directory accessible to all nodes.
|
||||
|
||||
**BF16 Versions:**
|
||||
|
||||
| Model | Hardware Requirement | Download |
|
||||
|-------|---------------------|----------|
|
||||
| Qwen3-0.6B | 1 Atlas A3 inference products (64GB × 16), 1 Atlas A2 inference products (64GB × 8) | [Download](https://modelers.cn/models/Modelers_Park/Qwen3-0.6B) |
|
||||
| Qwen3-1.7B | 1 Atlas A3 inference products (64GB × 16), 1 Atlas A2 inference products (64GB × 8) | [Download](https://modelers.cn/models/Modelers_Park/Qwen3-1.7B) |
|
||||
| Qwen3-4B | 1 Atlas A3 inference products (64GB × 16), 1 Atlas A2 inference products (64GB × 8) | [Download](https://modelers.cn/models/Modelers_Park/Qwen3-4B) |
|
||||
| Qwen3-8B | 1 Atlas A3 inference products (64GB × 16), 1 Atlas A2 inference products (64GB × 8) | [Download](https://modelers.cn/models/Modelers_Park/Qwen3-8B) |
|
||||
| Qwen3-14B | 1 Atlas A3 inference products (64GB × 16), 1 Atlas A2 inference products (64GB × 8) | [Download](https://modelers.cn/models/Modelers_Park/Qwen3-14B) |
|
||||
| Qwen3-32B | 1 Atlas A3 inference products (64GB × 16), 1 Atlas A2 inference products (64GB × 8) | [Download](https://modelers.cn/models/Modelers_Park/Qwen3-32B) |
|
||||
|
||||
**Quantized Versions for Atlas A2/A3 inference products:**
|
||||
|
||||
| Model | Quantization | Hardware Requirement | Download |
|
||||
|-------|-------------|---------------------|----------|
|
||||
| Qwen3-8B-W4A8 | W4A8 | 1 Atlas A3 inference products (64GB × 16) or 1 Atlas A2 inference products (64GB × 8) | [Download](https://www.modelscope.cn/models/vllm-ascend/Qwen3-8B-W4A8) |
|
||||
| Qwen3-32B-W4A4 | W4A4 | 1 Atlas A3 inference products (64GB × 16) or 1 Atlas A2 inference products (64GB × 8) | [Download](https://www.modelscope.cn/models/vllm-ascend/Qwen3-32B-W4A4) |
|
||||
| Qwen3-32B-W8A8 | W8A8 | 1 Atlas A3 inference products (64GB × 16) or 1 Atlas A2 inference products (64GB × 8) | [Download](https://www.modelscope.cn/models/vllm-ascend/Qwen3-32B-W8A8) |
|
||||
|
||||
**Quantized Versions for Atlas 300I DUO:**
|
||||
|
||||
| Model | Quantization | Hardware Requirement | Download |
|
||||
|-------|-------------|---------------------|----------|
|
||||
| Qwen3-8B-W8A8SC | W8A8SC | Atlas 300I DUO (TP1) | [Download](https://www.modelscope.cn/models/Eco-Tech/Qwen3-8B-w8a8sc-310-vllm) |
|
||||
| Qwen3-14B-W8A8SC | W8A8SC | Atlas 300I DUO (TP1) | [Download](https://www.modelscope.cn/models/Eco-Tech/Qwen3-14B-w8a8sc-310-vllm) |
|
||||
| Qwen3-32B-W8A8SC | W8A8SC | Atlas 300I DUO (TP4) | [Download](https://www.modelscope.cn/models/Eco-Tech/Qwen3-32B-w8a8sc-310-vllm) |
|
||||
|
||||
These are the recommended numbers of cards, which can be adjusted according to the actual situation.
|
||||
|
||||
### 3.2 Verify Multi-node Communication
|
||||
|
||||
If you need to deploy a multi-node environment, verify the multi-node communication according to [Verify Multi-node Communication Environment](../../installation.md#verify-multi-node-communication).
|
||||
|
||||
## 4 Installation
|
||||
|
||||
### 4.1 Docker Image Installation
|
||||
|
||||
You can use the official all-in-one Docker image for Qwen3 Dense models.
|
||||
|
||||
For Atlas 300I DUO, use `vllm-ascend:nightly-releases-v0.23.0-310p` (or a later `-310p` image).
|
||||
|
||||
**Docker Pull:**
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
docker pull quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
```
|
||||
|
||||
**Docker Run:**
|
||||
|
||||
Start the docker image on each node.
|
||||
|
||||
:::::{tab-set}
|
||||
|
||||
::::{tab-item} Atlas A3 inference products
|
||||
:sync: A3
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-a3
|
||||
|
||||
docker run --rm \
|
||||
--name vllm-ascend-env \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci8 \
|
||||
--device /dev/davinci9 \
|
||||
--device /dev/davinci10 \
|
||||
--device /dev/davinci11 \
|
||||
--device /dev/davinci12 \
|
||||
--device /dev/davinci13 \
|
||||
--device /dev/davinci14 \
|
||||
--device /dev/davinci15 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
:::{note}
|
||||
Atlas A3 inference products have 8 NPUs with dual-die design (16 chips total: `/dev/davinci[0-15]`).
|
||||
If you are on a shared machine, map only the chips you need (e.g., `/dev/davinci[0-7]` for NPU 0-3).
|
||||
:::
|
||||
|
||||
::::
|
||||
::::{tab-item} Atlas A2 inference products
|
||||
:sync: A2
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
|
||||
docker run --rm \
|
||||
--name vllm-ascend-env \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
# Use the vllm-ascend image
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-310p
|
||||
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=10g \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-p 8080:8080 \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
The default workdir is `/workspace`. vLLM and vLLM-Ascend are installed as Python packages in site-packages.
|
||||
|
||||
Installation Verification:
|
||||
After starting the container, run the following command to verify the installation:
|
||||
|
||||
```bash
|
||||
docker ps | grep vllm-ascend-env
|
||||
```
|
||||
|
||||
Expected result: The container is listed with status Up. You can also verify the vllm-ascend version inside the container:
|
||||
|
||||
```bash
|
||||
pip show vllm-ascend
|
||||
```
|
||||
|
||||
Expected result: The version information is displayed, matching the pulled image version.
|
||||
|
||||
### 4.2 Source Code Installation
|
||||
|
||||
If you prefer to build from source instead of using the Docker image, install vLLM-Ascend following the [Installation Guide](../../installation.md).
|
||||
|
||||
:::{note}
|
||||
For Atlas 300I DUO, source installation may pull in `triton` and `triton-ascend`. Uninstall them before running vLLM-Ascend on Atlas 300I DUO:
|
||||
|
||||
```bash
|
||||
pip uninstall -y triton-ascend triton
|
||||
```
|
||||
|
||||
:::
|
||||
|
||||
To verify the source installation:
|
||||
|
||||
```bash
|
||||
pip show vllm-ascend
|
||||
```
|
||||
|
||||
Expected result: The version information is displayed, confirming a successful installation.
|
||||
|
||||
:::{note}
|
||||
If deploying a multi-node environment, set up the environment on each node.
|
||||
:::
|
||||
|
||||
## 5 Online Service Deployment
|
||||
|
||||
### 5.1 Single-Node Online Deployment
|
||||
|
||||
Single-node deployment completes both Prefill and Decode within the same node, suitable for development, testing, and small-to-medium scale inference scenarios.
|
||||
|
||||
**Start the server:**
|
||||
> The following command is an example configuration. Adjust the parameters based on your actual scenario.
|
||||
|
||||
:::::{tab-set}
|
||||
::::{tab-item} Atlas A2 inference products / Atlas A3 inference products
|
||||
|
||||
Qwen3-32B-W8A8:
|
||||
|
||||
```bash
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
|
||||
vllm serve your_model_path \
|
||||
--served-model-name qwen3 \
|
||||
--trust-remote-code \
|
||||
--async-scheduling \
|
||||
--quantization ascend \
|
||||
--distributed-executor-backend mp \
|
||||
--tensor-parallel-size 4 \
|
||||
--max-model-len 5500 \
|
||||
--max-num-batched-tokens 40960 \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
|
||||
--port <port> \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--additional-config '{"enable_flashcomm1": true}'
|
||||
```
|
||||
|
||||
Qwen3-32B-W4A4:
|
||||
|
||||
```bash
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1
|
||||
export VLLM_USE_V1=1
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export HCCL_BUFFSIZE=1024
|
||||
vllm serve your_model_path \
|
||||
--port 8004 \
|
||||
--data-parallel-size 1 \
|
||||
--tensor-parallel-size 2 \
|
||||
--served-model-name qwen3 \
|
||||
--distributed-executor-backend "mp" \
|
||||
--max-model-len 40960 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--max-num-seqs 64 \
|
||||
--trust-remote-code \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--quantization ascend \
|
||||
--compilation-config '{"cudagraph_capture_sizes": [64]}' \
|
||||
--additional-config '{"enable_flashcomm1": true, "ascend_compilation_config": {"fuse_norm_quant": false}}'
|
||||
```
|
||||
|
||||
Qwen3-8B-W4A8:
|
||||
|
||||
```bash
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
vllm serve your_model_path \
|
||||
--served-model-name qwen3 \
|
||||
--max-model-len 4096 \
|
||||
--port 20001 \
|
||||
--additional-config '{"ascend_compilation_config": {"enable_npugraph_ex": false}}' \
|
||||
--quantization ascend
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
|
||||
Atlas 300I DUO uses the W8A8SC-310 quantized weights from the Eco-Tech official ModelScope repository. Keep an explicit `--max-model-len` for 310P deployment to avoid OOM caused by oversized attention mask allocation.
|
||||
|
||||
Qwen3-8B-W8A8SC:
|
||||
|
||||
```bash
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0
|
||||
|
||||
vllm serve Eco-Tech/Qwen3-8B-w8a8sc-310-vllm/TP1/Qwen3-8B-w8a8sc-310-vllm-tp1 \
|
||||
--host 127.0.0.1 \
|
||||
--port 8080 \
|
||||
--tensor-parallel-size 1 \
|
||||
--gpu-memory-utilization 0.90 \
|
||||
--max-num-seqs 32 \
|
||||
--served-model-name qwen3 \
|
||||
--dtype float16 \
|
||||
--additional-config '{"ascend_compilation_config": {"enable_npugraph_ex":false}}' \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY", "cudagraph_capture_sizes": [1,2,4,8,16,32]}' \
|
||||
--quantization ascend \
|
||||
--max-model-len 16384 \
|
||||
--no-enable-prefix-caching \
|
||||
--load-format sharded_state
|
||||
```
|
||||
|
||||
Qwen3-14B-W8A8SC:
|
||||
|
||||
```bash
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0
|
||||
|
||||
vllm serve Eco-Tech/Qwen3-14B-w8a8sc-310-vllm/TP1/Qwen3-14B-w8a8sc-310-vllm-tp1 \
|
||||
--host 127.0.0.1 \
|
||||
--port 8080 \
|
||||
--tensor-parallel-size 1 \
|
||||
--gpu-memory-utilization 0.90 \
|
||||
--max-num-seqs 16 \
|
||||
--served-model-name qwen3 \
|
||||
--dtype float16 \
|
||||
--additional-config '{"ascend_compilation_config": {"enable_npugraph_ex":false}}' \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY", "cudagraph_capture_sizes": [1,2,4,8,16]}' \
|
||||
--quantization ascend \
|
||||
--max-model-len 16384 \
|
||||
--no-enable-prefix-caching \
|
||||
--load-format sharded_state
|
||||
```
|
||||
|
||||
Qwen3-32B-W8A8SC:
|
||||
|
||||
```bash
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3
|
||||
|
||||
vllm serve Eco-Tech/Qwen3-32B-w8a8sc-310-vllm/TP4/Qwen3-32B-w8a8sc-310-vllm-tp4 \
|
||||
--host 127.0.0.1 \
|
||||
--port 8080 \
|
||||
--tensor-parallel-size 4 \
|
||||
--gpu-memory-utilization 0.90 \
|
||||
--max-num-seqs 32 \
|
||||
--served-model-name qwen3 \
|
||||
--dtype float16 \
|
||||
--additional-config '{"ascend_compilation_config": {"enable_npugraph_ex":false}}' \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY", "cudagraph_capture_sizes": [16,32]}' \
|
||||
--quantization ascend \
|
||||
--max-model-len 20480 \
|
||||
--no-enable-prefix-caching \
|
||||
--load-format sharded_state
|
||||
```
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
:::{note}
|
||||
|
||||
- [vLLM Serving Arguments documentation](https://docs.vllm.com.cn/en/latest/cli/serve/?h=block+size#arguments) — Additional parameter details for vLLM serve commands.
|
||||
- [Environment Variables](../../user_guide/configuration/env_vars.md) — Ascend-specific environment variables (`HCCL_*`, etc.).
|
||||
:::
|
||||
|
||||
**Service Verification:**
|
||||
|
||||
If the service starts successfully, the following startup log will be displayed:
|
||||
|
||||
```text
|
||||
(APIServer pid=<pid>) INFO: Started server process [<pid>]
|
||||
(APIServer pid=<pid>) INFO: Waiting for application startup.
|
||||
(APIServer pid=<pid>) INFO: Application startup complete.
|
||||
```
|
||||
|
||||
## 6 Functional Verification
|
||||
|
||||
After the service is started, the model can be invoked by sending a prompt.
|
||||
|
||||
**Chat Completions API:**
|
||||
|
||||
```bash
|
||||
curl http://localhost:<port>/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "qwen3",
|
||||
"messages": [
|
||||
{"role": "user", "content": "Give me a short introduction to large language models."}
|
||||
],
|
||||
"temperature": 0.6,
|
||||
"top_p": 0.95,
|
||||
"top_k": 20,
|
||||
"max_completion_tokens": 4096
|
||||
}'
|
||||
```
|
||||
|
||||
Expected result: HTTP 200 with a JSON response containing the `choices` field with generated text.
|
||||
|
||||
## 7 Accuracy Evaluation
|
||||
|
||||
### Using AISBench
|
||||
|
||||
For setup details, including installation, dataset download, and configuration, please refer to [Using AISBench](../../developer_guide/evaluation/using_ais_bench.md).
|
||||
|
||||
The following is an example configuration for the accuracy evaluation config file:
|
||||
|
||||
**Accuracy Evaluation Config File:**
|
||||
|
||||
```bash
|
||||
# Example configuration: benchmarks/ais_bench/benchmark/configs/models/vllm_api/vllm_api_general_chat.py
|
||||
from ais_bench.benchmark.models import VLLMCustomAPIChat
|
||||
from ais_bench.benchmark.utils.model_postprocessors import extract_non_reasoning_content
|
||||
|
||||
models = [
|
||||
dict(
|
||||
attr="service",
|
||||
type=VLLMCustomAPIChat,
|
||||
abbr='vllm-api-general-chat',
|
||||
path="your_model_path",
|
||||
model="qwen3",
|
||||
request_rate=0,
|
||||
retry=2,
|
||||
host_ip="127.0.0.1",
|
||||
host_port=2001,
|
||||
max_out_len=32768,
|
||||
batch_size=32,
|
||||
trust_remote_code=False,
|
||||
generation_kwargs=dict(
|
||||
temperature=0.6,
|
||||
top_k=20,
|
||||
top_p=0.95,
|
||||
),
|
||||
pred_postprocessor=dict(type=extract_non_reasoning_content)
|
||||
)
|
||||
]
|
||||
```
|
||||
|
||||
**Run the accuracy evaluation using the aime2025 dataset as an example:**
|
||||
|
||||
```bash
|
||||
ais_bench --models vllm_api_general_chat --datasets aime2025_gen_0_shot_chat_prompt --debug
|
||||
```
|
||||
|
||||
> The --models parameter value corresponds to the abbr field in the configuration file above. Adjust max_out_len, batch_size, and dataset tasks based on your scenario.
|
||||
|
||||
## 8 Performance Evaluation
|
||||
|
||||
### Using AISBench
|
||||
|
||||
For setup details, including installation, dataset download, and configuration, please refer to [Using AISBench](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation) for details.
|
||||
|
||||
The following is an example configuration for the accuracy evaluation config file:
|
||||
|
||||
```bash
|
||||
# Example configuration: benchmarks/ais_bench/benchmark/configs/models/vllm_api/vllm_api_stream_chat.py
|
||||
from ais_bench.benchmark.models import VLLMCustomAPIChat
|
||||
from ais_bench.benchmark.utils.postprocess.model_postprocessors import extract_non_reasoning_content
|
||||
|
||||
models = [
|
||||
dict(
|
||||
attr="service",
|
||||
type=VLLMCustomAPIChat,
|
||||
abbr="vllm-api-stream-chat",
|
||||
path="your_model_path",
|
||||
model="qwen3",
|
||||
stream=True,
|
||||
request_rate=0,
|
||||
use_timestamp=False,
|
||||
retry=2,
|
||||
host_ip="127.0.0.1",
|
||||
host_port=8004,
|
||||
max_out_len=1500,
|
||||
batch_size=90,
|
||||
trust_remote_code=False,
|
||||
generation_kwargs=dict(
|
||||
temperature=0,
|
||||
ignore_eos=True
|
||||
),
|
||||
)
|
||||
]
|
||||
```
|
||||
|
||||
**Run the performance evaluation using the GSM8K dataset as an example:**
|
||||
|
||||
```bash
|
||||
ais_bench --models vllm_api_stream_chat --datasets gsm8k_gen_0_shot_cot_str_perf --debug --summarizer default_perf --mode perf --num-prompts 360
|
||||
```
|
||||
|
||||
### Using vLLM Benchmark
|
||||
|
||||
Refer to [vLLM benchmark](https://docs.vllm.ai/en/latest/benchmarking/) for more details.
|
||||
|
||||
There are three `vllm bench` subcommands:
|
||||
|
||||
- `latency`: Benchmark the latency of a single batch of requests.
|
||||
- `serve`: Benchmark the online serving throughput.
|
||||
- `throughput`: Benchmark offline inference throughput.
|
||||
|
||||
Take `serve` as an example:
|
||||
|
||||
```shell
|
||||
vllm bench serve \
|
||||
--model your_model_path \
|
||||
--served-model-name qwen3 \
|
||||
--port <port> \
|
||||
--dataset-name random \
|
||||
--random-input 200 \
|
||||
--num-prompts 200 \
|
||||
--request-rate 1 \
|
||||
--save-result \
|
||||
--result-dir ./
|
||||
```
|
||||
|
||||
After several minutes, you will get the performance evaluation result.
|
||||
|
||||
## 9 Performance Tuning
|
||||
|
||||
### 9.1 Recommended Configurations
|
||||
|
||||
> **Note**: The following configurations are validated in specific test environments and are for reference only. The optimal configuration depends on factors such as maximum input/output length, prefix cache hit rate, precision requirements, and deployment machine ratios. It is recommended to refer to Section 9.2 for tuning based on actual conditions.
|
||||
|
||||
#### Table 1: Scenario Overview
|
||||
|
||||
| Scenario | Deployment Mode | *Total NPUs | Weight Version | Key Considerations |
|
||||
|----------|----------------|-------------|----------------|---------------------|
|
||||
| High Throughput | Single-Node (TP4) | 4 (Atlas A3 inference products) | W8A8 | 4-card TP maximizes concurrent request processing |
|
||||
| Long Context | Single-Node (TP4) | 4 (Atlas A3 inference products) | W8A8 | 4-card TP extends context window for long sequences |
|
||||
| Low Latency | Single-Node (TP8) | 8 (Atlas A3 inference products) | W8A8 | 8-card TP reduces per-token latency for interactive responses |
|
||||
|
||||
> `*Total NPUs` indicates the total number of NPUs used across all nodes.
|
||||
|
||||
#### Table 2: Detailed Node Configuration
|
||||
|
||||
| Scenario | Configuration | NPUs | TP | DP | FUSED_MC2 | EP Switch | Async Scheduling |
|
||||
|----------|---------------|-------|----|----|-------------|--------------|--------------|
|
||||
| High Throughput | Single-Node | 4 | 4 | 1 | Off | Off | On |
|
||||
| Long Context | Single-Node | 4 | 4 | 1 | Off | Off | On |
|
||||
| Low Latency | Single-Node | 8 | 8 | 1 | Off | Off | On |
|
||||
|
||||
For detailed parameter descriptions, please refer to the deployment examples in [Section 5](#5-online-service-deployment)
|
||||
|
||||
<u>High Throughput Configuration:</u>
|
||||
|
||||
```bash
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
|
||||
vllm serve your_model_path \
|
||||
--served-model-name qwen3 \
|
||||
--trust-remote-code \
|
||||
--distributed-executor-backend mp \
|
||||
--tensor-parallel-size 4 \
|
||||
--max-model-len 5500 \
|
||||
--max-num-batched-tokens 40960 \
|
||||
--no-enable-prefix-caching \
|
||||
--async-scheduling \
|
||||
--quantization ascend \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY","cudagraph_capture_sizes":[4,8,64,72,76,80,96,100,120,140,144,160,192,216,240,252,288,320,336,360,384,400,408,416,420,432,480,540,576,600]}' \
|
||||
--additional-config '{"weight_prefetch_config":{"enabled":true}, "enable_flashcomm1": true}' \
|
||||
--host <host_ip> \
|
||||
--port <port> \
|
||||
--block-size 128 \
|
||||
--gpu-memory-utilization 0.9
|
||||
```
|
||||
|
||||
<u>Long Context Configuration:</u>
|
||||
|
||||
```bash
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
|
||||
vllm serve your_model_path \
|
||||
--host <host_ip> \
|
||||
--port <port> \
|
||||
--served-model-name qwen3 \
|
||||
--trust-remote-code \
|
||||
--seed 1024 \
|
||||
--max-model-len 135000 \
|
||||
--max-num-batched-tokens 40960 \
|
||||
--tensor-parallel-size 4 \
|
||||
--distributed-executor-backend "mp" \
|
||||
--async-scheduling \
|
||||
--no-enable-prefix-caching \
|
||||
--speculative-config '{"method": "eagle3", "model":"your_eagle3_model_path", "num_speculative_tokens": 3}' \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
|
||||
--hf-overrides '{"rope_parameters": {"rope_type":"yarn","rope_theta":1000000,"factor":4,"original_max_position_embeddings":131072}}' \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--quantization ascend \
|
||||
--additional-config '{"enable_flashcomm1": true}'
|
||||
```
|
||||
|
||||
<u>Low Latency Configuration:</u>
|
||||
|
||||
```bash
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3,4,5,6,7
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
|
||||
vllm serve your_model_path \
|
||||
--served-model-name qwen3 \
|
||||
--trust-remote-code \
|
||||
--distributed-executor-backend mp \
|
||||
--tensor-parallel-size 8 \
|
||||
--max-model-len 5500 \
|
||||
--max-num-batched-tokens 40960 \
|
||||
--no-enable-prefix-caching \
|
||||
--async-scheduling \
|
||||
--quantization ascend \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY","cudagraph_capture_sizes":[1,2,4,8,16,32,64,72,76,80,96,100,120,140,144,160,192,216,240,252,288,320,336,360,384,400,408,416,420,432,480,540,576,600]}' \
|
||||
--speculative-config '{"method": "eagle3", "model":"your_eagle3_model_path", "enforce_eager": true, "num_speculative_tokens": 3}' \
|
||||
--port <port> \
|
||||
--block-size 128 \
|
||||
--gpu-memory-utilization 0.9
|
||||
```
|
||||
|
||||
### 9.2 Tuning Guidelines
|
||||
|
||||
#### 9.2.1 General Tuning Reference
|
||||
|
||||
Please refer to the [Public Performance Tuning Documentation](../../developer_guide/performance_and_debug/optimization_and_tuning.md) for tuning methods.
|
||||
Please refer to the [Feature Guide](../../user_guide/support_matrix/feature_matrix.md) for detailed feature descriptions.
|
||||
|
||||
## 10 FAQ
|
||||
|
||||
For common environment, installation, and general parameter issues, please refer to the [vLLM-Ascend FAQs](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html). This section only covers issues specific to Qwen3 Dense models.
|
||||
286
docs/source/tutorials/models/Qwen3-Embedding.md
Normal file
286
docs/source/tutorials/models/Qwen3-Embedding.md
Normal file
@@ -0,0 +1,286 @@
|
||||
# Qwen3-Embedding
|
||||
|
||||
## 1 Introduction
|
||||
|
||||
The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. Building upon the dense foundational models of the Qwen3 series, it provides a comprehensive range of text embeddings and reranking models in various sizes (0.6B, 4B, and 8B). This guide describes how to run the model with vLLM Ascend. Note that only vLLM Ascend 0.9.2rc1 and higher versions support the model.
|
||||
|
||||
## 2 Supported Features
|
||||
|
||||
Refer to [supported features](../../user_guide/support_matrix/supported_models.md) to get the model's supported feature matrix.
|
||||
|
||||
## 3 Prerequisites
|
||||
|
||||
### 3.1 Model Weight
|
||||
|
||||
- `Qwen3-Embedding-8B` [Download model weight](https://www.modelscope.cn/models/Qwen/Qwen3-Embedding-8B)
|
||||
- `Qwen3-Embedding-4B` [Download model weight](https://www.modelscope.cn/models/Qwen/Qwen3-Embedding-4B)
|
||||
- `Qwen3-Embedding-0.6B` [Download model weight](https://www.modelscope.cn/models/Qwen/Qwen3-Embedding-0.6B)
|
||||
|
||||
It is recommended to download the model weight to the shared directory of multiple nodes, such as `/root/.cache/`
|
||||
|
||||
## 4 Installation
|
||||
|
||||
### 4.1 Docker Image Installation
|
||||
|
||||
You can use our official docker image to run `Qwen3-Embedding` model directly.
|
||||
|
||||
Select an image based on your machine type and start the docker image on your node, refer to [using docker](../../installation.md#set-up-using-docker).
|
||||
|
||||
For Atlas 300I DUO, use `vllm-ascend:nightly-releases-v0.23.0-310p` (or a later `-310p` image).
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: install
|
||||
|
||||
::::{tab-item} A3 series
|
||||
:sync: A3 series
|
||||
|
||||
Start the docker image on each node.
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-a3
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--privileged=true \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} A2 series
|
||||
:sync: A2 series
|
||||
|
||||
Start the docker image on each node.
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--privileged=true \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
:sync: Atlas 300I DUO
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-310p
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--privileged=true \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
After a successful docker run, you can verify the running container service by executing the `docker ps` command.
|
||||
|
||||
### 4.2 Source Code Installation
|
||||
|
||||
If you don't want to use the docker image as above, you can also build all from source:
|
||||
|
||||
- Install `vllm-ascend` from source, refer to [installation](../../installation.md).
|
||||
|
||||
If you want to deploy multi-node environment, you need to set up environment on each node.
|
||||
|
||||
## 5 Online Service Deployment
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: Deployment
|
||||
|
||||
::::{tab-item} A3/A2 series
|
||||
:sync: A3/A2 series
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
vllm serve Qwen/Qwen3-Embedding-0.6B \
|
||||
--served-model-name Qwen/Qwen3-Embedding-0.6B \
|
||||
--runner pooling \
|
||||
--port 8000 \
|
||||
--max-model-len 1024
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
:sync: Atlas 300I DUO
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
vllm serve Qwen/Qwen3-Embedding-0.6B \
|
||||
--served-model-name Qwen/Qwen3-Embedding-0.6B \
|
||||
--compilation-config '{"cudagraph_capture_sizes": [1024,512]}' \
|
||||
--additional-config '{"ascend_compilation_config": {"fuse_norm_quant": false}}' \
|
||||
--runner pooling \
|
||||
--dtype float16 \
|
||||
--port 8000 \
|
||||
--max-model-len 1024
|
||||
```
|
||||
|
||||
Required Parameter Descriptions:
|
||||
|
||||
`--compilation-config` For Atlas 300I DUO, due to limited hardware streams, the size of cudagraph_capture_sizes is restricted.
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
Key Parameter Descriptions:
|
||||
|
||||
- `--max-model-len` represents the context length, which is the maximum value of the input plus output for a single request. For Atlas 300I DUO if automatic parsing resolves to a large context length, allocating this mask (O(max_model_len^2)) may exceed NPU memory and trigger OOM. Be sure to set an explicit and conservative value, such as --max-model-len 1024.
|
||||
|
||||
Common Issues Tip: If you encounter issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html) for troubleshooting.
|
||||
|
||||
## 6 Functional Verification
|
||||
|
||||
Once your server is started, you can verify by follow command:
|
||||
|
||||
Service Verification:
|
||||
|
||||
```bash
|
||||
curl -X POST http://localhost:8000/v1/embeddings -H "Content-Type: application/json" -d '{
|
||||
"input": [
|
||||
"The capital of China is Beijing.",
|
||||
"Gravity is a force that attracts two bodies towards each other. It gives weight to physical objects and is responsible for the movement of planets around the sun."
|
||||
]
|
||||
}'
|
||||
```
|
||||
|
||||
Expected Result:
|
||||
|
||||
The service returns HTTP 200 OK with a JSON response containing the `embedding` field. Example output:
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "embd-8136155c01e8411d",
|
||||
"object": "list",
|
||||
"created": 1784538286,
|
||||
"model": "Qwen/Qwen3-Embedding-0.6B",
|
||||
"data": [
|
||||
{
|
||||
"index": 0,
|
||||
"object": "embedding",
|
||||
"embedding": [
|
||||
-0.04725276678800583,-0.021066857501864433
|
||||
]
|
||||
},
|
||||
{
|
||||
"index": 1,
|
||||
"object": "embedding",
|
||||
"embedding": [
|
||||
-0.053165290504693985,-0.01480848714709282
|
||||
]
|
||||
}
|
||||
],
|
||||
"usage": {
|
||||
"prompt_tokens": 39,
|
||||
"total_tokens": 39,
|
||||
"completion_tokens": 0,
|
||||
"prompt_tokens_details": null
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
For more usage examples, please reference the [examples](https://github.com/vllm-project/vllm/tree/main/examples/pooling/embed)
|
||||
|
||||
## 7 Accuracy Evaluation
|
||||
|
||||
Here are two accuracy evaluation methods.
|
||||
|
||||
### Using MTEB
|
||||
|
||||
1. Refer to [MTEB](https://docs.mteb.org/) for details.
|
||||
|
||||
2. Run follow code to execute the accuracy evaluation.
|
||||
|
||||
```python
|
||||
|
||||
import os
|
||||
import mteb
|
||||
|
||||
from mteb.models.vllm_wrapper import VllmEncoderWrapper
|
||||
|
||||
if __name__ == "__main__":
|
||||
|
||||
data_path = "/home/data/mteb_data"
|
||||
os.environ["HF_DATASETS_CACHE"] = data_path
|
||||
os.environ["HF_ENDPOINT"] = "https://hf-mirror.com"
|
||||
|
||||
model = VllmEncoderWrapper(f"/root/.cache/Qwen3-Embedding-0.6B",
|
||||
revision="norm",
|
||||
dtype="float16",
|
||||
max_model_len=10240,
|
||||
)
|
||||
|
||||
cache = mteb.ResultCache("/home/data/mteb_data")
|
||||
tasks = mteb.get_tasks(tasks=["LeCaRDv2"])
|
||||
results = mteb.evaluate(model, tasks=tasks, cache=cache, encode_kwargs={"batch_size": 2}, overwrite_strategy="always")
|
||||
df = results.to_dataframe()
|
||||
print(df)
|
||||
```
|
||||
|
||||
3. After execution, you can get the result.
|
||||
|
||||
## 8 Performance Evaluation
|
||||
|
||||
### Using vLLM Benchmark
|
||||
|
||||
Run performance of `Qwen3-Embedding-0.6B` as an example.
|
||||
Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/cli/) for more details.
|
||||
|
||||
Take the `serve` as an example. Run the code as follows.
|
||||
|
||||
```bash
|
||||
vllm bench serve --model Qwen/Qwen3-Embedding-0.6B --backend openai-embeddings --port 8000 --dataset-name random --endpoint /v1/embeddings --random-input 200 --save-result --result-dir ./
|
||||
```
|
||||
|
||||
After about several minutes, you can get the performance evaluation result.
|
||||
|
||||
## 9 FAQ
|
||||
|
||||
For common environment, installation, and general parameter issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html).
|
||||
283
docs/source/tutorials/models/Qwen3-Next.md
Normal file
283
docs/source/tutorials/models/Qwen3-Next.md
Normal file
@@ -0,0 +1,283 @@
|
||||
# Qwen3-Next
|
||||
|
||||
## 1 Introduction
|
||||
|
||||
The Qwen3-Next model is a sparse MoE (Mixture of Experts) model with high sparsity. Compared to the MoE architecture of Qwen3, it has introduced key improvements in aspects such as the hybrid attention mechanism and multi-token prediction mechanism, enhancing the training and inference efficiency of the model under long contexts and large total parameter scales.
|
||||
|
||||
This document will present the core verification steps of the model, including supported features, environment preparation, as well as accuracy and performance evaluation. Qwen3-Next is currently using Triton Ascend, which is in the experimental phase. In subsequent versions, its performance related to stability and accuracy may change, and performance will be continuously optimized.
|
||||
|
||||
The `Qwen3-Next` model is first supported in `vllm-ascend:v0.10.2rc1` and can stably run in v0.16.0 and later version.
|
||||
|
||||
## 2 Supported Features
|
||||
|
||||
Refer to [Supported Features List](../../user_guide/support_matrix/supported_models.md) to get the model's supported feature matrix.
|
||||
|
||||
Refer to [Feature Guide](../../user_guide/feature_guide/index.md) to get the feature's configuration.
|
||||
|
||||
## 3 Prerequisites
|
||||
|
||||
### 3.1 Model Weight
|
||||
|
||||
`Qwen3-Next-80B-A3B-Instruct`: requires **8 cards in 1 Atlas 800 A3 (64GB × 16) node** or **8 cards in 1 Atlas 800 A2 (64GB × 8) node**. [Model Weight](https://www.modelscope.cn/models/Qwen/Qwen3-Next-80B-A3B-Instruct)
|
||||
|
||||
## 4 Installation
|
||||
|
||||
### 4.1 Docker Image Installation
|
||||
|
||||
Select an image based on your machine type and start the docker image on your node, refer to [using docker](../../installation.md#set-up-using-docker).
|
||||
|
||||
**A3 series:**
|
||||
|
||||
Start the docker image on each node.
|
||||
|
||||
```bash
|
||||
#!/bin/sh
|
||||
# Update the vllm-ascend image
|
||||
# For Atlas A2 machines:
|
||||
# export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
# For Atlas A3 machines:
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-a3
|
||||
docker run --rm \
|
||||
--shm-size=1g \
|
||||
--name vllm-ascend-qwen3 \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci8 \
|
||||
--device /dev/davinci9 \
|
||||
--device /dev/davinci10 \
|
||||
--device /dev/davinci11 \
|
||||
--device /dev/davinci12 \
|
||||
--device /dev/davinci13 \
|
||||
--device /dev/davinci14 \
|
||||
--device /dev/davinci15 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-p 8000:8000 \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
The Qwen3-Next is using [Triton Ascend](https://gitee.com/ascend/triton-ascend) which is currently experimental. In future versions, there may be behavioral changes related to stability, accuracy, and performance improvement.
|
||||
|
||||
**Installation Verification:**
|
||||
|
||||
```bash
|
||||
pip show vllm vllm-ascend
|
||||
```
|
||||
|
||||
Expected result: The version information for both packages is displayed, confirming a successful installation.
|
||||
|
||||
:::{note}
|
||||
If deploying a multi-node environment, set up the environment on each node.
|
||||
:::
|
||||
|
||||
For more details, please refer to the [Installation Guide](../../installation.md).
|
||||
|
||||
### 4.2 Source Code Installation
|
||||
|
||||
If you prefer not to use the Docker image, you can build from source. Install vLLM from source first:
|
||||
|
||||
1. Clone and install vLLM:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/vllm-project/vllm.git
|
||||
cd vllm
|
||||
pip install -e .
|
||||
```
|
||||
|
||||
2. Clone and install the vLLM-Ascend repository:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/vllm-project/vllm-ascend.git
|
||||
cd vllm-ascend
|
||||
pip install -e .
|
||||
```
|
||||
|
||||
**Installation Verification:**
|
||||
|
||||
```bash
|
||||
pip show vllm vllm-ascend
|
||||
```
|
||||
|
||||
Expected result: The version information for both packages is displayed, confirming a successful installation.
|
||||
|
||||
:::{note}
|
||||
If deploying a multi-node environment, set up the environment on each node.
|
||||
:::
|
||||
|
||||
For more details, please refer to the [Installation Guide](../../installation.md).
|
||||
|
||||
## 5 Online Service Deployment
|
||||
|
||||
### 5.1 Single-Node Online Deployment
|
||||
|
||||
Single-node deployment completes both Prefill and Decode within the same node. The model `Qwen3-Next-80B-A3B-Instruct` can be deployed on 1 Atlas 800 A3 (64GB × 16).
|
||||
|
||||
While a single-node setup supports all input/output scenarios, consider deploying multi-node clusters for optimal performance.
|
||||
|
||||
Startup Command:
|
||||
|
||||
```bash
|
||||
vllm serve Qwen/Qwen3-Next-80B-A3B-Instruct --served-model-name qwen3_next --tensor-parallel-size 4 --max-model-len 32768 --gpu-memory-utilization 0.8 --max-num-batched-tokens 4096 --compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}'
|
||||
```
|
||||
|
||||
If your service start successfully, you can see the info shown below:
|
||||
|
||||
```bash
|
||||
INFO: Started server process [2736]
|
||||
INFO: Waiting for application startup.
|
||||
INFO: Application startup complete.
|
||||
```
|
||||
|
||||
## 6 Functional Verification
|
||||
|
||||
Once your server is started, you can query the model with input prompts:
|
||||
|
||||
```bash
|
||||
curl http://localhost:8000/v1/chat/completions -H "Content-Type: application/json" -d '{
|
||||
"model": "qwen3_next",
|
||||
"messages": [
|
||||
{"role": "user", "content": "Who are you?"}
|
||||
],
|
||||
"temperature": 0.6,
|
||||
"top_p": 0.95,
|
||||
"top_k": 20,
|
||||
"max_completion_tokens": 32
|
||||
}'
|
||||
```
|
||||
|
||||
Expected result:
|
||||
|
||||
The service returns HTTP 200 OK with a JSON response containing the `choices` field. Example output (content truncated for brevity):
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "chatcmpl-9df13fd5e539af93",
|
||||
"object": "chat.completion",
|
||||
"created": 1780971952,
|
||||
"model": "qwen3_next",
|
||||
"choices": [
|
||||
{
|
||||
"index": 0,
|
||||
"message": {
|
||||
"role": "assistant",
|
||||
"content": "What do you know about me?\n\nHello! I am Qwen, a large-scale language model independently developed by the Tongyi Lab under Alibaba Group. I am...",
|
||||
"reasoning": "The user is asking for my thoughts on \"Who are you?\"...",
|
||||
"refusal": null,
|
||||
"annotations": null,
|
||||
"audio": null,
|
||||
"function_call": null
|
||||
},
|
||||
"logprobs": null,
|
||||
"finish_reason": "length",
|
||||
"stop_reason": null,
|
||||
"token_ids": null
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
## 7 Accuracy Evaluation
|
||||
|
||||
### Using AISBench
|
||||
|
||||
1. Refer to [Using AISBench](../../developer_guide/evaluation/using_ais_bench.md) for details.
|
||||
|
||||
2. After execution, you can get the result, here is the result of `Qwen3-Next-80B-A3B-Instruct` in `vllm-ascend:0.13.0rc1` for reference only.
|
||||
|
||||
| dataset | version | metric | mode | vllm-api-general-chat |
|
||||
|----- | ----- | ----- | ----- | -----|
|
||||
| gsm8k | - | accuracy | gen | 95.53 |
|
||||
|
||||
## 8 Performance Evaluation
|
||||
|
||||
### Using AISBench
|
||||
|
||||
Refer to [Using AISBench for performance evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation) for details.
|
||||
|
||||
### Using vLLM Benchmark
|
||||
|
||||
Run performance evaluation of `Qwen3-Next` as an example.
|
||||
|
||||
Refer to [vLLM Benchmark](https://docs.vllm.ai/en/latest/benchmarking/) for more details.
|
||||
|
||||
There are three `vllm bench` subcommands:
|
||||
|
||||
- `latency`: Benchmark the latency of a single batch of requests.
|
||||
- `serve`: Benchmark the online serving throughput.
|
||||
- `throughput`: Benchmark offline inference throughput.
|
||||
|
||||
Take the `serve` as an example. Run the code as follows.
|
||||
|
||||
```shell
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
vllm bench serve --model Qwen/Qwen3-Next-80B-A3B-Instruct --dataset-name random --random-input 200 --num-prompts 200 --request-rate 1 --save-result --result-dir ./
|
||||
```
|
||||
|
||||
After about several minutes, you can get the performance evaluation result.
|
||||
|
||||
The performance result is:
|
||||
|
||||
```bash
|
||||
Hardware: A3-752T, 2 node
|
||||
Deployment: TP4 + Full Decode Only
|
||||
Input/Output: 2k/2k
|
||||
Concurrency: 32
|
||||
Performance: 580tps, TPOT 54ms
|
||||
```
|
||||
|
||||
## 9 Performance Tuning
|
||||
|
||||
### 9.1 Recommended Configurations
|
||||
|
||||
> **Note**:
|
||||
>
|
||||
> - The following configurations are validated in specific test environments and are for reference only. The optimal configuration depends on factors such as maximum input/output length, prefix cache hit rate, precision requirements, and deployment machine ratios. It is recommended to refer to Section 9.2 for tuning based on actual conditions.
|
||||
>
|
||||
> - Qwen3-Next does not support TP>=16 now. Since this model has 16 query heads but only 2 key and value heads, GQA degenerates into MHA when TP >= 16. However, the FIA operator currently fails to function in MHA scenarios with a head dimension of 256 (which is the case for this model).
|
||||
|
||||
#### Table 1: Scenario Overview
|
||||
|
||||
|Scenario|Deployment Mode|*Total NPUs|Weight Version|Key Considerations|
|
||||
|--------|---------------|-----------|--------------|------------------|
|
||||
|High Throughput<br>(16k context)|Single-Node Mixed|2 (A3)|Qwen3-Next|Use tp2 for high-resolution text inputs|
|
||||
|Long Context<br>(128k, no prefix cache)|Single-Node Mixed|2 (A3)|Qwen3-Next|tp2 for high-resolution text inputs|
|
||||
|Long Context<br>(128k, with prefix cache)|Single-Node Mixed|2 (A3)|Qwen3-Next|tp2 for high-resolution text inputs|
|
||||
|Multimodal<br>(1080p)|Single-Node Mixed|2 (A3)|Qwen3-Next|tp2 for high-resolution visual inputs|
|
||||
|
||||
> `*Total NPUs` indicates the total number of NPUs used across all nodes. 1 node = 1 Atlas 800 A3 server (64GB × 16 NPUs).
|
||||
|
||||
#### Table 2: Detailed Node Configuration
|
||||
|
||||
|Scenario|Configuration|NPUs|TP|DP|Max Model Len|MTP Speculation Num|
|
||||
|--------|-------------|-----|--|--|-------------------|--------------------|
|
||||
|High Throughput / Low Latency (16k)|Server / Single Machine|2|1|1|~16k|3|
|
||||
|Long Context (128k, no cache)|Server / Single Machine|2|1|1|128k|3|
|
||||
|Long Context (128k, with cache)|Server / Single Machine|2|1|1|128k|3|
|
||||
|Multimodal (1080p)|Server / Single Machine|2|1|1|~16k|3|
|
||||
|
||||
> For complete startup commands and parameter descriptions, please refer to the deployment examples in [Chapter 5](#5-online-service-deployment).
|
||||
|
||||
### 9.2 Tuning Guidelines
|
||||
|
||||
#### 9.2.1 General Tuning Reference
|
||||
|
||||
Please refer to the [Public Performance Tuning Documentation](../../developer_guide/performance_and_debug/optimization_and_tuning.md) for tuning methods.
|
||||
|
||||
Please refer to the [Feature Matrix](../../user_guide/support_matrix/feature_matrix.md) for detailed feature descriptions.
|
||||
|
||||
## 10 FAQ
|
||||
|
||||
For common environment, installation, and general parameter issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html).
|
||||
594
docs/source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md
Normal file
594
docs/source/tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md
Normal file
@@ -0,0 +1,594 @@
|
||||
# Qwen3-Omni-30B-A3B-Thinking
|
||||
|
||||
## 1 Introduction
|
||||
|
||||
Qwen3-Omni is a native end-to-end multilingual omni-modal foundation model. It processes text, images, audio, and video, and delivers real-time streaming responses in both text and natural speech. We introduce several architectural upgrades to improve performance and efficiency. The Thinking model of Qwen3-Omni-30B-A3B, which contains the thinker component, is equipped with chain-of-thought reasoning and supports audio, video, and text input, with text output.
|
||||
|
||||
This document will show the main verification steps of the model, including supported features, feature configuration, environment preparation, single-node deployment, accuracy and performance evaluation.
|
||||
|
||||
The Qwen3-Omni-30B-A3B model is first supported in v0.12.0rc1. This document is validated and written based on vLLM-Ascend v0.22.1rc. All v0.22.1rc and later versions can run stably. To use the latest features, it is recommended to use the latest release candidate or official version.
|
||||
|
||||
## 2 Supported Features
|
||||
|
||||
Please refer to [Supported Features List](https://docs.vllm.ai/projects/ascend/zh-cn/latest/user_guide/support_matrix/supported_models.html) to get the model's supported feature matrix.
|
||||
|
||||
Please refer to [Feature Guide](https://docs.vllm.ai/projects/ascend/zh-cn/latest/user_guide/feature_guide/index.html) to get the feature's configuration.
|
||||
|
||||
## 3 Prerequisites
|
||||
|
||||
### 3.1 Model Weight
|
||||
|
||||
The following model variants are available. It is recommended to download the model weight to a shared directory accessible to all nodes.
|
||||
|
||||
| Model | Hardware Requirement | Download |
|
||||
| -------------------- | ------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------ |
|
||||
| Qwen3-Omni-30B-A3B (BF16) | Atlas 800I A3 (64G, 1\~2 cards)<br>Atlas 800I A2 (64G, 2\~4 cards) | [Download](https://www.modelscope.cn/models/Qwen/Qwen3-Omni-30B-A3B) |
|
||||
| Qwen3-Omni-30B-A3B-W8A8 | Atlas 800I A3 (64G, 1\~2 cards)<br>Atlas 800I A2 (64G, 2\~4 cards) | N/A|
|
||||
|
||||
The W8A8 quantized weights are not available for direct download, you can obtain them by quantizing the BF16 model using **msmodelslim**. Refer to the [Quantization Guide](../../user_guide/feature_guide/quantization.md) for details. All model paths in this document should be adjusted to your actual local paths.
|
||||
|
||||
These are the recommended numbers of cards, which can be adjusted according to the actual situation.
|
||||
|
||||
:::{note}
|
||||
Qwen3-Omni-30B-A3B-W8A8 adopts a hybrid quantization strategy (ordered by model structure):
|
||||
|
||||
- **Embedding layer**: BF16 (no quantization)
|
||||
- **Q/K normalization** (q_norm, k_norm): BF16
|
||||
- **Attention projections** (q/k/v/o_proj): Static W8A8 with pre-computed per-tensor scales
|
||||
- **MoE routing gate** (mlp.gate): BF16
|
||||
- **MoE expert projections** (gate/up/down_proj): Dynamic W8A8 where input scales are computed on-the-fly during inference
|
||||
:::
|
||||
|
||||
It is recommended to download the model weight to a shared directory across multiple nodes.
|
||||
|
||||
## 4 Installation
|
||||
|
||||
### 4.1 Docker Image Installation
|
||||
|
||||
You can use the official all-in-one Docker image for Qwen3-Omni MoE models.
|
||||
|
||||
**Docker Pull:**
|
||||
|
||||
```bash
|
||||
docker pull quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
```
|
||||
|
||||
**Docker Run:**
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: hardware
|
||||
|
||||
::::{tab-item} Atlas 800I A3
|
||||
:sync: a3
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
|
||||
docker run \
|
||||
--name vllm-ascend-env \
|
||||
--shm-size=128g \
|
||||
--net=host \
|
||||
--privileged=true \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci8 \
|
||||
--device /dev/davinci9 \
|
||||
--device /dev/davinci10 \
|
||||
--device /dev/davinci11 \
|
||||
--device /dev/davinci12 \
|
||||
--device /dev/davinci13 \
|
||||
--device /dev/davinci14 \
|
||||
--device /dev/davinci15 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /etc/hccn.conf:/etc/hccn.conf \
|
||||
-it -d $IMAGE bash
|
||||
```
|
||||
|
||||
:::{note}
|
||||
A3 has 8 NPUs with dual-die design (16 chips total: `/dev/davinci[0-15]`).
|
||||
If you are on a shared machine, map only the chips you need (e.g., `/dev/davinci[0-7]` for NPU 0-3).
|
||||
:::
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Atlas 800I A2
|
||||
:sync: a2
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
|
||||
docker run \
|
||||
--name vllm-ascend-env \
|
||||
--shm-size=128g \
|
||||
--net=host \
|
||||
--privileged=true \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /etc/hccn.conf:/etc/hccn.conf \
|
||||
-it -d $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
:::::
|
||||
|
||||
The default workdir is `/workspace`. vLLM and vLLM-Ascend are installed as Python packages in site-packages.
|
||||
|
||||
**Installation Verification:**
|
||||
|
||||
After starting the container, run the following command to verify the installation:
|
||||
|
||||
```bash
|
||||
docker ps | grep vllm-ascend-env
|
||||
```
|
||||
|
||||
Expected result: The container is listed with status `Up`. You can also verify the vllm-ascend version inside the container:
|
||||
|
||||
```bash
|
||||
pip show vllm-ascend
|
||||
```
|
||||
|
||||
Expected result: The version information is displayed, matching the pulled image version.
|
||||
|
||||
### 4.2 Source Code Installation
|
||||
|
||||
If you prefer not to use the Docker image, you can build from source. Install vLLM from source first:
|
||||
|
||||
1. Clone and install vLLM:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/vllm-project/vllm.git
|
||||
cd vllm
|
||||
pip install -e .
|
||||
```
|
||||
|
||||
2. Clone and install the vLLM-Ascend repository:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/vllm-project/vllm-ascend.git
|
||||
cd vllm-ascend
|
||||
pip install -e .
|
||||
```
|
||||
|
||||
**Installation Verification:**
|
||||
|
||||
```bash
|
||||
pip show vllm vllm-ascend
|
||||
```
|
||||
|
||||
Expected result: The version information for both packages is displayed, confirming a successful installation.
|
||||
|
||||
:::{note}
|
||||
If deploying a multi-node environment, set up the environment on each node.
|
||||
:::
|
||||
|
||||
Please install system dependencies.
|
||||
|
||||
```bash
|
||||
pip install qwen_omni_utils modelscope
|
||||
# Used for audio processing.
|
||||
apt-get update && apt-get install -y ffmpeg
|
||||
# Check the installation.
|
||||
ffmpeg -version
|
||||
```
|
||||
|
||||
Required to avoid HcclAllreduce failures caused by the default FFTS+ mode's stream and shape limitations.
|
||||
|
||||
```bash
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
```
|
||||
|
||||
## 5 Online Service Deployment
|
||||
|
||||
Because the model has fewer parameters, it doesn’t involve the PD separation scenario.
|
||||
|
||||
### 5.1 Single-Node Online Deployment
|
||||
|
||||
Single-node deployment completes both Prefill and Decode within the same node, suitable for development, testing, and small-to-medium scale inference scenarios. For the Qwen3-Omni-30B-A3B MoE model, Expert Parallelism (EP) is required to distribute experts across NPUs.
|
||||
|
||||
> The following command is an example configuration. Adjust the parameters based on your actual scenario.
|
||||
|
||||
**Atlas 800I A2/A3:**
|
||||
|
||||
```bash
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3
|
||||
export HCCL_OP_EXPANSION_MODE="AIV" # not needed on A2
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
|
||||
vllm serve your_model_path \
|
||||
--served-model-name qwen3-omni \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 100 \
|
||||
--max-model-len 40960 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--tensor-parallel-size 4 \
|
||||
--enable-expert-parallel \
|
||||
--quantization ascend \
|
||||
--distributed_executor_backend "mp" \
|
||||
--no-enable-prefix-caching \
|
||||
--async-scheduling \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
|
||||
--gpu-memory-utilization 0.95 \
|
||||
--additional-config '{"enable_flashcomm1": false, "weight_nz_mode": 2}' \
|
||||
--port 8000
|
||||
```
|
||||
|
||||
:::{note}
|
||||
|
||||
- `ASCEND_RT_VISIBLE_DEVICES`: must be set to the NPU chip IDs allocated to your environment (e.g., `0,1,2,3` for 4 chips).
|
||||
- `--port`: adjust to avoid conflicts with other services running on the same machine.
|
||||
- `--no-enable-prefix-caching`: disabled by default as prefix caching effectiveness for this model on Ascend NPUs has not been fully characterized. You can try enabling it to evaluate the cache hit rate for your workload.
|
||||
- `--quantization ascend`: required for W8A8 quantized models. Remove this parameter when using BF16 weights.
|
||||
|
||||
:::
|
||||
|
||||
:::{tip}
|
||||
For parameter details, refer to:
|
||||
|
||||
- [vLLM CLI documentation](https://docs.vllm.ai/en/stable/cli/) — standard serve parameters (`--host`, `--port`, `--max-model-len`, etc.)
|
||||
- [Environment Variables](../../user_guide/configuration/env_vars.md) — Ascend-specific environment variables (`HCCL_*`, etc.)
|
||||
- [Additional Configuration](../../user_guide/configuration/additional_config.md) — `--additional-config` format and options
|
||||
:::
|
||||
|
||||
**Service Verification:**
|
||||
|
||||
If the service starts successfully, the following startup log will be displayed:
|
||||
|
||||
```text
|
||||
(APIServer pid=<pid>) INFO: Started server process [<pid>]
|
||||
(APIServer pid=<pid>) INFO: Waiting for application startup.
|
||||
(APIServer pid=<pid>) INFO: Application startup complete.
|
||||
```
|
||||
|
||||
## 6 Functional Verification
|
||||
|
||||
Once your server is started, you can query the model with input prompts.
|
||||
|
||||
```bash
|
||||
curl http://localhost:8000/v1/chat/completions \
|
||||
-X POST \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "qwen3-omni",
|
||||
"messages": [
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{
|
||||
"type": "image_url",
|
||||
"image_url": {
|
||||
"url": "https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen3-Omni/demo/cars.jpg"
|
||||
}
|
||||
},
|
||||
{
|
||||
"type": "audio_url",
|
||||
"audio_url": {
|
||||
"url": "https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen3-Omni/demo/cough.wav"
|
||||
}
|
||||
},
|
||||
{
|
||||
"type": "video_url",
|
||||
"video_url": {
|
||||
"url": "https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen3-Omni/demo/draw.mp4"
|
||||
}
|
||||
|
||||
},
|
||||
{
|
||||
"type": "text",
|
||||
"text": "Analyze this audio, image, and video together."
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}'
|
||||
```
|
||||
|
||||
The output should resemble the following.
|
||||
|
||||
```bash
|
||||
{
|
||||
"id": "chatcmpl-8f3a1c2e4b5d4f7a9c0e1d2b3a4f5e6d",
|
||||
"object": "chat.completion",
|
||||
"created": 1782714240,
|
||||
"model": "Qwen/Qwen3-Omni-30B-A3B-Thinking",
|
||||
"choices": [
|
||||
{
|
||||
"index": 0,
|
||||
"message": {
|
||||
"role": "assistant",
|
||||
"reasoning_content": "Let me examine each modality. The image shows several cars parked in a row — looks like a street or lot, daytime. The audio is short and sharp, sounds like a human cough (a single dry cough). The video appears to show someone drawing/sketching on paper, a hand moving a pen. Now I need to connect them: there isn't a single obvious narrative linking cars + cough + drawing, so I should describe each and then note any plausible shared context (e.g., a person sketching cars while having a cold). I'll keep the synthesis honest rather than forcing a connection.",
|
||||
"content": "Here's an analysis of the three inputs together:\n\n**Image** — A row of parked cars photographed outdoors in daylight, likely on a street or in a parking area. The vehicles are of varying colors and models, shot at a slight angle so several are visible in perspective.\n\n**Audio** — A brief, dry human cough: a single forceful exhalation lasting under a second, with no speech or background music. It sounds like an isolated cough rather than a sustained coughing fit.\n\n**Video** — A close-up of a hand drawing/sketching with a pen on paper, the strokes building up a simple illustration over a few seconds.\n\n**Combined interpretation** — The three clips don't share an explicit storyline; they're distinct samples of vision, sound, and motion. If a connecting context is assumed, one plausible scene is a person sketching cars (the drawing in the video, the cars in the image) while momentarily coughing (the audio) — e.g., an artist working outdoors who has a cold. But strictly, each input stands on its own: a static photo of cars, a one-off cough sound, and a short hand-drawing clip.",
|
||||
"tool_calls": []
|
||||
},
|
||||
"logprobs": null,
|
||||
"finish_reason": "stop",
|
||||
"stop_reason": null
|
||||
}
|
||||
],
|
||||
"usage": {
|
||||
"prompt_tokens": 8423,
|
||||
"total_tokens": 8712,
|
||||
"completion_tokens": 289,
|
||||
"prompt_tokens_details": null
|
||||
},
|
||||
"prompt_logprobs": null
|
||||
}
|
||||
|
||||
```
|
||||
|
||||
Expected result: HTTP 200 with a JSON response containing the `choices` field with generated text.
|
||||
|
||||
## 7 Accuracy Evaluation
|
||||
|
||||
### Using EvalScope
|
||||
|
||||
As an example, take the `gsm8k` `omni_bench` `bbh` dataset as a test dataset, and run accuracy evaluation of `Qwen3-Omni-30B-A3B-Thinking` in online mode.
|
||||
|
||||
1. Refer to [Using evalscope](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/evaluation/using_evalscope.html#install-evalscope-using-pip) for `evalscope` installation.
|
||||
2. Run `evalscope` to execute the accuracy evaluation.
|
||||
|
||||
```bash
|
||||
evalscope eval \
|
||||
--model /root/.cache/modelscope/hub/models/Qwen/Qwen3-Omni-30B-A3B-Thinking \
|
||||
--api-url http://localhost:8000/v1 \
|
||||
--api-key EMPTY \
|
||||
--eval-type server \
|
||||
--datasets omni_bench, gsm8k, bbh \
|
||||
--dataset-args '{"omni_bench": { "extra_params": { "use_image": true, "use_audio": false}}}' \
|
||||
--eval-batch-size 1 \
|
||||
--generation-config '{"max_completion_tokens": 10000, "temperature": 0.6}' \
|
||||
--limit 100
|
||||
```
|
||||
|
||||
3. After execution, you can get the result, here is the result of `Qwen3-Omni-30B-A3B-Thinking` in vllm-ascend:0.13.0rc1 for reference only.
|
||||
|
||||
```bash
|
||||
+-----------------------------+------------+----------+----------+-------+---------+---------+
|
||||
| Model | Dataset | Metric | Subset | Num | Score | Cat.0 |
|
||||
+=============================+============+==========+==========+=======+=========+=========+
|
||||
| Qwen3-Omni-30B-A3B-Thinking | omni_bench | mean_acc | default | 100 | 0.44 | default |
|
||||
+-----------------------------+------------+----------+----------+-------+---------+---------+
|
||||
| Qwen3-Omni-30B-A3B-Thinking | gsm8k | mean_acc | main | 100 | 0.98 | default |
|
||||
+-----------------------------+-----------+----------+----------+-------+---------+---------+
|
||||
| Qwen3-Omni-30B-A3B-Thinking | bbh | mean_acc | OVERALL | 270 | 0.9148 | |
|
||||
+-----------------------------+------------+----------+----------+-------+---------+---------+
|
||||
```
|
||||
|
||||
## 8 Performance Evaluation
|
||||
|
||||
### Using vLLM Benchmark
|
||||
|
||||
Run performance evaluation of `Qwen3-Omni-30B-A3B-Thinking` as an example.
|
||||
Refer to [vLLM Benchmark](https://docs.vllm.ai/en/latest/benchmarking/) for more details.
|
||||
|
||||
There are three `vllm bench` subcommands:
|
||||
|
||||
- `latency`: Benchmark the latency of a single batch of requests.
|
||||
- `serve`: Benchmark the online serving throughput.
|
||||
- `throughput`: Benchmark offline inference throughput.
|
||||
|
||||
Take the `serve` as an example. Run the code as follows.
|
||||
|
||||
```bash
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
export MODEL=Qwen/Qwen3-Omni-30B-A3B-Thinking
|
||||
python3 -m vllm.entrypoints.openai.api_server --model $MODEL --tensor-parallel-size 2 --swap-space 16 --disable-log-stats --disable-log-request --load-format dummy
|
||||
|
||||
pip config set global.index-url https://mirrors.tuna.tsinghua.edu.cn/pypi/web/simple
|
||||
pip install -r vllm-ascend/benchmarks/requirements-bench.txt
|
||||
|
||||
vllm bench serve --model $MODEL --dataset-name random --random-input 200 --num-prompts 200 --request-rate 1 --save-result --result-dir ./
|
||||
```
|
||||
|
||||
After execution, you can get the result, here is the result of `Qwen3-Omni-30B-A3B-Thinking` in vllm-ascend:0.13.0rc1 for reference only.
|
||||
|
||||
```bash
|
||||
============ Serving Benchmark Result ============
|
||||
Successful requests: 200
|
||||
Failed requests: 0
|
||||
Request rate configured (RPS): 1.00
|
||||
Benchmark duration (s): 211.90
|
||||
Total input tokens: 40000
|
||||
Total generated tokens: 25600
|
||||
Request throughput (req/s): 0.94
|
||||
Output token throughput (tok/s): 120.81
|
||||
Peak output token throughput (tok/s): 216.00
|
||||
Peak concurrent requests: 24.00
|
||||
Total token throughput (tok/s): 309.58
|
||||
---------------Time to First Token----------------
|
||||
Mean TTFT (ms): 215.50
|
||||
Median TTFT (ms): 211.51
|
||||
P99 TTFT (ms): 317.18
|
||||
-----Time per Output Token (excl. 1st token)------
|
||||
Mean TPOT (ms): 98.96
|
||||
Median TPOT (ms): 99.19
|
||||
P99 TPOT (ms): 101.52
|
||||
---------------Inter-token Latency----------------
|
||||
Mean ITL (ms): 99.02
|
||||
Median ITL (ms): 96.10
|
||||
P99 ITL (ms): 176.02
|
||||
==================================================
|
||||
```
|
||||
|
||||
## 9 Performance Tuning
|
||||
|
||||
### 9.1 Recommended Configurations
|
||||
|
||||
> **Note**: The following configurations are validated in specific test environments and are for reference only. The optimal configuration depends on factors such as maximum input/output length, prefix cache hit rate, precision requirements, and deployment machine ratios. It is recommended to refer to Section 9.2 for tuning based on actual conditions.
|
||||
|
||||
#### Table 1: Scenario Overview
|
||||
|
||||
| Scenario | Deployment Mode | *Total NPUs | Weight Version | Key Considerations |
|
||||
| --------------- | ----------------- | ---------------- | -------------- | ---------------------------------------------------------------- |
|
||||
| High Throughput | Single-Node (TP1) | 1 (A3)<br>2 (A2) | W8A8 | Single-card deployment maximizes concurrent request processing |
|
||||
| Low Latency | Single-Node (TP4) | 2 (A3)<br>4 (A2) | W8A8 | Multi-card TP reduces per-token latency with expert parallelism |
|
||||
| Long Context | Single-Node (TP4) | 2 (A3)<br>4 (A2) | W8A8 | Reduces concurrent sequences to accommodate longer max-model-len |
|
||||
|
||||
> `*Total NPUs` indicates the total number of NPUs used across all nodes. On Atlas 800I A3, each NPU contains two dies (chips), so TP4 requires 4 chips = 2 NPUs.
|
||||
|
||||
#### Table 2: Detailed Node Configuration
|
||||
|
||||
| Scenario | NPUs | TP | max-model-len | max-num-seqs | FUSED_MC2 | EP | hf-overrides |
|
||||
| --------------- | ------ | --- | ------------- | ------------ | --------- | --- | ------------ |
|
||||
| High Throughput | 1 (A3) | 1 | 37364 | 100 | Off | Off | - |
|
||||
| Low Latency | 2 (A3) | 4 | 37364 | 100 | Off | On | - |
|
||||
| Long Context | 2 (A3) | 4 | 131072 | 14 | Off | On | YaRN |
|
||||
|
||||
> For detailed parameter descriptions, please refer to the deployment examples in Section 5.
|
||||
|
||||
**Low Latency Configuration:**
|
||||
|
||||
```shell
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
|
||||
vllm serve your_model_path \
|
||||
--served-model-name qwen3-omni \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 100 \
|
||||
--max-model-len 37364 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--tensor-parallel-size 4 \
|
||||
--enable-expert-parallel \
|
||||
--distributed_executor_backend "mp" \
|
||||
--no-enable-prefix-caching \
|
||||
--async-scheduling \
|
||||
--quantization ascend \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"enable_flashcomm1": false, "weight_nz_mode": 2}' \
|
||||
--gpu-memory-utilization 0.95 \
|
||||
--port 8000 \
|
||||
--speculative-config '{"method": "eagle3","model": "your_eagle3_model_path", "num_speculative_tokens": 3}'
|
||||
```
|
||||
|
||||
:::{tip}
|
||||
Example AISBench settings for this configuration:
|
||||
|
||||
- `request_rate`: 0
|
||||
- `batch_size`: 32
|
||||
- Input/Output length: 2048/2048 or 3500/1500
|
||||
:::
|
||||
|
||||
**High Throughput Configuration:**
|
||||
|
||||
```shell
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
|
||||
vllm serve your_model_path \
|
||||
--served-model-name qwen3-omni \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 100 \
|
||||
--max-model-len 37364 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--tensor-parallel-size 1 \
|
||||
--distributed_executor_backend "mp" \
|
||||
--no-enable-prefix-caching \
|
||||
--async-scheduling \
|
||||
--quantization ascend \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"weight_nz_mode": 2}' \
|
||||
--gpu-memory-utilization 0.95 \
|
||||
--port 8000 \
|
||||
--speculative-config '{"method": "eagle3","model": "your_eagle3_model_path", "num_speculative_tokens": 3}'
|
||||
```
|
||||
|
||||
:::{tip}
|
||||
Example AISBench settings for this configuration:
|
||||
|
||||
- `request_rate`: 0
|
||||
- `batch_size`: 32
|
||||
- Input/Output length: 2048/2048 or 3500/1500
|
||||
:::
|
||||
|
||||
**Long Context Configuration:**
|
||||
|
||||
```shell
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
|
||||
vllm serve your_model_path \
|
||||
--served-model-name qwen3-omni \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 14 \
|
||||
--max-model-len 131072 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--tensor-parallel-size 4 \
|
||||
--enable-expert-parallel \
|
||||
--distributed_executor_backend "mp" \
|
||||
--no-enable-prefix-caching \
|
||||
--async-scheduling \
|
||||
--quantization ascend \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"enable_flashcomm1": false, "weight_nz_mode": 2}' \
|
||||
--gpu-memory-utilization 0.95 \
|
||||
--port 8000 \
|
||||
--hf-overrides '{"rope_parameters": {"rope_type":"yarn","factor":4,"original_max_position_embeddings":32768}}'
|
||||
```
|
||||
|
||||
:::{tip}
|
||||
Example AISBench settings for this configuration:
|
||||
|
||||
- `request_rate`: 0
|
||||
- `batch_size`: 32
|
||||
- Input/Output length: 65536/1024 or 131072/1024
|
||||
:::
|
||||
|
||||
### 9.2 Tuning Guidelines
|
||||
|
||||
#### 9.2.1 General Tuning Reference
|
||||
|
||||
Please refer to the [Public Performance Tuning Documentation](../../developer_guide/performance_and_debug/optimization_and_tuning.md) for tuning methods.
|
||||
Please refer to the [Feature Guide](../../user_guide/support_matrix/feature_matrix.md) for detailed feature descriptions.
|
||||
|
||||
## 10 FAQ
|
||||
|
||||
For common environment, installation, and general parameter issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html).
|
||||
316
docs/source/tutorials/models/Qwen3-Reranker.md
Normal file
316
docs/source/tutorials/models/Qwen3-Reranker.md
Normal file
@@ -0,0 +1,316 @@
|
||||
# Qwen3-Reranker
|
||||
|
||||
## 1 Introduction
|
||||
|
||||
The Qwen3 Reranker model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. Building upon the dense foundational models of the Qwen3 series, it provides a comprehensive range of text embeddings and reranking models in various sizes (0.6B, 4B, and 8B). This guide describes how to run the model with vLLM Ascend. Note that only 0.9.2rc1 and higher versions of vLLM Ascend support the model.
|
||||
|
||||
## 2 Supported Features
|
||||
|
||||
Refer to [supported features](../../user_guide/support_matrix/supported_models.md) to get the model's supported feature matrix.
|
||||
|
||||
## 3 Prerequisites
|
||||
|
||||
### 3.1 Model Weight
|
||||
|
||||
- `Qwen3-Reranker-8B` [Download model weight](https://www.modelscope.cn/models/Qwen/Qwen3-Reranker-8B)
|
||||
- `Qwen3-Reranker-4B` [Download model weight](https://www.modelscope.cn/models/Qwen/Qwen3-Reranker-4B)
|
||||
- `Qwen3-Reranker-0.6B` [Download model weight](https://www.modelscope.cn/models/Qwen/Qwen3-Reranker-0.6B)
|
||||
|
||||
It is recommended to download the model weight to the shared directory of multiple nodes, such as `/root/.cache/`
|
||||
|
||||
## 4 Installation
|
||||
|
||||
### 4.1 Docker Image Installation
|
||||
|
||||
You can use our official docker image to run `Qwen3-Reranker` model directly.
|
||||
|
||||
Select an image based on your machine type and start the docker image on your node, refer to [using docker](../../installation.md#set-up-using-docker).
|
||||
|
||||
For Atlas 300I DUO, use `vllm-ascend:nightly-releases-v0.23.0-310p` (or a later `-310p` image).
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: install
|
||||
|
||||
::::{tab-item} A3 series
|
||||
:sync: A3 series
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-a3
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--privileged=true \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} A2 series
|
||||
:sync: A2 series
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--privileged=true \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
:sync: Atlas 300I DUO
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-310p
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--privileged=true \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
After a successful docker run, you can verify the running container service by executing the `docker ps` command.
|
||||
|
||||
### 4.2 Source Code Installation
|
||||
|
||||
If you don't want to use the docker image as above, you can also build all from source:
|
||||
|
||||
- Install `vllm-ascend` from source, refer to [installation](../../installation.md).
|
||||
|
||||
If you want to deploy multi-node environment, you need to set up environment on each node.
|
||||
|
||||
## 5 Online Service Deployment
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: Deployment
|
||||
|
||||
::::{tab-item} A3/A2 series
|
||||
:sync: A3/A2 series
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
vllm serve Qwen/Qwen3-Reranker-0.6B \
|
||||
--served-model-name Qwen/Qwen3-Reranker-0.6B \
|
||||
--runner pooling \
|
||||
--hf_overrides '{"architectures": ["Qwen3VLForSequenceClassification"],"classifier_from_token": ["no", "yes"],"is_original_qwen3_reranker": true}' \
|
||||
--port 8000 \
|
||||
--max-model-len 1024
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
:sync: Atlas 300I DUO
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
vllm serve Qwen/Qwen3-Reranker-0.6B \
|
||||
--served-model-name Qwen/Qwen3-Reranker-0.6B \
|
||||
--runner pooling \
|
||||
--hf_overrides '{"architectures": ["Qwen3VLForSequenceClassification"],"classifier_from_token": ["no", "yes"],"is_original_qwen3_reranker": true}' \
|
||||
--compilation-config '{"cudagraph_capture_sizes": [1024,512]}' \
|
||||
--additional-config '{"ascend_compilation_config": {"fuse_norm_quant": false}}' \
|
||||
--dtype float16 \
|
||||
--port 8000 \
|
||||
--max-model-len 1024
|
||||
```
|
||||
|
||||
Required Parameter Descriptions:
|
||||
|
||||
`--compilation-config` For Atlas 300I DUO, due to limited hardware streams, the size of cudagraph_capture_sizes is restricted.
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
Key Parameter Descriptions:
|
||||
|
||||
- `--max-model-len` represents the context length, which is the maximum value of the input plus output for a single request. For Atlas 300I DUO if automatic parsing resolves to a large context length, allocating this mask (O(max_model_len^2)) may exceed NPU memory and trigger OOM. Be sure to set an explicit and conservative value, such as --max-model-len 1024.
|
||||
|
||||
Common Issues Tip: If you encounter issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html) for troubleshooting.
|
||||
|
||||
## 6 Functional Verification
|
||||
|
||||
Once your server is started, you can verify by follow command:
|
||||
|
||||
Service Verification:
|
||||
|
||||
```python
|
||||
import requests
|
||||
|
||||
url = "http://127.0.0.1:8000/v1/rerank"
|
||||
|
||||
# Please use the query_template and document_template to format the query and
|
||||
# document for better reranker results.
|
||||
|
||||
prefix = '<|im_start|>system\nJudge whether the Document meets the requirements based on the Query and the Instruct provided. Note that the answer can only be "yes" or "no".<|im_end|>\n<|im_start|>user\n'
|
||||
suffix = "<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n"
|
||||
|
||||
query_template = "{prefix}<Instruct>: {instruction}\n<Query>: {query}\n"
|
||||
document_template = "<Document>: {doc}{suffix}"
|
||||
|
||||
instruction = (
|
||||
"Given a web search query, retrieve relevant passages that answer the query"
|
||||
)
|
||||
|
||||
query = "What is the capital of China?"
|
||||
|
||||
documents = [
|
||||
"The capital of China is Beijing.",
|
||||
"Gravity is a force that attracts two bodies towards each other. It gives weight to physical objects and is responsible for the movement of planets around the sun.",
|
||||
]
|
||||
|
||||
documents = [
|
||||
document_template.format(doc=doc, suffix=suffix) for doc in documents
|
||||
]
|
||||
|
||||
response = requests.post(url,
|
||||
json={
|
||||
"query": query_template.format(prefix=prefix, instruction=instruction, query=query),
|
||||
"documents": documents,
|
||||
}).json()
|
||||
|
||||
print(response)
|
||||
```
|
||||
|
||||
Expected Result:
|
||||
|
||||
The service returns HTTP 200 OK with a JSON response containing the `relevance_score` field. Example output:
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "score-xxx",
|
||||
"model": "Qwen/Qwen3-Reranker-0.6B",
|
||||
"usage": {
|
||||
"prompt_tokens": 193,
|
||||
"total_tokens": 193
|
||||
},
|
||||
"results": [
|
||||
{
|
||||
"index": 0,
|
||||
"document": {
|
||||
"text": "<Document>: The capital of China is Beijing.<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n",
|
||||
"multi_modal": null
|
||||
},
|
||||
"relevance_score": 0.9994981288909912
|
||||
},
|
||||
{
|
||||
"index": 1,
|
||||
"document": {
|
||||
"text": "<Document>: Gravity is a force that attracts two bodies towards each other. It gives weight to physical objects and is responsible for the movement of planets around the sun.<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n",
|
||||
"multi_modal": null
|
||||
},
|
||||
"relevance_score": 0.00000506485957885161
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
For more usage examples, please reference the [examples](https://github.com/vllm-project/vllm/tree/main/examples/pooling/score)
|
||||
|
||||
## 7 Accuracy Evaluation
|
||||
|
||||
Here are two accuracy evaluation methods.
|
||||
|
||||
### Using MTEB
|
||||
|
||||
1. Refer to [MTEB](https://docs.mteb.org/) for details.
|
||||
|
||||
2. Run follow code to execute the accuracy evaluation.
|
||||
|
||||
```python
|
||||
|
||||
import os
|
||||
|
||||
from mteb.models.vllm_wrapper import VllmCrossEncoderWrapper
|
||||
|
||||
if __name__ == "__main__":
|
||||
import mteb
|
||||
|
||||
data_path = "/home/data/mteb_data"
|
||||
os.environ["HF_DATASETS_CACHE"] = data_path
|
||||
os.environ["HF_ENDPOINT"] = "https://hf-mirror.com"
|
||||
|
||||
model = VllmCrossEncoderWrapper(f"/home/data/Qwen3-Reranker-0.6B",
|
||||
revision="norm",
|
||||
dtype="float16",
|
||||
enforce_eager=True,
|
||||
max_model_len=10240,
|
||||
hf_overrides={"architectures": ["Qwen3VLForSequenceClassification"],"classifier_from_token": ["no", "yes"],"is_original_qwen3_reranker": True})
|
||||
|
||||
cache = mteb.ResultCache("/home/data/mteb_data")
|
||||
tasks = mteb.get_tasks(
|
||||
task_types=["Reranking"],
|
||||
languages=["zho"]
|
||||
)
|
||||
tasks = mteb.get_tasks(tasks=["MultiLongDocReranking"])
|
||||
results = mteb.evaluate(model, tasks=tasks, cache=cache, overwrite_strategy="always")
|
||||
print(results)
|
||||
|
||||
```
|
||||
|
||||
3. After execution, you can get the result.
|
||||
|
||||
## 8 Performance Evaluation
|
||||
|
||||
### Using vLLM Benchmark
|
||||
|
||||
Run performance of `Qwen3-Reranker-0.6B` as an example.
|
||||
Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/cli/) for more details.
|
||||
|
||||
Take the `serve` as an example. Run the code as follows.
|
||||
|
||||
```bash
|
||||
vllm bench serve --model Qwen/Qwen3-Reranker-0.6B --backend vllm-rerank --port 8000 --dataset-name random-rerank --endpoint /v1/rerank --random-input 200 --save-result --result-dir ./
|
||||
```
|
||||
|
||||
After about several minutes, you can get the performance evaluation result.
|
||||
|
||||
## 9 FAQ
|
||||
|
||||
For common environment, installation, and general parameter issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html).
|
||||
754
docs/source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md
Normal file
754
docs/source/tutorials/models/Qwen3-VL-235B-A22B-Instruct.md
Normal file
@@ -0,0 +1,754 @@
|
||||
# Qwen3-VL-235B-A22B-Instruct
|
||||
|
||||
## 1 Introduction
|
||||
|
||||
Qwen3-VL-235B-A22B-Instruct is a large-scale sparse MoE vision-language model in the Qwen3-VL family. It is designed for multimodal chat, image understanding, multi-image reasoning, OCR-like visual question answering, and long-context generation.
|
||||
|
||||
This document describes the main validation steps for the model, including supported features, prerequisites, installation, single-node online deployment, multi-node deployment, Prefill-Decode (PD) disaggregation, functional verification, accuracy and performance evaluation, performance tuning, and FAQs.
|
||||
|
||||
The `Qwen3-VL-235B-A22B-Instruct` tutorial was introduced in the vLLM-Ascend validation cycle around `v0.12.0`. Use the current `vllm-ascend` documentation image placeholder or a later release for the examples below.
|
||||
|
||||
## 2 Supported Features
|
||||
|
||||
Refer to [Supported Features List](../../user_guide/support_matrix/supported_features.md) to get the model's supported feature matrix.
|
||||
|
||||
Refer to [Feature Guide](../../user_guide/feature_guide/index.md) to get the feature's configuration.
|
||||
|
||||
## 3 Prerequisites
|
||||
|
||||
### 3.1 Model Weight
|
||||
|
||||
- `Qwen3-VL-235B-A22B-Instruct` (BF16 version): requires 1 Atlas 800 A3 (64G x 16) node or 2 Atlas 800 A2 (64G x 8) nodes. [Model Weight](https://www.modelscope.cn/models/Qwen/Qwen3-VL-235B-A22B-Instruct/).
|
||||
- `Qwen3-VL-235B-A22B-Instruct-w8a8-QuaRot` (quantized version used by single-node validation): requires 1 Atlas 800 A3 (64G x 16) node. [Model Weight](https://www.modelscope.cn/models/Eco-Tech/Qwen3-VL-235B-A22B-Instruct-w8a8-QuaRot).
|
||||
- `Qwen3-VL-235B-A22B-Instruct-w8a8-mxfp8` (quantized version): requires 1 Ascend 950DT (96G x 8) node. [Model Weight](https://modelscope.cn/models/Eco-Tech/Qwen3-VL-235B-A22B-Instruct-w8a8-mxfp8).
|
||||
|
||||
It is recommended to download the model weight to a shared directory across multiple nodes.
|
||||
|
||||
### 3.2 Verify Multi-node Communication (Optional)
|
||||
|
||||
If you want to deploy the model in a multi-node environment, verify the communication environment according to [verify multi-node communication environment](../../installation.md#verify-multi-node-communication).
|
||||
|
||||
## 4 Installation
|
||||
|
||||
### 4.1 Docker Image Installation
|
||||
|
||||
Select an image based on your machine type and start the docker image on your node, refer to [using docker](../../installation.md#set-up-using-docker).
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: install
|
||||
|
||||
::::{tab-item} Ascend 950DT series
|
||||
:sync: 950dt
|
||||
|
||||
Start the docker image on your each node.
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-#TODO
|
||||
export NAME=vllm-ascend
|
||||
|
||||
docker run --rm \
|
||||
--name $NAME \
|
||||
--net=host \
|
||||
--shm-size=1g \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/hisi_hdc \
|
||||
--device /dev/ummu \
|
||||
--device /dev/uburma \
|
||||
-v /usr/local/Ascend/driver:/usr/local/Ascend/driver \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /etc/hccl_rootinfo.json:/etc/hccl_rootinfo.json \
|
||||
-v /etc/hixlep/:/etc/hixlep/ \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-v /usr/local/sbin:/usr/local/sbin \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/sbin/npu-smi:/usr/local/sbin/npu-smi \
|
||||
-v /usr/lib64:/usr/lib64 \
|
||||
-itd $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} A3 series
|
||||
:sync: A3
|
||||
|
||||
Start the docker image on each node.
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-a3
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=512g \
|
||||
--net=host \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci8 \
|
||||
--device /dev/davinci9 \
|
||||
--device /dev/davinci10 \
|
||||
--device /dev/davinci11 \
|
||||
--device /dev/davinci12 \
|
||||
--device /dev/davinci13 \
|
||||
--device /dev/davinci14 \
|
||||
--device /dev/davinci15 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /etc/hccn.conf:/etc/hccn.conf \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} A2 series
|
||||
:sync: A2
|
||||
|
||||
Start the docker image on each node.
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=512g \
|
||||
--net=host \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /etc/hccn.conf:/etc/hccn.conf \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
After a successful docker run, you can verify the running container service by executing the `docker ps` command.
|
||||
|
||||
### 4.2 Source Code Installation
|
||||
|
||||
If you prefer not to use the Docker image, you can build from source. Install vLLM from source first:
|
||||
|
||||
1. Clone and install vLLM:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/vllm-project/vllm.git
|
||||
cd vllm
|
||||
pip install -e .
|
||||
```
|
||||
|
||||
2. Clone and install the vLLM-Ascend repository:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/vllm-project/vllm-ascend.git
|
||||
cd vllm-ascend
|
||||
pip install -e .
|
||||
```
|
||||
|
||||
**Installation Verification:**
|
||||
|
||||
```bash
|
||||
pip show vllm vllm-ascend
|
||||
```
|
||||
|
||||
Expected result: The version information for both packages is displayed, confirming a successful installation.
|
||||
|
||||
:::{note}
|
||||
If deploying a multi-node environment, set up the environment on each node.
|
||||
:::
|
||||
|
||||
For more details, please refer to the [Installation Guide](../../installation.md).
|
||||
|
||||
## 5 Online Service Deployment
|
||||
|
||||
### 5.1 Single-Node Online Deployment
|
||||
|
||||
Single-node deployment runs both Prefill and Decode on the same node. The W8A8 version needs `--quantization ascend`.
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: deploy
|
||||
|
||||
::::{tab-item} Ascend 950DT series
|
||||
:sync: 950dt
|
||||
|
||||
Run the following script to execute online inference on 1 Ascend 950DT (96G x 8). The quantized version (`Qwen3-VL-235B-A22B-Instruct-w8a8-mxfp8`) can be deployed on a single Ascend 950DT node.
|
||||
|
||||
```shell
|
||||
#!/bin/sh
|
||||
|
||||
# Load model from ModelScope to speed up download.
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
|
||||
# Reduce memory fragmentation and avoid out-of-memory errors.
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export HCCL_BUFFSIZE=400
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=100
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export VLLM_ASCEND_ENABLE_FLASHCOMM1=1
|
||||
|
||||
vllm serve Eco-Tech/Qwen3-VL-235B-A22B-Instruct-w8a8-mxfp8 \
|
||||
--host 0.0.0.0 \
|
||||
--port 8000 \
|
||||
--distributed-executor-backend mp \
|
||||
--data-parallel-size 1 \
|
||||
--tensor-parallel-size 8 \
|
||||
--enable-expert-parallel \
|
||||
--seed 1024 \
|
||||
--quantization ascend \
|
||||
--served-model-name qwen3-vl-235b \
|
||||
--max-num-seqs 32 \
|
||||
--max-model-len 32768 \
|
||||
--max-num-batched-tokens 8192 \
|
||||
--trust-remote-code \
|
||||
--no-enable-prefix-caching \
|
||||
--mm-processor-cache-gb 0 \
|
||||
--limit-mm-per-prompt.image 1 \
|
||||
--limit-mm-per-prompt.video 0 \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}'
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} A3 series
|
||||
:sync: A3
|
||||
|
||||
Run the following script to start online serving on 1 Atlas 800 A3 (64G x 16) node. The W8A8 example is suitable for functional validation and image-only online serving.
|
||||
|
||||
```shell
|
||||
#!/bin/sh
|
||||
|
||||
# Load model from ModelScope to speed up download.
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
|
||||
# Reduce memory fragmentation and avoid out-of-memory errors.
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export HCCL_BUFFSIZE=1536
|
||||
export OMP_NUM_THREADS=1
|
||||
export OMP_PROC_BIND=false
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export VLLM_ASCEND_ENABLE_FLASHCOMM1=1
|
||||
export VLLM_ASCEND_ENABLE_FUSED_MC2=1
|
||||
export VLLM_ASCEND_BALANCE_SCHEDULING=1
|
||||
|
||||
vllm serve Eco-Tech/Qwen3-VL-235B-A22B-Instruct-w8a8-QuaRot \
|
||||
--host 0.0.0.0 \
|
||||
--port 8000 \
|
||||
--served-model-name qwen3-vl-235b \
|
||||
--quantization ascend \
|
||||
--data-parallel-size 4 \
|
||||
--tensor-parallel-size 4 \
|
||||
--enable-expert-parallel \
|
||||
--seed 1024 \
|
||||
--max-num-seqs 32 \
|
||||
--max-model-len 32768 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--trust-remote-code \
|
||||
--gpu-memory-utilization 0.92 \
|
||||
--no-enable-prefix-caching \
|
||||
--mm-processor-cache-gb 0 \
|
||||
--limit-mm-per-prompt.image 1 \
|
||||
--limit-mm-per-prompt.video 0 \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY","cudagraph_capture_sizes":[1,2,4,8,16,24,32]}'
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} A2 series
|
||||
:sync: A2
|
||||
|
||||
For W8A8 deployment on A2, 2 Atlas 800 A2 (64G x 8) nodes are required. Refer to [Section 5.2](#52-multi-node-deployment-with-mp-recommended-for-bf16) for multi-node MP deployment.
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
Common Issues Tip: If you encounter issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html) for troubleshooting.
|
||||
|
||||
**Key parameters:**
|
||||
|
||||
- `--data-parallel-size 4` and `--tensor-parallel-size 4` map the 16 NPUs on one A3 node into four DP groups, each with TP4.
|
||||
- `--enable-expert-parallel` enables expert parallelism for MoE layers. Do not mix MoE tensor parallelism and expert parallelism in the same MoE layer.
|
||||
- `--max-model-len` is the maximum input plus output length for a single request. Multimodal inputs consume text tokens and visual tokens, so increase it only when enough KV cache is available.
|
||||
- `--max-num-seqs` is the maximum number of active requests scheduled by each DP group. For performance tests, set `--max-num-seqs * --data-parallel-size` greater than or equal to the test concurrency.
|
||||
- `--max-num-batched-tokens` is the maximum number of tokens processed in one scheduler step. A larger value can improve prefill efficiency but consumes more activation memory.
|
||||
- `--gpu-memory-utilization` controls how much HBM vLLM can use to calculate KV cache capacity. A higher value increases KV cache size but can trigger OOM if runtime memory is higher than the profile run.
|
||||
- `--quantization ascend` enables Ascend quantization for the W8A8 model. Remove this option when deploying the BF16 model.
|
||||
- `--limit-mm-per-prompt.image 1` and `--limit-mm-per-prompt.video 0` reserve multimodal capacity for one image per request and disable video inputs to save memory.
|
||||
- `--mm-processor-cache-gb 0` disables the multimodal processor cache. Increase it only when your workload benefits from reused media preprocessing and you have enough host memory.
|
||||
- `--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}'` enables full decode ACLGraph replay to reduce dispatch overhead.
|
||||
|
||||
### 5.2 Multi-Node Deployment with MP (Recommended for BF16)
|
||||
|
||||
Multi-node MP deployment uses vLLM data parallelism across nodes and tensor parallelism within each node. It is recommended for the BF16 model on 2 Atlas 800 A2 (64G x 8) nodes, or for long-context validation where one node does not provide enough HBM headroom.
|
||||
|
||||
Assume you have 2 Atlas 800 A2 nodes and want to deploy `Qwen3-VL-235B-A22B-Instruct-w8a8-QuaRot` across them. Replace `nic_name`, `local_ip`, and `node0_ip` with the actual network interface and IP addresses in your environment.
|
||||
|
||||
Run the following script on node 0.
|
||||
|
||||
```shell
|
||||
#!/bin/sh
|
||||
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
|
||||
# Get these values through ifconfig.
|
||||
# nic_name is the network interface name corresponding to local_ip.
|
||||
nic_name="xxxx"
|
||||
local_ip="xxxx"
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
|
||||
vllm serve Eco-Tech/Qwen3-VL-235B-A22B-Instruct-w8a8-QuaRot \
|
||||
--host 0.0.0.0 \
|
||||
--port 8000 \
|
||||
--quantization ascend \
|
||||
--data-parallel-size 2 \
|
||||
--api-server-count 2 \
|
||||
--data-parallel-size-local 1 \
|
||||
--data-parallel-address $local_ip \
|
||||
--data-parallel-rpc-port 13389 \
|
||||
--seed 1024 \
|
||||
--served-model-name qwen3-vl-235b \
|
||||
--tensor-parallel-size 8 \
|
||||
--enable-expert-parallel \
|
||||
--max-num-seqs 16 \
|
||||
--max-model-len 262144 \
|
||||
--max-num-batched-tokens 4096 \
|
||||
--trust-remote-code \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--no-enable-prefix-caching \
|
||||
--mm-processor-cache-gb 0 \
|
||||
--limit-mm-per-prompt.image 1 \
|
||||
--limit-mm-per-prompt.video 0 \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"enable_cpu_binding":true,"enable_flashcomm1":true}'
|
||||
```
|
||||
|
||||
Common Issues Tip: If node 1 cannot join the service or HCCL initialization times out, refer to [verify multi-node communication environment](../../installation.md#verify-multi-node-communication) and [Public FAQs](../../faqs.md). Make sure the network interface names, IP addresses, and RPC ports are consistent across nodes.
|
||||
|
||||
Run the following script on node 1.
|
||||
|
||||
```shell
|
||||
#!/bin/sh
|
||||
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
|
||||
# Get these values through ifconfig.
|
||||
# nic_name is the network interface name corresponding to local_ip.
|
||||
nic_name="xxxx"
|
||||
local_ip="xxxx"
|
||||
|
||||
# The value of node0_ip must be consistent with local_ip on node 0.
|
||||
node0_ip="xxxx"
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
|
||||
vllm serve Eco-Tech/Qwen3-VL-235B-A22B-Instruct-w8a8-QuaRot \
|
||||
--host 0.0.0.0 \
|
||||
--port 8000 \
|
||||
--quantization ascend \
|
||||
--headless \
|
||||
--data-parallel-size 2 \
|
||||
--data-parallel-size-local 1 \
|
||||
--data-parallel-start-rank 1 \
|
||||
--data-parallel-address $node0_ip \
|
||||
--data-parallel-rpc-port 13389 \
|
||||
--seed 1024 \
|
||||
--tensor-parallel-size 8 \
|
||||
--served-model-name qwen3-vl-235b \
|
||||
--max-num-seqs 16 \
|
||||
--max-model-len 262144 \
|
||||
--max-num-batched-tokens 4096 \
|
||||
--enable-expert-parallel \
|
||||
--trust-remote-code \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--no-enable-prefix-caching \
|
||||
--mm-processor-cache-gb 0 \
|
||||
--limit-mm-per-prompt.image 1 \
|
||||
--limit-mm-per-prompt.video 0 \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"enable_cpu_binding":true,"enable_flashcomm1":true}'
|
||||
```
|
||||
|
||||
If the service starts successfully, the following information is displayed on node 0:
|
||||
|
||||
```shell
|
||||
INFO: Started server process [44610]
|
||||
INFO: Waiting for application startup.
|
||||
INFO: Application startup complete.
|
||||
INFO: Started server process [44611]
|
||||
INFO: Waiting for application startup.
|
||||
INFO: Application startup complete.
|
||||
```
|
||||
|
||||
**Key parameters for MP deployment:**
|
||||
|
||||
- `--data-parallel-size` is the global DP size across all nodes. In the example, 2 DP ranks are used.
|
||||
- `--data-parallel-size-local` is the number of DP ranks on the current node. In the example, each A2 node has 1 local DP rank.
|
||||
- `--data-parallel-start-rank` is the first DP rank on the current node. Node 0 starts from 0 by default, and node 1 starts from 1.
|
||||
- `--data-parallel-address` must point to the master DP node. Use node 0 `local_ip` on node 0 and `node0_ip` on other nodes.
|
||||
- `--data-parallel-rpc-port` is the DP RPC port. Use the same value on all nodes and ensure the port is available.
|
||||
- `--api-server-count` controls how many API server processes are started on the master node.
|
||||
- `--headless` starts a worker node without exposing an API server. Use it on non-master nodes.
|
||||
- `--tensor-parallel-size 8` maps one TP group to the 8 NPUs on each A2 node.
|
||||
- `HCCL_IF_IP`, `GLOO_SOCKET_IFNAME`, `TP_SOCKET_IFNAME`, and `HCCL_SOCKET_IFNAME` bind HCCL, Gloo, and TP communication to the selected network.
|
||||
|
||||
### 5.3 Multi-Node PD Separation Deployment
|
||||
|
||||
PD disaggregation separates Prefill and Decode into different service groups. Prefill nodes process large prompt chunks, Decode nodes serve token generation, and a proxy forwards requests between them. This mode is suitable for production serving scenarios where prefill and decode resource ratios need to be tuned separately.
|
||||
|
||||
We recommend using Mooncake for deployment. Refer to [Mooncake](../features/pd_disaggregation_mooncake_multi_node.md) for the general PD disaggregation workflow and request forwarding setup.
|
||||
|
||||
The following example matches the validated A3 two-node topology for `Qwen3-VL-235B-A22B-Instruct-w8a8-QuaRot`:
|
||||
|
||||
- 1 Prefill node: 1 Atlas 800 A3 (64G x 16), DP2 + TP8 + EP.
|
||||
- 1 Decode node: 1 Atlas 800 A3 (64G x 16), DP4 + TP4 + EP + full decode ACLGraph.
|
||||
|
||||
#### 5.3.1 Prefill Node
|
||||
|
||||
Create `run_p.sh` on the prefill node.
|
||||
|
||||
```shell
|
||||
#!/bin/bash
|
||||
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
|
||||
vllm serve Eco-Tech/Qwen3-VL-235B-A22B-Instruct-w8a8-QuaRot \
|
||||
--host 0.0.0.0 \
|
||||
--port 8080 \
|
||||
--quantization ascend \
|
||||
--data-parallel-size 2 \
|
||||
--data-parallel-size-local 2 \
|
||||
--tensor-parallel-size 8 \
|
||||
--seed 1024 \
|
||||
--served-model-name qwen3-vl-235b \
|
||||
--enable-expert-parallel \
|
||||
--max-num-seqs 32 \
|
||||
--max-model-len 8192 \
|
||||
--max-num-batched-tokens 8192 \
|
||||
--trust-remote-code \
|
||||
--no-enable-prefix-caching \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--kv-transfer-config \
|
||||
'{"kv_connector":"MooncakeConnectorV1",
|
||||
"kv_role":"kv_producer",
|
||||
"kv_port":"30000",
|
||||
"kv_connector_extra_config":{
|
||||
"prefill":{"dp_size":2,"tp_size":8},
|
||||
"decode":{"dp_size":4,"tp_size":4}
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
Common Issues Tip: If the prefill service is not ready for a long time, check whether the model path is shared, all 16 NPUs are visible, and the Mooncake `kv_port` is available.
|
||||
|
||||
#### 5.3.2 Decode Node
|
||||
|
||||
Create `run_d.sh` on the decode node.
|
||||
|
||||
```shell
|
||||
#!/bin/bash
|
||||
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
|
||||
vllm serve Eco-Tech/Qwen3-VL-235B-A22B-Instruct-w8a8-QuaRot \
|
||||
--host 0.0.0.0 \
|
||||
--port 8080 \
|
||||
--quantization ascend \
|
||||
--data-parallel-size 4 \
|
||||
--data-parallel-size-local 4 \
|
||||
--tensor-parallel-size 4 \
|
||||
--seed 1024 \
|
||||
--served-model-name qwen3-vl-235b \
|
||||
--enable-expert-parallel \
|
||||
--max-num-seqs 32 \
|
||||
--max-model-len 8192 \
|
||||
--max-num-batched-tokens 8192 \
|
||||
--trust-remote-code \
|
||||
--no-enable-prefix-caching \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \
|
||||
--kv-transfer-config \
|
||||
'{"kv_connector":"MooncakeConnectorV1",
|
||||
"kv_role":"kv_consumer",
|
||||
"kv_port":"30200",
|
||||
"kv_connector_extra_config":{
|
||||
"prefill":{"dp_size":2,"tp_size":8},
|
||||
"decode":{"dp_size":4,"tp_size":4}
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
**Key parameters for PD disaggregation:**
|
||||
|
||||
- Prefill uses `--data-parallel-size 2`, `--data-parallel-size-local 2`, and `--tensor-parallel-size 8`.
|
||||
- Decode uses `--data-parallel-size 4`, `--data-parallel-size-local 4`, and `--tensor-parallel-size 4`.
|
||||
- `--max-num-batched-tokens` is set to 8192 on both sides in this validation topology. Increase the prefill value only if activation memory is sufficient.
|
||||
- `--kv-transfer-config` sets the Mooncake connector. `kv_role` is `kv_producer` on prefill and `kv_consumer` on decode.
|
||||
- `kv_connector_extra_config.prefill.dp_size/tp_size` and `decode.dp_size/tp_size` must match the actual global DP and TP layout.
|
||||
- `--no-enable-prefix-caching` disables prefix caching. For PD disaggregation, first validate the service without prefix caching before enabling additional cache features.
|
||||
- `--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}'` is recommended on decode nodes to reduce decode dispatch overhead.
|
||||
|
||||
Common Issues Tip: If you encounter issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html) for troubleshooting.
|
||||
|
||||
Service Verification:
|
||||
|
||||
```shell
|
||||
curl http://<server_ip>:<port>/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "qwen3-vl-235b",
|
||||
"messages": [
|
||||
{
|
||||
"role": "user",
|
||||
"content": "Who are you?"
|
||||
}
|
||||
],
|
||||
"max_tokens": 256,
|
||||
"temperature": 0
|
||||
}'
|
||||
```
|
||||
|
||||
Expected Result:
|
||||
|
||||
The service returns HTTP 200 OK with a JSON response containing the `choices` field.
|
||||
|
||||
## 6 Functional Verification
|
||||
|
||||
After the server is started, send a request to verify basic multimodal functionality. For single-node and MP deployment, use the API endpoint on node 0. For PD disaggregation, use the proxy endpoint from the Mooncake deployment guide.
|
||||
|
||||
```shell
|
||||
curl http://<server_ip>:<port>/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "qwen3-vl-235b",
|
||||
"messages": [
|
||||
{"role": "system", "content": "You are a helpful assistant."},
|
||||
{"role": "user", "content": [
|
||||
{"type": "image_url", "image_url": {"url": "https://modelscope.oss-cn-beijing.aliyuncs.com/resource/qwen.png"}},
|
||||
{"type": "text", "text": "What is the text in the illustration?"}
|
||||
]}
|
||||
],
|
||||
"max_completion_tokens": 100,
|
||||
"temperature": 0
|
||||
}'
|
||||
```
|
||||
|
||||
Expected result: the HTTP status is 200 and the JSON response contains a `choices` field with generated text, for example text similar to `TONGYI Qwen`.
|
||||
|
||||
## 7 Accuracy Evaluation
|
||||
|
||||
### Using AISBench
|
||||
|
||||
1. Refer to [Using AISBench](../../developer_guide/evaluation/using_ais_bench.md) for details.
|
||||
|
||||
2. After execution, you can get the result.
|
||||
|
||||
| dataset | version | metric | mode | vllm-api-general-chat |
|
||||
| ------- | ------- | ------ | ---- | --------------------- |
|
||||
| textvqa-lite | - | accuracy | gen | 83 |
|
||||
| aime2024 | - | accuracy | gen | 93 |
|
||||
|
||||
## 8 Performance Evaluation
|
||||
|
||||
### 8.1 Using AISBench
|
||||
|
||||
Refer to [Using AISBench for performance evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation) for details. For multimodal performance, use a dataset with image payloads, such as TextVQA-style requests, instead of random text-only prompts.
|
||||
|
||||
### 8.2 Using vLLM Benchmark
|
||||
|
||||
Run performance evaluation of `Qwen3-VL-235B-A22B-Instruct` as an example. Refer to [vLLM benchmark](https://docs.vllm.ai/en/latest/benchmarking/) for more details.
|
||||
|
||||
There are three `vllm bench` subcommands:
|
||||
|
||||
- `latency`: benchmark the latency of a single batch of requests.
|
||||
- `serve`: benchmark online serving throughput.
|
||||
- `throughput`: benchmark offline inference throughput.
|
||||
|
||||
Take `serve` as an example:
|
||||
|
||||
```shell
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
|
||||
vllm bench serve \
|
||||
--model Eco-Tech/Qwen3-VL-235B-A22B-Instruct-w8a8-QuaRot \
|
||||
--served-model-name qwen3-vl-235b \
|
||||
--dataset-name random \
|
||||
--random-input 200 \
|
||||
--num-prompts 200 \
|
||||
--request-rate 1 \
|
||||
--save-result \
|
||||
--result-dir ./
|
||||
```
|
||||
|
||||
After several minutes, you can get the performance evaluation result. This random benchmark is useful for serving pipeline validation; use AISBench or a custom multimodal dataset for image-token performance.
|
||||
|
||||
## 9 Performance Tuning
|
||||
|
||||
### 9.1 Recommended Configurations
|
||||
|
||||
> **Note**: The following configurations are validated in specific test environments and are for reference only. The optimal configuration depends on hardware type, maximum input/output length, image resolution, request concurrency, prefix cache hit rate, quantization, and prefill/decode ratio. Tune the parameters in Section 9.2 based on your actual workload.
|
||||
|
||||
#### Table 1: Scenario Overview
|
||||
|
||||
| Scenario | Deployment Mode | *Total NPUs | Weight Version | Key Considerations |
|
||||
| -------- | --------------- | ---------- | -------------- | ------------------ |
|
||||
| Functional validation | Single-node online serving | 16 A3 NPUs | W8A8 | Use shorter context, disable video, and set `--mm-processor-cache-gb 0` to reduce memory pressure. |
|
||||
| Long context | Multi-node MP | 16 A3 NPUs | W8A8 | Use TP across each node and DP across nodes. Lower image count or context length if OOM occurs. |
|
||||
| Low latency | 1P1D PD disaggregation | 32 A3 NPUs | W8A8 | Separate prefill and decode resources and enable full decode ACLGraph on decode nodes. |
|
||||
|
||||
> `*Total NPUs` indicates the total number of NPUs used across all nodes. 1 node = 1 Atlas 800 A3 server (64G × 16 NPUs).
|
||||
|
||||
#### Table 2: Detailed Node Configuration
|
||||
|
||||
| Scenario | Node Role | NPUs | TP | DP | Max Num Seqs | Max Model Len | Max Num Batched Tokens | Prefix Cache | Main Optimizations |
|
||||
| -------- | --------- | ---- | -- | -- | ------------ | ------------- | ---------------------- | ------------ | ------------------ |
|
||||
| Functional validation | Single node | 16 | 4 | 4 | 32 | 32768 | 16384 | Off | W8A8, FullGraph, FlashComm1, Fused MC2 |
|
||||
| Long context | MP node | 8 per node | 8 | 1 per node, 2 global | 16 per DP | 262144 | 4096 | Off | FullGraph, FlashComm1, CPU binding |
|
||||
| Low latency | Prefill node | 16 | 8 | 2 | 32 | 8192 | 8192 | Off | Mooncake KV producer, EP |
|
||||
| Low latency | Decode node | 16 | 4 | 4 | 32 | 8192 | 8192 | Off | Mooncake KV consumer, FullGraph, EP |
|
||||
|
||||
> For complete startup commands and parameter descriptions, please refer to the deployment examples in [Chapter 5](#5-online-service-deployment).
|
||||
|
||||
### 9.2 Tuning Guidelines
|
||||
|
||||
#### 9.2.1 General Tuning Reference
|
||||
|
||||
Please refer to the [Public Performance Tuning Documentation](../../developer_guide/performance_and_debug/optimization_and_tuning.md) for tuning methods.
|
||||
|
||||
Please refer to the [Feature Guide](../../user_guide/support_matrix/feature_matrix.md) for detailed feature descriptions.
|
||||
|
||||
#### 9.2.2 Recommended tuning order
|
||||
|
||||
1. Set the deployment topology first. Use single-node deployment for validation, MP deployment for simple multi-node serving, and PD disaggregation when prefill and decode need different resource ratios.
|
||||
2. Choose the maximum context length with `--max-model-len`. Multimodal requests consume KV cache for both text tokens and visual tokens, so reduce image resolution, image count, `--max-num-seqs`, or context length if OOM occurs.
|
||||
3. Tune multimodal limits. Use `--limit-mm-per-prompt.image` and `--limit-mm-per-prompt.video` to match your request shape. Disable video with `--limit-mm-per-prompt.video 0` for image-only services.
|
||||
4. Tune `--max-num-batched-tokens`. Larger values usually improve prefill throughput but increase activation memory. Decode-heavy workloads usually need smaller values.
|
||||
5. Tune `--max-num-seqs` according to service concurrency. Requests above this value wait in the queue and the waiting time is counted in TTFT and TPOT.
|
||||
6. Tune `--gpu-memory-utilization`. Increase it to provide more KV cache, but leave headroom for runtime memory fluctuation, image preprocessing, and expert imbalance.
|
||||
7. Tune ACLGraph capture. `FULL_DECODE_ONLY` is recommended for decode. If you set `cudagraph_capture_sizes` manually, include common decode batch sizes.
|
||||
|
||||
### 9.3 Model-Specific Optimizations
|
||||
|
||||
| Optimization | Enablement | Benefit | Notes |
|
||||
| ------------ | ---------- | ------- | ----- |
|
||||
| Multimodal prompt limits | `--limit-mm-per-prompt.image`, `--limit-mm-per-prompt.video` | Avoids reserving memory for unused media types. | Disable video for image-only serving. |
|
||||
| Multimodal processor cache | `--mm-processor-cache-gb` | Caches processed media features when repeated media appears. | Set to 0 for memory-constrained validation. |
|
||||
| Full decode ACLGraph | `--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}'` | Reduces operator dispatch overhead and stabilizes decode performance. | Recommended for decode-heavy serving. |
|
||||
| FlashComm1 | `VLLM_ASCEND_ENABLE_FLASHCOMM1=1` or `--additional-config '{"enable_flashcomm1":true}'` | Reduces communication overhead in large TP and high-concurrency scenarios. | May not help low-concurrency workloads. |
|
||||
| Fused MC2 | `VLLM_ASCEND_ENABLE_FUSED_MC2=1` | Enables MoE fused operators to improve MoE efficiency. | Compare with disabled state if accuracy or performance regresses. |
|
||||
| Prefix caching | `--enable-prefix-caching` | Improves repeated-prefix workloads. | Validate HBM usage first. For PD, start with prefix caching disabled. |
|
||||
| Asynchronous scheduling | `--async-scheduling` | Can improve high-concurrency throughput. | Disable and compare for latency-sensitive workloads. |
|
||||
| PD disaggregation | `--kv-transfer-config` | Separates prefill and decode resources. | Ensure producer/consumer DP and TP sizes match the actual topology. |
|
||||
|
||||
## 10 FAQ
|
||||
|
||||
For common environment, installation, and general parameter issues, refer to [Public FAQs](../../faqs.md). This section only covers model-specific issues for Qwen3-VL-235B-A22B-Instruct.
|
||||
|
||||
### Q1: Why does the service report OOM during startup or soon after accepting requests?
|
||||
|
||||
**Phenomenon:** The service fails during profile run, or it starts successfully but reports OOM when real traffic arrives.
|
||||
|
||||
**Cause:** Qwen3-VL-235B-A22B-Instruct has high weight, KV cache, and multimodal preprocessing memory requirements. Large `--max-model-len`, large `--max-num-seqs`, large `--max-num-batched-tokens`, high image resolution, too many images per prompt, or high `--gpu-memory-utilization` can leave insufficient HBM headroom.
|
||||
|
||||
**Solution:** Use the W8A8 model with `--quantization ascend` when possible, lower `--max-model-len`, lower `--max-num-seqs`, lower `--max-num-batched-tokens`, lower image/video limits, or reduce `--gpu-memory-utilization`. Keep `PYTORCH_NPU_ALLOC_CONF=expandable_segments:True`.
|
||||
|
||||
### Q2: Why does multi-node MP deployment hang during initialization?
|
||||
|
||||
**Phenomenon:** One node waits for other ranks, HCCL initialization times out, or the headless node exits.
|
||||
|
||||
**Cause:** Network interface names, IP addresses, DP ranks, or RPC ports are inconsistent across nodes.
|
||||
|
||||
**Solution:** Verify multi-node communication first. Ensure `HCCL_IF_IP`, `GLOO_SOCKET_IFNAME`, `TP_SOCKET_IFNAME`, and `HCCL_SOCKET_IFNAME` match the selected NIC. Ensure all nodes use the same `--data-parallel-rpc-port`, non-master nodes use `--headless`, and `--data-parallel-start-rank` does not overlap.
|
||||
|
||||
### Q3: Why is video disabled in the image-only examples?
|
||||
|
||||
**Phenomenon:** The service reserves more memory than expected, or startup OOM occurs even when requests only contain images.
|
||||
|
||||
**Cause:** Allowing video inputs can reserve memory for long visual embeddings and preprocessing paths that are not needed by image-only workloads.
|
||||
|
||||
**Solution:** Use `--limit-mm-per-prompt.video 0` for image-only serving. Enable video only when your workload needs it, and lower `--max-model-len` or request concurrency if needed.
|
||||
|
||||
### Q4: Why does enabling prefix caching not improve performance?
|
||||
|
||||
**Phenomenon:** Prefix caching is enabled, but throughput or latency does not improve.
|
||||
|
||||
**Cause:** Prefix caching only helps when requests share reusable prefixes. Random prompts, unique images, or low cache hit rates may add memory pressure without visible gains.
|
||||
|
||||
**Solution:** Enable prefix caching for repeated-prefix workloads. For random benchmark datasets, memory-constrained long-context workloads, or PD validation, compare with `--no-enable-prefix-caching`.
|
||||
|
||||
### Q5: Why does PD disaggregation fail to transfer KV cache?
|
||||
|
||||
**Phenomenon:** Requests reach the proxy or prefill service, but decode nodes do not produce output or report KV transfer errors.
|
||||
|
||||
**Cause:** The Mooncake connector ports, producer/consumer roles, or `kv_connector_extra_config` DP/TP sizes do not match the actual topology.
|
||||
|
||||
**Solution:** Check `kv_role`, `kv_port`, and the prefill/decode DP/TP sizes on all nodes. Start from the validated topology in Section 5.4, then change one dimension at a time.
|
||||
499
docs/source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md
Normal file
499
docs/source/tutorials/models/Qwen3-VL-30B-A3B-Instruct.md
Normal file
@@ -0,0 +1,499 @@
|
||||
# Qwen3-VL-30B-A3B-Instruct
|
||||
|
||||
## 1 Introduction
|
||||
|
||||
Qwen3-VL-30B-A3B-Instruct is a sparse MoE vision-language model in the Qwen3-VL family, with about 30B total parameters and about 3B activated parameters per token. It is suitable for image understanding, video understanding, multimodal dialogue, and long-context online serving on Ascend hardware.
|
||||
|
||||
This document describes the main validation steps for the model, including supported features, prerequisites, installation, image and video online deployment, offline inference, functional verification, accuracy and performance evaluation, performance tuning, and FAQs.
|
||||
|
||||
The `Qwen3-VL-30B-A3B-Instruct` tutorial was introduced for the `vllm-ascend` `v0.13.0` validation cycle. Use `v0.13.0` or later for this model. The examples below use the version placeholder configured by the documentation build system.
|
||||
|
||||
## 2 Supported Features
|
||||
|
||||
Refer to [Supported Features List](../../user_guide/support_matrix/supported_features.md) to get the model's supported feature matrix.
|
||||
|
||||
Refer to [Feature Guide](../../user_guide/feature_guide/index.md) to get the feature's configuration.
|
||||
|
||||
## 3 Prerequisites
|
||||
|
||||
### 3.1 Model Weight
|
||||
|
||||
- `Qwen3-VL-30B-A3B-Instruct` (BF16 version): requires 1 Atlas 800 A3 (64G x 16) node or 1 Atlas 800 A2 (64G x 8) node. [Model Weight](https://www.modelscope.cn/models/Qwen/Qwen3-VL-30B-A3B-Instruct).
|
||||
- `Qwen3-VL-30B-A3B-Instruct-w8a8-mxfp8` (quantized version): requires 1 Ascend 950DT (96G x 8) node. [Model Weight](https://modelscope.cn/models/Eco-Tech/Qwen3-VL-30B-A3B-Instruct-w8a8-mxfp8).
|
||||
|
||||
It is recommended to download the model weight to a shared directory across multiple nodes.
|
||||
|
||||
## 4 Installation
|
||||
|
||||
### 4.1 Docker Image Installation
|
||||
|
||||
Select an image based on your machine type and start the docker image on your node, refer to [using docker](../../installation.md#set-up-using-docker).
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: install
|
||||
|
||||
::::{tab-item} Ascend 950DT series
|
||||
:sync: 950dt
|
||||
|
||||
Start the docker image on your each node.
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-#TODO
|
||||
export NAME=vllm-ascend
|
||||
|
||||
docker run --rm \
|
||||
--name $NAME \
|
||||
--net=host \
|
||||
--shm-size=1g \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/hisi_hdc \
|
||||
--device /dev/ummu \
|
||||
--device /dev/uburma \
|
||||
-v /usr/local/Ascend/driver:/usr/local/Ascend/driver \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /etc/hccl_rootinfo.json:/etc/hccl_rootinfo.json \
|
||||
-v /etc/hixlep/:/etc/hixlep/ \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-v /usr/local/sbin:/usr/local/sbin \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/sbin/npu-smi:/usr/local/sbin/npu-smi \
|
||||
-v /usr/lib64:/usr/lib64 \
|
||||
-itd $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} A3 series
|
||||
:sync: A3
|
||||
|
||||
Start the docker image on each node.
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-a3
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=512g \
|
||||
--net=host \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci8 \
|
||||
--device /dev/davinci9 \
|
||||
--device /dev/davinci10 \
|
||||
--device /dev/davinci11 \
|
||||
--device /dev/davinci12 \
|
||||
--device /dev/davinci13 \
|
||||
--device /dev/davinci14 \
|
||||
--device /dev/davinci15 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /etc/hccn.conf:/etc/hccn.conf \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} A2 series
|
||||
:sync: A2
|
||||
|
||||
Start the docker image on each node.
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=512g \
|
||||
--net=host \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /etc/hccn.conf:/etc/hccn.conf \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
After a successful docker run, you can verify the running container service by executing the `docker ps` command.
|
||||
|
||||
### 4.2 Source Code Installation
|
||||
|
||||
If you prefer not to use the Docker image, you can build from source. Install vLLM from source first:
|
||||
|
||||
1. Clone and install vLLM:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/vllm-project/vllm.git
|
||||
cd vllm
|
||||
pip install -e .
|
||||
```
|
||||
|
||||
2. Clone and install the vLLM-Ascend repository:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/vllm-project/vllm-ascend.git
|
||||
cd vllm-ascend
|
||||
pip install -e .
|
||||
```
|
||||
|
||||
**Installation Verification:**
|
||||
|
||||
```bash
|
||||
pip show vllm vllm-ascend
|
||||
```
|
||||
|
||||
Expected result: The version information for both packages is displayed, confirming a successful installation.
|
||||
|
||||
:::{note}
|
||||
If deploying a multi-node environment, set up the environment on each node.
|
||||
:::
|
||||
|
||||
For more details, please refer to the [Installation Guide](../../installation.md).
|
||||
|
||||
## 5 Online Service Deployment
|
||||
|
||||
### 5.1 Single-Node Online Deployment
|
||||
|
||||
Single-node deployment runs both Prefill and Decode on the same node. The following examples are suitable for image-only online serving.
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: deploy
|
||||
|
||||
::::{tab-item} Ascend 950DT series
|
||||
:sync: 950dt
|
||||
|
||||
Run the following script to execute online inference on 1 Ascend 950DT (96G x 8). The quantized version (`Qwen3-VL-30B-A3B-Instruct-w8a8-mxfp8`) can be deployed on a single Ascend 950DT node. The W8A8 version needs `--quantization ascend`.
|
||||
|
||||
```shell
|
||||
#!/bin/sh
|
||||
|
||||
# Load model from ModelScope to speed up download.
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
|
||||
# Reduce memory fragmentation and avoid out-of-memory errors.
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export HCCL_BUFFSIZE=400
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=100
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export VLLM_ASCEND_ENABLE_FLASHCOMM1=1
|
||||
|
||||
vllm serve Eco-Tech/Qwen3-VL-30B-A3B-Instruct-w8a8-mxfp8 \
|
||||
--host 0.0.0.0 \
|
||||
--port 8000 \
|
||||
--distributed-executor-backend mp \
|
||||
--data-parallel-size 1 \
|
||||
--tensor-parallel-size 8 \
|
||||
--enable-expert-parallel \
|
||||
--seed 1024 \
|
||||
--quantization ascend \
|
||||
--served-model-name qwen3-vl-30b \
|
||||
--max-num-seqs 32 \
|
||||
--max-model-len 32768 \
|
||||
--max-num-batched-tokens 8192 \
|
||||
--trust-remote-code \
|
||||
--no-enable-prefix-caching \
|
||||
--mm-processor-cache-gb 0 \
|
||||
--limit-mm-per-prompt.image 1 \
|
||||
--limit-mm-per-prompt.video 0 \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}'
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} A3 series
|
||||
:sync: A3
|
||||
|
||||
Run the following script to start image-only serving on 1 Atlas 800 A3 (64G x 16) node.
|
||||
|
||||
```shell
|
||||
#!/bin/sh
|
||||
|
||||
# Load model from ModelScope to speed up download.
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
|
||||
# Reduce memory fragmentation and avoid out-of-memory errors.
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export OMP_NUM_THREADS=1
|
||||
export OMP_PROC_BIND=false
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export VLLM_ASCEND_ENABLE_FLASHCOMM1=1
|
||||
export VLLM_ASCEND_ENABLE_FUSED_MC2=1
|
||||
|
||||
vllm serve Qwen/Qwen3-VL-30B-A3B-Instruct \
|
||||
--host 0.0.0.0 \
|
||||
--port 8000 \
|
||||
--served-model-name qwen3-vl-30b \
|
||||
--data-parallel-size 1 \
|
||||
--tensor-parallel-size 2 \
|
||||
--enable-expert-parallel \
|
||||
--seed 1024 \
|
||||
--max-num-seqs 32 \
|
||||
--max-model-len 32768 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--no-enable-prefix-caching \
|
||||
--mm-processor-cache-gb 0 \
|
||||
--limit-mm-per-prompt.image 1 \
|
||||
--limit-mm-per-prompt.video 0 \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY","cudagraph_capture_sizes":[1,2,4,8,16,24,32]}'
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} A2 series
|
||||
:sync: A2
|
||||
|
||||
The A3 series script above also works on 1 Atlas 800 A2 (64G x 8) node.
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
Key Parameter Descriptions:
|
||||
|
||||
- `--tensor-parallel-size 2` maps the model to two NPUs. Increase TP only after validating memory, communication, and throughput on your hardware.
|
||||
- `--enable-expert-parallel` enables expert parallelism for MoE layers. Do not mix MoE tensor parallelism and expert parallelism in the same MoE layer.
|
||||
- `--max-model-len` is the maximum input plus output length for a single request. By default, the model can support long context, but `128000` is a practical validation value for many image/video workloads.
|
||||
- `--max-num-seqs` is the maximum number of active requests scheduled by each DP group. Video requests consume more memory, so the video example uses a smaller value.
|
||||
- `--max-num-batched-tokens` is the maximum number of tokens processed in one scheduler step. A larger value can improve prefill efficiency but consumes more activation memory.
|
||||
- `--gpu-memory-utilization` controls how much HBM vLLM can use to calculate KV cache capacity. Increase it only after confirming the service is stable.
|
||||
- `--limit-mm-per-prompt.video 0` disables video inputs and saves memory for image-only serving.
|
||||
- `--allowed-local-media-path /media` allows requests to use local files such as `file:///media/test.mp4`.
|
||||
- `--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}'` enables full decode ACLGraph replay to reduce dispatch overhead.
|
||||
|
||||
Common Issues Tip: If you encounter issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html) for troubleshooting.
|
||||
|
||||
Service Verification:
|
||||
|
||||
```shell
|
||||
curl http://<server_ip>:<port>/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "qwen3-vl-30b",
|
||||
"messages": [
|
||||
{
|
||||
"role": "user",
|
||||
"content": "Who are you?"
|
||||
}
|
||||
],
|
||||
"max_tokens": 256,
|
||||
"temperature": 0
|
||||
}'
|
||||
```
|
||||
|
||||
Expected Result:
|
||||
|
||||
The service returns HTTP 200 OK with a JSON response containing the `choices` field.
|
||||
|
||||
## 6 Functional Verification
|
||||
|
||||
After the server is started, send a request to verify basic multimodal functionality.
|
||||
|
||||
```shell
|
||||
curl http://<server_ip>:<port>/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "qwen3-vl-30b",
|
||||
"messages": [
|
||||
{"role": "system", "content": "You are a helpful assistant."},
|
||||
{"role": "user", "content": [
|
||||
{"type": "image_url", "image_url": {"url": "https://modelscope.oss-cn-beijing.aliyuncs.com/resource/qwen.png"}},
|
||||
{"type": "text", "text": "What is the text in the illustration?"}
|
||||
]}
|
||||
],
|
||||
"max_completion_tokens": 100,
|
||||
"temperature": 0
|
||||
}'
|
||||
```
|
||||
|
||||
Expected result: the HTTP status is 200 and the JSON response contains a `choices` field with generated text, for example text similar to `TONGYI Qwen`.
|
||||
|
||||
## 7 Accuracy Evaluation
|
||||
|
||||
### Using AISBench
|
||||
|
||||
1. Refer to [Using AISBench](../../developer_guide/evaluation/using_ais_bench.md) for details.
|
||||
|
||||
2. After execution, you can get the result.
|
||||
|
||||
| dataset | version | metric | mode | result |
|
||||
| ------- | ------- | ------ | ---- | ------ |
|
||||
| mmmu_val | - | acc,none | gen | 0.58 |
|
||||
|
||||
## 8 Performance Evaluation
|
||||
|
||||
### 8.1 Using AISBench
|
||||
|
||||
Refer to [Using AISBench for performance evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation) for details. For image or video performance, use a dataset with real multimodal payloads instead of random text-only prompts.
|
||||
|
||||
### 8.2 Using vLLM Benchmark
|
||||
|
||||
Run performance evaluation of `Qwen3-VL-30B-A3B-Instruct` as an example. Refer to [vLLM benchmark](https://docs.vllm.ai/en/latest/benchmarking/) for more details.
|
||||
|
||||
There are three `vllm bench` subcommands:
|
||||
|
||||
- `latency`: benchmark the latency of a single batch of requests.
|
||||
- `serve`: benchmark online serving throughput.
|
||||
- `throughput`: benchmark offline inference throughput.
|
||||
|
||||
Take `serve` as an example:
|
||||
|
||||
```shell
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
|
||||
vllm bench serve \
|
||||
--model Qwen/Qwen3-VL-30B-A3B-Instruct \
|
||||
--served-model-name qwen3-vl-30b \
|
||||
--dataset-name random \
|
||||
--random-input 200 \
|
||||
--num-prompts 200 \
|
||||
--request-rate 1 \
|
||||
--save-result \
|
||||
--result-dir ./
|
||||
```
|
||||
|
||||
After several minutes, you can get the performance evaluation result. This random benchmark is useful for serving pipeline validation; use AISBench or a custom multimodal dataset for image/video-token performance.
|
||||
|
||||
## 9 Performance Tuning
|
||||
|
||||
### 9.1 Recommended Configurations
|
||||
|
||||
> **Note**: The following configurations are validated in specific test environments and are for reference only. The optimal configuration depends on hardware type, image resolution, video length, maximum input/output length, request concurrency, prefix cache hit rate, and prefill/decode ratio. Tune the parameters in Section 9.2 based on your actual workload.
|
||||
|
||||
#### Table 1: Scenario Overview
|
||||
|
||||
| Scenario | Deployment Mode | *Total NPUs | Weight Version | Key Considerations |
|
||||
| -------- | --------------- | ---------- | -------------- | ------------------ |
|
||||
| Image-only serving | Single-node online serving | 2 or more NPUs | BF16 | Disable video, tune context length, and keep enough KV cache for visual tokens. |
|
||||
| Video serving | Single-node online serving | 2 or more NPUs | BF16 | Use local media paths, lower concurrency, and reduce video length or frame sampling if OOM occurs. |
|
||||
| Functional graph validation | Single-node PP | 2 NPUs | BF16 | Use shorter context and explicit capture sizes to validate full decode ACLGraph behavior. |
|
||||
|
||||
> `*Total NPUs` indicates the total number of NPUs used across all nodes. 1 node = 1 Atlas 800 A3 server (64G × 16 NPUs).
|
||||
|
||||
#### Table 2: Detailed Node Configuration
|
||||
|
||||
| Scenario | Node Role | NPUs | TP | PP | Max Num Seqs | Max Model Len | Max Num Batched Tokens | Prefix Cache | Main Optimizations |
|
||||
| -------- | --------- | ---- | -- | -- | ------------ | ------------- | ---------------------- | ------------ | ------------------ |
|
||||
| Image-only serving | Single node | 2 or more | 2 | 1 | 16 | 128000 | 4096 | Workload dependent | FullGraph, EP, video disabled |
|
||||
| Video serving | Single node | 2 or more | 2 | 1 | 8 | 128000 | 4096 | Workload dependent | FullGraph, EP, local media path |
|
||||
| Graph validation | Single node | 2 | 1 | 2 | Tune by test | 4096 | 1024 | Off | FullGraph capture sizes |
|
||||
|
||||
> For complete startup commands and parameter descriptions, please refer to the deployment examples in [Chapter 5](#5-online-service-deployment).
|
||||
|
||||
### 9.2 Tuning Guidelines
|
||||
|
||||
#### 9.2.1 General Tuning Reference
|
||||
|
||||
Please refer to the [Public Performance Tuning Documentation](../../developer_guide/performance_and_debug/optimization_and_tuning.md) for tuning methods.
|
||||
|
||||
Please refer to the [Feature Guide](../../user_guide/support_matrix/feature_matrix.md) for detailed feature descriptions.
|
||||
|
||||
#### 9.2.2 Recommended tuning order
|
||||
|
||||
1. Start from image-only serving. Add video only after the image path is stable.
|
||||
2. Choose the maximum context length with `--max-model-len`. Multimodal requests consume KV cache for both text tokens and visual tokens, so reduce image resolution, video length, request concurrency, or context length if OOM occurs.
|
||||
3. Tune multimodal limits. Use `--limit-mm-per-prompt.image` and `--limit-mm-per-prompt.video` to match your request shape.
|
||||
4. Tune `--max-num-batched-tokens`. Larger values usually improve prefill throughput but increase activation memory. Video-heavy workloads usually need conservative values.
|
||||
5. Tune `--max-num-seqs` according to service concurrency. Video requests are more memory intensive than image requests, so start with a smaller value.
|
||||
6. Tune `--gpu-memory-utilization`. Increase it to provide more KV cache, but leave headroom for runtime memory fluctuation and media preprocessing.
|
||||
7. Tune ACLGraph capture. `FULL_DECODE_ONLY` is recommended for decode. If you set `cudagraph_capture_sizes` manually, include common decode batch sizes.
|
||||
|
||||
### 9.3 Model-Specific Optimizations
|
||||
|
||||
| Optimization | Enablement | Benefit | Notes |
|
||||
| ------------ | ---------- | ------- | ----- |
|
||||
| Multimodal prompt limits | `--limit-mm-per-prompt.image`, `--limit-mm-per-prompt.video` | Avoids reserving memory for unused media types. | Disable video for image-only serving. |
|
||||
| Local media access | `--allowed-local-media-path /media` | Avoids slow network video downloads during serving. | Use `file:///media/...` in requests. |
|
||||
| Full decode ACLGraph | `--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}'` | Reduces operator dispatch overhead and stabilizes decode performance. | Recommended for decode-heavy serving. |
|
||||
| Expert parallelism | `--enable-expert-parallel` | Improves MoE serving throughput. | Do not mix MoE tensor parallelism and expert parallelism in the same MoE layer. |
|
||||
| Prefix caching | `--enable-prefix-caching` | Improves repeated-prefix workloads. | Random prompts or unique media may not benefit. |
|
||||
| Asynchronous scheduling | `--async-scheduling` | Can improve high-concurrency throughput. | Disable and compare for latency-sensitive workloads. |
|
||||
| Pipeline parallel validation | `--pipeline-parallel-size 2` | Provides another two-card validation layout. | Use shorter context and lower batch tokens for functional tests. |
|
||||
|
||||
## 10 FAQ
|
||||
|
||||
For common environment, installation, and general parameter issues, refer to [Public FAQs](../../faqs.md). This section only covers model-specific issues for Qwen3-VL-30B-A3B-Instruct.
|
||||
|
||||
### Q1: Why does the service report OOM during startup?
|
||||
|
||||
**Phenomenon:** The service fails during profile run or exits before accepting requests.
|
||||
|
||||
**Cause:** Long context, high image resolution, video inputs, large `--max-num-seqs`, large `--max-num-batched-tokens`, or high `--gpu-memory-utilization` can leave insufficient HBM headroom.
|
||||
|
||||
**Solution:** Start with image-only serving, set `--limit-mm-per-prompt.video 0`, reduce `--max-model-len`, lower `--max-num-seqs`, lower `--max-num-batched-tokens`, or reduce `--gpu-memory-utilization`. Keep `PYTORCH_NPU_ALLOC_CONF=expandable_segments:True`.
|
||||
|
||||
### Q2: Why is video disabled in the image-only command?
|
||||
|
||||
**Phenomenon:** The service reserves more memory than expected even when requests only contain images.
|
||||
|
||||
**Cause:** Allowing video inputs can reserve memory for long visual embeddings and preprocessing paths.
|
||||
|
||||
**Solution:** Use `--limit-mm-per-prompt.video 0` for image-only serving. Enable video only when the workload needs it.
|
||||
|
||||
### Q3: Why does the video request fail with a local file path?
|
||||
|
||||
**Phenomenon:** The request reports that the file is not allowed or cannot be found.
|
||||
|
||||
**Cause:** The server can only access local media paths that are mounted into the container and allowed by `--allowed-local-media-path`.
|
||||
|
||||
**Solution:** Mount the host media directory to `/media`, start the server with `--allowed-local-media-path /media`, and use a request URL like `file:///media/test.mp4`.
|
||||
|
||||
### Q4: Why does enabling prefix caching not improve performance?
|
||||
|
||||
**Phenomenon:** Prefix caching is enabled, but throughput or latency does not improve.
|
||||
|
||||
**Cause:** Prefix caching only helps when requests share reusable prefixes. Unique images, unique videos, or random prompts may add memory pressure without visible gains.
|
||||
|
||||
**Solution:** Enable prefix caching for repeated-prefix workloads. For random benchmarks or memory-constrained video workloads, compare with prefix caching disabled.
|
||||
|
||||
### Q5: Why does multimodal accuracy evaluation fail to insert image tokens?
|
||||
|
||||
**Phenomenon:** Evaluation fails because image placeholders cannot be found in the prompt.
|
||||
|
||||
**Cause:** Qwen3-VL multimodal tasks rely on the model chat template to insert image placeholder tokens before multimodal processing.
|
||||
|
||||
**Solution:** Enable chat template application in the evaluation configuration. For lm_eval-based multimodal tasks, set `apply_chat_template` to true.
|
||||
283
docs/source/tutorials/models/Qwen3-VL-Embedding.md
Normal file
283
docs/source/tutorials/models/Qwen3-VL-Embedding.md
Normal file
@@ -0,0 +1,283 @@
|
||||
# Qwen3-VL-Embedding
|
||||
|
||||
## 1 Introduction
|
||||
|
||||
The Qwen3-VL-Embedding and Qwen3-VL-Reranker model series are the latest additions to the Qwen family, built upon the recently open-sourced and powerful Qwen3-VL foundation model. Specifically designed for multimodal information retrieval and cross-modal understanding, this suite accepts diverse inputs including text, images, screenshots, and videos, as well as inputs containing a mixture of these modalities. This guide describes how to run the model with vLLM Ascend.
|
||||
|
||||
## 2 Supported Features
|
||||
|
||||
Refer to [supported features](../../user_guide/support_matrix/supported_models.md) to get the model's supported feature matrix.
|
||||
|
||||
## 3 Prerequisites
|
||||
|
||||
### 3.1 Model Weight
|
||||
|
||||
- `Qwen3-VL-Embedding-8B` [Download model weight](https://www.modelscope.cn/models/Qwen/Qwen3-VL-Embedding-8B)
|
||||
- `Qwen3-VL-Embedding-2B` [Download model weight](https://www.modelscope.cn/models/Qwen/Qwen3-VL-Embedding-2B)
|
||||
|
||||
It is recommended to download the model weight to the shared directory of multiple nodes, such as `/root/.cache/`
|
||||
|
||||
## 4 Installation
|
||||
|
||||
### 4.1 Docker Image Installation
|
||||
|
||||
You can use our official docker image to run `Qwen3-VL-Embedding` model directly.
|
||||
|
||||
Select an image based on your machine type and start the docker image on your node, refer to [using docker](../../installation.md#set-up-using-docker).
|
||||
|
||||
For Atlas 300I DUO, use `vllm-ascend:nightly-releases-v0.23.0-310p` (or a later `-310p` image).
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: install
|
||||
|
||||
::::{tab-item} A3 series
|
||||
:sync: A3 series
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-a3
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--privileged=true \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} A2 series
|
||||
:sync: A2 series
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--privileged=true \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
:sync: Atlas 300I DUO
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-310p
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--privileged=true \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
After a successful docker run, you can verify the running container service by executing the `docker ps` command.
|
||||
|
||||
### 4.2 Source Code Installation
|
||||
|
||||
If you don't want to use the docker image as above, you can also build all from source:
|
||||
|
||||
- Install `vllm-ascend` from source, refer to [installation](../../installation.md).
|
||||
|
||||
If you want to deploy multi-node environment, you need to set up environment on each node.
|
||||
|
||||
## 5 Online Service Deployment
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: Deployment
|
||||
|
||||
::::{tab-item} A3/A2 series
|
||||
:sync: A3/A2 series
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
vllm serve Qwen/Qwen3-VL-Embedding-2B \
|
||||
--served-model-name Qwen/Qwen3-VL-Embedding-2B \
|
||||
--runner pooling \
|
||||
--port 8000 \
|
||||
--max-model-len 1024
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
:sync: Atlas 300I DUO
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
vllm serve Qwen/Qwen3-VL-Embedding-2B \
|
||||
--served-model-name Qwen/Qwen3-VL-Embedding-2B \
|
||||
--compilation-config '{"cudagraph_capture_sizes": [1024,512]}' \
|
||||
--additional-config '{"ascend_compilation_config": {"fuse_norm_quant": false}}' \
|
||||
--runner pooling \
|
||||
--dtype float16 \
|
||||
--port 8000 \
|
||||
--max-model-len 1024
|
||||
```
|
||||
|
||||
Required Parameter Descriptions:
|
||||
|
||||
`--compilation-config` For Atlas 300I DUO, due to limited hardware streams, the size of cudagraph_capture_sizes is restricted.
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
Key Parameter Descriptions:
|
||||
|
||||
- `--max-model-len` represents the context length, which is the maximum value of the input plus output for a single request. For Atlas 300I DUO if automatic parsing resolves to a large context length, allocating this mask (O(max_model_len^2)) may exceed NPU memory and trigger OOM. Be sure to set an explicit and conservative value, such as --max-model-len 1024.
|
||||
|
||||
Common Issues Tip: If you encounter issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html) for troubleshooting.
|
||||
|
||||
## 6 Functional Verification
|
||||
|
||||
Once your server is started, you can verify by follow command:
|
||||
|
||||
Service Verification:
|
||||
|
||||
```bash
|
||||
curl -X POST http://localhost:8000/v1/embeddings -H "Content-Type: application/json" -d '{
|
||||
"input": [
|
||||
"The capital of China is Beijing.",
|
||||
"Gravity is a force that attracts two bodies towards each other. It gives weight to physical objects and is responsible for the movement of planets around the sun."
|
||||
]
|
||||
}'
|
||||
```
|
||||
|
||||
Expected Result:
|
||||
|
||||
The service returns HTTP 200 OK with a JSON response containing the `embedding` field. Example output:
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "embd-8136155c01e8411d",
|
||||
"object": "list",
|
||||
"created": 1784538286,
|
||||
"model": "Qwen/Qwen3-VL-Embedding-2B",
|
||||
"data": [
|
||||
{
|
||||
"index": 0,
|
||||
"object": "embedding",
|
||||
"embedding": [
|
||||
-0.028474265709519386,
|
||||
-0.02678542211651802
|
||||
]
|
||||
},
|
||||
{
|
||||
"index": 1,
|
||||
"object": "embedding",
|
||||
"embedding": [
|
||||
-0.016785264015197754,
|
||||
-0.003787524998188019
|
||||
]
|
||||
}
|
||||
],
|
||||
"usage": {
|
||||
"prompt_tokens": 39,
|
||||
"total_tokens": 39,
|
||||
"completion_tokens": 0,
|
||||
"prompt_tokens_details": null
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
For more usage examples, please reference the [examples](https://github.com/vllm-project/vllm/tree/main/examples/pooling/embed)
|
||||
|
||||
## 7 Accuracy Evaluation
|
||||
|
||||
Here are two accuracy evaluation methods.
|
||||
|
||||
### Using MTEB
|
||||
|
||||
1. Refer to [MTEB](https://docs.mteb.org/) for details.
|
||||
|
||||
2. Run follow code to execute the accuracy evaluation.
|
||||
|
||||
```python
|
||||
|
||||
import os
|
||||
import mteb
|
||||
|
||||
from mteb.models.vllm_wrapper import VllmEncoderWrapper
|
||||
|
||||
if __name__ == "__main__":
|
||||
|
||||
data_path = "/home/data/mteb_data"
|
||||
os.environ["HF_DATASETS_CACHE"] = data_path
|
||||
os.environ["HF_ENDPOINT"] = "https://hf-mirror.com"
|
||||
|
||||
model = VllmEncoderWrapper(f"/root/.cache/Qwen3-VL-Embedding-2B",
|
||||
revision="norm",
|
||||
dtype="float16",
|
||||
max_model_len=10240,
|
||||
)
|
||||
|
||||
cache = mteb.ResultCache("/home/data/mteb_data")
|
||||
tasks = mteb.get_tasks(tasks=["LeCaRDv2"])
|
||||
results = mteb.evaluate(model, tasks=tasks, cache=cache, encode_kwargs={"batch_size": 2}, overwrite_strategy="always")
|
||||
df = results.to_dataframe()
|
||||
print(df)
|
||||
```
|
||||
|
||||
3. After execution, you can get the result.
|
||||
|
||||
## 8 Performance Evaluation
|
||||
|
||||
### Using vLLM Benchmark
|
||||
|
||||
Run performance of `Qwen3-VL-Embedding-2B` as an example.
|
||||
Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/cli/) for more details.
|
||||
|
||||
Take the `serve` as an example. Run the code as follows.
|
||||
|
||||
```bash
|
||||
vllm bench serve --model Qwen/Qwen3-VL-Embedding-2B --backend openai-embeddings --port 8000 --dataset-name random --endpoint /v1/embeddings --random-input 200 --save-result --result-dir ./
|
||||
```
|
||||
|
||||
After about several minutes, you can get the performance evaluation result.
|
||||
|
||||
## 9 FAQ
|
||||
|
||||
For common environment, installation, and general parameter issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html).
|
||||
320
docs/source/tutorials/models/Qwen3-VL-Reranker.md
Normal file
320
docs/source/tutorials/models/Qwen3-VL-Reranker.md
Normal file
@@ -0,0 +1,320 @@
|
||||
# Qwen3-VL-Reranker
|
||||
|
||||
## 1 Introduction
|
||||
|
||||
The Qwen3-VL-Embedding and Qwen3-VL-Reranker model series are the latest additions to the Qwen family, built upon the recently open-sourced and powerful Qwen3-VL foundation model. Specifically designed for multimodal information retrieval and cross-modal understanding, this suite accepts diverse inputs including text, images, screenshots, and videos, as well as inputs containing a mixture of these modalities. This guide describes how to run the model with vLLM Ascend.
|
||||
|
||||
## 2 Supported Features
|
||||
|
||||
Refer to [supported features](../../user_guide/support_matrix/supported_models.md) to get the model's supported feature matrix.
|
||||
|
||||
## 3 Prerequisites
|
||||
|
||||
### 3.1 Model Weight
|
||||
|
||||
- `Qwen3-VL-Reranker-8B` [Download model weight](https://www.modelscope.cn/models/Qwen/Qwen3-VL-Reranker-8B)
|
||||
- `Qwen3-VL-Reranker-2B` [Download model weight](https://www.modelscope.cn/models/Qwen/Qwen3-VL-Reranker-2B)
|
||||
|
||||
It is recommended to download the model weight to the shared directory of multiple nodes, such as `/root/.cache/`
|
||||
|
||||
## 4 Installation
|
||||
|
||||
### 4.1 Docker Image Installation
|
||||
|
||||
You can use our official docker image to run `Qwen3-VL-Reranker` model directly.
|
||||
|
||||
Select an image based on your machine type and start the docker image on your node, refer to [using docker](../../installation.md#set-up-using-docker).
|
||||
|
||||
For Atlas 300I DUO, use `vllm-ascend:nightly-releases-v0.23.0-310p` (or a later `-310p` image).
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: install
|
||||
|
||||
::::{tab-item} A3 series
|
||||
:sync: A3 series
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-a3
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--privileged=true \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} A2 series
|
||||
:sync: A2 series
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--privileged=true \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
:sync: Atlas 300I DUO
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-310p
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--privileged=true \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
After a successful docker run, you can verify the running container service by executing the `docker ps` command.
|
||||
|
||||
### 4.2 Source Code Installation
|
||||
|
||||
If you don't want to use the docker image as above, you can also build all from source:
|
||||
|
||||
- Install `vllm-ascend` from source, refer to [installation](../../installation.md).
|
||||
|
||||
If you want to deploy multi-node environment, you need to set up environment on each node.
|
||||
|
||||
## 5 Online Service Deployment
|
||||
|
||||
### 5.1 Chat Template
|
||||
|
||||
The Qwen3-VL-Reranker model requires a specific chat template for proper formatting. Create a file named `qwen3_vl_reranker.jinja` with the following content:
|
||||
|
||||
```jinja
|
||||
<|im_start|>system
|
||||
Judge whether the Document meets the requirements based on the Query and the Instruct provided. Note that the answer can only be "yes" or "no".<|im_end|>
|
||||
<|im_start|>user
|
||||
<Instruct>: {{
|
||||
messages
|
||||
| selectattr("role", "eq", "system")
|
||||
| map(attribute="content")
|
||||
| first
|
||||
| default("Given a search query, retrieve relevant candidates that answer the query.")
|
||||
}}<Query>:{{
|
||||
messages
|
||||
| selectattr("role", "eq", "query")
|
||||
| map(attribute="content")
|
||||
| first
|
||||
}}
|
||||
<Document>:{{
|
||||
messages
|
||||
| selectattr("role", "eq", "document")
|
||||
| map(attribute="content")
|
||||
| first
|
||||
}}<|im_end|>
|
||||
<|im_start|>assistant
|
||||
|
||||
```
|
||||
|
||||
Save this file to a location of your choice (e.g., `./qwen3_vl_reranker.jinja`).
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: Deployment
|
||||
|
||||
::::{tab-item} A3/A2 series
|
||||
:sync: A3/A2 series
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
vllm serve Qwen/Qwen3-VL-Reranker-2B \
|
||||
--served-model-name Qwen/Qwen3-VL-Reranker-2B \
|
||||
--runner pooling \
|
||||
--hf_overrides '{"architectures": ["Qwen3VLForSequenceClassification"],"classifier_from_token": ["no", "yes"],"is_original_qwen3_reranker": true}' \
|
||||
--chat-template ./qwen3_vl_reranker.jinja \
|
||||
--port 8000 \
|
||||
--max-model-len 1024
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
:sync: Atlas 300I DUO
|
||||
|
||||
Start the docker image on each node.
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
vllm serve Qwen/Qwen3-VL-Reranker-2B \
|
||||
--served-model-name Qwen/Qwen3-VL-Reranker-2B \
|
||||
--runner pooling \
|
||||
--hf_overrides '{"architectures": ["Qwen3VLForSequenceClassification"],"classifier_from_token": ["no", "yes"],"is_original_qwen3_reranker": true}' \
|
||||
--chat-template ./qwen3_vl_reranker.jinja \
|
||||
--compilation-config '{"cudagraph_capture_sizes": [1024,512]}' \
|
||||
--additional-config '{"ascend_compilation_config": {"fuse_norm_quant": false}}' \
|
||||
--dtype float16 \
|
||||
--port 8000 \
|
||||
--max-model-len 1024
|
||||
```
|
||||
|
||||
Required Parameter Descriptions:
|
||||
|
||||
`--compilation-config` For Atlas 300I DUO, due to limited hardware streams, the size of cudagraph_capture_sizes is restricted.
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
Key Parameter Descriptions:
|
||||
|
||||
- `--max-model-len` represents the context length, which is the maximum value of the input plus output for a single request. For Atlas 300I DUO if automatic parsing resolves to a large context length, allocating this mask (O(max_model_len^2)) may exceed NPU memory and trigger OOM. Be sure to set an explicit and conservative value, such as --max-model-len 1024.
|
||||
|
||||
Common Issues Tip: If you encounter issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html) for troubleshooting.
|
||||
|
||||
## 6 Functional Verification
|
||||
|
||||
Once your server is started, you can verify by follow command:
|
||||
|
||||
Service Verification:
|
||||
|
||||
```bash
|
||||
curl http://localhost:8000/v1/rerank \
|
||||
-X POST \
|
||||
-d '{"query":"What is the capital of China?", "documents": ["The capital of China is Beijing.", "Gravity is a force that attracts two bodies towards each other. It gives weight to physical objects and is responsible for the movement of planets around the sun."]}' \
|
||||
-H 'Content-Type: application/json'
|
||||
```
|
||||
|
||||
Expected Result:
|
||||
|
||||
The service returns HTTP 200 OK with a JSON response containing the `relevance_score` field. Example output:
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "score-xxxxx",
|
||||
"model": "Qwen/Qwen3-VL-Reranker-2B",
|
||||
"usage": {
|
||||
"prompt_tokens": 179,
|
||||
"total_tokens": 179
|
||||
},
|
||||
"results": [
|
||||
{
|
||||
"index": 0,
|
||||
"document": {
|
||||
"text": "The capital of China is Beijing.",
|
||||
"multi_modal": null
|
||||
},
|
||||
"relevance_score": 0.7209711670875549
|
||||
},
|
||||
{
|
||||
"index": 1,
|
||||
"document": {
|
||||
"text": "Gravity is a force that attracts two bodies towards each other. It gives weight to physical objects and is responsible for the movement of planets around the sun.",
|
||||
"multi_modal": null
|
||||
},
|
||||
"relevance_score": 0.18871910870075226
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
For more usage examples, please reference the [examples](https://github.com/vllm-project/vllm/tree/main/examples/pooling/score)
|
||||
|
||||
## 7 Accuracy Evaluation
|
||||
|
||||
Here are two accuracy evaluation methods.
|
||||
|
||||
### Using MTEB
|
||||
|
||||
1. Refer to [MTEB](https://docs.mteb.org/) for details.
|
||||
|
||||
2. Run follow code to execute the accuracy evaluation.
|
||||
|
||||
```python
|
||||
|
||||
import os
|
||||
|
||||
from mteb.models.vllm_wrapper import VllmCrossEncoderWrapper
|
||||
|
||||
if __name__ == "__main__":
|
||||
import mteb
|
||||
|
||||
data_path = "/home/data/mteb_data"
|
||||
os.environ["HF_DATASETS_CACHE"] = data_path
|
||||
os.environ["HF_ENDPOINT"] = "https://hf-mirror.com"
|
||||
|
||||
model = VllmCrossEncoderWrapper(f"/home/data/Qwen3-VL-Reranker-2B",
|
||||
revision="norm",
|
||||
dtype="float16",
|
||||
enforce_eager=True,
|
||||
max_model_len=10240,
|
||||
hf_overrides={"architectures": ["Qwen3VLForSequenceClassification"],"classifier_from_token": ["no", "yes"],"is_original_qwen3_reranker": True})
|
||||
|
||||
cache = mteb.ResultCache("/home/data/mteb_data")
|
||||
tasks = mteb.get_tasks(
|
||||
task_types=["Reranking"],
|
||||
languages=["zho"]
|
||||
)
|
||||
tasks = mteb.get_tasks(tasks=["MultiLongDocReranking"])
|
||||
results = mteb.evaluate(model, tasks=tasks, cache=cache, overwrite_strategy="always")
|
||||
print(results)
|
||||
|
||||
```
|
||||
|
||||
3. After execution, you can get the result.
|
||||
|
||||
## 8 Performance Evaluation
|
||||
|
||||
### Using vLLM Benchmark
|
||||
|
||||
Run performance of `Qwen3-VL-Reranker-2B` as an example.
|
||||
Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/cli/) for more details.
|
||||
|
||||
Take the `serve` as an example. Run the code as follows.
|
||||
|
||||
```bash
|
||||
vllm bench serve --model Qwen/Qwen3-VL-Reranker-2B --backend vllm-rerank --port 8000 --dataset-name random-rerank --endpoint /v1/rerank --random-input 200 --save-result --result-dir ./
|
||||
```
|
||||
|
||||
After about several minutes, you can get the performance evaluation result.
|
||||
|
||||
## 9 FAQ
|
||||
|
||||
For common environment, installation, and general parameter issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html).
|
||||
947
docs/source/tutorials/models/Qwen3.5-27B-Qwen3.6-27B.md
Normal file
947
docs/source/tutorials/models/Qwen3.5-27B-Qwen3.6-27B.md
Normal file
@@ -0,0 +1,947 @@
|
||||
# Qwen3.5-27B/Qwen3.6-27B
|
||||
|
||||
## 1 Introduction
|
||||
|
||||
Qwen3.5-27B and Qwen3.6-27B are dense hybrid Mamba-Transformer language models in the Qwen3.5/Qwen3.6 family, integrating breakthroughs in architectural efficiency, reinforcement learning scale, and global accessibility. They share the same hybrid attention design (GDN + full attention), so deployment on Ascend NPUs follows the same pattern for both models. They are suitable for general-purpose text generation tasks such as dialogue, content creation, and code generation running on Ascend NPUs.
|
||||
|
||||
This document will demonstrate the main validation steps for the models, including supported features, feature configuration, environment preparation, single-node and multi-node deployment, as well as accuracy and performance evaluation.
|
||||
|
||||
It is **strongly recommended to use the latest release candidate (rc) version or the latest official version** of `vllm-ascend`. As a minimum-version requirement, `Qwen3.5-27B` is first supported in `vllm-ascend:v0.17.0rc1`, and `Qwen3.6-27B` is first supported in `vllm-ascend:v0.18.0rc1`. Support for Atlas 300I DUO and Ascend950DT series starts from `vllm-ascend:v0.23.0rc1`.
|
||||
|
||||
## 2 Supported Features
|
||||
|
||||
Please refer to the [Supported Features List](../../user_guide/support_matrix/supported_models.md) for the model support matrix.
|
||||
|
||||
Please refer to the [Feature Guide](../../user_guide/feature_guide/index.md) for feature configuration information.
|
||||
|
||||
## 3 Prerequisites
|
||||
|
||||
### 3.1 Model Weight
|
||||
|
||||
**Qwen3.5-27B**
|
||||
|
||||
- `Qwen3.5-27B` (BF16 version): requires 1 Atlas 800 A3 (64GB × 16) node or 1 Atlas 800 A2 (64GB × 8) node or Atlas 300I DUO. [Download model weight](https://www.modelscope.cn/models/Qwen/Qwen3.5-27B)
|
||||
- `Qwen3.5-27B-w8a8` (Quantized version): requires 1 Atlas 800 A3 (64GB × 16) node or 1 Atlas 800 A2 (64GB × 8) node or Atlas 300I DUO. [Download model weight](https://www.modelscope.cn/models/Eco-Tech/Qwen3.5-27B-w8a8-mtp)
|
||||
|
||||
**Qwen3.6-27B**
|
||||
|
||||
- `Qwen3.6-27B` (BF16 version): requires 1 Ascend 950DT(96GB × 8) node or 1 Atlas 800 A3 (64GB × 16) node or 1 Atlas 800 A2 (64GB × 8) node or Atlas 300I DUO. [Download model weight](https://www.modelscope.cn/models/Qwen/Qwen3.6-27B)
|
||||
- `Qwen3.6-27B-w8a8` (Quantized version): requires 1 Atlas 800 A3 (64GB × 16) node or 1 Atlas 800 A2 (64GB × 8) node or Atlas 300I DUO. [Download model weight](https://www.modelscope.cn/models/Eco-Tech/Qwen3.6-27B-w8a8)
|
||||
- `Qwen3.6-27B-w8a8-mxfp8` (Quantized version): requires 1 Ascend950DT series (96GB × 8) node. [Download model weight](https://www.modelscope.cn/models/Eco-Tech/Qwen3.6-27B-w8a8-mxfp8)
|
||||
|
||||
It is recommended to download the model weight to the shared directory of multiple nodes, such as `/root/.cache/`.
|
||||
|
||||
### 3.2 Verify Multi-node Communication
|
||||
|
||||
If you want to deploy multi-node environment, you need to verify multi-node communication according to [verify multi-node communication environment](../../installation.md#verify-multi-node-communication).
|
||||
|
||||
## 4 Installation
|
||||
|
||||
### 4.1 Docker Image Installation
|
||||
|
||||
Select an image based on your machine type and start the docker image on your node, refer to [using docker](../../installation.md#set-up-using-docker).
|
||||
|
||||
It is **recommended to use the latest release candidate (rc) version or the latest official version** of the `vllm-ascend` image to ensure the best compatibility and access to the latest features. As a minimum-version requirement, use `vllm-ascend:v0.17.0rc1` (or a later version) for `Qwen3.5-27B`, and `vllm-ascend:v0.18.0rc1` (or a later version) for `Qwen3.6-27B`. For `Qwen3.6-27B` on Atlas 800 A3, please use the matching `v0.18.0rc1-a3` (or a later `-a3`) image. For Atlas 300I DUO, use `vllm-ascend:nightly-releases-v0.23.0-310p` (or a later `-310p` image).
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: install
|
||||
|
||||
::::{tab-item} A3 series
|
||||
:sync: A3
|
||||
|
||||
Start the docker image on each node.
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-a3
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci8 \
|
||||
--device /dev/davinci9 \
|
||||
--device /dev/davinci10 \
|
||||
--device /dev/davinci11 \
|
||||
--device /dev/davinci12 \
|
||||
--device /dev/davinci13 \
|
||||
--device /dev/davinci14 \
|
||||
--device /dev/davinci15 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} A2 series
|
||||
:sync: A2
|
||||
|
||||
Start the docker image on each node.
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
:sync: atlas
|
||||
|
||||
Start the docker image on each node.
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-310p
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-p 8080:8080 \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Ascend950DT series
|
||||
:sync: 950dt
|
||||
|
||||
Start the docker image on each node.
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-a5
|
||||
export NAME=vllm-ascend
|
||||
|
||||
docker run --rm \
|
||||
--name $NAME \
|
||||
--net=host \
|
||||
--shm-size=1g \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/hisi_hdc \
|
||||
--device /dev/ummu \
|
||||
--device /dev/uburma \
|
||||
-v /usr/local/Ascend/driver:/usr/local/Ascend/driver \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /etc/hccl_rootinfo.json:/etc/hccl_rootinfo.json \
|
||||
-v /etc/hixlep/:/etc/hixlep/ \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-v /usr/local/sbin:/usr/local/sbin \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/sbin/npu-smi:/usr/local/sbin/npu-smi \
|
||||
-v /usr/lib64:/usr/lib64 \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
After a successful docker run, you can verify the running container service by executing the `docker ps` command. The expected result is that the container `vllm-ascend` is listed with status `Up`, confirming the docker installation is successful.
|
||||
|
||||
### 4.2 Source Code Installation
|
||||
|
||||
If you don't want to use the docker image as above, you can also build all from source:
|
||||
|
||||
1. Clone the repository and install `vllm-ascend` from source:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/vllm-project/vllm-ascend.git
|
||||
cd vllm-ascend
|
||||
pip install -e .
|
||||
```
|
||||
|
||||
For the complete installation steps, refer to [installation](../../installation.md).
|
||||
|
||||
````{note}
|
||||
On Atlas 300I DUO, you may need to uninstall `triton-ascend` and `triton` to avoid dependency conflicts:
|
||||
|
||||
```bash
|
||||
pip uninstall -y triton-ascend triton
|
||||
```
|
||||
````
|
||||
|
||||
2. If you want to deploy a multi-node environment, you need to set up the environment on each node.
|
||||
|
||||
To verify the source code installation, run the following command and confirm the displayed version matches the one you installed:
|
||||
|
||||
```bash
|
||||
pip show vllm-ascend
|
||||
```
|
||||
|
||||
Expected result: The version information of `vllm-ascend` is displayed, confirming a successful installation.
|
||||
|
||||
## 5 Online Service Deployment
|
||||
|
||||
### 5.1 Single-Node Online Deployment
|
||||
|
||||
Single-node deployment completes both Prefill and Decode within the same node, suitable for development, testing, and medium-scale inference scenarios. The `Qwen3.5-27B`, `Qwen3.5-27B-w8a8`, `Qwen3.6-27B`, and `Qwen3.6-27B-w8a8` models can all be deployed on 1 Atlas 800 A3 (64GB × 16), 1 Atlas 800 A2 (64GB × 8). On Atlas 300I DUO, at least 2 devices are required. The `Qwen3.6-27B`, and `Qwen3.6-27B-w8a8-mxfp8` models can all be deployed on 1 Ascend950DT series (96GB × 8). The quantized versions need to start with the `--quantization ascend` parameter.
|
||||
|
||||
Both `Qwen3.5-27B` and `Qwen3.6-27B` share the same MTP head design, so the `qwen3_5_mtp` speculative decoding method can be used for both.
|
||||
|
||||
:::::::{tab-set}
|
||||
|
||||
::::::{tab-item} Atlas 800 A3 / Atlas 800 A2
|
||||
|
||||
The following examples are for Atlas 800 A3 / Atlas 800 A2. Quantized versions need `--quantization ascend`.
|
||||
|
||||
:::::{tab-set}
|
||||
|
||||
::::{tab-item} Qwen3.5-27B-w8a8
|
||||
|
||||
Startup Command:
|
||||
|
||||
```bash
|
||||
#!/bin/sh
|
||||
# Load model from ModelScope to speed up download
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
# To reduce memory fragmentation and avoid out of memory
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export HCCL_BUFFSIZE=512
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
|
||||
vllm serve Eco-Tech/Qwen3.5-27B-w8a8-mtp \
|
||||
--host 0.0.0.0 \
|
||||
--port 8000 \
|
||||
--data-parallel-size 1 \
|
||||
--tensor-parallel-size 2 \
|
||||
--seed 1024 \
|
||||
--quantization ascend \
|
||||
--served-model-name qwen3.5 \
|
||||
--max-num-seqs 32 \
|
||||
--max-model-len 133000 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--trust-remote-code \
|
||||
--gpu-memory-utilization 0.90 \
|
||||
--no-enable-prefix-caching \
|
||||
--speculative-config '{"method": "qwen3_5_mtp", "num_speculative_tokens": 3, "enforce_eager": true}' \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"enable_cpu_binding":true}' \
|
||||
--async-scheduling
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Qwen3.6-27B-w8a8
|
||||
|
||||
Startup Command:
|
||||
|
||||
```bash
|
||||
#!/bin/sh
|
||||
# Load model from ModelScope to speed up download
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
# To reduce memory fragmentation and avoid out of memory
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export HCCL_BUFFSIZE=512
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
|
||||
vllm serve Eco-Tech/Qwen3.6-27B-w8a8 \
|
||||
--host 0.0.0.0 \
|
||||
--port 8000 \
|
||||
--data-parallel-size 1 \
|
||||
--tensor-parallel-size 2 \
|
||||
--seed 1024 \
|
||||
--quantization ascend \
|
||||
--served-model-name qwen3.6 \
|
||||
--max-num-seqs 32 \
|
||||
--max-model-len 262144 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--trust-remote-code \
|
||||
--gpu-memory-utilization 0.90 \
|
||||
--no-enable-prefix-caching \
|
||||
--speculative-config '{"method": "qwen3_5_mtp", "num_speculative_tokens": 3, "enforce_eager": true}' \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"enable_cpu_binding":true}' \
|
||||
--async-scheduling
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
:::::
|
||||
|
||||
Key Parameter Descriptions:
|
||||
|
||||
- `--data-parallel-size 1` and `--tensor-parallel-size 2` are common settings for data parallelism (DP) and tensor parallelism (TP) sizes.
|
||||
- `--max-model-len` represents the context length, which is the maximum value of the input plus output for a single request. The Qwen3.6-27B model supports up to 262144.
|
||||
- `--max-num-seqs` indicates the maximum number of requests that each DP group is allowed to process. If the number of requests sent to the service exceeds this limit, the excess requests will remain in a waiting state and will not be scheduled. Note that the time spent in the waiting state is also counted in metrics such as TTFT and TPOT. Therefore, when testing performance, it is generally recommended that `--max-num-seqs` * `--data-parallel-size` >= the actual total concurrency.
|
||||
- `--max-num-batched-tokens` represents the maximum number of tokens that the model can process in a single step. Currently, vLLM v1 scheduling enables ChunkPrefill/SplitFuse by default, which means:
|
||||
- (1) If the input length of a request is greater than `--max-num-batched-tokens`, it will be divided into multiple rounds of computation according to `--max-num-batched-tokens`;
|
||||
- (2) Decode requests are prioritized for scheduling, and prefill requests are scheduled only if there is available capacity.
|
||||
- Generally, if `--max-num-batched-tokens` is set to a larger value, the overall latency will be lower, but the pressure on HBM memory (activation value usage) will be greater.
|
||||
- `--gpu-memory-utilization` represents the proportion of HBM that vLLM will use for actual inference. Its essential function is to calculate the available kv_cache size. During the warm-up phase (referred to as profile run in vLLM), vLLM records the peak HBM memory usage during an inference process with an input size of `--max-num-batched-tokens`. The available kv_cache size is then calculated as: `--gpu-memory-utilization` * HBM size - peak HBM memory usage. Therefore, the larger the value of `--gpu-memory-utilization`, the more kv_cache can be used. However, since the HBM memory usage during the warm-up phase may differ from that during actual inference (e.g., due to uneven EP load), setting `--gpu-memory-utilization` too high may lead to OOM (Out of Memory) issues during actual inference. The default value is `0.9`.
|
||||
- `--no-enable-prefix-caching` indicates that prefix caching is disabled. The current implementation of hybrid kv cache for Qwen3.5-27B / Qwen3.6-27B may result in a very large effective `block_size` when prefix caching is enabled (e.g., 2048), which means any prefix shorter than `block_size` will never be cached. If your workload has many short repeated prefixes, consider keeping prefix caching disabled. For related issues, see the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html).
|
||||
- `--quantization ascend` indicates that quantization is used. To disable quantization, remove this option.
|
||||
- `--speculative-config` uses `qwen3_5_mtp` for both `Qwen3.5-27B` and `Qwen3.6-27B` because they share the same MTP head design.
|
||||
- `--compilation-config` contains configurations related to the aclgraph graph mode. The most significant configurations are `"cudagraph_mode"` and `"cudagraph_capture_sizes"`, which have the following meanings:
|
||||
- `"cudagraph_mode"`: represents the specific graph mode. Currently, `"PIECEWISE"` and `"FULL_DECODE_ONLY"` are supported. The graph mode is mainly used to reduce the cost of operator dispatch. Currently, `"FULL_DECODE_ONLY"` is recommended.
|
||||
- `"cudagraph_capture_sizes"`: represents different levels of graph modes. The default value is `[1, 2, 4, 8, 16, 24, 32, 40,..., --max-num-seqs]`. In the graph mode, the input for graphs at different levels is fixed, and inputs between levels are automatically padded to the next level. Currently, the default setting is recommended. Only in some scenarios is it necessary to set this separately to achieve optimal performance.
|
||||
|
||||
::::::
|
||||
|
||||
::::::{tab-item} Atlas 300I DUO
|
||||
|
||||
Currently only the **TP** scenario is supported. Choose **TP=2** or **TP=4** according to the available devices. Replace `MODEL_PATH` with a ModelScope model id or a local directory path. The quantized versions need to start with the `--quantization ascend` parameter.
|
||||
|
||||
:::::{tab-set}
|
||||
|
||||
::::{tab-item} Qwen3.5-27B-w8a8
|
||||
|
||||
Startup Command:
|
||||
|
||||
```bash
|
||||
#!/bin/sh
|
||||
# Load model from ModelScope to speed up download
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
|
||||
# Model weight path; can be a ModelScope model id (e.g., Eco-Tech/Qwen3.5-27B-w8a8-mtp) or a local directory path
|
||||
export MODEL_PATH=Eco-Tech/Qwen3.5-27B-w8a8-mtp
|
||||
|
||||
vllm serve $MODEL_PATH \
|
||||
--host 127.0.0.1 \
|
||||
--port 1025 \
|
||||
--tensor-parallel-size 4 \
|
||||
--served-model-name qwen3.5 \
|
||||
--max-num-seqs 128 \
|
||||
--max-model-len 16384 \
|
||||
--trust-remote-code \
|
||||
--gpu-memory-utilization 0.90 \
|
||||
--mamba-ssm-cache-dtype float16 \
|
||||
--dtype float16 \
|
||||
--speculative-config '{"method": "qwen3_5_mtp","num_speculative_tokens":1}' \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY", "cudagraph_capture_sizes": [2,8]}' \
|
||||
--additional-config '{"ascend_compilation_config": {"enable_npugraph_ex": false}}'
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Qwen3.6-27B-w8a8
|
||||
|
||||
Startup Command:
|
||||
|
||||
```bash
|
||||
#!/bin/sh
|
||||
# Load model from ModelScope to speed up download
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
|
||||
# Model weight path; can be a ModelScope model id (e.g., Eco-Tech/Qwen3.6-27B-w8a8) or a local directory path
|
||||
export MODEL_PATH=Eco-Tech/Qwen3.6-27B-w8a8
|
||||
|
||||
vllm serve $MODEL_PATH \
|
||||
--host 127.0.0.1 \
|
||||
--port 1025 \
|
||||
--tensor-parallel-size 4 \
|
||||
--served-model-name qwen3.6 \
|
||||
--max-num-seqs 128 \
|
||||
--max-model-len 16384 \
|
||||
--trust-remote-code \
|
||||
--gpu-memory-utilization 0.90 \
|
||||
--mamba-ssm-cache-dtype float16 \
|
||||
--dtype float16 \
|
||||
--speculative-config '{"method": "qwen3_5_mtp","num_speculative_tokens":1}' \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY", "cudagraph_capture_sizes": [2,8]}' \
|
||||
--additional-config '{"ascend_compilation_config": {"enable_npugraph_ex": false}}'
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
:::::
|
||||
|
||||
Key Parameter Descriptions:
|
||||
|
||||
- `--tensor-parallel-size` sets the tensor parallel size. Choose **TP=2** or **TP=4** according to the available devices.
|
||||
- `--max-model-len` represents the context length, which is the maximum value of the input plus output for a single request. Configure it based on the actual workload and available memory; with **TP=4**, the Qwen3.6-27B model supports up to 262144.
|
||||
- `--max-num-seqs` indicates the maximum number of concurrent requests. Configure it as needed—setting it too high may cause OOM.
|
||||
- `--gpu-memory-utilization` represents the proportion of HBM that vLLM will use for actual inference. Configure this value according to the actual device memory; setting it too high may cause OOM. The default value is `0.9`.
|
||||
- `--mamba-ssm-cache-dtype` sets the data type of the Mamba SSM cache. On Atlas 300I DUO, only `float16` is supported.
|
||||
- `--dtype float16` must be set on Atlas 300I DUO. These devices only support the FP16 data type.
|
||||
- `--speculative-config` uses `qwen3_5_mtp` for both `Qwen3.5-27B` and `Qwen3.6-27B` because they share the same MTP head design. On Atlas 300I DUO, it is recommended to set `num_speculative_tokens` to `1`.
|
||||
- `--compilation-config` contains configurations related to the aclgraph graph mode. The most significant configurations are `"cudagraph_mode"` and `"cudagraph_capture_sizes"`, which have the following meanings:
|
||||
- `"cudagraph_mode"`: represents the specific graph mode. Currently, `"PIECEWISE"` and `"FULL_DECODE_ONLY"` are supported. The graph mode is mainly used to reduce the cost of operator dispatch. Currently, `"FULL_DECODE_ONLY"` is recommended.
|
||||
- `"cudagraph_capture_sizes"`: represents different levels of graph modes. When tensor parallelism (TP) is enabled, hardware event-id constraints allow at most two capture sizes (for example, `[1, 8]`). With MTP enabled, calculate each capture size as `n * (num_speculative_tokens + 1)`, where `n` is a capture size for the deployment without MTP. For example, when `num_speculative_tokens` is `1`, the non-MTP sizes `[1,4]` become `[2,8]`.
|
||||
- `--additional-config` with `"ascend_compilation_config": {"enable_npugraph_ex": false}` is required on Atlas 300I DUO because `enable_npugraph_ex` is not supported on this platform.
|
||||
|
||||
::::::
|
||||
|
||||
::::::{tab-item} Ascend950DT series
|
||||
|
||||
The following example is for Ascend950DT series. Quantized versions need `--quantization ascend`.
|
||||
|
||||
:::::{tab-set}
|
||||
|
||||
::::{tab-item} Qwen3.6-27B-w8a8-mxfp8
|
||||
|
||||
Startup Command:
|
||||
|
||||
```bash
|
||||
#!/bin/sh
|
||||
# Load model from ModelScope to speed up download
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
# To reduce memory fragmentation and avoid out of memory
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
# Size of the shared buffer (in MB) used by HCCL for NPU-to-NPU collective communication
|
||||
export HCCL_BUFFSIZE=512
|
||||
# Whether OpenMP threads are bound to specific CPU cores
|
||||
export OMP_PROC_BIND=false
|
||||
# Number of OpenMP threads available for parallel regions
|
||||
export OMP_NUM_THREADS=1
|
||||
# Enables the Ascend task queue for asynchronous operator dispatch
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
|
||||
# Model weight path; can be a ModelScope model id (e.g., Eco-Tech/Qwen3.6-27B-w8a8-mxfp8) or a local directory path
|
||||
export MODEL_PATH=Eco-Tech/Qwen3.6-27B-w8a8-mxfp8
|
||||
|
||||
vllm serve $MODEL_PATH \
|
||||
--host 0.0.0.0 \
|
||||
--port 8000 \
|
||||
--data-parallel-size 1 \
|
||||
--tensor-parallel-size 1 \
|
||||
--seed 1024 \
|
||||
--quantization ascend \
|
||||
--served-model-name qwen3.6 \
|
||||
--max-num-seqs 32 \
|
||||
--max-model-len 262144 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--trust-remote-code \
|
||||
--gpu-memory-utilization 0.90 \
|
||||
--no-enable-prefix-caching \
|
||||
--speculative-config '{"method": "qwen3_5_mtp", "num_speculative_tokens": 3, "enforce_eager": true}' \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"enable_cpu_binding":true}' \
|
||||
--async-scheduling
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
:::::
|
||||
|
||||
Key Parameter Descriptions:
|
||||
|
||||
- `--data-parallel-size 1` and `--tensor-parallel-size 1` are common settings for data parallelism (DP) and tensor parallelism (TP) sizes.
|
||||
- `--max-model-len` represents the context length, which is the maximum value of the input plus output for a single request. The Qwen3.6-27B model supports up to 262144.
|
||||
- `--max-num-seqs` indicates the maximum number of requests that each DP group is allowed to process. If the number of requests sent to the service exceeds this limit, the excess requests will remain in a waiting state and will not be scheduled. Note that the time spent in the waiting state is also counted in metrics such as TTFT and TPOT. Therefore, when testing performance, it is generally recommended that `--max-num-seqs` * `--data-parallel-size` >= the actual total concurrency.
|
||||
- `--max-num-batched-tokens` represents the maximum number of tokens that the model can process in a single step. Currently, vLLM v1 scheduling enables ChunkPrefill/SplitFuse by default, which means:
|
||||
- (1) If the input length of a request is greater than `--max-num-batched-tokens`, it will be divided into multiple rounds of computation according to `--max-num-batched-tokens`;
|
||||
- (2) Decode requests are prioritized for scheduling, and prefill requests are scheduled only if there is available capacity.
|
||||
- Generally, if `--max-num-batched-tokens` is set to a larger value, the overall latency will be lower, but the pressure on HBM memory (activation value usage) will be greater.
|
||||
- `--gpu-memory-utilization` represents the proportion of HBM that vLLM will use for actual inference. Its essential function is to calculate the available kv_cache size. During the warm-up phase (referred to as profile run in vLLM), vLLM records the peak HBM memory usage during an inference process with an input size of `--max-num-batched-tokens`. The available kv_cache size is then calculated as: `--gpu-memory-utilization` * HBM size - peak HBM memory usage. Therefore, the larger the value of `--gpu-memory-utilization`, the more kv_cache can be used. However, since the HBM memory usage during the warm-up phase may differ from that during actual inference (e.g., due to uneven EP load), setting `--gpu-memory-utilization` too high may lead to OOM (Out of Memory) issues during actual inference. The default value is `0.9`.
|
||||
- `--no-enable-prefix-caching` indicates that prefix caching is disabled. The current implementation of hybrid kv cache for Qwen3.6-27B may result in a very large effective `block_size` when prefix caching is enabled (e.g., 2048), which means any prefix shorter than `block_size` will never be cached. If your workload has many short repeated prefixes, consider keeping prefix caching disabled. For related issues, see the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html).
|
||||
- `--quantization ascend` indicates that quantization is used. To disable quantization, remove this option.
|
||||
- `--speculative-config` uses `qwen3_5_mtp` for `Qwen3.6-27B` because it shares the same MTP head design as `Qwen3.5-27B`.
|
||||
- `--compilation-config` contains configurations related to the aclgraph graph mode. The most significant configurations are `"cudagraph_mode"` and `"cudagraph_capture_sizes"`, which have the following meanings:
|
||||
- `"cudagraph_mode"`: represents the specific graph mode. Currently, `"PIECEWISE"` and `"FULL_DECODE_ONLY"` are supported. The graph mode is mainly used to reduce the cost of operator dispatch. Currently, `"FULL_DECODE_ONLY"` is recommended.
|
||||
- `"cudagraph_capture_sizes"`: represents different levels of graph modes. The default value is `[1, 2, 4, 8, 16, 24, 32, 40,..., --max-num-seqs]`. In the graph mode, the input for graphs at different levels is fixed, and inputs between levels are automatically padded to the next level. Currently, the default setting is recommended. Only in some scenarios is it necessary to set this separately to achieve optimal performance.
|
||||
|
||||
::::::
|
||||
|
||||
:::::::
|
||||
|
||||
Common Issues Tip: If you encounter issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html) for troubleshooting.
|
||||
|
||||
Service Verification:
|
||||
|
||||
If the service starts successfully, the following startup log will be displayed:
|
||||
|
||||
```text
|
||||
(APIServer pid=<pid>) INFO: Started server process [<pid>]
|
||||
(APIServer pid=<pid>) INFO: Waiting for application startup.
|
||||
(APIServer pid=<pid>) INFO: Application startup complete.
|
||||
```
|
||||
|
||||
For functional testing (e.g., `completions` and `chat.completions` curl examples with expected responses), please refer to [Section 6](#6-functional-verification).
|
||||
|
||||
### 5.2 Multi-Node PD Separation Deployment
|
||||
|
||||
```{note}
|
||||
Multi-node PD separation deployment is **not supported** on Atlas 300I DUO.
|
||||
```
|
||||
|
||||
For high-concurrency production scenarios, multi-node PD (Prefill-Decode) separation can be used to scale the service. The recommended approach is to use Mooncake for deployment: [Mooncake Multi-Node PD Disaggregation Guide](../features/pd_disaggregation_mooncake_multi_node.md).
|
||||
|
||||
In the standard single-node deployment mode, Prefill (prompt processing) and Decode (token generation) tasks run on the same set of NPUs. This can lead to two issues:
|
||||
|
||||
1. **Prefill preemption interrupts Decode**: Prefill is a compute-intensive task that processes the entire input context at once, while Decode generates tokens one by one. When a new user request arrives, its Prefill phase can preempt and interrupt ongoing Decode tasks, causing jitter and higher time-per-output-token (TPOT) latency.
|
||||
2. **Inflexible resource allocation**: Prefill and Decode have fundamentally different computational characteristics — Prefill is compute-bound and memory-bandwidth-intensive, while Decode is memory-bandwidth-bound. Running them on the same hardware forces a compromise that satisfies neither optimally.
|
||||
|
||||
PD (Prefill-Decode) separation addresses these issues by running Prefill and Decode on dedicated node groups, each configured independently:
|
||||
|
||||
- **Prefill nodes** focus on high-throughput prompt processing, optimized for compute and communication (e.g., enabling FlashComm for Allreduce acceleration).
|
||||
- **Decode nodes** focus on low-latency token generation, optimized for memory bandwidth (e.g., enabling async-scheduling and full-decode aclgraph).
|
||||
|
||||
For `Qwen3.5-27B-w8a8` and `Qwen3.6-27B-w8a8`, a typical **1P1D** configuration requires **2 Atlas 800 A3 (64GB × 16) nodes** (1 Prefill node + 1 Decode node), with **TP=2** and **DP=8** on each node, which fully utilizes all 16 NPUs of an Atlas A3. The example below uses `Qwen3.5-27B-w8a8`; for `Qwen3.6-27B-w8a8`, replace the model path with `Eco-Tech/Qwen3.6-27B-w8a8` and adjust `--served-model-name` to `qwen3.6` (and `--max-model-len` to 262144 if needed).
|
||||
|
||||
> **Why TP=2 + DP=8 (DP-first strategy)?** The `Qwen3.5-27B-w8a8` (and `Qwen3.6-27B-w8a8`) model is only ~30 GB, which easily fits in a single NPU (each NPU has 64 GB HBM). **TP > 1 is mainly needed for models that do not fit in one NPU.** For a 27 B model, `TP=2` is sufficient to balance operator-dispatch overhead across NPUs, while **maximizing DP** keeps all 16 NPUs of an Atlas A3 busy with independent request batches, fully utilizing the hardware. This **DP-first parallelism strategy** is the standard practice for small dense models (e.g., Qwen3.5-27B, Qwen3.6-27B, Llama-3-8B) and has been validated by the [Qwen3.5-27B B200 benchmark](https://thenextgentechinsider.com/pulse/qwen-35-27b-delivers-11m-tokenssecond-on-nvidia-b200-cluster), where switching from TP=8 to DP=8 lifted per-node throughput from 9.5k to 95k tokens/s.
|
||||
>
|
||||
> **Note**: Since `Qwen3.5-27B` and `Qwen3.6-27B` fit in a single node, multi-node PD separation is only recommended for high-concurrency production deployments. For the Mooncake deployment specifics, please refer to the [Mooncake Multi-Node PD Disaggregation Guide](../features/pd_disaggregation_mooncake_multi_node.md).
|
||||
|
||||
To run the vllm-ascend Prefill-Decode Disaggregation service, you need to:
|
||||
|
||||
- Deploy a `launch_online_dp.py` script and a `run_dp_template.sh` script on each node;
|
||||
- Deploy a `load_balance_proxy_server_example.py` script on the prefill master node to forward requests.
|
||||
|
||||
1. `launch_online_dp.py` is used to launch external dp vllm servers.
|
||||
[launch_online_dp.py](https://github.com/vllm-project/vllm-ascend/blob/main/examples/external_online_dp/launch_online_dp.py)
|
||||
|
||||
Parameter descriptions:
|
||||
|
||||
|Parameter|Type|Required|Default|Description|
|
||||
|---------|----|--------|-------|-----------|
|
||||
|`--dp-size`|int|Yes|-|Data parallel size (total number of DP ranks across all nodes).|
|
||||
|`--tp-size`|int|No|1|Tensor parallel size within each DP rank.|
|
||||
|`--dp-size-local`|int|No|(same as `--dp-size`)|Number of DP ranks on the current node. If not set, defaults to `--dp-size`.|
|
||||
|`--dp-rank-start`|int|No|0|Starting rank offset for data parallel ranks on this node.|
|
||||
|`--dp-address`|str|Yes|-|IP address of the data parallel master node (node 0).|
|
||||
|`--dp-rpc-port`|str|No|12345|RPC port for data parallel master communication.|
|
||||
|`--vllm-start-port`|int|No|9000|Starting port for each vLLM engine instance on this node. Each DP rank's engine port = `vllm_start_port` + local rank index.|
|
||||
|
||||
2. Prefill Node 0 `run_dp_template.sh` script. You can get the template in the repository's examples: [run_dp_template.sh](https://github.com/vllm-project/vllm-ascend/blob/main/examples/external_online_dp/run_dp_template.sh).
|
||||
|
||||
```shell
|
||||
# nic_name is the network interface name corresponding to local_ip of the current node
|
||||
nic_name="xxx"
|
||||
local_ip="141.xx.xx.1"
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
|
||||
# [Optional] jemalloc
|
||||
# jemalloc is for better performance, if `libjemalloc.so` is installed on your machine, you can turn it on.
|
||||
# export LD_PRELOAD=/usr/lib/aarch64-linux-gnu/libjemalloc.so.2:$LD_PRELOAD
|
||||
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export LD_LIBRARY_PATH=/usr/local/Ascend/ascend-toolkit/latest/python/site-packages/mooncake:$LD_LIBRARY_PATH
|
||||
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export VLLM_ASCEND_ENABLE_FLASHCOMM1=1
|
||||
export ASCEND_RT_VISIBLE_DEVICES=$1
|
||||
|
||||
vllm serve Eco-Tech/Qwen3.5-27B-w8a8-mtp \
|
||||
--host 0.0.0.0 \
|
||||
--port $2 \
|
||||
--data-parallel-size $3 \
|
||||
--data-parallel-rank $4 \
|
||||
--data-parallel-address $5 \
|
||||
--data-parallel-rpc-port $6 \
|
||||
--tensor-parallel-size $7 \
|
||||
--seed 1024 \
|
||||
--quantization ascend \
|
||||
--served-model-name qwen3.5 \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 4 \
|
||||
--max-model-len 32768 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--no-enable-prefix-caching \
|
||||
--gpu-memory-utilization 0.95 \
|
||||
--enforce-eager \
|
||||
--speculative-config '{"method": "qwen3_5_mtp", "num_speculative_tokens": 3, "enforce_eager": true}' \
|
||||
--additional-config '{"enable_cpu_binding":true}' \
|
||||
--kv-transfer-config \
|
||||
'{"kv_connector": "MooncakeConnectorV1",
|
||||
"kv_role": "kv_producer",
|
||||
"kv_port": "30000",
|
||||
"engine_id": "0",
|
||||
"kv_connector_extra_config": {
|
||||
"prefill": {
|
||||
"dp_size": 8,
|
||||
"tp_size": 2
|
||||
},
|
||||
"decode": {
|
||||
"dp_size": 8,
|
||||
"tp_size": 2
|
||||
}
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
3. Decode Node 0 `run_dp_template.sh` script. You can get the template in the repository's examples: [run_dp_template.sh](https://github.com/vllm-project/vllm-ascend/blob/main/examples/external_online_dp/run_dp_template.sh).
|
||||
|
||||
```shell
|
||||
# nic_name is the network interface name corresponding to local_ip of the current node
|
||||
nic_name="xxx"
|
||||
local_ip="141.xx.xx.2"
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
|
||||
# [Optional] jemalloc
|
||||
# jemalloc is for better performance, if `libjemalloc.so` is installed on your machine, you can turn it on.
|
||||
# export LD_PRELOAD=/usr/lib/aarch64-linux-gnu/libjemalloc.so.2:$LD_PRELOAD
|
||||
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=1
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export LD_LIBRARY_PATH=/usr/local/Ascend/ascend-toolkit/latest/python/site-packages/mooncake:$LD_LIBRARY_PATH
|
||||
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export ASCEND_RT_VISIBLE_DEVICES=$1
|
||||
|
||||
vllm serve Eco-Tech/Qwen3.5-27B-w8a8-mtp \
|
||||
--host 0.0.0.0 \
|
||||
--port $2 \
|
||||
--data-parallel-size $3 \
|
||||
--data-parallel-rank $4 \
|
||||
--data-parallel-address $5 \
|
||||
--data-parallel-rpc-port $6 \
|
||||
--tensor-parallel-size $7 \
|
||||
--seed 1024 \
|
||||
--quantization ascend \
|
||||
--served-model-name qwen3.5 \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 16 \
|
||||
--max-model-len 32768 \
|
||||
--max-num-batched-tokens 2048 \
|
||||
--no-enable-prefix-caching \
|
||||
--gpu-memory-utilization 0.91 \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"recompute_scheduler_enable":true,"enable_cpu_binding":true}' \
|
||||
--async-scheduling \
|
||||
--speculative-config '{"method": "qwen3_5_mtp", "num_speculative_tokens": 3, "enforce_eager": true}' \
|
||||
--kv-transfer-config \
|
||||
'{"kv_connector": "MooncakeConnectorV1",
|
||||
"kv_role": "kv_consumer",
|
||||
"kv_port": "30200",
|
||||
"engine_id": "1",
|
||||
"kv_connector_extra_config": {
|
||||
"prefill": {
|
||||
"dp_size": 8,
|
||||
"tp_size": 2
|
||||
},
|
||||
"decode": {
|
||||
"dp_size": 8,
|
||||
"tp_size": 2
|
||||
}
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
Key Parameter Descriptions:
|
||||
|
||||
- `VLLM_ASCEND_ENABLE_FLASHCOMM1=1`: enables the Allreduce communication optimization on prefill nodes, which reduces the communication overhead of long-context prefill.
|
||||
- `recompute_scheduler_enable: true`: enables the recomputation scheduler. When the KV Cache of the decode node is insufficient, requests will be sent to the prefill node to recompute the KV Cache. In the PD separation scenario, enable this configuration only on decode nodes.
|
||||
- `--async-scheduling` (on decode nodes): enables asynchronous scheduling, which can reduce TPOT for high-concurrency decode workloads.
|
||||
- `--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}'` (on decode nodes): enables the full-decode aclgraph mode, which significantly reduces scheduling latency on the decode side.
|
||||
|
||||
4. Run server for each node:
|
||||
|
||||
```shell
|
||||
# p0 (Prefill node 0)
|
||||
python launch_online_dp.py --dp-size 8 --tp-size 2 --dp-size-local 8 --dp-rank-start 0 --dp-address 141.xx.xx.1 --dp-rpc-port 12321 --vllm-start-port 7100
|
||||
# d0 (Decode node 0)
|
||||
python launch_online_dp.py --dp-size 8 --tp-size 2 --dp-size-local 8 --dp-rank-start 0 --dp-address 141.xx.xx.2 --dp-rpc-port 12321 --vllm-start-port 7100
|
||||
```
|
||||
|
||||
5. Run the proxy server on the prefill master node.
|
||||
|
||||
You can get the proxy program in the repository's examples: [load_balance_proxy_server_example.py](https://github.com/vllm-project/vllm-ascend/blob/main/examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py).
|
||||
|
||||
Note: Since each node has 8 DP ranks (with `--vllm-start-port 7100` + local rank index, occupying ports 7100-7107), you need to list all 8 ports for each node in the proxy command:
|
||||
|
||||
```shell
|
||||
python load_balance_proxy_server_example.py \
|
||||
--port 1999 \
|
||||
--host 141.xx.xx.1 \
|
||||
--prefiller-hosts \
|
||||
141.xx.xx.1 \
|
||||
141.xx.xx.1 \
|
||||
141.xx.xx.1 \
|
||||
141.xx.xx.1 \
|
||||
141.xx.xx.1 \
|
||||
141.xx.xx.1 \
|
||||
141.xx.xx.1 \
|
||||
141.xx.xx.1 \
|
||||
--prefiller-ports \
|
||||
7100 7101 7102 7103 7104 7105 7106 7107 \
|
||||
--decoder-hosts \
|
||||
141.xx.xx.2 \
|
||||
141.xx.xx.2 \
|
||||
141.xx.xx.2 \
|
||||
141.xx.xx.2 \
|
||||
141.xx.xx.2 \
|
||||
141.xx.xx.2 \
|
||||
141.xx.xx.2 \
|
||||
141.xx.xx.2 \
|
||||
--decoder-ports \
|
||||
7100 7101 7102 7103 7104 7105 7106 7107 \
|
||||
```
|
||||
|
||||
Deployment Verification:
|
||||
|
||||
After the PD separation service is fully started, send a request through the proxy port on the prefill master node to verify that Prefill and Decode nodes are working correctly together:
|
||||
|
||||
```bash
|
||||
curl http://<proxy_node0_ip>:1999/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "qwen3.5",
|
||||
"messages": [
|
||||
{"role": "user", "content": "The future of AI is"}
|
||||
],
|
||||
"max_tokens": 1024,
|
||||
"temperature": 1.0,
|
||||
"top_p": 0.95
|
||||
}'
|
||||
```
|
||||
|
||||
> **Note**: For `Qwen3.6-27B-w8a8`, change the `model` field above to `"qwen3.6"` and the `--served-model-name` of the Prefill/Decode nodes to `qwen3.6`.
|
||||
|
||||
Expected Result: The proxy returns HTTP 200 OK. The JSON response contains the `choices` field with the generated text, confirming that Prefill nodes have successfully processed the prompt and Decode nodes have generated the response.
|
||||
|
||||
Common Issues Tip: If you encounter issues with PD separation deployment, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html) for troubleshooting.
|
||||
|
||||
## 6 Functional Verification
|
||||
|
||||
After the service is started, the model can be invoked by sending a prompt. Two API interfaces are supported: `completions` and `chat.completions`. Use the `--served-model-name` you configured (`qwen3.5` for `Qwen3.5-27B` or `qwen3.6` for `Qwen3.6-27B`).
|
||||
|
||||
**Completions API:**
|
||||
|
||||
```bash
|
||||
curl http://localhost:8000/v1/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "qwen3.5",
|
||||
"prompt": "The future of AI is",
|
||||
"max_tokens": 50,
|
||||
"temperature": 0
|
||||
}'
|
||||
```
|
||||
|
||||
> **Note**: For `Qwen3.6-27B`, set `"model": "qwen3.6"` in the request body.
|
||||
|
||||
**Chat Completions API:**
|
||||
|
||||
```bash
|
||||
curl http://localhost:8000/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "qwen3.5",
|
||||
"messages": [
|
||||
{"role": "user", "content": "The future of AI is"}
|
||||
],
|
||||
"max_completion_tokens": 1024,
|
||||
"temperature": 0.7,
|
||||
"top_p": 0.95
|
||||
}'
|
||||
```
|
||||
|
||||
> **Note**: For `Qwen3.6-27B`, set `"model": "qwen3.6"` in the request body.
|
||||
|
||||
Expected Result: The service returns HTTP 200 OK. The JSON response contains the `choices` field with generated text. Example output for the completions API (content truncated for brevity):
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "cmpl-xxxxxxxxxxxxx",
|
||||
"object": "text_completion",
|
||||
"created": 1780971952,
|
||||
"model": "qwen3.5",
|
||||
"choices": [
|
||||
{
|
||||
"index": 0,
|
||||
"text": "The future of AI is a rapidly evolving landscape with breakthroughs in natural language understanding, multimodal reasoning, and autonomous agents. As models grow more capable and efficient...",
|
||||
"logprobs": null,
|
||||
"finish_reason": "length"
|
||||
}
|
||||
],
|
||||
"usage": {
|
||||
"prompt_tokens": 4,
|
||||
"total_tokens": 54,
|
||||
"completion_tokens": 50
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## 7 Accuracy Evaluation
|
||||
|
||||
Here are two accuracy evaluation methods.
|
||||
|
||||
### Using AISBench
|
||||
|
||||
1. Refer to [Using AISBench](../../developer_guide/evaluation/using_ais_bench.md) for details.
|
||||
|
||||
2. After execution, you can get the result. Here is the result of `Qwen3.5-27B-w8a8` in `vllm-ascend:v0.17.0rc1` for reference only. The accuracy result of `Qwen3.6-27B-w8a8` can be obtained in the same way and is not listed here.
|
||||
|
||||
| dataset | version | metric | mode | vllm-api-general-chat |
|
||||
|----- | ----- | ----- | ----- | -----|
|
||||
| gsm8k | - | accuracy | gen | 96.74 |
|
||||
|
||||
### Using Language Model Evaluation Harness
|
||||
|
||||
Using the `gsm8k` dataset as an example test dataset, run the accuracy evaluation for `Qwen3.5-27B-w8a8` in online mode.
|
||||
|
||||
1. For `lm_eval` installation, please refer to [Using lm_eval](../../developer_guide/evaluation/using_lm_eval.md).
|
||||
2. Run `lm_eval` to execute the accuracy evaluation.
|
||||
|
||||
```shell
|
||||
# For Qwen3.5-27B-w8a8
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
vllm serve Eco-Tech/Qwen3.5-27B-w8a8-mtp \
|
||||
--served-model-name qwen3.5 \
|
||||
--trust-remote-code \
|
||||
--quantization ascend \
|
||||
--tensor-parallel-size 2 \
|
||||
--max-model-len 133000 \
|
||||
--max-num-seqs 32 \
|
||||
--gpu-memory-utilization 0.90 \
|
||||
--no-enable-prefix-caching
|
||||
|
||||
# Run lm_eval in another terminal
|
||||
lm_eval \
|
||||
--model local-completions \
|
||||
--model_args model=qwen3.5,base_url=http://127.0.0.1:8000/v1/completions,tokenized_requests=False,trust_remote_code=True \
|
||||
--tasks gsm8k \
|
||||
--output_path ./
|
||||
```
|
||||
|
||||
## 8 Performance Evaluation
|
||||
|
||||
### Using AISBench
|
||||
|
||||
Refer to [Using AISBench for performance evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation) for details.
|
||||
|
||||
### Using vLLM Benchmark
|
||||
|
||||
Run performance evaluation of `Qwen3.5-27B-w8a8` or `Qwen3.6-27B-w8a8` as an example.
|
||||
|
||||
Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/) for more details.
|
||||
|
||||
There are three `vllm bench` subcommands:
|
||||
|
||||
- `latency`: Benchmark the latency of a single batch of requests.
|
||||
- `serve`: Benchmark the online serving throughput.
|
||||
- `throughput`: Benchmark offline inference throughput.
|
||||
|
||||
Take the `serve` as an example. Run the code as follows.
|
||||
|
||||
```bash
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
# For Qwen3.5-27B-w8a8:
|
||||
vllm bench serve --model Eco-Tech/Qwen3.5-27B-w8a8-mtp --dataset-name random --random-input 200 --num-prompts 200 --request-rate 1 --save-result --result-dir ./
|
||||
# For Qwen3.6-27B-w8a8:
|
||||
vllm bench serve --model Eco-Tech/Qwen3.6-27B-w8a8 --dataset-name random --random-input 200 --num-prompts 200 --request-rate 1 --save-result --result-dir ./
|
||||
```
|
||||
|
||||
After about several minutes, you can get the performance evaluation result.
|
||||
|
||||
## 9 Performance Tuning
|
||||
|
||||
### 9.1 Recommended Configurations
|
||||
|
||||
> **Note**: The following configurations are validated in specific test environments and are for reference only. The optimal configuration depends on factors such as maximum input/output length, prefix cache hit rate, precision requirements, and deployment machine ratios. It is recommended to refer to [Section 9.2](#92-tuning-guidelines) for tuning based on actual conditions.
|
||||
>
|
||||
> **Parallelism Strategy**: `Qwen3.5-27B-w8a8` and `Qwen3.6-27B-w8a8` are only ~30 GB and easily fit in a single NPU (64 GB HBM per NPU). Following the **DP-first** principle, **TP=2 is the recommended default** for most scenarios, and the remaining NPUs should be allocated to DP for parallel request batches. **TP=8 is only recommended for ultra-long context (256k+) scenarios**, where it shards the KV cache across 8 NPUs to maximize the available context window per rank. For `Qwen3.6-27B-w8a8`, you can also raise `--max-model-len` up to 262144 in the same TP/DP layout.
|
||||
>
|
||||
> **Atlas 300I DUO**: Currently only the TP scenario is supported. Choose **TP=2** or **TP=4** according to the available devices. With **TP=4**, `--max-model-len` can support **128k** and **256k** long-sequence scenarios; configure `--max-num-seqs` as needed—setting it too high may cause OOM.
|
||||
>
|
||||
> **Ascend950DT series**: The `Qwen3.6-27B-w8a8-mxfp8` model weight easily fits in a single NPU (96 GB HBM per NPU). Following the **DP-first** principle, **TP=1 is the recommended default** for most scenarios, and the remaining NPUs should be allocated to DP for parallel request batches. For `Qwen3.6-27B-w8a8-mxfp8`, `--max-model-len` can support up to **262144** in the same TP=1 + DP=8 layout.
|
||||
|
||||
#### Table 1: Scenario Overview
|
||||
|
||||
| Scenario | Deployment Mode | *Total NPUs | Weight Version | Key Considerations |
|
||||
|----------|----------------|-------------|----------------|---------------------|
|
||||
| High Throughput<br>(128k context) | Single-Node (A2) | 8 (A2) | Qwen3.5-27B-w8a8 / Qwen3.6-27B-w8a8 | TP=2 + DP=4 fully utilizes all 8 NPUs for parallel request batches |
|
||||
| High Throughput<br>(128k context) | Single-Node (A3) | 16 (A3) | Qwen3.5-27B-w8a8 / Qwen3.6-27B-w8a8 | TP=2 + DP=8 fully utilizes all 16 NPUs for parallel request batches |
|
||||
| Low Latency<br>(128k context) | Single-Node (A3) | 16 (A3) | Qwen3.5-27B-w8a8 / Qwen3.6-27B-w8a8 | TP=2 + DP=8 reduces per-layer Allreduce overhead for small interactive batches |
|
||||
| Long Context<br>(256k+ context) | Single-Node (A3) | 16 (A3) | Qwen3.5-27B-w8a8 / Qwen3.6-27B-w8a8 | TP=8 + DP=2 shards the KV cache across 8 NPUs to maximize the available context window |
|
||||
| High Throughput<br>(128k context) | Single-Node (A5DT) | 8 (A5DT) | Qwen3.6-27B-w8a8-mxfp8 | TP=1 + DP=8 fully utilizes all 8 NPUs for parallel request batches |
|
||||
| Long Context<br>(256k+ context) | Single-Node (A5DT) | 8 (A5DT) | Qwen3.6-27B-w8a8-mxfp8 | TP=1 + DP=8 maximizes the available context window while keeping all 8 NPUs busy |
|
||||
|
||||
> `*Total NPUs` indicates the total number of NPUs used across all nodes. 1 Atlas 800 A3 node = 16 NPUs, 1 Atlas 800 A2 node = 8 NPUs, 1 Ascend950DT series node = 8 NPUs.
|
||||
|
||||
#### Table 2: Detailed Node Configuration
|
||||
|
||||
| Scenario | Configuration | NPUs | TP | DP | Max Num Seqs | Max Num Batched Tokens | Max Model Len | MTP Speculation Num | Async Scheduling |
|
||||
|----------|---------------|-------|----|----|----|-------------|--------------------|---------------------|------------------|
|
||||
| High Throughput (128k) | Single-Node (A2) | 8 | 2 | 4 | 32 | 16384 | 133000 | 3 | On |
|
||||
| High Throughput (128k) | Single-Node (A3) | 16 | 2 | 8 | 32 | 16384 | 133000 | 3 | On |
|
||||
| Low Latency (128k) | Single-Node (A3) | 16 | 2 | 8 | 4 | 4096 | 133000 | 3 | On |
|
||||
| Long Context (256k+) | Single-Node (A3) | 16 | 8 | 2 | 8 | 8192 | 262144 | 3 | On |
|
||||
| High Throughput (128k) | Single-Node (A5DT) | 8 | 1 | 8 | 32 | 16384 | 133000 | 3 | On |
|
||||
| Long Context (256k+) | Single-Node (A5DT) | 8 | 1 | 8 | 32 | 16384 | 262144 | 3 | On |
|
||||
|
||||
> For complete startup commands and parameter descriptions, please refer to the deployment examples in [Chapter 5](#5-online-service-deployment).
|
||||
|
||||
### 9.2 Tuning Guidelines
|
||||
|
||||
#### 9.2.1 General Tuning Reference
|
||||
|
||||
Please refer to the [Public Performance Tuning Documentation](../../developer_guide/performance_and_debug/optimization_and_tuning.md) for tuning methods.
|
||||
Please refer to the [Feature Guide](../../user_guide/support_matrix/feature_matrix.md) for detailed feature descriptions.
|
||||
|
||||
## 10 FAQ
|
||||
|
||||
For common environment, installation, and general parameter issues, please refer to the [vLLM-Ascend Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html).
|
||||
1035
docs/source/tutorials/models/Qwen3.5-397B-A17B.md
Normal file
1035
docs/source/tutorials/models/Qwen3.5-397B-A17B.md
Normal file
File diff suppressed because it is too large
Load Diff
418
docs/source/tutorials/models/Qwen3.5-Dense.md
Normal file
418
docs/source/tutorials/models/Qwen3.5-Dense.md
Normal file
@@ -0,0 +1,418 @@
|
||||
# Qwen3.5-Dense (Qwen3.5-2B/4B/9B)
|
||||
|
||||
## 1 Introduction
|
||||
|
||||
Qwen3.5-2B, Qwen3.5-4B, and Qwen3.5-9B are dense hybrid Mamba-Transformer language models in the Qwen3.5 family. They share the same hybrid attention design (GDN + full attention) and are suitable for general-purpose text generation tasks such as dialogue, content creation, and code generation.
|
||||
|
||||
This document describes deployment and verification of these models on **Atlas 300I DUO** and **Atlas 200I Pro**, including environment preparation, Docker installation, single-node online deployment, functional verification, and tuning notes.
|
||||
|
||||
It is **strongly recommended to use the latest release candidate (rc) version or the latest official version** of `vllm-ascend`. Support for Qwen3.5-2B/4B/9B on Atlas 300I DUO and Atlas 200I Pro starts from `vllm-ascend:v0.23.0rc1`.
|
||||
|
||||
## 2 Supported Features
|
||||
|
||||
Please refer to the [Supported Features List](../../user_guide/support_matrix/supported_models.md) for the model support matrix.
|
||||
|
||||
Please refer to the [Feature Guide](../../user_guide/feature_guide/index.md) for feature configuration information.
|
||||
|
||||
## 3 Prerequisites
|
||||
|
||||
### 3.1 Model Weight
|
||||
|
||||
| Model | Version | Hardware Requirement | Download |
|
||||
|-------|---------|----------------------|----------|
|
||||
| Qwen3.5-2B | FP16 | Atlas 300I DUO or Atlas 200I Pro | [Download](https://www.modelscope.cn/models/Qwen/Qwen3.5-2B) |
|
||||
| Qwen3.5-4B | FP16 | Atlas 300I DUO or Atlas 200I Pro | [Download](https://www.modelscope.cn/models/Qwen/Qwen3.5-4B) |
|
||||
| Qwen3.5-9B | FP16 | Atlas 300I DUO or Atlas 200I Pro | [Download](https://www.modelscope.cn/models/Qwen/Qwen3.5-9B) |
|
||||
|
||||
It is recommended to download the model weight to a local directory such as `/root/.cache/` or `/home/data/`.
|
||||
|
||||
## 4 Installation
|
||||
|
||||
### 4.1 Docker Image Installation
|
||||
|
||||
Select an image based on your machine type and start the docker image on your node, refer to [using docker](../../installation.md#set-up-using-docker).
|
||||
|
||||
It is **recommended to use the latest release candidate (rc) version or the latest official version** of the `vllm-ascend` image. For Atlas 300I DUO and Atlas 200I Pro on Ubuntu, use `vllm-ascend:nightly-releases-v0.23.0-310p` (or a later `-310p` image). For Atlas 200I Pro on openEuler, use `vllm-ascend:nightly-releases-v0.23.0-310p-openeuler` (or a later `-310p-openeuler` image).
|
||||
|
||||
:::::::{tab-set}
|
||||
|
||||
::::::{tab-item} Atlas 300I DUO
|
||||
|
||||
Start the docker image on each node.
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
export IMAGE=m.daocloud.io/quay.io/ascend/vllm-ascend:|vllm_ascend_version|-310p
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-p 8080:8080 \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::::
|
||||
|
||||
::::::{tab-item} Atlas 200I Pro
|
||||
|
||||
Start the docker image on each node. Adjust `--device=/dev/davinci0` according to the NPU ID you want to use.
|
||||
|
||||
:::::{tab-set}
|
||||
|
||||
::::{tab-item} Ubuntu 24.04
|
||||
:selected:
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-310p
|
||||
|
||||
docker run --rm \
|
||||
--privileged \
|
||||
--name vllm-ascend \
|
||||
--shm-size=10g \
|
||||
--device=/dev/davinci0:/dev/davinci0 \
|
||||
--device=/dev/davinci_manager \
|
||||
--device=/dev/ascend_manager \
|
||||
--device=/dev/user_config \
|
||||
-v /etc/sys_version.conf:/etc/sys_version.conf \
|
||||
-v /etc/ld.so.conf.d/mind_so.conf:/etc/ld.so.conf.d/mind_so.conf \
|
||||
-v /etc/hdcBasic.cfg:/etc/hdcBasic.cfg \
|
||||
-v /var/dmp_daemon:/var/dmp_daemon \
|
||||
-v /usr/lib64/libmmpa.so:/usr/lib64/libmmpa.so \
|
||||
-v /usr/lib64/libcrypto.so.1.1:/usr/lib64/libcrypto.so.1.1 \
|
||||
-v /usr/local/sbin/npu-smi:/usr/local/sbin/npu-smi \
|
||||
-v /usr/lib64/libstackcore.so:/usr/lib64/libstackcore.so \
|
||||
-v /usr/lib/aarch64-linux-gnu/libyaml-0.so.2:/usr/lib64/libyaml-0.so.2 \
|
||||
-v /etc/slog.conf:/etc/slog.conf \
|
||||
-v /var/slogd:/var/slogd \
|
||||
-v /usr/local/Ascend/driver/lib64:/usr/local/Ascend/driver/lib64 \
|
||||
-v /usr/lib64/libtensorflow.so:/usr/lib64/libtensorflow.so \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-p 8080:8080 \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} openEuler 24.03
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-310p-openeuler
|
||||
|
||||
docker run --rm \
|
||||
--privileged \
|
||||
--name vllm-ascend \
|
||||
--shm-size=10g \
|
||||
--device=/dev/davinci0:/dev/davinci0 \
|
||||
--device=/dev/davinci_manager \
|
||||
--device=/dev/ascend_manager \
|
||||
--device=/dev/user_config \
|
||||
-v /etc/sys_version.conf:/etc/sys_version.conf \
|
||||
-v /etc/ld.so.conf.d/mind_so.conf:/etc/ld.so.conf.d/mind_so.conf \
|
||||
-v /etc/hdcBasic.cfg:/etc/hdcBasic.cfg \
|
||||
-v /var/dmp_daemon:/var/dmp_daemon \
|
||||
-v /usr/lib64/libsemanage.so.2:/usr/lib64/libsemanage.so.2 \
|
||||
-v /usr/lib64/libmmpa.so:/usr/lib64/libmmpa.so \
|
||||
-v /usr/lib64/libcrypto.so.1.1:/usr/lib64/libcrypto.so.1.1 \
|
||||
-v /usr/lib64/libyaml-0.so.2.0.9:/usr/lib64/libyaml-0.so.2 \
|
||||
-v /usr/local/sbin/npu-smi:/usr/local/sbin/npu-smi \
|
||||
-v /usr/lib64/libstackcore.so:/usr/lib64/libstackcore.so \
|
||||
-v /etc/slog.conf:/etc/slog.conf \
|
||||
-v /var/slogd:/var/slogd \
|
||||
-v /usr/local/Ascend/driver/lib64:/usr/local/Ascend/driver/lib64 \
|
||||
-v /usr/lib64/libtensorflow.so:/usr/lib64/libtensorflow.so \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-p 8080:8080 \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
::::::
|
||||
:::::::
|
||||
|
||||
After a successful docker run, you can verify the running container service by executing the `docker ps` command. The expected result is that the container `vllm-ascend` is listed with status `Up`, confirming the docker installation is successful.
|
||||
|
||||
### 4.2 Source Code Installation
|
||||
|
||||
If you don't want to use the docker image as above, you can also build all from source:
|
||||
|
||||
1. Clone the repository and install `vllm-ascend` from source:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/vllm-project/vllm-ascend.git
|
||||
cd vllm-ascend
|
||||
pip install -e .
|
||||
```
|
||||
|
||||
For the complete installation steps, refer to [installation](../../installation.md).
|
||||
|
||||
````{note}
|
||||
On Atlas 300I DUO and Atlas 200I Pro, you may need to uninstall `triton-ascend` and `triton` to avoid dependency conflicts:
|
||||
|
||||
```bash
|
||||
pip uninstall -y triton-ascend triton
|
||||
```
|
||||
````
|
||||
|
||||
To verify the source code installation, run the following command and confirm the displayed version matches the one you installed:
|
||||
|
||||
```bash
|
||||
pip show vllm-ascend
|
||||
```
|
||||
|
||||
Expected result: The version information of `vllm-ascend` is displayed, confirming a successful installation.
|
||||
|
||||
## 5 Online Service Deployment
|
||||
|
||||
### 5.1 Single-Node Online Deployment
|
||||
|
||||
Single-node deployment completes both Prefill and Decode within the same node. `Qwen3.5-2B`, `Qwen3.5-4B`, and `Qwen3.5-9B` can be deployed on Atlas 300I DUO or Atlas 200I Pro.
|
||||
|
||||
> **Parallelism note**: These platforms currently support the **TP** scenario. Choose **TP=1** or **TP=2** according to the available devices. On Atlas 200I Pro with a single visible NPU, use **TP=1**.
|
||||
|
||||
The following examples use FP16 weights from ModelScope. Replace `MODEL_PATH` with your local directory if needed.
|
||||
|
||||
:::::{tab-set}
|
||||
|
||||
::::{tab-item} Qwen3.5-2B
|
||||
|
||||
Startup Command:
|
||||
|
||||
```bash
|
||||
#!/bin/sh
|
||||
# Load model from ModelScope to speed up download
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
|
||||
# Model weight path; can be a ModelScope model id or a local directory path
|
||||
export MODEL_PATH=Qwen/Qwen3.5-2B
|
||||
|
||||
vllm serve $MODEL_PATH \
|
||||
--host 127.0.0.1 \
|
||||
--port 1025 \
|
||||
--tensor-parallel-size 1 \
|
||||
--served-model-name qwen3.5 \
|
||||
--max-num-seqs 32 \
|
||||
--max-model-len 16384 \
|
||||
--trust-remote-code \
|
||||
--gpu-memory-utilization 0.90 \
|
||||
--mamba-ssm-cache-dtype float16 \
|
||||
--dtype float16 \
|
||||
--speculative-config '{"method": "qwen3_5_mtp","num_speculative_tokens":1}' \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY", "cudagraph_capture_sizes": [2,4,8,16]}' \
|
||||
--additional-config '{"ascend_compilation_config": {"enable_npugraph_ex": false}}'
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Qwen3.5-4B
|
||||
|
||||
Startup Command:
|
||||
|
||||
```bash
|
||||
#!/bin/sh
|
||||
# Load model from ModelScope to speed up download
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
|
||||
# Model weight path; can be a ModelScope model id or a local directory path
|
||||
export MODEL_PATH=Qwen/Qwen3.5-4B
|
||||
|
||||
vllm serve $MODEL_PATH \
|
||||
--host 127.0.0.1 \
|
||||
--port 1025 \
|
||||
--tensor-parallel-size 1 \
|
||||
--served-model-name qwen3.5 \
|
||||
--max-num-seqs 32 \
|
||||
--max-model-len 16384 \
|
||||
--trust-remote-code \
|
||||
--gpu-memory-utilization 0.90 \
|
||||
--mamba-ssm-cache-dtype float16 \
|
||||
--dtype float16 \
|
||||
--speculative-config '{"method": "qwen3_5_mtp","num_speculative_tokens":1}' \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY", "cudagraph_capture_sizes": [2,4,8,16]}' \
|
||||
--additional-config '{"ascend_compilation_config": {"enable_npugraph_ex": false}}'
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Qwen3.5-9B
|
||||
|
||||
Startup Command:
|
||||
|
||||
```bash
|
||||
#!/bin/sh
|
||||
# Load model from ModelScope to speed up download
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
|
||||
# Model weight path; can be a ModelScope model id or a local directory path
|
||||
export MODEL_PATH=Qwen/Qwen3.5-9B
|
||||
|
||||
vllm serve $MODEL_PATH \
|
||||
--host 127.0.0.1 \
|
||||
--port 1025 \
|
||||
--tensor-parallel-size 1 \
|
||||
--served-model-name qwen3.5 \
|
||||
--max-num-seqs 32 \
|
||||
--max-model-len 16384 \
|
||||
--trust-remote-code \
|
||||
--gpu-memory-utilization 0.90 \
|
||||
--mamba-ssm-cache-dtype float16 \
|
||||
--dtype float16 \
|
||||
--speculative-config '{"method": "qwen3_5_mtp","num_speculative_tokens":1}' \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY", "cudagraph_capture_sizes": [2,4,8,16]}' \
|
||||
--additional-config '{"ascend_compilation_config": {"enable_npugraph_ex": false}}'
|
||||
```
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
Key Parameter Descriptions:
|
||||
|
||||
- `--tensor-parallel-size` sets the tensor parallel size. Prefer **TP=1** on Atlas 200I Pro. On Atlas 300I DUO, **TP=1** and **TP=2** are both supported; choose according to the available devices.
|
||||
- `--max-model-len` represents the context length (input plus output for a single request). On Atlas 300I DUO and Atlas 200I Pro, configure this value according to the actual device memory; setting it too high may cause OOM.
|
||||
- `--max-num-seqs` indicates the maximum number of requests that can be processed concurrently. On Atlas 300I DUO and Atlas 200I Pro, configure this value according to the actual device memory; setting it too high may cause OOM.
|
||||
- `--gpu-memory-utilization` represents the proportion of HBM that vLLM will use for actual inference. On Atlas 300I DUO and Atlas 200I Pro, configure this value according to the actual device memory; setting it too high may cause OOM. The default value is `0.9`.
|
||||
- `--dtype float16` must be set on Atlas 300I DUO and Atlas 200I Pro. These devices only support the FP16 data type.
|
||||
- `--mamba-ssm-cache-dtype` sets the data type of the Mamba SSM cache. On Atlas 300I DUO and Atlas 200I Pro, only `float16` is supported.
|
||||
- `--speculative-config` uses `qwen3_5_mtp` for Qwen3.5 Dense models that include an MTP head. It is recommended to set `num_speculative_tokens` to `1`.
|
||||
- `--compilation-config` contains configurations related to the aclgraph graph mode:
|
||||
- `"cudagraph_mode"`: `"FULL_DECODE_ONLY"` is recommended.
|
||||
- `"cudagraph_capture_sizes"`: when tensor parallelism (TP) is enabled, hardware event-id constraints allow at most two capture sizes (for example, `[1, 8]`). With MTP enabled, calculate each capture size as `n * (num_speculative_tokens + 1)`, where `n` is a capture size for the deployment without MTP. For example, when `num_speculative_tokens` is `1`, the non-MTP sizes `[1,2,4,8]` become `[2,4,8,16]`.
|
||||
- `--additional-config` with `"ascend_compilation_config": {"enable_npugraph_ex": false}` is required because `enable_npugraph_ex` is not supported on these platforms.
|
||||
|
||||
Common Issues Tip: If you encounter issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html) for troubleshooting.
|
||||
|
||||
Service Verification:
|
||||
|
||||
If the service starts successfully, the following startup log will be displayed:
|
||||
|
||||
```text
|
||||
(APIServer pid=<pid>) INFO: Started server process [<pid>]
|
||||
(APIServer pid=<pid>) INFO: Waiting for application startup.
|
||||
(APIServer pid=<pid>) INFO: Application startup complete.
|
||||
```
|
||||
|
||||
## 6 Functional Verification
|
||||
|
||||
After the service is started, the model can be invoked by sending a prompt. Two API interfaces are supported: `completions` and `chat.completions`. Use the `--served-model-name` you configured (for example, `qwen3.5`). If you used `--port 1025` or `-p 8080:8080`, adjust the URL accordingly.
|
||||
|
||||
**Completions API:**
|
||||
|
||||
```bash
|
||||
curl http://127.0.0.1:1025/v1/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "qwen3.5",
|
||||
"prompt": "The future of AI is",
|
||||
"max_completion_tokens": 50,
|
||||
"temperature": 0
|
||||
}'
|
||||
```
|
||||
|
||||
**Chat Completions API:**
|
||||
|
||||
```bash
|
||||
curl http://127.0.0.1:1025/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "qwen3.5",
|
||||
"messages": [
|
||||
{"role": "user", "content": "The future of AI is"}
|
||||
],
|
||||
"max_completion_tokens": 1024,
|
||||
"temperature": 0.7,
|
||||
"top_p": 0.95
|
||||
}'
|
||||
```
|
||||
|
||||
Expected Result: The service returns HTTP 200 OK. The JSON response contains the `choices` field with generated text.
|
||||
|
||||
## 7 Accuracy Evaluation
|
||||
|
||||
### Using AISBench
|
||||
|
||||
1. Refer to [Using AISBench](../../developer_guide/evaluation/using_ais_bench.md) for details.
|
||||
|
||||
2. After execution, you can get the result. Here are the accuracy results of `Qwen3.5-2B`, `Qwen3.5-4B`, and `Qwen3.5-9B` on Atlas 300I DUO for reference only.
|
||||
|
||||
**Accuracy Evaluation Config File:**
|
||||
|
||||
```python
|
||||
# Example configuration: benchmarks/ais_bench/benchmark/configs/models/vllm_api/vllm_api_general_chat.py
|
||||
from ais_bench.benchmark.models import VLLMCustomAPIChat
|
||||
from ais_bench.benchmark.utils.model_postprocessors import extract_non_reasoning_content
|
||||
|
||||
models = [
|
||||
dict(
|
||||
attr="service",
|
||||
type=VLLMCustomAPIChat,
|
||||
abbr="vllm-api-general-chat",
|
||||
path="your_model_path",
|
||||
model="qwen3.5",
|
||||
request_rate=0,
|
||||
retry=2,
|
||||
host_ip="127.0.0.1",
|
||||
host_port=1025,
|
||||
max_out_len=4096,
|
||||
batch_size=16,
|
||||
trust_remote_code=False,
|
||||
generation_kwargs=dict(
|
||||
temperature=0.0,
|
||||
ignore_eos=False,
|
||||
chat_template_kwargs = {"enable_thinking": False},
|
||||
),
|
||||
pred_postprocessor=dict(type=extract_non_reasoning_content)
|
||||
)
|
||||
]
|
||||
```
|
||||
|
||||
| Model | dataset | version | metric | mode | vllm-api-general-chat |
|
||||
|-------|---------|---------|--------|------|------------------------|
|
||||
| Qwen3.5-2B | gsm8k | - | accuracy | gen | 77.71 |
|
||||
| Qwen3.5-2B | textvqa | - | accuracy | gen | 76.09 |
|
||||
| Qwen3.5-4B | gsm8k | - | accuracy | gen | 93.18 |
|
||||
| Qwen3.5-4B | textvqa | - | accuracy | gen | 79.08 |
|
||||
| Qwen3.5-9B | gsm8k | - | accuracy | gen | 95.30 |
|
||||
| Qwen3.5-9B | textvqa | - | accuracy | gen | 82.33 |
|
||||
|
||||
## 8 Performance Evaluation
|
||||
|
||||
### Using AISBench
|
||||
|
||||
Refer to [Using AISBench for performance evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation) for details.
|
||||
|
||||
## 9 Performance Tuning
|
||||
|
||||
### 9.1 Recommended Configurations
|
||||
|
||||
> **Note**: The following configurations are for reference only. The optimal configuration depends on model size, maximum input/output length, and actual device memory.
|
||||
>
|
||||
> **Atlas 300I DUO / Atlas 200I Pro**: Currently only the TP scenario is supported. Prefer **TP=1** on Atlas 200I Pro. On Atlas 300I DUO, **TP=1** and **TP=2** are both supported; choose according to the available devices. Configure `--max-model-len`, `--max-num-seqs`, and `--gpu-memory-utilization` based on the actual device memory; setting them too high may cause OOM.
|
||||
|
||||
### 9.2 Tuning Guidelines
|
||||
|
||||
Please refer to the [Public Performance Tuning Documentation](../../developer_guide/performance_and_debug/optimization_and_tuning.md) for tuning methods.
|
||||
|
||||
Please refer to the [Feature Guide](../../user_guide/support_matrix/feature_matrix.md) for detailed feature descriptions.
|
||||
|
||||
## 10 FAQ
|
||||
|
||||
For common environment, installation, and general parameter issues, please refer to the [vLLM-Ascend Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html).
|
||||
420
docs/source/tutorials/models/Qwen3.6-35B-A3B.md
Normal file
420
docs/source/tutorials/models/Qwen3.6-35B-A3B.md
Normal file
@@ -0,0 +1,420 @@
|
||||
# Qwen3.6-35B-A3B
|
||||
|
||||
## 1 Introduction
|
||||
|
||||
Qwen3.6-35B-A3B is a sparse MoE model in the Qwen3.6 family, with 35B total parameters and about 3B activated parameters per token. It uses the hybrid attention architecture used by Qwen3.5-style models, and is suitable for long-context online serving on Ascend hardware.
|
||||
|
||||
This document describes the main validation steps for the model, including supported features, prerequisites, installation, single-node online deployment, functional verification, accuracy and performance evaluation, performance tuning, and FAQs.
|
||||
|
||||
The `Qwen3.6-35B-A3B` model is first supported in `vllm-ascend:v0.18.0rc1`. Use `v0.18.0rc1` or later for this model. The examples below use the version placeholder configured by the documentation build system.
|
||||
|
||||
## 2 Supported Features
|
||||
|
||||
Refer to [supported features](../../user_guide/support_matrix/supported_features.md) to get the model's supported feature matrix, including BF16, W8A8 quantization, chunked prefill, automatic prefix caching, asynchronous scheduling, tensor parallelism, expert parallelism, and ACLGraph support.
|
||||
|
||||
Refer to [feature guide](../../user_guide/feature_guide/index.md) to get feature configuration details.
|
||||
|
||||
## 3 Prerequisites
|
||||
|
||||
### 3.1 Model Weight
|
||||
|
||||
- `Qwen3.6-35B-A3B` (BF16 version): requires 1 Atlas A3 inference products (64G x 16) node, 1 Atlas A2 inference products (64G x 8) node, or Atlas 300I DUO. [Download model weight](https://www.modelscope.cn/models/Qwen/Qwen3.6-35B-A3B).
|
||||
- `Qwen3.6-35B-A3B-w8a8` (quantized version): requires 1 Atlas A3 inference products (64G x 16) node, 1 Atlas A2 inference products (64G x 8) node, or Atlas 300I DUO. [Download model weight](https://www.modelscope.cn/models/Eco-Tech/Qwen3.6-35B-A3B-w8a8).
|
||||
|
||||
It is recommended to download the model weight to `/root/.cache/`.
|
||||
|
||||
## 4 Installation
|
||||
|
||||
### 4.1 Docker Image Installation
|
||||
|
||||
Select an image based on your machine type. For example, use `quay.io/ascend/vllm-ascend:|vllm_ascend_version|` for Atlas A2 inference products, `quay.io/ascend/vllm-ascend:|vllm_ascend_version|-a3` for Atlas A3 inference products, and `quay.io/ascend/vllm-ascend:|vllm_ascend_version|-310p` for Atlas 300I DUO.
|
||||
|
||||
For Atlas 300I DUO, use `vllm-ascend:nightly-releases-v0.23.0-310p` (or a later `-310p` image).
|
||||
|
||||
Refer to [using docker](../../installation.md#set-up-using-docker) for the complete installation guide.
|
||||
|
||||
:::::{tab-set}
|
||||
|
||||
::::{tab-item} Atlas A3 inference products
|
||||
:sync: A3
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
# Download the model weight to /root/.cache in advance.
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-a3
|
||||
export NAME=vllm-ascend
|
||||
|
||||
docker run --rm \
|
||||
--name $NAME \
|
||||
--net=host \
|
||||
--shm-size=1g \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci8 \
|
||||
--device /dev/davinci9 \
|
||||
--device /dev/davinci10 \
|
||||
--device /dev/davinci11 \
|
||||
--device /dev/davinci12 \
|
||||
--device /dev/davinci13 \
|
||||
--device /dev/davinci14 \
|
||||
--device /dev/davinci15 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} Atlas A2 inference products
|
||||
:sync: A2
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
# Download the model weight to /root/.cache in advance.
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
export NAME=vllm-ascend
|
||||
|
||||
docker run --rm \
|
||||
--name $NAME \
|
||||
--net=host \
|
||||
--shm-size=1g \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
# Use the vllm-ascend image
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-310p
|
||||
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=10g \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-p 8080:8080 \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
After entering the container, verify that vLLM and vLLM-Ascend can be imported:
|
||||
|
||||
```shell
|
||||
python -c "import vllm, vllm_ascend; print('vllm and vllm_ascend are ready')"
|
||||
```
|
||||
|
||||
### 4.2 Source Code Installation
|
||||
|
||||
You can also build and install `vllm-ascend` from source. Refer to [set up using python](../../installation.md#set-up-using-python).
|
||||
|
||||
:::{note}
|
||||
For Atlas 300I DUO, source installation may pull in `triton` and `triton-ascend`. Uninstall them before running vLLM-Ascend on Atlas 300I DUO:
|
||||
|
||||
```bash
|
||||
pip uninstall -y triton-ascend triton
|
||||
```
|
||||
|
||||
:::
|
||||
|
||||
## 5 Online Service Deployment
|
||||
|
||||
### 5.1 Single-Node Online Deployment
|
||||
|
||||
Single-node deployment runs both Prefill and Decode on the same node. `Qwen3.6-35B-A3B-w8a8` can be deployed on 1 Atlas A3 inference products (64G x 16) or 1 Atlas A2 inference products (64G x 8), or Atlas 300I DUO. The W8A8 version needs `--quantization ascend`.
|
||||
|
||||
:::::{tab-set}
|
||||
::::{tab-item} Atlas A2 inference products / Atlas A3 inference products
|
||||
|
||||
Run the following script to execute online inference with up to 262144 context length on 1 Atlas A3 inference products (64G x 16).
|
||||
|
||||
```shell
|
||||
#!/bin/sh
|
||||
|
||||
# Load model from ModelScope to speed up download.
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
|
||||
# Reduce memory fragmentation and avoid out-of-memory errors.
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export OMP_NUM_THREADS=1
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
|
||||
sysctl -w vm.swappiness=0
|
||||
sysctl -w kernel.numa_balancing=0
|
||||
sysctl kernel.sched_migration_cost_ns=50000
|
||||
export LD_PRELOAD=/usr/lib/aarch64-linux-gnu/libjemalloc.so.2:$LD_PRELOAD
|
||||
|
||||
vllm serve Eco-Tech/Qwen3.6-35B-A3B-w8a8 \
|
||||
--host 0.0.0.0 \
|
||||
--port 8000 \
|
||||
--data-parallel-size 1 \
|
||||
--tensor-parallel-size 2 \
|
||||
--enable-expert-parallel \
|
||||
--seed 1024 \
|
||||
--quantization ascend \
|
||||
--served-model-name qwen3.6 \
|
||||
--max-num-seqs 128 \
|
||||
--max-model-len 262144 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--trust-remote-code \
|
||||
--gpu-memory-utilization 0.90 \
|
||||
--enable-prefix-caching \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \
|
||||
--additional-config '{"enable_cpu_binding":true, "enable_flashcomm1":true, "multistream_overlap_shared_expert": true}' \
|
||||
--async-scheduling
|
||||
```
|
||||
|
||||
**Key parameters:**
|
||||
|
||||
- `--data-parallel-size 1` and `--tensor-parallel-size 2` set DP and TP for the default single-node serving example.
|
||||
- `--enable-expert-parallel` enables expert parallelism for MoE layers. Do not mix MoE tensor parallelism and expert parallelism in the same MoE layer.
|
||||
- `--max-model-len` is the maximum input plus output length for a single request. Increase it only when enough KV cache is available.
|
||||
- `--max-num-seqs` is the maximum number of active requests scheduled by each DP group. For performance tests, set `--max-num-seqs * --data-parallel-size` greater than or equal to the test concurrency.
|
||||
- `--max-num-batched-tokens` is the maximum number of tokens processed in one scheduler step. A larger value can improve prefill efficiency but consumes more activation memory.
|
||||
- `--gpu-memory-utilization` controls how much HBM vLLM can use to calculate KV cache capacity. A higher value increases KV cache size but can trigger OOM if runtime memory is higher than the profile run.
|
||||
- `--enable-prefix-caching` enables prefix caching. For long-context serving, monitor memory usage because prefix caching can increase KV cache pressure.
|
||||
- `--quantization ascend` enables Ascend quantization for the W8A8 model. Remove this option when deploying the BF16 model.
|
||||
- `--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}'` enables full decode ACLGraph replay to reduce dispatch overhead.
|
||||
- `--additional-config` enables Ascend-specific optimizations. `enable_flashcomm1` enables FlashComm1, `multistream_overlap_shared_expert` overlaps shared expert computation, and `enable_cpu_binding` enables Ascend-native CPU binding.
|
||||
- `--async-scheduling` enables asynchronous scheduling, which can improve high-concurrency throughput.
|
||||
|
||||
::::
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
|
||||
```shell
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1
|
||||
|
||||
vllm serve Eco-Tech/Qwen3.6-35B-A3B-w8a8 \
|
||||
--host 127.0.0.1 \
|
||||
--port 8080 \
|
||||
--tensor-parallel-size 2 \
|
||||
--gpu-memory-utilization 0.90 \
|
||||
--max-num-seqs 16 \
|
||||
--served-model-name qwen3.6 \
|
||||
--dtype float16 \
|
||||
--additional-config '{"ascend_compilation_config": {"enable_npugraph_ex":false}}' \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY", "cudagraph_capture_sizes": [1,8]}' \
|
||||
--quantization ascend \
|
||||
--max-model-len 20480 \
|
||||
--no-enable-prefix-caching
|
||||
```
|
||||
|
||||
**Key parameters:**
|
||||
|
||||
- `--tensor-parallel-size 2` maps the model across two Atlas inference devices. Adjust it together with `ASCEND_RT_VISIBLE_DEVICES` according to the available devices and memory.
|
||||
- `--dtype float16` is used for Atlas 300I DUO to match the Atlas inference execution path.
|
||||
- `--max-num-seqs 16` limits concurrent active requests to reduce KV cache and graph capture pressure on Atlas 300I DUO.
|
||||
- `--gpu-memory-utilization` controls KV cache capacity. Reduce it if startup or runtime requests report OOM.
|
||||
- `--additional-config` with `"ascend_compilation_config": {"enable_npugraph_ex": false}` is required because `enable_npugraph_ex` is not supported on Atlas 300I DUO.
|
||||
- `--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY", "cudagraph_capture_sizes": [1,2,4,8,16]}'` enables decode ACLGraph replay and explicitly limits capture sizes for Atlas 300I DUO.
|
||||
- `--no-enable-prefix-caching` is the default recommendation for this Atlas 300I DUO example to reduce memory pressure.
|
||||
- `--quantization ascend` enables Ascend quantization for the W8A8 model. Remove this option when deploying the BF16 model.
|
||||
- To enable MTP speculative decoding, use --speculative_config '{"method": "mtp", "num_speculative_tokens": 1}'. We recommend setting num_speculative_tokens to 1.
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
Common Issues Tip: If the service fails to start, HBM is insufficient, or requests are not scheduled as expected, refer to [FAQs](../../faqs.md) first, and then check the model-specific FAQ in Section 10.
|
||||
|
||||
## 6 Functional Verification
|
||||
|
||||
After the server is started, send a request to verify basic model functionality.
|
||||
|
||||
```shell
|
||||
curl http://<server_ip>:<port>/v1/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "qwen3.6",
|
||||
"prompt": "The future of AI is",
|
||||
"max_tokens": 50,
|
||||
"temperature": 0
|
||||
}'
|
||||
```
|
||||
|
||||
Expected result: the HTTP status is 200 and the JSON response contains a `choices` field with generated text.
|
||||
|
||||
## 7 Accuracy Evaluation
|
||||
|
||||
Here are two accuracy evaluation methods.
|
||||
|
||||
### 7.1 Using AISBench
|
||||
|
||||
Refer to [Using AISBench](../../developer_guide/evaluation/using_ais_bench.md) for details. After execution, you can get the accuracy result of `Qwen3.6-35B-A3B-w8a8`.
|
||||
|
||||
| dataset | version | metric | mode | vllm-api-general-chat |
|
||||
| ------- | ------- | ------ | ---- | --------------------- |
|
||||
| mmmu | - | accuracy | gen | 83.3 |
|
||||
| gpqa | - | accuracy | gen | 83.3 |
|
||||
|
||||
### 7.2 Using Language Model Evaluation Harness
|
||||
|
||||
Refer to [Using lm_eval](../../developer_guide/evaluation/using_lm_eval.md) for installation and usage details. When using online serving, set `base_url` to the endpoint started in Section 5.
|
||||
|
||||
```shell
|
||||
lm_eval \
|
||||
--model local-completions \
|
||||
--model_args model=qwen3.6,base_url=http://127.0.0.1:8000/v1/completions,tokenized_requests=False,trust_remote_code=True \
|
||||
--tasks gsm8k \
|
||||
--output_path ./
|
||||
```
|
||||
|
||||
## 8 Performance Evaluation
|
||||
|
||||
### 8.1 Using AISBench
|
||||
|
||||
Refer to [Using AISBench for performance evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation) for details.
|
||||
|
||||
### 8.2 Using vLLM Benchmark
|
||||
|
||||
Run performance evaluation of `Qwen3.6-35B-A3B-w8a8` as an example. Refer to [vLLM benchmark](https://docs.vllm.ai/en/latest/benchmarking/) for more details.
|
||||
|
||||
There are three `vllm bench` subcommands:
|
||||
|
||||
- `latency`: benchmark the latency of a single batch of requests.
|
||||
- `serve`: benchmark online serving throughput.
|
||||
- `throughput`: benchmark offline inference throughput.
|
||||
|
||||
Take `serve` as an example:
|
||||
|
||||
```shell
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
|
||||
vllm bench serve \
|
||||
--model Eco-Tech/Qwen3.6-35B-A3B-w8a8 \
|
||||
--served-model-name qwen3.6 \
|
||||
--dataset-name random \
|
||||
--random-input 200 \
|
||||
--num-prompts 200 \
|
||||
--request-rate 1 \
|
||||
--save-result \
|
||||
--result-dir ./
|
||||
```
|
||||
|
||||
After several minutes, you can get the performance evaluation result.
|
||||
|
||||
## 9 Performance Tuning
|
||||
|
||||
### 9.1 Recommended Configurations
|
||||
|
||||
The following configurations are validated in specific test environments and are for reference only. The optimal configuration depends on hardware type, maximum input/output length, request concurrency, prefix cache hit rate, and quantization. Tune the parameters in Section 9.2 based on your actual workload.
|
||||
|
||||
| Scenario | Deployment Mode | Total NPUs | Weight Version | Key Considerations |
|
||||
| -------- | --------------- | ---------- | -------------- | ------------------ |
|
||||
| Long context | Single-node online serving | 2 or more NPUs | W8A8 | Use larger `--max-model-len` and reserve enough KV cache. Lower `--max-num-seqs` if OOM occurs. |
|
||||
| High throughput | Single-node online serving | 8 or more NPUs | W8A8 | Increase local DP groups within one node and tune `--max-num-batched-tokens`. |
|
||||
| Low latency | Single-node online serving | 2 or more NPUs | W8A8 | Use smaller `--max-num-batched-tokens`, full decode ACLGraph, and disable speculative decoding by default. |
|
||||
|
||||
| Scenario | Node Role | NPUs | TP | DP | Max Num Seqs | Max Model Len | Max Num Batched Tokens | Prefix Cache | Main Optimizations |
|
||||
| -------- | --------- | ---- | -- | -- | ------------ | ------------- | ---------------------- | ------------ | ------------------ |
|
||||
| Long context | Single node | 2 or more | 2 | 1 | 128 | 262144 | 16384 | On | FullGraph, FlashComm1, shared expert overlap, CPU binding |
|
||||
| High throughput | Single node | 8 or more | 2 | 4 or more | 32 per DP | 65536 | 8192 | On | FullGraph, FlashComm1, async scheduling, shared expert overlap |
|
||||
| Low latency | Single node | 2 or more | 2 | 1 | Tune by concurrency | 32768 or 65536 | 1024 to 4096 | Workload dependent | FullGraph, CPU binding, speculative decoding disabled |
|
||||
|
||||
### 9.2 Tuning Guidelines
|
||||
|
||||
Refer to [public performance tuning documentation](../../developer_guide/performance_and_debug/optimization_and_tuning.md) for general tuning methods, and refer to [feature matrix](../../user_guide/support_matrix/feature_matrix.md) for feature descriptions.
|
||||
|
||||
Recommended tuning order:
|
||||
|
||||
1. Use single-node deployment. If more throughput is required, increase local DP groups within the same node.
|
||||
2. Choose the maximum context length with `--max-model-len`. Long context increases KV cache usage, so reduce `--max-num-seqs` or `--gpu-memory-utilization` if OOM occurs.
|
||||
3. Tune `--max-num-batched-tokens`. Larger values usually improve prefill throughput but increase activation memory. Decode-heavy workloads usually need smaller values.
|
||||
4. Tune `--max-num-seqs` according to service concurrency. Requests above this value wait in the queue and the waiting time is counted in TTFT and TPOT.
|
||||
5. Tune `--gpu-memory-utilization`. Increase it to provide more KV cache, but leave headroom for runtime memory fluctuation and expert imbalance.
|
||||
6. Tune ACLGraph capture. `FULL_DECODE_ONLY` is recommended for decode. If you set `cudagraph_capture_sizes` manually, include common decode batch sizes. With FlashComm1, use capture sizes that are multiples of TP size.
|
||||
|
||||
### 9.3 Model-Specific Optimizations
|
||||
|
||||
| Optimization | Enablement | Benefit | Notes |
|
||||
| ------------ | ---------- | ------- | ----- |
|
||||
| Hybrid attention support | Enabled by model implementation | Supports Qwen3.6 long-context inference. | Tune context length based on KV cache capacity. |
|
||||
| Full decode ACLGraph | `--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}'` | Reduces operator dispatch overhead and stabilizes decode performance. | Recommended for decode-heavy serving. |
|
||||
| FlashComm1 | `--additional-config '{"enable_flashcomm1": true}'` | Reduces communication overhead in TP and high-concurrency scenarios. | May not help low-concurrency workloads. |
|
||||
| Shared expert overlap | `--additional-config '{"multistream_overlap_shared_expert": true}'` | Overlaps shared expert computation in MoE workloads. | Recommended for throughput scenarios. |
|
||||
| Asynchronous scheduling | `--async-scheduling` | Improves high-concurrency throughput by using non-blocking scheduling. | Disable it and compare if the workload is latency-sensitive. |
|
||||
| Prefix caching | `--enable-prefix-caching` | Improves repeated-prefix workloads. | Monitor HBM usage for long-context workloads. |
|
||||
| Qwen3.6 MTP speculative decoding | `--speculative-config '{"method": "qwen3_5_mtp", "num_speculative_tokens": 3, "enforce_eager": true}'` | Can improve decode throughput when stable and accepted tokens are high. | Validate stability, TTFT, TPOT, and throughput for your workload. |
|
||||
|
||||
## 10 FAQ
|
||||
|
||||
For common environment, installation, and general parameter issues, refer to [FAQs](../../faqs.md). This section only covers model-specific issues for Qwen3.6-35B-A3B.
|
||||
|
||||
### Q1: Why does the service report OOM during startup or soon after accepting requests?
|
||||
|
||||
**Phenomenon:** The service fails during profile run, or it starts successfully but reports OOM when real traffic arrives.
|
||||
|
||||
**Cause:** Qwen3.6 long-context serving consumes a large KV cache. Large `--max-model-len`, large `--max-num-seqs`, large `--max-num-batched-tokens`, or high `--gpu-memory-utilization` can leave insufficient HBM headroom.
|
||||
|
||||
**Solution:** Use the W8A8 model with `--quantization ascend` when possible, lower `--max-model-len`, lower `--max-num-seqs`, lower `--max-num-batched-tokens`, or reduce `--gpu-memory-utilization`. Keep `PYTORCH_NPU_ALLOC_CONF=expandable_segments:True`.
|
||||
|
||||
### Q2: Why does enabling prefix caching not improve performance?
|
||||
|
||||
**Phenomenon:** Prefix caching is enabled, but throughput or latency does not improve.
|
||||
|
||||
**Cause:** Prefix caching only helps when requests share reusable prefixes. For random prompts or low cache hit rates, it may add memory pressure without visible gains.
|
||||
|
||||
**Solution:** Enable prefix caching for repeated-prefix workloads. For random benchmark datasets or memory-constrained long-context workloads, compare with `--no-enable-prefix-caching`.
|
||||
|
||||
### Q3: How should I tune async scheduling for Qwen3.6?
|
||||
|
||||
**Phenomenon:** Throughput improves in high-concurrency scenarios, but some latency-sensitive workloads may not benefit.
|
||||
|
||||
**Cause:** Asynchronous scheduling reduces blocking overhead, but the benefit depends on concurrency, prompt/output length, and graph capture shape.
|
||||
|
||||
**Solution:** Use `--async-scheduling` for high-throughput serving. For low-latency serving, compare TTFT and TPOT with and without this option.
|
||||
193
docs/source/tutorials/models/gpt-oss-120b.md
Normal file
193
docs/source/tutorials/models/gpt-oss-120b.md
Normal file
@@ -0,0 +1,193 @@
|
||||
# gpt-oss-120b
|
||||
|
||||
## Introduction
|
||||
|
||||
gpt-oss-120b and gpt-oss-20b are two open-weight reasoning models that push the frontier of accuracy and inference cost. The models use an efficient mixture-of-experts transformer architecture and are trained using large-scale distillation and reinforcement learning. We optimize the models to have strong agentic capabilities (deep research browsing, python tool use, and support for developer-provided functions), all while using a rendered chat format that enables clear instruction following and role delineation. Both models achieve strong results on benchmarks ranging from mathematics, coding, and safety. We release the model weights, inference implementations, tool environments, and tokenizers under an Apache 2.0 license to enable broad use and further research.
|
||||
|
||||
## Supported Features
|
||||
|
||||
Refer to [supported features](../../user_guide/support_matrix/supported_models.md) to get the model's supported feature matrix.
|
||||
|
||||
Refer to [feature guide](../../user_guide/feature_guide/index.md) to get the feature's configuration.
|
||||
|
||||
## Environment Preparation
|
||||
|
||||
### Model Weight
|
||||
|
||||
- `gpt-oss-120b`(bf16 version): require 1 Atlas 800 A3 (64GB × 16) nodes or 1 Atlas 800 A2 (64GB × 8) nodes. [Download model weight](https://huggingface.co/unsloth/gpt-oss-120b-BF16)
|
||||
|
||||
### Installation
|
||||
|
||||
You can use our official docker image for supporting gpt-oss-120b-bf16 models.
|
||||
Currently, we provide the all-in-one images. [Download images](https://quay.io/repository/ascend/vllm-ascend?tab=tags)
|
||||
|
||||
#### Docker Pull (by tag)
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
docker pull quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
|
||||
```
|
||||
|
||||
#### Docker run
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
# Update --device according to your device (Atlas A2: /dev/davinci[0-7] Atlas A3:/dev/davinci[0-15]).
|
||||
# Update the vllm-ascend image according to your environment.
|
||||
# Note you should download the weight to /root/.cache in advance.
|
||||
# For Atlas A2 machines:
|
||||
# export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
# For Atlas A3 machines:
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-a3
|
||||
docker run --rm \
|
||||
--name vllm-ascend-env \
|
||||
--shm-size=1g \
|
||||
--net=host \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
The default workdir is `/workspace`, vLLM and vLLM Ascend code are placed in `/vllm-workspace` and installed in [development mode](https://setuptools.pypa.io/en/latest/userguide/development_mode.html) (`pip install -e`) to help developers apply changes immediately without requiring a new installation.
|
||||
|
||||
In addition, if you don't want to use the docker image as above, you can also build all from source:
|
||||
|
||||
- Install `vllm-ascend` from source, refer to [installation](../../installation.md).
|
||||
|
||||
## Deployment
|
||||
|
||||
### Troubleshooting
|
||||
|
||||
Run into
|
||||
|
||||
"openai_harmony.HarmonyError: error downloading or loading vocab file: failed to download or load vocab error"
|
||||
|
||||
Solution: This is caused by a bug in openai_harmony code. This can be worked around by downloading the tiktoken encoding files in advance and setting the TIKTOKEN_ENCODINGS_BASE environment variable. See this [GitHub](https://github.com/openai/harmony/issues/35) issue for more information.
|
||||
|
||||
```bash
|
||||
mkdir -p tiktoken_encodings
|
||||
wget -O tiktoken_encodings/o200k_base.tiktoken "https://openaipublic.blob.core.windows.net/encodings/o200k_base.tiktoken"
|
||||
wget -O tiktoken_encodings/cl100k_base.tiktoken "https://openaipublic.blob.core.windows.net/encodings/cl100k_base.tiktoken"
|
||||
export TIKTOKEN_ENCODINGS_BASE=${PWD}/tiktoken_encodings
|
||||
```
|
||||
|
||||
### Single-node Deployment
|
||||
|
||||
`gpt-oss-120b` can both be deployed on 1 Atlas 800 A3(64GB × 16), 1 Atlas 800 A2(64GB × 8).
|
||||
|
||||
Run the following script to execute online inference.
|
||||
|
||||
```shell
|
||||
#!/bin/sh
|
||||
# Load model from ModelScope to speed up download
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
# To reduce memory fragmentation and avoid out of memory
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export HCCL_BUFFSIZE=512
|
||||
export NPU_MEMORY_FRACTION=0.95
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3
|
||||
export OMP_PROC_BIND=false
|
||||
export VLLM_USE_V1=1
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export OMP_NUM_THREADS=1
|
||||
export TIKTOKEN_ENCODINGS_BASE=${PWD}/tiktoken_encodings
|
||||
|
||||
vllm serve unsloth/gpt-oss-120b-BF16 \
|
||||
--served-model-name gpt-oss-120b-bf16 \
|
||||
--port 8000 \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 4 \
|
||||
--gpu-memory-utilization 0.90 \
|
||||
--tensor-parallel-size 4 \
|
||||
--max-model-len 4096 \
|
||||
--max-num-batched-tokens 4096 \
|
||||
--enable-expert-parallel \
|
||||
--compilation_config '{"cudagraph_mode": "FULL_DECODE_ONLY", "cudagraph_capture_sizes":[1,2,3,4]}'
|
||||
```
|
||||
|
||||
The parameters are explained as follows:
|
||||
|
||||
- `--tensor-parallel-size` are common settings for tensor parallelism (TP) sizes.
|
||||
- `--max-model-len` represents the context length, which is the maximum value of the input plus output for a single request.
|
||||
- `--max-num-seqs` indicates the maximum number of requests that each DP group is allowed to process. If the number of requests sent to the service exceeds this limit, the excess requests will remain in a waiting state and will not be scheduled. Note that the time spent in the waiting state is also counted in metrics such as TTFT and TPOT. Therefore, when testing performance, it is generally recommended that `--max-num-seqs` * `--data-parallel-size` >= the actual total concurrency.
|
||||
- `--max-num-batched-tokens` represents the maximum number of tokens that the model can process in a single step. Currently, vLLM v1 scheduling enables ChunkPrefill/SplitFuse by default, which means:
|
||||
- (1) If the input length of a request is greater than `--max-num-batched-tokens`, it will be divided into multiple rounds of computation according to `--max-num-batched-tokens`;
|
||||
- (2) Decode requests are prioritized for scheduling, and prefill requests are scheduled only if there is available capacity.
|
||||
- Generally, if `--max-num-batched-tokens` is set to a larger value, the overall latency will be lower, but the pressure on GPU memory (activation value usage) will be greater.
|
||||
- `--gpu-memory-utilization` represents the proportion of HBM that vLLM will use for actual inference. Its essential function is to calculate the available kv_cache size. During the warm-up phase (referred to as profile run in vLLM), vLLM records the peak GPU memory usage during an inference process with an input size of `--max-num-batched-tokens`. The available kv_cache size is then calculated as: `--gpu-memory-utilization` * HBM size - peak GPU memory usage. Therefore, the larger the value of `--gpu-memory-utilization`, the more kv_cache can be used. However, since the GPU memory usage during the warm-up phase may differ from that during actual inference (e.g., due to uneven EP load), setting `--gpu-memory-utilization` too high may lead to OOM (Out of Memory) issues during actual inference. The default value is `0.9`.
|
||||
- `--compilation-config` contains configurations related to the aclgraph graph mode. The most significant configurations are "cudagraph_mode" and "cudagraph_capture_sizes", which have the following meanings:
|
||||
"cudagraph_mode": represents the specific graph mode. Currently, "PIECEWISE" and "FULL_DECODE_ONLY" are supported. The graph mode is mainly used to reduce the cost of operator dispatch. Currently, "FULL_DECODE_ONLY" is recommended.
|
||||
- "cudagraph_capture_sizes": represents different levels of graph modes. The default value is [1, 2, 4, 8, 16, 24, 32, 40,..., `--max-num-seqs`]. In the graph mode, the input for graphs at different levels is fixed, and inputs between levels are automatically padded to the next level. Currently, the default setting is recommended. Only in some scenarios is it necessary to set this separately to achieve optimal performance.
|
||||
|
||||
## Functional Verification
|
||||
|
||||
Once your server is started, you can query the model with input prompts:
|
||||
|
||||
```shell
|
||||
curl http://localhost:8000/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "gpt-oss-120b-bf16",
|
||||
"messages": [{"role":"user", "content":"who are you"}]
|
||||
}'
|
||||
```
|
||||
|
||||
## Accuracy Evaluation
|
||||
|
||||
Here are two accuracy evaluation methods.
|
||||
|
||||
### Using AISBench
|
||||
|
||||
1. Refer to [Using AISBench](../../developer_guide/evaluation/using_ais_bench.md) for details.
|
||||
|
||||
2. After execution, you can get the result, here is the result of `gpt-oss-120b-bf16` for reference only.
|
||||
|
||||
| dataset | version | metric | mode | vllm-api-general-chat |
|
||||
|----- | ----- | ----- | ----- | -----|
|
||||
| mmlu | - | accuracy | gen | 89.50 |
|
||||
|
||||
## Performance
|
||||
|
||||
### Using AISBench
|
||||
|
||||
Refer to [Using AISBench for performance evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation) for details.
|
||||
|
||||
### Using vLLM Benchmark
|
||||
|
||||
Run performance evaluation of `gpt-oss-120b-BF16` as an example.
|
||||
|
||||
Refer to [vllm benchmark](https://docs.vllm.ai/en/latest/benchmarking/) for more details.
|
||||
|
||||
There are three `vllm bench` subcommands:
|
||||
|
||||
- `latency`: Benchmark the latency of a single batch of requests.
|
||||
- `serve`: Benchmark the online serving throughput.
|
||||
- `throughput`: Benchmark offline inference throughput.
|
||||
|
||||
Take the `serve` as an example. Run the code as follows.
|
||||
|
||||
```shell
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
vllm bench serve --model unsloth/gpt-oss-120b-BF16 --dataset-name random --random-input 200 --num-prompts 200 --request-rate 1 --save-result --result-dir ./
|
||||
```
|
||||
|
||||
After about several minutes, you can get the performance evaluation result.
|
||||
49
docs/source/tutorials/models/index.md
Normal file
49
docs/source/tutorials/models/index.md
Normal file
@@ -0,0 +1,49 @@
|
||||
# Model Tutorials
|
||||
|
||||
This section provides tutorials for different models of vLLM Ascend.
|
||||
|
||||
:::{toctree}
|
||||
:caption: Model Tutorials
|
||||
:maxdepth: 1
|
||||
Qwen3-Dense.md
|
||||
Qwen-VL-Dense.md
|
||||
Qwen3-30B-A3B.md
|
||||
Qwen3-235B-A22B.md
|
||||
Qwen3-VL-30B-A3B-Instruct.md
|
||||
Qwen3-VL-235B-A22B-Instruct.md
|
||||
Qwen3-Coder-30B-A3B.md
|
||||
Qwen3-Embedding.md
|
||||
Qwen3-VL-Embedding.md
|
||||
Qwen3-Reranker.md
|
||||
Qwen3-VL-Reranker.md
|
||||
Qwen3-Next.md
|
||||
Qwen3-Omni-30B-A3B-Thinking.md
|
||||
Qwen3.5-27B-Qwen3.6-27B.md
|
||||
Qwen3.5-Dense.md
|
||||
Qwen3.5-397B-A17B.md
|
||||
Qwen3.6-35B-A3B.md
|
||||
DeepSeek-V3.1.md
|
||||
DeepSeek-V3.2.md
|
||||
DeepSeek-V4-Flash.md
|
||||
DeepSeek-V4-Pro.md
|
||||
DeepSeek-R1.md
|
||||
DeepSeekOCR2.md
|
||||
GLM4.x.md
|
||||
GLM5.md
|
||||
GLM5.2.md
|
||||
Kimi-K2-Thinking.md
|
||||
Kimi-K2.5.md
|
||||
Kimi-K2.6.md
|
||||
Kimi-K3.md
|
||||
PaddleOCR-VL.md
|
||||
MiniMax-M2.md
|
||||
Hunyuan-A13B-Instruct.md
|
||||
Hy3-preview.md
|
||||
Minitron-8B-Base.md
|
||||
LLaVA-OneVision-Qwen2-0.5B-OV.md
|
||||
gpt-oss-120b.md
|
||||
Mixtral-8x7B-Instruct-v0.1.md
|
||||
Qwen3-ASR-1.7B.md
|
||||
Qwen2.5-Math-RM-72B.md
|
||||
InternVL3.5
|
||||
:::
|
||||
Reference in New Issue
Block a user