295
docs/source/_templates/Model-Deployment-Tutorial-Template.md
Normal file
295
docs/source/_templates/Model-Deployment-Tutorial-Template.md
Normal file
@@ -0,0 +1,295 @@
|
||||
# Technical Documentation Template for Deployment Tutorials Based on the XXX Model
|
||||
|
||||
<p align="center">
|
||||
<a href="Model-Deployment-Tutorial-Template.md"><b>English</b></a> | <a href="Model-Deployment-Tutorial-Template.zh.md"><b>中文</b></a>
|
||||
</p>
|
||||
|
||||
This template is based on deployment tutorials for models such as DeepSeek-V3.2 and Qwen-VL-Dense, and is intended to serve as a reference for technical documentation writing. Users can systematically construct relevant technical documentation by following the guidelines provided in this template.
|
||||
|
||||
## 1 Introduction
|
||||
|
||||
**Content Writing Requirements:**
|
||||
|
||||
- Provide a one-sentence description of the model's basic architecture, core features, and primary application scenarios.
|
||||
- Provide a one-sentence description of the document's purpose and the objectives to be achieved.
|
||||
- Specify the version of vLLM-Ascend used in the document and the version support status of the model.
|
||||
|
||||
**Example 1: Model Introduction**
|
||||
|
||||
DeepSeek-V3.2 is a sparse attention model. Its core architecture is similar to that of DeepSeek-V3.1, but it employs a sparse attention mechanism, aiming to explore and validate optimization solutions for training and inference efficiency in long-context scenarios.
|
||||
|
||||
**Example 2: Document Purpose**
|
||||
|
||||
This document will demonstrate the primary validation steps for the model, including supported features, feature configuration, environment preparation, single-node and multi-node deployment, as well as accuracy and performance evaluation.
|
||||
|
||||
**Example 3: Version Information**
|
||||
|
||||
This document is validated and written based on **vLLM-Ascend v0.13.0**. The current model (XXX) is fully supported in this version, and all **v0.13.0 and later versions** can run stably. To use the latest features (e.g., PD separation, MTP), it is recommended to use the latest release candidate or official version.
|
||||
|
||||
## 2 Supported Features
|
||||
|
||||
This section introduces the features supported by the model, including supported hardware, quantization methods, data parallelism, long-sequence features, etc.
|
||||
|
||||
**Content Writing Requirements:**
|
||||
|
||||
- Present the support status of models and features in a table format.
|
||||
- Or provide cross-references with jump links (recommended).
|
||||
|
||||
**Example 1: Feature Support List**
|
||||
|
||||
| Model Name | Support Status | Remarks | BF16 | Supported Hardware | W8A8 | Chunked Prefill | Automatic Prefix Caching | LoRA | Speculative Decoding | Asynchronous Scheduling | Tensor Parallelism | Pipeline Parallelism | Expert Parallelism | Data Parallelism | Prefill-Decode Separation | Segmented ACL Graph Execution | Full ACL Graph Execution | Max Model Length | MLP Weight Prefetch | Documentation |
|
||||
| ------ | ---------- | ------ | ------ | ---------- | ------ | ------------ | -------------- | ------ | ---------- | ---------- | ---------- | ------------ | ---------- | ---------- | ------------------- | ----------- | ----------- | ------------- | ------------- | ---------- |
|
||||
| DeepSeek V3/3.1 | ✅ | | ✅ | Atlas 800I A2:<br>Minimum card requirement: xx | ✅ | ✅ | ✅ | | ✅ | | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 240k | | [DeepSeek-V3.1](../../tutorials/models/DeepSeek-V3.1.md) |
|
||||
| DeepSeek V3.2 | ✅ | | ✅ | Atlas 800I A2:<br>Minimum card requirement: xx | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 160k | ✅ | [DeepSeek-V3.2](../../tutorials/models/DeepSeek-V3.2.md) |
|
||||
| Qwen3 | ✅ | | ✅ | Atlas 800I A2:<br>Minimum card requirement: xx | ✅ | ✅ | ✅ | | | ✅ | ✅ | | | ✅ | | ✅ | ✅ | 128k | ✅ | [Qwen3-Dense](../../tutorials/models/Qwen3-Dense.md) |
|
||||
|
||||
>**Note**: This is a simplified example. Please refer to the complete feature matrix for the full table.
|
||||
|
||||
**Example 2: Reference Citation**
|
||||
|
||||
Please refer to the [Supported Features List](../user_guide/support_matrix/supported_models.md) for the model support matrix.
|
||||
|
||||
Please refer to the [Feature Guide](../user_guide/feature_guide/index.md) for feature configuration information.
|
||||
|
||||
## 3 Prerequisites
|
||||
|
||||
### 3.1 Model Weight
|
||||
|
||||
**Content Writing Requirements:** Describe the hardware resources, software environment, and model files required for deployment.
|
||||
|
||||
**Example:**
|
||||
|
||||
- `DeepSeek-V3.2-Exp-W8A8` (Quantized version): requires 1 Atlas 800 A3 (64G × 16) node or 2 Atlas 800 A2 (64G × 8) nodes. [Model Weight](https://www.modelscope.cn/models/vllm-ascend/DeepSeek-V3.2-Exp-W8A8)
|
||||
- `DeepSeek-V3.2-w8a8` (Quantized version): requires 1 Atlas 800 A3 (64G × 16) node or 2 Atlas 800 A2 (64G × 8) nodes. [Model Weight](https://www.modelscope.cn/models/vllm-ascend/DeepSeek-V3.2-W8A8/)
|
||||
|
||||
It is recommended to download the model weight to a shared directory across multiple nodes.
|
||||
|
||||
### 3.2 Verify Multi-node Communication (Optional)
|
||||
|
||||
**Example:**
|
||||
|
||||
If multi-node deployment is required, please follow the [Verify Multi-node Communication Environment](../installation.md#verify-multi-node-communication) guide for communication verification.
|
||||
|
||||
## 4 Installation
|
||||
|
||||
**Content Writing Requirements:**
|
||||
|
||||
- Provide specific installation steps and commands (parameters should be explained with meaning, value range, units, etc.).
|
||||
- **Version Number Writing Specification:** Prefer using placeholders (values are centrally configured). If a fixed value is used and it differs from the documented validation version, a comment MUST be added stating: "Please replace with your actual version."
|
||||
- Provide verification commands and expected status: guide users to check the installation result by executing commands (e.g., docker ps), specifying success criteria such as status codes or output characteristics.
|
||||
- When content involves multiple hardware series (e.g., A3/A2), the `tab-set` markup syntax must be used to present them in separate tabs,and the tabs should be arranged with the newest models first.
|
||||
|
||||
### 4.1 Docker Image Installation
|
||||
|
||||
**Example:**
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: install
|
||||
|
||||
::::{tab-item} A3 series
|
||||
:sync: A3
|
||||
|
||||
```bash
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-a3
|
||||
docker run ...
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} A2 series
|
||||
:sync: A2
|
||||
|
||||
```bash
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
docker run ...
|
||||
```
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
### 4.2 Source Code Installation
|
||||
|
||||
**Example:** Omitted
|
||||
|
||||
## 5 Online Service Deployment
|
||||
|
||||
### 5.1 Single-Node Online Deployment
|
||||
|
||||
**Content Writing Requirements:**
|
||||
|
||||
- Describe the architectural characteristics and applicable scenarios of single-node deployment.
|
||||
- Provide startup command templates and key parameter descriptions.
|
||||
- Provide service verification methods (e.g., curl commands) and expected results, specifying success indicators (e.g., 200 OK).
|
||||
- When content involves multiple hardware series (e.g., A3/A2), the `tab-set` markup syntax must be used to present them in separate tabs,and the tabs should be arranged with the newest models first.
|
||||
- Below the startup command, provide guidance on common issues; if already described in the public FAQ, a direct link may be provided.
|
||||
|
||||
**Example:**
|
||||
|
||||
Single-node deployment completes both Prefill and Decode within the same node, suitable for XXX scenarios.
|
||||
|
||||
Startup Command:
|
||||
|
||||
```bash
|
||||
# Omitted
|
||||
```
|
||||
|
||||
Common Issues Tip: If you encounter XXX issues, please refer to the [Public FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html) for troubleshooting.
|
||||
|
||||
Service Verification:
|
||||
|
||||
```bash
|
||||
# Omitted
|
||||
```
|
||||
|
||||
Expected Result: Omitted (fill in according to actual output).
|
||||
|
||||
### 5.2 Multi-Node PD Separation Deployment
|
||||
|
||||
**Content Writing Requirements:**
|
||||
|
||||
- Describe the principles of PD separation architecture and applicable scenarios.
|
||||
- Provide startup procedures, key configurations, and **deployment verification instructions**, and indicate performance metrics.
|
||||
- Below the startup command, provide guidance on common issues; if already described in the public FAQ, a direct link may be provided.
|
||||
- When content involves multiple hardware series (e.g., A3/A2), the `tab-set` markup syntax must be used to present them in separate tabs,and the tabs should be arranged with the newest models first.
|
||||
|
||||
**Example:** Omitted
|
||||
|
||||
### 5.3 Special Deployment Modes (Optional)
|
||||
|
||||
**Content Writing Requirements:**
|
||||
|
||||
- If the model features non‑standard deployment modes (e.g., offline batch processing for embedding models, low‑latency online serving for reranker models), the corresponding deployment solutions must be explicitly documented.
|
||||
- Section 5.1 and 5.2 above can be referenced for extension.
|
||||
|
||||
## 6 Functional Verification
|
||||
|
||||
**Content Writing Requirements:**
|
||||
|
||||
- Guide users on how to test the basic functionality of the model through simple interface calls after the service is started.
|
||||
- Provide expected results, specifying success indicators (e.g., HTTP 200, JSON response containing a choices field).
|
||||
|
||||
**Example:**
|
||||
|
||||
After the service is started, the model can be invoked by sending a prompt:
|
||||
|
||||
```shell
|
||||
curl http://<node0_ip>:<port>/v1/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "deepseek_v3.2",
|
||||
"prompt": "The future of AI is",
|
||||
"max_tokens": 50,
|
||||
"temperature": 0
|
||||
}'
|
||||
```
|
||||
|
||||
Expected Result: Omitted (fill in according to actual output).
|
||||
|
||||
## 7 Accuracy Evaluation
|
||||
|
||||
**Content Writing Requirements:** Introduce standardized methods and tools for evaluating model output quality (accuracy). Two accuracy evaluation methods are provided below as examples; alternatively, provide direct links to existing documentation.
|
||||
|
||||
### Using AISBench
|
||||
|
||||
For details, please refer to [Using AISBench](../developer_guide/evaluation/using_ais_bench.md).
|
||||
|
||||
### Using Language Model Evaluation Harness
|
||||
|
||||
Using the `gsm8k` dataset as an example test dataset, run the accuracy evaluation for `DeepSeek-V3.2-W8A8` in online mode.
|
||||
|
||||
1. For `lm_eval` installation, please refer to [Using lm_eval](../developer_guide/evaluation/using_lm_eval.md).
|
||||
2. Run `lm_eval` to execute the accuracy evaluation.
|
||||
|
||||
```shell
|
||||
lm_eval \
|
||||
--model local-completions \
|
||||
--model_args model=/root/.cache/Eco-Tech/DeepSeek-V3.2-w8a8-mtp-QuaRot,base_url=http://127.0.0.1:8000/v1/completions,tokenized_requests=False,trust_remote_code=True \
|
||||
--tasks gsm8k \
|
||||
--output_path ./
|
||||
```
|
||||
|
||||
## 8 Performance Evaluation
|
||||
|
||||
Omitted. Requirements are the same as for Accuracy Evaluation.
|
||||
|
||||
## 9 Performance Tuning
|
||||
|
||||
### 9.1 Recommended Configurations
|
||||
|
||||
**Content Writing Requirements:**
|
||||
|
||||
Provide recommended configurations for three typical scenarios (long context, low latency, high throughput). Clearly state that the configurations are not globally optimal and guide users to perform tuning based on their actual circumstances.
|
||||
|
||||
**Example:**
|
||||
|
||||
> **Note**: The following configurations are validated in specific test environments and are for reference only. The optimal configuration depends on factors such as maximum input/output length, prefix cache hit rate, precision requirements, and deployment machine ratios. It is recommended to refer to Section 9.2 for tuning based on actual conditions.
|
||||
|
||||
#### Table 1: Scenario Overview
|
||||
|
||||
| Scenario | Deployment Mode | *Total NPUs | Weight Version | Key Considerations |
|
||||
|----------|----------------|-------------|----------------|------------------------|
|
||||
| High Throughput<br>(32K context → 1K output) | 1P1D deployment | 16 (A3) | glm5.1w4a8 | For short-sequence high throughput, try adjusting xxx parameters |
|
||||
| Long Context | | | | |
|
||||
| Low Latency | | | | |
|
||||
|
||||
> `*Total NPUs` indicates the total number of NPUs used across all nodes.
|
||||
|
||||
#### Table 2: Detailed Node Configuration
|
||||
|
||||
| Scenario | Configuration | NPUs | TP | DP | Max Num Seqs | Max Num Batched Tokens | Max Model Len | MTP Speculation Num | FUSED_MC2 | EP Switch | FC+CP Switch | Async Scheduling |
|
||||
|----------|---------------|-------|----|----|----|-------------|--------------------|---------------------|-----------|-----------|--------------|------------------|
|
||||
| High Throughput (32K→1K) | Server-P Node / Single Machine | 8 | 8 | 2 | 32 | 4096 | 30k | 3 | Off | On | On | On |
|
||||
| High Throughput (32K→1K) | Server-D Node | 8 | 2 | 8 | 8 | 4096 | 30k | 12 | Off | On | Off | On |
|
||||
| Long Context | Server-P Node / Single Machine | | | | | | | | | | | |
|
||||
| Long Context | Server-D Node | | | | | | | | | | | |
|
||||
| Low Latency | Server-P Node / Single Machine | | | | | | | | | | | |
|
||||
| Low Latency | Server-D Node | | | | | | | | | | | |
|
||||
|
||||
> For complete startup commands and parameter descriptions, please refer to the deployment examples in Chapter 5.
|
||||
|
||||
### 9.2 Tuning Guidelines
|
||||
|
||||
#### 9.2.1 General Tuning Reference
|
||||
|
||||
**Content Writing Requirements:**
|
||||
|
||||
If no special tuning is involved, directly provide a feature combination table and a link to the public performance tuning documentation.
|
||||
|
||||
**Example:**
|
||||
|
||||
Please refer to the [Public Performance Tuning Documentation](../../developer_guide/performance_and_debug/optimization_and_tuning.md) for tuning methods.
|
||||
Please refer to the [Feature Guide](../../user_guide/support_matrix/feature_matrix.md) for detailed feature descriptions.
|
||||
|
||||
#### 9.2.2 Model-Specific Optimizations (Optional)
|
||||
|
||||
**Documentation Requirements:**
|
||||
|
||||
If the model has specific optimizations, summarize the key optimization techniques and tuning experience for this model.
|
||||
|
||||
**Example:**
|
||||
|
||||
#### Optimizations Enabled by Default
|
||||
|
||||
The following optimizations are enabled by default and require no additional configuration:
|
||||
|
||||
| Optimization Technique | Technical Principle | Performance Benefit |
|
||||
| --------- | --------- | --------- |
|
||||
| Rope Optimization | The cos_sin_cache and indexing operations of positional encoding are executed only in the first layer, and subsequent layers reuse them directly | Reduces redundant computation during the decoding phase, accelerating inference |
|
||||
| AddRMSNormQuant Fusion | Merges address-wise multi-scale normalization and quantization operations into a single operator | Optimizes memory access patterns, improving computational efficiency |
|
||||
| Zero-like Elimination | Removes unnecessary zero-tensor operations in Attention forward pass | Reduces memory footprint, improves matrix operation efficiency |
|
||||
| FullGraph Optimization | Captures and replays the entire decoding graph at once using `compilation_config={"cudagraph_mode":"FULL_DECODE_ONLY"}` | Significantly reduces scheduling latency, stabilizes multi-device performance |
|
||||
|
||||
#### Optimizations That Require Explicit Enabling
|
||||
|
||||
| Optimization Technique | Applicable Scenarios | Enablement Method | Technical Principle | Precautions |
|
||||
| --------------------- | -------------------- | ----------------- | ------------------- | ----------- |
|
||||
| FlashComm_v1 | High-concurrency, Tensor Parallelism (TP) scenarios | `export VLLM_ASCEND_ENABLE_FLASHCOMM1=1` | Decomposes traditional Allreduce into Reduce-Scatter and All-Gather, reducing RMSNorm computation dimensions | Threshold protection: Only takes effect when the actual number of tokens exceeds the threshold to avoid performance degradation in low-concurrency scenarios |
|
||||
| Matmul-ReduceScatter Fusion | Large-scale distributed environments | Automatically enabled after enabling FlashComm_v1 | Fuses matrix multiplication and Reduce-Scatter operations to achieve pipelined parallel processing | Same as FlashComm_v1, has threshold protection |
|
||||
| Weight Prefetch | MLP-intensive scenarios (Dense models) | `export VLLM_ASCEND_ENABLE_PREFETCH_MLP=1` | Utilizes vector computation time to prefetch MLP weights into L2 cache in advance | Requires coordination with prefetch buffer size adjustment |
|
||||
|
||||
## 10 FAQ
|
||||
|
||||
**Content Writing Requirements:**
|
||||
|
||||
- Add a note at the beginning of the section: For common environment, installation, and general parameter issues, please refer to the [Public FAQs](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html); this chapter only covers model-specific issues.
|
||||
- For **model-specific issues**, provide the following elements: problem phenomenon description, cause analysis, and solution measures.
|
||||
296
docs/source/_templates/Model-Deployment-Tutorial-Template.zh.md
Normal file
296
docs/source/_templates/Model-Deployment-Tutorial-Template.zh.md
Normal file
@@ -0,0 +1,296 @@
|
||||
# 基于XXX模型部署教程的技术文档模板
|
||||
|
||||
<p align="center">
|
||||
<a href="Model-Deployment-Tutorial-Template.md"><b>English</b></a> | <a href="Model-Deployment-Tutorial-Template.zh.md"><b>中文</b></a>
|
||||
</p>
|
||||
|
||||
本模板基于DeepSeek-V3.2、Qwen-VL-Dense等部署教程,旨在为技术文档撰写提供参考。使用者可遵循模板指引,系统性完成相关技术文档的构建工作。
|
||||
|
||||
## 1 简介
|
||||
|
||||
**资料写作要求:**
|
||||
|
||||
- 一句话介绍模型的基本架构、核心特性及主要应用场景。
|
||||
- 一句话写清楚文档要干什么,要达成的目的。
|
||||
- 说明文档使用的vLLM-Ascend版本及模型的版本支持情况。
|
||||
|
||||
**示例1:模型介绍**
|
||||
|
||||
DeepSeek-V3.2 是一种稀疏注意力模型。其主要架构与 DeepSeek-V3.1 类似,但采用了稀疏注意力机制,旨在探索和验证在长上下文场景下训练和推理效率的优化方案。
|
||||
|
||||
**示例2:文档目的**
|
||||
|
||||
本文档将展示模型的主要验证步骤,包括支持的功能、功能配置、环境准备、单节点和多节点部署、准确性和性能评估。
|
||||
|
||||
**示例3:版本信息**
|
||||
|
||||
本文档基于 **vLLM-Ascend v0.13.0** 版本进行验证和编写。当前模型(XXX)在该版本中已完整支持,**v0.13.0 及更高版本**均可稳定运行。如需使用最新特性(如PD分离、MTP等),建议使用最新的候选版本或正式版本。
|
||||
|
||||
## 2 支持的特性
|
||||
|
||||
介绍该模型支持的特性,包括支持的硬件、量化方式、数据并行、长序列特性等。
|
||||
|
||||
**资料写作要求:**
|
||||
|
||||
- 采用表格形式,呈现模型和特性的支持情况。
|
||||
- 或提供可跳转的交叉引用(推荐)。
|
||||
|
||||
**示例1:特性支持列表**
|
||||
|
||||
| 模型名称 | 支持状态 | 备注 | BF16 | 支持的硬件 | W8A8 | 分块预填充 | 自动前缀缓存 | LoRA | 推测解码 | 异步调度 | 张量并行 | 流水线并行 | 专家并行 | 数据并行 | Prefill-Decode分离 | 分段式ACL图执行 | 整图ACL图执行 | 最大模型长度 | MLP权重预取 | 文档 |
|
||||
| ------ | ---------- | ------ | ------ | ---------- | ------ | ------------ | -------------- | ------ | ---------- | ---------- | ---------- | ------------ | ---------- | ---------- | ------------------- |----------- | ----------- | ------------- | ------------- | ---------- |
|
||||
| DeepSeek V3/3.1 | ✅ | | ✅ | Atlas 800I A2:<br>最低卡数要求为xx | ✅ | ✅ | ✅ | | ✅ | | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 240k | | [DeepSeek-V3.1](../../tutorials/models/DeepSeek-V3.1.md) |
|
||||
| DeepSeek V3.2 | ✅ | | ✅ | Atlas 800I A2:<br>最低卡数要求为xx | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 160k | ✅ | [DeepSeek-V3.2](../../tutorials/models/DeepSeek-V3.2.md)|
|
||||
| Qwen3 | ✅ | | ✅ | Atlas 800I A2:<br>最低卡数要求为xx | ✅ | ✅ | ✅ | | | ✅ | ✅ | | | ✅ | | ✅ | ✅ | 128k | ✅ | [Qwen3-Dense](../../tutorials/models/Qwen3-Dense.md) |
|
||||
|
||||
>**注意**:此为简化示例,完整表格请参考完整特性矩阵。
|
||||
|
||||
**示例2:引用**
|
||||
|
||||
请参考[支持的模型](../user_guide/support_matrix/supported_models.md),获取模型支持的功能矩阵。
|
||||
|
||||
请参考[特性指南](../user_guide/feature_guide/index.md)获取功能配置信息。
|
||||
|
||||
## 3 前置准备
|
||||
|
||||
### 3.1 模型权重
|
||||
|
||||
**资料写作要求:** 说明部署所需的硬件资源、软件环境和模型文件。
|
||||
|
||||
**示例:**
|
||||
|
||||
- `DeepSeek-V3.2-Exp-W8A8`(量化版):需要 1 台 Atlas 800 A3(64G × 16)节点或 2 台 Atlas 800 A2(64G × 8)节点。 [模型权重](https://www.modelscope.cn/models/vllm-ascend/DeepSeek-V3.2-Exp-W8A8)
|
||||
- `DeepSeek-V3.2-w8a8`(量化版):需要 1 台 Atlas 800 A3(64G × 16)节点或 2 台 Atlas 800 A2(64G × 8)节点。 [模型权重](https://www.modelscope.cn/models/vllm-ascend/DeepSeek-V3.2-W8A8/)
|
||||
|
||||
建议将模型权重下载至多节点共享目录。
|
||||
|
||||
### 3.2 验证多节点通信(可选)
|
||||
|
||||
**示例:**
|
||||
|
||||
若需部署多节点环境,请依据[验证多节点通信环境](../installation.md#verify-multi-node-communication)指南进行通信验证。
|
||||
|
||||
## 4 安装
|
||||
|
||||
**资料写作要求:**
|
||||
|
||||
- 提供具体的安装步骤与命令(参数需解释含义、取值范围、单位等)。
|
||||
- 版本号书写规范:优先使用占位符(值统一配置);若使用固定值且该值与文档验证版本不一致,须加注释“请按实际版本替换”。
|
||||
- 提供验证命令及预期状态:指导用户通过执行命令(如 docker ps)检查安装结果,说明成功时的状态码或输出特征。
|
||||
- 当涉及多硬件系列(如 A3/A2 系列)时,须使用`tab-set`标记语法进行分标签呈现,标签顺序按新机型优先排列。
|
||||
|
||||
### 4.1 Docker镜像安装
|
||||
|
||||
**示例:**
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: install
|
||||
|
||||
::::{tab-item} A3 series
|
||||
:sync: A3
|
||||
|
||||
```bash
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-a3
|
||||
docker run ...
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} A2 series
|
||||
:sync: A2
|
||||
|
||||
```bash
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
docker run ...
|
||||
```
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
### 4.2 源码安装
|
||||
|
||||
**示例:** 略
|
||||
|
||||
## 5 在线服务化部署(Online service deployment)
|
||||
|
||||
### 5.1 单机在线部署
|
||||
|
||||
**资料写作要求:**
|
||||
|
||||
- 说明单机部署的架构特点与适用场景
|
||||
- 提供启动命令模板和关键参数说明
|
||||
- 提供服务验证方法(如 curl 命令)及预期结果,说明成功特征(如 200 OK)。
|
||||
- 当涉及多硬件系列(如 A3/A2 系列)时,须使用`tab-set`标记语法进行分标签呈现,标签顺序按新机型优先排列。
|
||||
- 在启动命令下方提供常见问题指引,如公共FAQ中已有描述可直接链接呈现。
|
||||
|
||||
**示例:**
|
||||
|
||||
单机部署将Prefill与Decode在同一节点内完成,适用于XXX场景。
|
||||
|
||||
启动命令:
|
||||
|
||||
```bash
|
||||
# 略
|
||||
```
|
||||
|
||||
常见问题提示:如遇xxx问题,请参考[公共FAQ](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html)进行检查。
|
||||
|
||||
服务验证:
|
||||
|
||||
```bash
|
||||
# 略
|
||||
```
|
||||
|
||||
预期结果:略(按实际输出书写即可)。
|
||||
|
||||
### 5.2 多机PD分离部署
|
||||
|
||||
**资料写作要求:**
|
||||
|
||||
- 说明PD分离架构的原理与适用场景。
|
||||
- 提供启动流程、关键配置及**部署验证说明**,并注明性能指标。
|
||||
- 在启动命令下方提供常见问题指引,如公共FAQ中已有描述可直接链接呈现。
|
||||
- 当涉及多硬件系列(如 A3/A2 系列)时,须使用`tab-set`标记语法进行分标签呈现,标签顺序按新机型优先排列。
|
||||
|
||||
**示例:** 略
|
||||
|
||||
### 5.3 特殊部署形态(可选)
|
||||
|
||||
**资料写作要求:**
|
||||
|
||||
- 若模型存在非标准部署形态(如embedding模型的离线批处理、reranker模型的低延迟在线服务等),需在文档中明确体现对应部署方案。
|
||||
- 可参考本章5.1和5.2节进行扩展。
|
||||
|
||||
## 6 功能验证
|
||||
|
||||
**资料写作要求:**
|
||||
|
||||
- 指导用户如何在服务启动后,通过简单接口测试模型的基本功能是否正常。
|
||||
- 提供预期结果,说明成功特征(如 HTTP 200、返回包含 choices 字段的 JSON)。
|
||||
|
||||
**示例:**
|
||||
|
||||
服务启动后,即可通过发送提示词来调用模型:
|
||||
|
||||
```shell
|
||||
curl http://<node0_ip>:<port>/v1/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "deepseek_v3.2",
|
||||
"prompt": "The future of AI is",
|
||||
"max_tokens": 50,
|
||||
"temperature": 0
|
||||
}'
|
||||
```
|
||||
|
||||
预期结果:略(按实际输出书写即可)。
|
||||
|
||||
## 7 精度评估
|
||||
|
||||
**资料写作要求:** 介绍评估模型输出质量(精度)的标准化方法及工具,以下提供两种精度评估方法作为示例;或直接链接现有文档进行呈现。
|
||||
|
||||
### AISBench的使用
|
||||
|
||||
详情请参考[Using AISBench](../developer_guide/evaluation/using_ais_bench.md)。
|
||||
|
||||
### Language Model Evaluation Harness的使用
|
||||
|
||||
以`gsm8k`数据集作为测试数据集为例,在线模式下运行`DeepSeek-V3.2-W8A8`的精度评估。
|
||||
|
||||
1. `lm_eval`安装请参考[Using lm_eval](../developer_guide/evaluation/using_lm_eval.md)。
|
||||
2. 运行`lm_eval`执行精度评估。
|
||||
|
||||
```shell
|
||||
lm_eval \
|
||||
--model local-completions \
|
||||
--model_args model=/root/.cache/Eco-Tech/DeepSeek-V3.2-w8a8-mtp-QuaRot,base_url=http://127.0.0.1:8000/v1/completions,tokenized_requests=False,trust_remote_code=True \
|
||||
--tasks gsm8k \
|
||||
--output_path ./
|
||||
```
|
||||
|
||||
## 8 性能评估
|
||||
|
||||
略,要求同精度评估
|
||||
|
||||
## 9 性能调优
|
||||
|
||||
### 9.1推荐配置
|
||||
|
||||
**资料写作要求:**
|
||||
|
||||
提供模型在三个典型场景下的推荐配置(长序列、低时延、高吞吐),需明确说明配置的非全局最优性,并引导用户根据实际进行调优。
|
||||
|
||||
**示例:**
|
||||
|
||||
> **说明**:以下配置基于特定测试环境验证,仅作参考。实际最优配置取决于最大输入输出长度、前缀缓存命中率、精度要求、部署机器配比等因素,建议根据实际参考9.2章节进行调优。
|
||||
|
||||
#### 表1:场景概览
|
||||
|
||||
| 场景 | 部署形态 | *总卡数 | 权重版本 | 场景要点 |
|
||||
|------|------|---------|----------|----------|
|
||||
| 高吞吐<br>(32K推1K) | 1P1D部署 | 16(A3) | glm5.1w4a8 | 短序列高吞吐情况下,尝试调整xxx参数 |
|
||||
| 长序列 | | | | |
|
||||
| 低时延 | | | | |
|
||||
|
||||
> `*总卡数` 表示所有节点使用的 NPU 总数。
|
||||
|
||||
#### 表2:节点详细配置
|
||||
|
||||
| 场景 | 配置 | 卡数 | TP | DP | 最大序列数 | 最大批量Token数 | 最大上下文 | MTP投机数 | FUSED_MC2 | EP开关 | FC+CP开关 | 异步调度 |
|
||||
|------|------|------|----|----|----|------|----------|---------|---------------|--------|-------|------|
|
||||
| 高吞吐(32K推1K) | 服务端-P节点/单机 | 8 | 8 | 2 | 32 | 4096 | 30k | 3 | 关 | 开 | 开 | 开 |
|
||||
| 高吞吐(32K推1K) | 服务端-D节点 | 8 | 2 | 8 | 8 | 4096 | 30k | 12 | 关 | 开 | 关 | 开 |
|
||||
| 长序列 | 服务端-P节点/单机 | | | | | | | | | | | |
|
||||
| 长序列 | 服务端-D节点 | | | | | | | | | | | |
|
||||
| 低时延 | 服务端-P节点/单机 | | | | | | | | | | | |
|
||||
| 低时延 | 服务端-D节点 | | | | | | | | | | | |
|
||||
|
||||
> 完整启动命令及参数含义请参考第5章部署示例。
|
||||
|
||||
### 9.2 调优思路
|
||||
|
||||
#### 9.2.1 通用调优参考
|
||||
|
||||
**资料写作要求:**
|
||||
|
||||
若不涉及特殊调优可直接给出特性叠加表和公共性能调优文档链接供参考。
|
||||
|
||||
**示例:**
|
||||
|
||||
请参考[公共性能调优文档](../../developer_guide/performance_and_debug/optimization_and_tuning.md)获得调优方法。
|
||||
请参考[特性指南](../../user_guide/support_matrix/feature_matrix.md)获得详细特性说明。
|
||||
|
||||
#### 9.2.2 模型特有优化(可选)
|
||||
|
||||
**资料写作要求:**
|
||||
|
||||
若该模型存在特有优化,需总结针对该模型的关键优化技术和调参经验。
|
||||
|
||||
**示例:**
|
||||
|
||||
#### 默认启用的优化
|
||||
|
||||
以下优化默认启用,无需额外配置:
|
||||
|
||||
| 优化技术 | 技术原理 | 性能收益 |
|
||||
| --------- | --------- | --------- |
|
||||
| Rope优化 | 位置编码的cos_sin_cache及索引操作仅在第一层执行,后续层直接复用 | 减少解码阶段重复计算,加速推理 |
|
||||
| AddRMSNormQuant融合 | 将逐地址多尺度归一化与量化操作合并为单算子 | 优化内存访问模式,提升计算效率 |
|
||||
| Zero-like Elimination | 移除Attention前向中的非必要零张量操作 | 减少内存占用,提高矩阵运算效率 |
|
||||
| FullGraph优化 | 通过`compilation_config={"cudagraph_mode":"FULL_DECODE_ONLY"}`将整个解码图一次性捕获重放 | 显著降低调度延迟,稳定多设备性能 |
|
||||
|
||||
#### 需显式开启的优化
|
||||
|
||||
| 优化技术 | 适用场景 | 启用方式 | 技术原理 | 注意事项 |
|
||||
| --------- | --------- | --------- | --------- | --------- |
|
||||
| FlashComm_v1 | 大并发、张量并行(TP)场景 | `export VLLM_ASCEND_ENABLE_FLASHCOMM1=1` | 将传统Allreduce分解为Reduce-Scatter和All-Gather,减少RMSNorm计算维度 | 阈值保护:仅当实际token数超过阈值时生效,避免小并发场景性能倒退|
|
||||
| Matmul-ReduceScatter融合 | 大型分布式环境 | 启用FlashComm_v1后自动开启 | 将矩阵乘法与Reduce-Scatter操作融合,实现流水线并行处理 | 同FlashComm_v1,有阈值保护 |
|
||||
| 权重预取 | MLP密集型场景(Dense模型)| `export VLLM_ASCEND_ENABLE_PREFETCH_MLP=1` | 利用向量计算时间,提前将MLP权重加载到L2 Cache | 需配合预取缓冲区大小调整 |
|
||||
| 异步调度 | 大规模模型、高并发场景 | `--async-scheduling` | 非阻塞任务调度,提升并发处理能力 | 与FullGraph优化协同使用 |
|
||||
|
||||
## 10 FAQ
|
||||
|
||||
**资料写作要求:**
|
||||
|
||||
- 在章节开头添加说明:常见环境、安装、通用参数问题请参考[公共FAQs](https://docs.vllm.ai/projects/ascend/en/latest/faqs.html);本章仅收录本模型特有疑难问题。
|
||||
- 针对**本模型特有疑难问题** ,提供以下要素:问题现象描述、原因分析、解决措施。
|
||||
@@ -54,5 +54,5 @@
|
||||
</style>
|
||||
|
||||
<div class="notification-bar">
|
||||
<p>You are viewing the latest developer preview docs. <a href="https://vllm-ascend.readthedocs.io/en/v0.9.1-dev">Click here</a> to view docs for the latest stable release(v0.9.1).</p>
|
||||
<p>You are viewing the stable release (v0.23.0) documentation. <a href="https://docs.vllm.ai/projects/ascend/en/latest/">Click here</a> to view the latest developer preview documentation.</p>
|
||||
</div>
|
||||
Reference in New Issue
Block a user