@@ -1,6 +1,45 @@
|
||||
# Additional Configuration
|
||||
|
||||
additional configuration is a mechanism provided by vLLM to allow plugins to control inner behavior by their own. vLLM Ascend uses this mechanism to make the project more flexible.
|
||||
Additional configuration is a mechanism provided by vLLM to allow plugins to control internal behavior by themselves. VLLM Ascend uses this mechanism to make the project more flexible.
|
||||
|
||||
## Migration Guide
|
||||
|
||||
Starting from [PR #9064](https://github.com/vllm-project/vllm-ascend/pull/9064), vLLM Ascend is migrating **10 environment variables** to `--additional-config`.
|
||||
|
||||
### Important Notice
|
||||
|
||||
- **Current Support**: Both environment variables and `--additional-config` are supported during the transition period
|
||||
- **Recommendation**: Use `--additional-config` for new deployments and migrate existing configurations
|
||||
- **Future Plan**: Environment variables will be **removed** in a future release; only `--additional-config` will be supported
|
||||
|
||||
### Quick Reference
|
||||
|
||||
| Environment Variable | Config Key | Type Conversion |
|
||||
|---------------------|------------|-----------------|
|
||||
| `VLLM_ASCEND_BALANCE_SCHEDULING` | `enable_balance_scheduling` | `"1"` → `true`, `"0"` → `false` |
|
||||
| `VLLM_ASCEND_ENABLE_FLASHCOMM1` | `enable_flashcomm1` | `"1"` → `true`, `"0"` → `false` |
|
||||
| `VLLM_ASCEND_ENABLE_MATMUL_ALLREDUCE` | `enable_matmul_allreduce` | `"1"` → `true`, `"0"` → `false` |
|
||||
| `VLLM_ASCEND_FLASHCOMM2_PARALLEL_SIZE` | `enable_flashcomm2_parallel_size` | Integer (unchanged) |
|
||||
| `MSMONITOR_USE_DAEMON` | `msmonitor_use_daemon` | `"1"` → `true`, `"0"` → `false` |
|
||||
| `VLLM_ASCEND_ENABLE_MLAPO` | `enable_mlapo` | `"1"` → `true`, `"0"` → `false` |
|
||||
| `VLLM_ASCEND_ENABLE_NZ` | `weight_nz_mode` | Integer (unchanged, field name changed) |
|
||||
| `VLLM_ASCEND_ENABLE_FUSED_MC2` | `enable_fused_mc2` | Integer (unchanged) |
|
||||
| `VLLM_ASCEND_FUSION_OP_TRANSPOSE_KV_CACHE_BY_BLOCK` | `enable_transpose_kv_cache_by_block` | `"1"` → `true`, `"0"` → `false` |
|
||||
|
||||
### Example Migration
|
||||
|
||||
**Before (environment variable):**
|
||||
|
||||
```bash
|
||||
export VLLM_ASCEND_ENABLE_FLASHCOMM1=1
|
||||
vllm serve Qwen/Qwen3-8B
|
||||
```
|
||||
|
||||
**After (additional-config):**
|
||||
|
||||
```bash
|
||||
vllm serve Qwen/Qwen3-8B --additional-config='{"enable_flashcomm1": true}'
|
||||
```
|
||||
|
||||
## How to use
|
||||
|
||||
@@ -22,75 +61,156 @@ LLM(model="Qwen/Qwen3-8B", additional_config={"config_key":"config_value"})
|
||||
|
||||
### Configuration options
|
||||
|
||||
The following table lists the additional configuration options available in vLLM Ascend:
|
||||
The following table lists additional configuration options available in vLLM Ascend:
|
||||
|
||||
| Name | Type | Default | Description |
|
||||
|-------------------------------| ---- |------|-----------------------------------------------------------------------------------------------|
|
||||
| `torchair_graph_config` | dict | `{}` | The config options for torchair graph mode |
|
||||
| `ascend_scheduler_config` | dict | `{}` | The config options for ascend scheduler |
|
||||
| `refresh` | bool | `false` | Whether to refresh global ascend config content. This value is usually used by rlhf or ut/e2e test case. |
|
||||
| `expert_map_path` | str | `None` | When using expert load balancing for the MOE model, an expert map path needs to be passed in. |
|
||||
| `enable_prefetch` | bool | `False` | Whether to enable weight prefetch. |
|
||||
| `kv_cache_dtype` | str | `None` | When using the kv cache quantization method, kv cache dtype needs to be set, currently only int8 is supported. |
|
||||
| `enable_shared_expert_dp` | bool | `False` | When the shared expert in DP, it has better performance but consumes more memory. Currently only DeepSeek series models are supported to use. |
|
||||
| `lmhead_tensor_parallel_size` | int | `None` | The custom tensor parallel size of lmhead. |
|
||||
| `oproj_tensor_parallel_size` | int | `None` | The custom tensor parallel size of oproj. |
|
||||
| `multistream_overlap_shared_expert`| bool | `False` | Whether to enable multistream shared expert. This option only takes effects on moe models with shared experts. |
|
||||
| `dynamic_eplb` | bool | `False` | Whether to enable dynamic eplb |
|
||||
|`num_iterations_eplb_update`| int | `400` | Forward iterations when eplb would begin |
|
||||
|`gate_eplb`| bool | `False` | Whether to enale eplb only once. |
|
||||
|`num_wait_worker_iterations`| int | `30` | The forward iterations when eplb worker will finish cpu task. In our test default value 30 would cover most cases. |
|
||||
|`expert_map_record_path`| str | `None` | When dynamic eplb is completed, save the current expert load heatmap to the specified path. |
|
||||
|`init_redundancy_expert`| int | `0` |Specify redundant experts during initialization.|
|
||||
| Name | Type | Default | Description |
|
||||
|-------------------------------------|------|---------|-----------------------------------------------------------------------------------------------------------|
|
||||
| `xlite_graph_config` | dict | `{}` | Configuration options for Xlite graph mode |
|
||||
| `weight_prefetch_config` | dict | `{}` | Configuration options for weight prefetch |
|
||||
| `finegrained_tp_config` | dict | `{}` | Configuration options for module tensor parallelism |
|
||||
| `ascend_compilation_config` | dict | `{}` | Configuration options for ascend compilation |
|
||||
| `eplb_config` | dict | `{}` | Configuration options for eplb |
|
||||
| `refresh` | bool | `false` | Whether to refresh global Ascend configuration content. This is usually used by rlhf or ut/e2e test case. |
|
||||
| `dump_config` | dict | `None` | Inline msprobe dump configuration. vLLM-Ascend will materialize it to a temporary JSON file and pass that file to the debugger. |
|
||||
| `dump_config_path` | str | `None` | Configuration file path for msprobe dump (compatible legacy option). |
|
||||
| `enable_async_exponential` | bool | `False` | Whether to enable asynchronous exponential overlap. To enable asynchronous exponential, set this config to True. |
|
||||
| `enable_shared_expert_dp` | bool | `False` | When the expert is shared in DP, it delivers better performance but consumes more memory. |
|
||||
| `multistream_overlap_shared_expert` | bool | `False` | Whether to enable multi-stream shared expert. This option only takes effect on MoE models with shared experts. |
|
||||
| `multistream_overlap_gate` | bool | `False` | Whether to enable multi-stream overlap gate. This option only takes effect on MoE models with shared experts. |
|
||||
| `recompute_scheduler_enable` | bool | `False` | Whether to enable the recompute scheduler. **Only valid on PD-disaggregated D nodes** (`kv_role` is `kv_consumer`). **Do not enable on P nodes or in PD-mixed mode** (no `kv_transfer_config`, `kv_role` is `kv_producer`, or `kv_role` is `kv_both`); startup will fail with a clear error. |
|
||||
| `enable_cpu_binding` | bool | `True` | Enables Ascend-native CPU binding on ARM servers. Set to `False` to disable. See [CPU Binding](../feature_guide/cpu_binding.md). |
|
||||
| `enable_sleep_mode_extra_cleanup` | bool | `False` | Enables extra sleep-mode cleanup for RL workloads, including HCCL process-group release and ACL graph workspace cleanup. Disabled by default because wakeup may need to restore HCCL and recapture ACL graphs. |
|
||||
| `SLO_limits_for_dynamic_batch` | int | `-1` | SLO limits for dynamic batch. This is new scheduler to support dynamic batch feature |
|
||||
| `pa_shape_list` | list | `[]` | The custom shape list of page attention ops. |
|
||||
| `enable_kv_nz` | bool | `False` | Whether to enable KV cache NZ layout. This option only takes effects on models using MLA (e.g., DeepSeek). |
|
||||
| `layer_sharding` | dict | `{}` | Configuration options for Layer Sharding Linear. Layer Sharding can only be enabled in PD-disaggregated's P node. |
|
||||
| `enable_sparse_sfa_c8` | bool | `False` | Whether to enable the packed C8 KV cache for Sparse Flash Attention in DSA models (e.g., DeepSeek V3.2 and GLM5). This option is independent of `enable_sparse_li_c8`. SFA prefill context parallelism and Ascend 950 DCP are not supported. |
|
||||
| `enable_sparse_li_c8` | bool | `False` | Whether to enable the C8 key and scale caches for LightningIndexer in DSA models. This option is independent of `enable_sparse_sfa_c8` and only applies to eligible indexer layers from the model quantization config. SFA prefill context parallelism and Ascend 950 DCP are not supported. |
|
||||
| `c8_enable_reshape_optim` | bool | `False` | Whether to use the StoreKVBlock operator to accelerate LightningIndexer C8 cache writes. `enable_sparse_li_c8` must also be enabled. In the PD separation scenario, only the P node is enabled. |
|
||||
| `enable_mc2_hierarchy_comm` | bool | `False` | Enable dispatch/combine op inter-node communication by RoCE. |
|
||||
| `enable_prefill_mc2` | bool | `False` | Whether to reserve mc2_token_capacity for prefill batches. When enabled, `max_num_batched_tokens` is used to calculate the mc2_token_capacity instead of the decode-only capacity. In this scenario, the recommended maximum value of `max_num_batched_tokens` is `tp_size * 512`. This is a temporary switch; once MC2 operators are complete for all scenarios, this switch will be removed and MC2 will be enabled by default. |
|
||||
| `mega_moe_max_tokens` | int | `65536` | Per-rank token capacity after dispatch in the mega moe (dispatch_ffn_combine) fused operator. When load imbalance causes a rank to receive more tokens than this limit, the excess tokens are dropped and skipped from computation, degrading accuracy. Do not set this too large: workspace memory scales linearly with this value. |
|
||||
| `profiling_chunk_config` | dict | `{}` | Configuration options for dynamic chunked pipeline parallel. See [Dynamic Chunked Pipeline Parallel](../feature_guide/dynamic_chunk_pipeline_parallel.md) for details. |
|
||||
| `enable_balance_scheduling` | bool | `False` | Whether to enable balance scheduling. Can also be configured via the `VLLM_ASCEND_BALANCE_SCHEDULING` environment variable during the migration period. |
|
||||
| `enable_flashcomm1` | bool | `False` | Whether to enable FlashComm1 optimization. Can also be configured via the `VLLM_ASCEND_ENABLE_FLASHCOMM1` environment variable during the migration period. |
|
||||
| `enable_matmul_allreduce` | bool | `False` | Whether to enable matmul allreduce optimization. Can also be configured via the `VLLM_ASCEND_ENABLE_MATMUL_ALLREDUCE` environment variable during the migration period. |
|
||||
| `flashcomm2_parallel_size` | int | `0` | FlashComm2 parallel size. Can also be configured via the `VLLM_ASCEND_FLASHCOMM2_PARALLEL_SIZE` environment variable during the migration period. |
|
||||
| `msmonitor_use_daemon` | bool | `False` | Whether to use daemon mode for msmonitor. Can also be configured via the `MSMONITOR_USE_DAEMON` environment variable during the migration period. |
|
||||
| `enable_mlapo` | bool | `True` | Whether to enable MLAPO (Model Layer-wise Adaptive Parallel Optimization). Can also be configured via the `VLLM_ASCEND_ENABLE_MLAPO` environment variable during the migration period. |
|
||||
| `weight_nz_mode` | int | `1` | Weight NZ mode. Can also be configured via the `VLLM_ASCEND_ENABLE_NZ` environment variable during the migration period. |
|
||||
| `enable_fused_mc2` | int | `0` | Fused MC2 configuration. Can also be configured via the `VLLM_ASCEND_ENABLE_FUSED_MC2` environment variable during the migration period. |
|
||||
| `enable_transpose_kv_cache_by_block`| bool | `True` | Whether to enable transpose KV cache by block. Can also be configured via the `VLLM_ASCEND_FUSION_OP_TRANSPOSE_KV_CACHE_BY_BLOCK` environment variable during the migration period. |
|
||||
| `enable_dsa_cp` | bool | `False` | Whether to enable dsa_cp for DeepSeek V3.2, DeepSeek V4, and other models with the same architecture. This feature depends on FlashComm1. Please ensure that FlashComm1 is enabled before enabling this feature.|
|
||||
| `rejection_sampler_config` | dict | `{}` | Configuration options for rejection sampler (block verify and entropy verify). |
|
||||
| `multistream_dsv4_dsa_overlap` | bool | `True` | Whether to enable dsa multi-stream overlap for DeepSeek V4. |
|
||||
| `enable_reduce_sample` | bool | `False` | Whether to enable reduce sample optimization to reduce communication and computation overheads in the tensor parallelism scenario. When enabled, logits are kept partitioned across TP ranks and only the small set of top-k candidate values/indices is communicated, instead of performing a full-vocabulary all-to-all/all-gather. **Note**: This is an experimental feature. **Limitations**: (1) Not supported on PD-disaggregated scenario. (2) Must be disabled when sampling logprobs are requested. When reduce sample is enabled, logprobs are silently computed over partitioned logits instead of the full vocabulary, producing incorrect logprob values and top-k rankings. (3) Cannot be enabled together with lmhead TP.|
|
||||
|
||||
The details of each config option are as follows:
|
||||
The details of each configuration option are as follows:
|
||||
|
||||
**torchair_graph_config**
|
||||
**xlite_graph_config**
|
||||
|
||||
| Name | Type | Default | Description |
|
||||
| ---- | ---- | ------- | ----------- |
|
||||
| `enabled` | bool | `False` | Whether to enable torchair graph mode. Currently only DeepSeek series models and PanguProMoE are supported to use torchair graph mode |
|
||||
| `mode` | str | `None` | When using reduce-overhead mode for torchair, mode needs to be set |
|
||||
| `enable_multistream_mla`| bool | `False` | Whether to put vector ops of MLA to another stream. This option only takes effects on models using MLA (e.g., DeepSeek). |
|
||||
| `enable_view_optimize` | bool | `True` | Whether to enable torchair view optimization |
|
||||
| `enable_frozen_parameter` | bool | `True` | Whether to fix the memory address of weights during inference to reduce the input address refresh time during graph execution. |
|
||||
| `use_cached_graph` | bool | `False` | Whether to use cached graph |
|
||||
| `graph_batch_sizes` | list[int] | `[]` | The batch size for torchair graph cache |
|
||||
| `graph_batch_sizes_init` | bool | `False` | Init graph batch size dynamically if `graph_batch_sizes` is empty |
|
||||
| `enable_kv_nz`| bool | `False` | Whether to enable kvcache NZ layout. This option only takes effects on models using MLA (e.g., DeepSeek). |
|
||||
| `enabled` | bool | `False` | Whether to enable Xlite graph mode. Currently only Llama, Qwen dense series models, and Qwen3-VL are supported. |
|
||||
| `full_mode` | bool | `False` | Whether to enable Xlite for both the prefill and decode stages. By default, Xlite is only enabled for the decode stage. |
|
||||
|
||||
**ascend_scheduler_config**
|
||||
**weight_prefetch_config**
|
||||
|
||||
| Name | Type | Default | Description |
|
||||
|------------------|------|-------------------------------------------------------------|------------------------------------|
|
||||
| `enabled` | bool | `False` | Whether to enable weight prefetch. |
|
||||
| `prefetch_ratio` | dict | `{"attn": {"qkv": 1.0, "o": 1.0}, "moe": {"gate_up": 0.8}, "mlp": { "gate_up": 1.0, "down": 1.0}}` | Prefetch ratio of each weight. |
|
||||
|
||||
**finegrained_tp_config**
|
||||
|
||||
| Name | Type | Default | Description |
|
||||
| ---- | ---- | ------- | ----------- |
|
||||
| `enabled` | bool | `False` | Whether to enable ascend scheduler for V1 engine|
|
||||
| `enable_pd_transfer` | bool | `False` | Whether to enable pd transfer. When using it, decode is started only when prefill of all requests is done. This option only takes effects on offline inference. |
|
||||
| `decode_max_num_seqs` | int | `0` | Whether to change max_num_seqs of decode phase when enable pd transfer. This option only takes effects when enable_pd_transfer is True. |
|
||||
| `max_long_partial_prefills` | Union[int, float] | `float('inf')` | the maximum number of prompts longer than long_prefill_token_threshold that will be prefilled concurrently. |
|
||||
| `long_prefill_token_threshold` | Union[int, float] | `float('inf')` | a request is considered long if the prompt is longer than this number of tokens. |
|
||||
| `lmhead_tensor_parallel_size` | int | `0` | The custom tensor parallel size of lm_head. |
|
||||
| `oproj_tensor_parallel_size` | int | `0` | The custom tensor parallel size of o_proj. |
|
||||
| `embedding_tensor_parallel_size` | int | `0` | The custom tensor parallel size of embedding. |
|
||||
| `mlp_tensor_parallel_size` | int | `0` | The custom tensor parallel size of mlp. |
|
||||
|
||||
ascend_scheduler_config also support the options from [vllm scheduler config](https://docs.vllm.ai/en/stable/api/vllm/config.html#vllm.config.SchedulerConfig). For example, you can add `enable_chunked_prefill: True` to ascend_scheduler_config as well.
|
||||
**ascend_compilation_config**
|
||||
|
||||
| Name | Type | Default | Description |
|
||||
| ---- | ---- | ------- | ----------- |
|
||||
| `enable_npugraph_ex` | bool | `True` | Whether to enable npugraph_ex backend. |
|
||||
| `enable_static_kernel` | bool | `False` | Whether to enable static kernel. Suitable for scenarios where shape changes are minimal and some time is available for static kernel compilation. |
|
||||
| `fuse_norm_quant` | bool | `True` | Whether to enable fuse_norm_quant pass. |
|
||||
| `fuse_qknorm_rope` | bool | `True` | Whether to enable fuse_qknorm_rope pass. If Triton is not in the environment, set it to False. |
|
||||
| `fuse_allreduce_rms` | bool | `False` | Whether to enable fuse_allreduce_rms pass. It's set to False because of conflict with SP. |
|
||||
| `fuse_muls_add` | bool | `True` | Whether to enable fuse_muls_add pass.|
|
||||
|
||||
**eplb_config**
|
||||
|
||||
| Name | Type | Default | Description |
|
||||
| ---- | ---- | ------- | ----------- |
|
||||
| `dynamic_eplb` | bool| `False`| Whether to enable dynamic EPLB. |
|
||||
| `expert_map_path` | str | `None` | When using expert load balancing for an MoE model, an expert map path needs to be passed in.|
|
||||
| `expert_heat_collection_interval`| int | `400` | Forward iterations when EPLB begins. |
|
||||
| `algorithm_execution_interval` | int | `30` | The forward iterations when the EPLB worker will finish CPU tasks. |
|
||||
| `expert_map_record_path` | str | `None` | Save the expert load calculation results to a new expert table in the specified directory.|
|
||||
| `num_redundant_experts` | int | `0` | Specify redundant experts during initialization. |
|
||||
| `eplb_policy_type` | int | `2` | EPLB balancing policy: `0`=Random, `1`=DefaultEplb (open-source algorithm), `2`=SwiftBalanceEplb (optimized for low-bandwidth), `3`=FlashLB (statistical method with sliding windows). |
|
||||
| `eplb_heat_collection_stage` | str | `"all"`| Stage to collect EPLB heat: `"prefill"` collects only during prefill, `"decode"` collects only during decode, `"all"` collects during both stages. In PD colocation scenarios, prefill and decode requests may produce different expert workloads. Selectively collecting heat on one stage can reduce expert imbalance more effectively. |
|
||||
|
||||
**profiling_chunk_config**
|
||||
|
||||
| Name | Type | Default | Description |
|
||||
| ---- | ---- | ------- | ----------- |
|
||||
| `enabled` | bool | `False` | Whether to enable dynamic chunked pipeline parallel. Requires `pipeline-parallel-size > 1`. |
|
||||
| `smooth_factor` | float | `1.0` | Smoothing factor (0 < x ≤ 1.0). Higher values trust the dynamic prediction more; `0.0` disables dynamic adjustment. |
|
||||
| `min_chunk` | int | `4096` | Minimum chunk size for dynamic calculation. Should be smaller than `max-num-batched-tokens`. |
|
||||
| `need_timing` | bool | True | Enable/disable Online Calibration |
|
||||
| `max_fit_chunk` | int | 30 | Number of chunk-time data for Online Calibration |
|
||||
|
||||
**rejection_sampler_config**
|
||||
|
||||
> **Note**: Both block verify and entropy verify improve speculative decoding performance (higher acceptance rate, lower latency) at the cost of reduced sampling precision. A larger `posterior_alpha` makes the adjustment more aggressive — it further lowers the acceptance threshold for high-entropy tokens, improving throughput but degrading output quality. Users should tune these parameters based on their specific model weights and application scenario to find the right trade-off between performance and precision.
|
||||
|
||||
| Name | Type | Default | Description |
|
||||
| ---- | ---- | ------- | ----------- |
|
||||
| `enable_block_verify` | bool | `False` | Whether to enable block verify mode. Block verify evaluates all draft tokens as a block using cumulative probability products, which can improve acceptance rate. |
|
||||
| `enable_entropy_verify` | bool | `False` | Whether to enable entropy verify mode. Entropy verify adjusts the acceptance threshold based on the entropy of the target distribution — higher entropy (uncertain) tokens get a lower threshold (easier to accept), while lower entropy (confident) tokens get a stricter threshold. |
|
||||
| `posterior_threshold` | float | `0.95` | Upper bound for the entropy-adjusted acceptance threshold. Must be in (0, 1]. The effective threshold is `min(exp(-entropy * posterior_alpha), posterior_threshold)`. |
|
||||
| `posterior_alpha` | float | `0.4` | Scaling factor for entropy in the threshold computation. Must be >= 0. Higher values make the threshold more sensitive to entropy — high-entropy tokens become much easier to accept, improving performance but reducing precision. |
|
||||
|
||||
### Example
|
||||
|
||||
An example of additional configuration is as follows:
|
||||
|
||||
```
|
||||
```python
|
||||
{
|
||||
"torchair_graph_config": {
|
||||
"weight_prefetch_config": {
|
||||
"enabled": True,
|
||||
"use_cached_graph": True,
|
||||
"graph_batch_sizes": [1, 2, 4, 8],
|
||||
"graph_batch_sizes_init": False,
|
||||
"enable_kv_nz": False
|
||||
"prefetch_ratio": {
|
||||
"attn": {
|
||||
"qkv": 1.0,
|
||||
"o": 1.0,
|
||||
},
|
||||
"moe": {
|
||||
"gate_up": 0.8
|
||||
},
|
||||
"mlp": {
|
||||
"gate_up": 1.0,
|
||||
"down": 1.0
|
||||
}
|
||||
},
|
||||
},
|
||||
"ascend_scheduler_config": {
|
||||
"enabled": True,
|
||||
"enable_chunked_prefill": True,
|
||||
"max_long_partial_prefills": 1,
|
||||
"long_prefill_token_threshold": 4096,
|
||||
"finegrained_tp_config": {
|
||||
"lmhead_tensor_parallel_size": 8,
|
||||
"oproj_tensor_parallel_size": 8,
|
||||
"embedding_tensor_parallel_size": 8,
|
||||
"mlp_tensor_parallel_size": 8,
|
||||
},
|
||||
"enable_kv_nz": False,
|
||||
"multistream_overlap_shared_expert": True,
|
||||
"refresh": False,
|
||||
"rejection_sampler_config": {
|
||||
"enable_block_verify": True,
|
||||
"enable_entropy_verify": True,
|
||||
"posterior_threshold": 0.95,
|
||||
"posterior_alpha": 0.4,
|
||||
},
|
||||
"refresh": False
|
||||
}
|
||||
```
|
||||
|
||||
@@ -2,6 +2,8 @@
|
||||
|
||||
vllm-ascend uses the following environment variables to configure the system:
|
||||
|
||||
**Note:** Some environment variables are being migrated to `--additional-config` options. These environment variables are still supported during the migration period, and it is recommended to use `--additional-config` for new deployments. See [Additional Configuration](additional_config.md) for details.
|
||||
|
||||
:::{literalinclude} ../../../../vllm_ascend/envs.py
|
||||
:language: python
|
||||
:start-after: begin-env-vars-definition
|
||||
|
||||
8
docs/source/user_guide/deployment_guide/index.md
Normal file
@@ -0,0 +1,8 @@
|
||||
# Deployment Guide
|
||||
|
||||
:::{toctree}
|
||||
:caption: Deployment Guide
|
||||
:maxdepth: 1
|
||||
using_volcano_kthena
|
||||
using_mindie_motor
|
||||
:::
|
||||
@@ -0,0 +1,9 @@
|
||||
# Deploy vLLM-Ascend with MindIE-Motor
|
||||
|
||||
## 1. Overview
|
||||
|
||||
[MindIE-Motor](https://gitcode.com/Ascend/MindIE-Motor) provides one-click deployment for **prefill–decode (PD) disaggregation** and **PD aggregation** on Ascend NPUs with vLLM-Ascend. It uses **high-performance scheduling and load balancing**, together with **RAS (Reliability, Availability and Serviceability) capabilities**, to build inference services that are fast and highly stable.
|
||||
|
||||
## 2. Getting Started
|
||||
|
||||
For quick deployment instructions, refer to the [MindIE-Motor Quick Start](https://gitcode.com/Ascend/MindIE-Motor/blob/master/docs/zh/user_guide/README.md).
|
||||
433
docs/source/user_guide/deployment_guide/using_volcano_kthena.md
Normal file
@@ -0,0 +1,433 @@
|
||||
# Using Volcano Kthena
|
||||
|
||||
This guide shows how to run **prefill–decode (PD) disaggregation** on Huawei Ascend NPUs using **vLLM-Ascend**, with [**Kthena**](https://kthena.volcano.sh/) handling orchestration on Kubernetes. About vLLM support with Kthena, please refer to [Deploy vLLM with Kthena](https://docs.vllm.ai/en/latest/deployment/integrations/kthena/).
|
||||
|
||||
---
|
||||
|
||||
## 1. What is Prefill–Decode Disaggregation?
|
||||
|
||||
Large language model inference naturally splits into two phases:
|
||||
|
||||
- **Prefill**
|
||||
- Processes input tokens and builds the key–value (KV) cache.
|
||||
- Batch-friendly, high-throughput, well-suited to parallel NPU execution.
|
||||
- **Decode**
|
||||
- Consumes the KV cache to generate output tokens.
|
||||
- Latency-sensitive, memory-intensive, more sequential.
|
||||
|
||||
From the client's perspective, this still looks like a single Chat / Completions endpoint.
|
||||
|
||||
---
|
||||
|
||||
## 2. Deploy on Kubernetes with Kthena
|
||||
|
||||
[Kthena](https://kthena.volcano.sh/) is a Kubernetes-native LLM inference platform that transforms how organizations deploy and manage Large Language Models in production. Built with declarative model lifecycle management and intelligent request routing, it provides high-performance and enterprise-grade scalability for LLM inference workloads. In this example, we use three key Custom Resource Definitions (CRDs):
|
||||
|
||||
- `ModelServing` — defines the workloads (prefill and decode roles).
|
||||
- `ModelServer` — manages PD groupings and internal routing.
|
||||
- `ModelRoute` — exposes a stable model endpoint.
|
||||
|
||||
This section uses the `deepseek-ai/DeepSeek-V2-Lite` example, but you can swap in any model supported by vLLM-Ascend.
|
||||
|
||||
### 2.1 Prerequisites
|
||||
|
||||
- Kubernetes cluster with Ascend NPU nodes:
|
||||
|
||||
The resources corresponding to different NPU Drivers may vary slightly. For example:
|
||||
|
||||
- If using [MindCluster](https://gitee.com/ascend/mind-cluster#https://gitee.com/link?target=https%3A%2F%2Fgitcode.com%2FAscend%2Fmind-cluster), please use `huawei.com/Ascend310P` or `huawei.com/Ascend910`.
|
||||
|
||||
- If running on CCE (Cloud Container Engine) of Huawei Cloud and the [CCE AI Suite Plugin (Ascend NPU)](https://support.huaweicloud.com/intl/en-us/usermanual-cce/cce_10_0239.html) is installed, please use `huawei.com/ascend-310` or `huawei.com/ascend-1980`.
|
||||
|
||||
- Kthena installed. Please follow the [Kthena installation guide](https://kthena.volcano.sh/docs/getting-started/installation).
|
||||
|
||||
### 2.2 Deploy Prefill-Decode Disaggregated DeepSeek-V2-Lite on Kubernetes
|
||||
|
||||
A concrete example is provided in Kthena as [prefill-decode-disaggregation.yaml](https://github.com/volcano-sh/kthena/blob/main/examples/model-serving/prefill-decode-disaggregation.yaml)
|
||||
|
||||
Deploy it with the command below:
|
||||
|
||||
```bash
|
||||
kubectl apply -f https://raw.githubusercontent.com/volcano-sh/kthena/refs/heads/main/examples/model-serving/prefill-decode-disaggregation.yaml
|
||||
```
|
||||
|
||||
or
|
||||
|
||||
```bash
|
||||
cat << EOF | kubectl apply -f -
|
||||
apiVersion: workload.serving.volcano.sh/v1alpha1
|
||||
kind: ModelServing
|
||||
metadata:
|
||||
name: deepseek-v2-lite
|
||||
namespace: dev
|
||||
spec:
|
||||
schedulerName: volcano
|
||||
replicas: 1
|
||||
recoveryPolicy: ServingGroupRecreate
|
||||
template:
|
||||
restartGracePeriodSeconds: 60
|
||||
roles:
|
||||
- name: prefill
|
||||
replicas: 1
|
||||
entryTemplate:
|
||||
spec:
|
||||
initContainers:
|
||||
- name: downloader
|
||||
imagePullPolicy: Always
|
||||
image: ghcr.io/volcano-sh/downloader:latest
|
||||
args:
|
||||
- --source
|
||||
- deepseek-ai/DeepSeek-V2-Lite
|
||||
- --output-dir
|
||||
- /mnt/cache/deepseek-ai/DeepSeek-V2-Lite/
|
||||
volumeMounts:
|
||||
- name: models
|
||||
mountPath: /mnt/cache/deepseek-ai/DeepSeek-V2-Lite/
|
||||
containers:
|
||||
- name: runtime
|
||||
image: ghcr.io/volcano-sh/runtime:latest
|
||||
ports:
|
||||
- containerPort: 8100
|
||||
args:
|
||||
- --port
|
||||
- "8100"
|
||||
- --engine
|
||||
- vllm
|
||||
- --pod
|
||||
- $(POD_NAME).$(NAMESPACE)
|
||||
- --model
|
||||
- deepseek-v2-lite
|
||||
- --engine-base-url
|
||||
- http://localhost:8000
|
||||
- name: vllm
|
||||
image: ghcr.io/volcano-sh/kthena-engine:vllm-ascend_v0.10.1rc1_mooncake_v0.3.5
|
||||
ports:
|
||||
- containerPort: 8000
|
||||
env:
|
||||
- name: HF_HUB_OFFLINE
|
||||
value: "1"
|
||||
- name: HCCL_IF_IP
|
||||
valueFrom:
|
||||
fieldRef:
|
||||
fieldPath: status.podIP
|
||||
- name: GLOO_SOCKET_IFNAME
|
||||
value: eth0
|
||||
- name: TP_SOCKET_IFNAME
|
||||
value: eth0
|
||||
- name: HCCL_SOCKET_IFNAME
|
||||
value: eth0
|
||||
- name: VLLM_LOGGING_LEVEL
|
||||
value: DEBUG
|
||||
- name: AscendRealDevices
|
||||
valueFrom:
|
||||
fieldRef:
|
||||
fieldPath: metadata.annotations['huawei.com/AscendReal']
|
||||
args:
|
||||
- "/mnt/cache/deepseek-ai/DeepSeek-V2-Lite/"
|
||||
- "--served-model-name"
|
||||
- "deepseek-ai/DeepSeekV2"
|
||||
- "--tensor-parallel-size"
|
||||
- "2"
|
||||
- "--gpu-memory-utilization"
|
||||
- "0.8"
|
||||
- "--max-model-len"
|
||||
- "8192"
|
||||
- "--max-num-batched-tokens"
|
||||
- "8192"
|
||||
- "--trust-remote-code"
|
||||
- "--enforce-eager"
|
||||
- "--kv-transfer-config"
|
||||
- '{"kv_connector":"MooncakeConnectorV1","kv_buffer_device":"npu","kv_role":"kv_producer","kv_parallel_size":1,"kv_port":"20001","kv_rank":0,"kv_connector_extra_config":{"prefill":{"dp_size":2,"tp_size":2},"decode":{"dp_size":2,"tp_size":2}}}'
|
||||
imagePullPolicy: Always
|
||||
resources:
|
||||
limits:
|
||||
cpu: "8"
|
||||
memory: 64Gi
|
||||
huawei.com/ascend-1980: "4"
|
||||
requests:
|
||||
cpu: "8"
|
||||
memory: 64Gi
|
||||
huawei.com/ascend-1980: "4"
|
||||
readinessProbe:
|
||||
initialDelaySeconds: 5
|
||||
periodSeconds: 5
|
||||
failureThreshold: 3
|
||||
httpGet:
|
||||
path: /health
|
||||
port: 8000
|
||||
livenessProbe:
|
||||
initialDelaySeconds: 900
|
||||
periodSeconds: 5
|
||||
failureThreshold: 3
|
||||
httpGet:
|
||||
path: /health
|
||||
port: 8000
|
||||
volumeMounts:
|
||||
- name: models
|
||||
mountPath: /mnt/cache/deepseek-ai/DeepSeek-V2-Lite/
|
||||
readOnly: true
|
||||
- name: hccn-config
|
||||
mountPath: /etc/hccn.conf
|
||||
readOnly: true
|
||||
- name: shared-memory-volume
|
||||
mountPath: /dev/shm
|
||||
volumes:
|
||||
- name: models
|
||||
hostPath:
|
||||
path: /mnt/cache/deepseek-ai/DeepSeek-V2-Lite/
|
||||
type: DirectoryOrCreate
|
||||
- name: hccn-config
|
||||
hostPath:
|
||||
path: /etc/hccn.conf
|
||||
type: File
|
||||
- name: shared-memory-volume
|
||||
emptyDir:
|
||||
sizeLimit: 256Mi
|
||||
medium: Memory
|
||||
- name: decode
|
||||
replicas: 1
|
||||
entryTemplate:
|
||||
spec:
|
||||
initContainers:
|
||||
- name: downloader
|
||||
imagePullPolicy: Always
|
||||
image: ghcr.io/volcano-sh/downloader:latest
|
||||
args:
|
||||
- --source
|
||||
- deepseek-ai/DeepSeek-V2-Lite
|
||||
- --output-dir
|
||||
- /mnt/cache/deepseek-ai/DeepSeek-V2-Lite/
|
||||
volumeMounts:
|
||||
- name: models
|
||||
mountPath: /mnt/cache/deepseek-ai/DeepSeek-V2-Lite/
|
||||
containers:
|
||||
- name: vllm
|
||||
image: ghcr.io/volcano-sh/kthena-engine:vllm-ascend_v0.10.1rc1_mooncake_v0.3.5
|
||||
ports:
|
||||
- containerPort: 8000
|
||||
env:
|
||||
- name: HF_HUB_OFFLINE
|
||||
value: "1"
|
||||
- name: HCCL_IF_IP
|
||||
valueFrom:
|
||||
fieldRef:
|
||||
fieldPath: status.podIP
|
||||
- name: GLOO_SOCKET_IFNAME
|
||||
value: eth0
|
||||
- name: TP_SOCKET_IFNAME
|
||||
value: eth0
|
||||
- name: HCCL_SOCKET_IFNAME
|
||||
value: eth0
|
||||
- name: VLLM_LOGGING_LEVEL
|
||||
value: DEBUG
|
||||
- name: AscendRealDevices
|
||||
valueFrom:
|
||||
fieldRef:
|
||||
fieldPath: metadata.annotations['huawei.com/AscendReal']
|
||||
args:
|
||||
- "/mnt/cache/deepseek-ai/DeepSeek-V2-Lite/"
|
||||
- "--served-model-name"
|
||||
- "deepseek-ai/DeepSeekV2"
|
||||
- "--tensor-parallel-size"
|
||||
- "2"
|
||||
- "--gpu-memory-utilization"
|
||||
- "0.8"
|
||||
- "--max-model-len"
|
||||
- "8192"
|
||||
- "--max-num-batched-tokens"
|
||||
- "16384"
|
||||
- "--trust-remote-code"
|
||||
- "--no-enable-prefix-caching"
|
||||
- "--enforce-eager"
|
||||
- "--kv-transfer-config"
|
||||
- '{"kv_connector":"MooncakeConnectorV1","kv_buffer_device":"npu","kv_role":"kv_consumer","kv_parallel_size":1,"kv_port":"20002","kv_rank":1,"kv_connector_extra_config":{"prefill":{"dp_size":2,"tp_size":2},"decode":{"dp_size":2,"tp_size":2}}}'
|
||||
imagePullPolicy: Always
|
||||
resources:
|
||||
limits:
|
||||
cpu: "8"
|
||||
memory: 64Gi
|
||||
huawei.com/ascend-1980: "4"
|
||||
requests:
|
||||
cpu: "8"
|
||||
memory: 64Gi
|
||||
huawei.com/ascend-1980: "4"
|
||||
readinessProbe:
|
||||
initialDelaySeconds: 5
|
||||
periodSeconds: 5
|
||||
failureThreshold: 3
|
||||
httpGet:
|
||||
path: /health
|
||||
port: 8000
|
||||
livenessProbe:
|
||||
initialDelaySeconds: 900
|
||||
periodSeconds: 5
|
||||
failureThreshold: 3
|
||||
httpGet:
|
||||
path: /health
|
||||
port: 8000
|
||||
volumeMounts:
|
||||
- name: models
|
||||
mountPath: /mnt/cache/deepseek-ai/DeepSeek-V2-Lite/
|
||||
readOnly: true
|
||||
- name: hccn-config
|
||||
mountPath: /etc/hccn.conf
|
||||
readOnly: true
|
||||
- name: shared-memory-volume
|
||||
mountPath: /dev/shm
|
||||
volumes:
|
||||
- name: models
|
||||
hostPath:
|
||||
path: /mnt/cache/deepseek-ai/DeepSeek-V2-Lite/
|
||||
type: DirectoryOrCreate
|
||||
- name: hccn-config
|
||||
hostPath:
|
||||
path: /etc/hccn.conf
|
||||
type: File
|
||||
- name: shared-memory-volume
|
||||
emptyDir:
|
||||
sizeLimit: 256Mi
|
||||
medium: Memory
|
||||
EOF
|
||||
```
|
||||
|
||||
You should see Pods such as:
|
||||
|
||||
- `deepseek-v2-lite-0-prefill-0-0`
|
||||
- `deepseek-v2-lite-0-decode-0-0`
|
||||
|
||||
To enable the LLM access, we still need to configure the routing layer with `ModelServer` and `ModelRoute`.
|
||||
|
||||
### 2.3 ModelServer: PD Group Management
|
||||
|
||||
The `ModelServer` resource:
|
||||
|
||||
- Selects the `ModelServing` workloads via labels.
|
||||
- Groups prefill and decode Pods into PD pairs.
|
||||
- Configures KV connector details and timeouts.
|
||||
- Exposes an internal gRPC/HTTP interface.
|
||||
|
||||
Create ModelServer with the command below:
|
||||
|
||||
```bash
|
||||
kubectl apply -f https://raw.githubusercontent.com/volcano-sh/kthena/refs/heads/main/examples/kthena-router/ModelServer-prefill-decode-disaggregation.yaml
|
||||
```
|
||||
|
||||
or
|
||||
|
||||
```bash
|
||||
cat << EOF | kubectl apply -f -
|
||||
apiVersion: networking.serving.volcano.sh/v1alpha1
|
||||
kind: ModelServer
|
||||
metadata:
|
||||
name: deepseek-v2
|
||||
namespace: dev
|
||||
spec:
|
||||
kvConnector:
|
||||
type: nixl
|
||||
workloadSelector:
|
||||
matchLabels:
|
||||
modelserving.volcano.sh/name: deepseek-v2-lite
|
||||
pdGroup:
|
||||
groupKey: "modelserving.volcano.sh/group-name"
|
||||
prefillLabels:
|
||||
modelserving.volcano.sh/role: prefill
|
||||
decodeLabels:
|
||||
modelserving.volcano.sh/role: decode
|
||||
workloadPort:
|
||||
port: 8000
|
||||
model: "deepseek-ai/DeepSeekV2"
|
||||
inferenceEngine: "vLLM"
|
||||
trafficPolicy:
|
||||
timeout: 10s
|
||||
EOF
|
||||
```
|
||||
|
||||
### 2.4 ModelRoute: User-Facing Endpoint
|
||||
|
||||
The `ModelRoute` resource maps a model name (e.g., `"deepseek-ai/DeepSeekV2"`) to the `ModelServer`.
|
||||
|
||||
Example manifest:
|
||||
|
||||
```bash
|
||||
cat << EOF | kubectl apply -f -
|
||||
apiVersion: networking.serving.volcano.sh/v1alpha1
|
||||
kind: ModelRoute
|
||||
metadata:
|
||||
name: deepseek-v2
|
||||
namespace: dev
|
||||
spec:
|
||||
modelName: "deepseek-ai/DeepSeekV2"
|
||||
rules:
|
||||
- name: "default"
|
||||
targetModels:
|
||||
- modelServerName: "deepseek-v2"
|
||||
EOF
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. Verification
|
||||
|
||||
### 3.1 Check Workloads
|
||||
|
||||
Confirm that prefill and decode Pods are up:
|
||||
|
||||
```bash
|
||||
kubectl get modelserving deepseek-v2-lite -n dev -o yaml | grep status -A 10
|
||||
|
||||
kubectl get pod -n dev -owide \
|
||||
-l modelserving.volcano.sh/name=deepseek-v2-lite
|
||||
```
|
||||
|
||||
You should see both roles in `Running` and `Ready` state.
|
||||
|
||||
### 3.2 Test the Chat Endpoint
|
||||
|
||||
Once routing is configured, you can send a test request to the Kthena-router:
|
||||
|
||||
```bash
|
||||
|
||||
export ENDPOINT=$(kubectl get svc kthena-router -n kthena-system --output=jsonpath='{.status.loadBalancer.ingress[0].ip}:{.spec.ports[0].port}')
|
||||
|
||||
curl --location "http://${ENDPOINT}/v1/chat/completions" \
|
||||
--header "Content-Type: application/json" \
|
||||
--data '{
|
||||
"model": "deepseek-ai/DeepSeekV2",
|
||||
"messages": [
|
||||
{
|
||||
"role": "user",
|
||||
"content": "Where is the capital of China?"
|
||||
}
|
||||
],
|
||||
"stream": false
|
||||
}'
|
||||
```
|
||||
|
||||
A successful JSON response confirms that:
|
||||
|
||||
- The prefill and decode services are both running on Ascend NPUs.
|
||||
- KV transfer between them is working.
|
||||
- The Kthena routing layer is correctly fronting the vLLM-Ascend plugin.
|
||||
|
||||
---
|
||||
|
||||
## 4. Cleanup
|
||||
|
||||
To remove the deployment:
|
||||
|
||||
```bash
|
||||
# 1. Remove user-facing routing
|
||||
kubectl delete modelroute deepseek-v2 -n dev
|
||||
|
||||
# 2. Remove internal server
|
||||
kubectl delete modelserver deepseek-v2 -n dev
|
||||
|
||||
# 3. Remove workloads
|
||||
kubectl delete modelserving deepseek-v2-lite -n dev
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 5. Summary
|
||||
|
||||
For more advanced features, please refer to the [Kthena website](https://kthena.volcano.sh/).
|
||||
@@ -0,0 +1,82 @@
|
||||
# AI QoS Feature
|
||||
|
||||
## Background
|
||||
|
||||
In the inference scenario, there are different types of traffic, such as operator delivery, collective communication, and KVCache. Such traffic are transmitted through the network and affect each other, increasing the inference latency.
|
||||
|
||||
For example, in the Agentic AI era, as the context length continues to increase, the size of the KVCache also gradually grows. To conserve HBM usage, the approach of offloading KVCache to DDR is adopted to enhance inference TPS. At the same time, to maximize the utilization of computing power, a pipeline orchestration method using computation to mask KVCache is commonly employed. This method involves prefetching the next layer's KVCache during the current layer's computation/communication to reduce overall latency. However, this approach introduces a traffic conflict issue between the KVCache and the operator delivery/collective communication, leading to increased inference latency and impacting the SLO.
|
||||
|
||||

|
||||
|
||||
As shown in the preceding figure, traffic conflicts occur on the UB switch when intra-node device-to-device (D2D) traffic, intra-node host-to-device (H2D) traffic, and inter-node D2D traffic are transmitted.
|
||||
|
||||
## Introduction
|
||||
|
||||
When different types of traffic conflict with each other, the Virtual Lane (VL) can be used to isolate the traffic at the UB switch and perform differentiated scheduling between the VLs. This helps to: (1) isolate the VLs of different types of traffic to prevent congestion from spreading; (2) perform differentiated scheduling for different types of traffic.
|
||||
|
||||
As shown in the following figure, different types of traffic are mapped to different VLs to isolate the traffic. In addition, the priority of each VL is set and the strict priority (SP) scheduling mode is used. When different types of traffic reach the UB switch at the same time, the traffic in the VL with the high priority is scheduled first, and then the traffic in the VL with the middle priority is scheduled. This process repeats until all the traffic is scheduled. In this way, differentiated scheduling is implemented for different types of traffic.
|
||||
|
||||

|
||||
|
||||
Different traffic is transmitted through different channels. Therefore, the AI QoS solution implements isolation and differentiated scheduling of different traffic to meet service requirements by (1) setting priorities for different NPU channels on the host, (2) establishing the mapping between the NPU channel priority and the VL of the UB switch, and (3) performing differentiated scheduling among different VLs of the UB switch based on the priority.
|
||||
|
||||
## Build AI QoS Module
|
||||
|
||||
Build and install the AI QoS extension before using `tools/ai_qos.py`.
|
||||
The DSMI include‑file (dsmi_common_interface.h) and library‑file (libdrvdsmi_host.so) paths are environment-dependent. Locate the paths on your machine first, then replace `YOUR_DSMI_INCLUDE_DIR` and `YOUR_DSMI_LIBRARY_FILE` in the command (for example, `/usr/local/Ascend/driver/include` and `/usr/local/Ascend/driver/lib64/driver/libdrvdsmi_host.so`).
|
||||
|
||||
In most deployments, these commands are executed inside a container. When creating the container, make sure the DSMI header/library directories are mounted into the container filesystem; otherwise CMake cannot find the files.
|
||||
|
||||
Run the following commands from the vLLM-Ascend repository root:
|
||||
|
||||
```bash
|
||||
cmake -S tools/ai_qos -B tools/ai_qos/build \
|
||||
-DCMAKE_BUILD_TYPE=Release \
|
||||
-DCMAKE_INSTALL_PREFIX=${PWD}/vllm_ascend \
|
||||
-DDSMI_INCLUDE_DIR=YOUR_DSMI_INCLUDE_DIR \
|
||||
-DDSMI_LIBRARY=YOUR_DSMI_LIBRARY_FILE
|
||||
cmake --build tools/ai_qos/build -j
|
||||
cmake --install tools/ai_qos/build
|
||||
```
|
||||
|
||||
## Usage Instruction
|
||||
|
||||
The AI QoS feature supports two modes: Auto and Manual. Enter the vLLM-Ascend installation directory and run the following command before running the inference job:
|
||||
|
||||
### 1) Auto mode
|
||||
|
||||
`python tools/ai_qos.py`
|
||||
|
||||
AI QoS auto mode automatically classifies the priorities of different types of traffic and generates QoS tags. It also prints the UB switch configuration. You can copy the outputs and log in to the UB switch to configure the QoS configurations of UB switch. This configuration will overwrite the current QoS configuration on the UB switch. If there is any existing QoS configuration, please back it up in advance.
|
||||
|
||||
### 2) Manual mode
|
||||
|
||||
python tools/ai_qos.py --mode manual --AIV_D2D {priority} --AIV_H2D {priority} --SDMA_D2D {priority} --SDMA_H2D {priority} --PCIEDMA_H2D {priority}
|
||||
|
||||
AI QoS manual mode calculates the QoS tags of traffic based on the priority of different types of traffic set by users, and generates and prints the UB switch configuration. You can copy the outputs and log in to the UB switch to configure the QoS configurations of UB switch. This configuration will overwrite the current QoS configuration on the UB switch. If there is any existing QoS configuration, please back it up in advance.
|
||||
|
||||
In manual mode, you can specify the priority of only one type of traffic. The parameters are described as follows:
|
||||
|
||||
| Name | Type | Default | Description |
|
||||
| ----------------- | ---- | ------------------------------------------------------------ | ------------------------------------------------------------ |
|
||||
| mode | str | auto | The mode of AI QoS, default mode is "auto", another mode is "manual", some parameters need to be configured if you choose "manual" mode. |
|
||||
| AIV_D2D,AIV_H2D,SDMA_D2D,SDMA_H2D,PCIEDMA_H2D | str | AIV_D2D: high,<br />AIV_H2D: high,<br />SDMA_D2D: high,<br />SDMA_H2D: low,<br />PCIEDMA_H2D: high | Parameters for "manual" mode, determined the QoS priority of different types of traffic. <br />The default configuration is the same as "auto" mode. <br />Typical traffic types are as follows for reference: AIV_D2D: AIV-based Device-to-Device communication, such as dispatch and combine.<br /> AIV_H2D: AIV-based Operator Delivery.<br /> SDMA_D2D: SDMA-based Device-to-Device communication, such as Allreduce and Allgather.<br />SDMA_H2D: SDMA-based Host-to-Device/Device-to-Host communication, such as KVCache offloading and prefetching.<br />PCIEDMA_H2D: PCIe DMA-based Operator Delivery. <br /> You can change the priority of different types of traffic, with "high/middle/low" options available.Due to hardware restrictions, "PCIEDMA_H2D" only supports "high/low" priority. |
|
||||
|
||||
**How to disable AI QoS**:
|
||||
|
||||
```bash
|
||||
python tools/ai_qos.py unset
|
||||
```
|
||||
|
||||
The command for disabling the AI QoS feature on the UB Switch will be printed on the screen. Please log in to the UB Switch and execute the command printed on the screen to complete the feature disabling.
|
||||
|
||||
## Usage Constraints
|
||||
|
||||
Due to underlying driver limitations, the QoS configurations for AIV_H2D and AIV_D2D do not take effect currently. Once the required adaptation capabilities are added in a future driver release, this feature will be delivered through a module upgrade.
|
||||
|
||||
The AI QoS feature supports the Atlas 800T A3 server and Atlas 900 A3 SuperPoD cluster. It must be used in privileged containers and requires the following software versions:
|
||||
|
||||
| Software | Matched Version |
|
||||
| :----------: | :--------------------------------------: |
|
||||
| Ascend HDK | 25.5.2 or later |
|
||||
| UB Switch | LingQu Computing Network 1.5.1 or later |
|
||||
109
docs/source/user_guide/feature_guide/Fine_grained_TP.md
Normal file
@@ -0,0 +1,109 @@
|
||||
# Fine-Grained Tensor Parallelism (Fine-grained TP)
|
||||
|
||||
## Overview
|
||||
|
||||
Fine-Grained Tensor Parallelism (Fine-grained TP) extends standard tensor parallelism by enabling **independent tensor-parallel sizes for different model components**. Instead of applying a single global `tensor_parallel_size` to all layers, Fine-grained TP allows users to configure separate TP sizes for key modules—such as embedding, language model head (lm_head), attention output projection (o_proj), and MLP blocks—via the `finegrained_tp_config` parameter.
|
||||
|
||||
This capability supports heterogeneous parallelism strategies within a single model, providing finer control over weight distribution, memory layout, and communication patterns across devices. The feature is compatible with standard dense transformer architectures and integrates seamlessly into vLLM’s serving pipeline.
|
||||
|
||||
---
|
||||
|
||||
## Benefits of Fine-grained TP
|
||||
|
||||
Fine-Grained Tensor Parallelism delivers two primary performance advantages through targeted weight sharding:
|
||||
|
||||
- **Reduced Per-Device Memory Footprint**:
|
||||
Fine-grained TP shards large weight matrices (e.g., LM Head, o_proj) across devices, lowering peak memory usage and enabling larger batches or deployment on memory-limited hardware—without quantization.
|
||||
|
||||
- **Faster Memory Access in GEMMs**:
|
||||
In decode-heavy workloads, GEMM performance is often memory-bound. Weight sharding reduces per-device weight fetch volume, cutting DRAM traffic and improving bandwidth efficiency—especially for latency-sensitive layers like LM Head and o_proj.
|
||||
|
||||
Together, these effects allow practitioners to better balance memory, communication, and compute—particularly in high-concurrency serving scenarios—while maintaining compatibility with standard dense transformer models.
|
||||
|
||||
---
|
||||
|
||||
## Supported Scenarios
|
||||
|
||||
### Models
|
||||
|
||||
Fine-grained TP is **model-agnostic** and supports all standard dense transformer architectures, including Llama, Qwen, DeepSeek (base/dense variants), and others.
|
||||
|
||||
### Component & Execution Mode Support
|
||||
|
||||
| TP config | Eager | Graph | Hybrid | Prefill | Decode |
|
||||
| ------------- | ----- | ----- | ------ | ------- | ------ |
|
||||
| **embedding** | ✅ | ✅ | ✅ | ✅ | ✅ |
|
||||
| **o_proj** | ❌ | ✅ | ❌ | ❌ | ✅ |
|
||||
| **mlp** | ✅ | ✅ | ✅ | ✅ | ✅ |
|
||||
| **LMhead** | ✅ | ✅ | ✅ | ✅ | ✅ |
|
||||
|
||||
> ⚠️ Note:
|
||||
>
|
||||
> - `o_proj` TP is only supported in Graph mode during Decode, because dummy_run in eager mode will not trigger o_proj.
|
||||
> - `mlp` TP supports dense models, or dense layers in MoE models. For example, the first three dense layers of DeepSeek-R1.
|
||||
|
||||
### Configuration Limit
|
||||
|
||||
The Fine-Grained TP size for any component must:
|
||||
|
||||
- Be **≤ the data-parallel (DP) size**, and
|
||||
- **Evenly divide the DP size** (i.e., `dp_size % tp_size == 0`) to ensure valid device assignment and communication grouping.
|
||||
|
||||
> ⚠️ Violating these constraints will result in runtime errors or undefined behavior.
|
||||
|
||||
---
|
||||
|
||||
## How to Use Fine-grained TP
|
||||
|
||||
### Configuration Format
|
||||
|
||||
Fine-grained TP is controlled via the `finegrained_tp_config` field inside `--additional-config`.
|
||||
|
||||
```bash
|
||||
--additional-config '{
|
||||
"finegrained_tp_config": {
|
||||
"embedding_tensor_parallel_size": 8,
|
||||
"lmhead_tensor_parallel_size": 8,
|
||||
"oproj_tensor_parallel_size": 8,
|
||||
"mlp_tensor_parallel_size": 8
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
### Example Usage
|
||||
|
||||
```bash
|
||||
vllm serve deepseek-ai/DeepSeek-R1 \
|
||||
--data-parallel-size 16 \
|
||||
--tensor-parallel-size 1 \
|
||||
--enable-expert-parallel \
|
||||
--additional-config '{
|
||||
"finegrained_tp_config": {
|
||||
"embedding_tensor_parallel_size": 8,
|
||||
"lmhead_tensor_parallel_size": 8,
|
||||
"mlp_tensor_parallel_size": 8
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Experimental Results
|
||||
|
||||
To evaluate the effectiveness of fine-grained TP in large-scale service scenarios, we use the model **DeepSeek-R1-W8A8**, deploy PD separated decode instances in an environment of 32 cards Ascend Atlas A2 inference products*64G (A2), with parallel configuration as DP32+EP32, and fine-grained TP size of 8; the performance data is as follows.
|
||||
|
||||
| Module | Memory Savings | TPOT Impact (batch=24) |
|
||||
| ---------------- | -------------- | ------------------------- |
|
||||
| o_proj TP = 8 | 5.8 GB | **+1.5 ms** (degradation) |
|
||||
| LM head TP = 8 | 1.51 GB | **−1.2 ms** (improvement) |
|
||||
| FFN TP = 8 | 0.9 GB | **−1.0 ms** (improvement) |
|
||||
| Embedding TP = 8 | 1.51 GB | **−1.0 ms** (improvement) |
|
||||
| **Total** | **9.72 GB** | — |
|
||||
|
||||
- We achieved significant gains in terms of high memory capacity on a single card, as well as the benefits of TPOT.
|
||||
|
||||
---
|
||||
|
||||
## ✅ Deployment Recommendations
|
||||
|
||||
Fine-grained TP is the **most effective** in the **decode instance** of PD separation, where models are typically deployed in all-DP mode. In this setup, sharding weight-heavy layers reduces redundant storage and memory pressure.
|
||||
147
docs/source/user_guide/feature_guide/batch_invariance.md
Normal file
@@ -0,0 +1,147 @@
|
||||
# Batch Invariance
|
||||
|
||||
```{note}
|
||||
Batch invariance is currently in beta. Some features are still under active development.
|
||||
Track progress and planned improvements at [tracking issue #5487](https://github.com/vllm-project/vllm-ascend/issues/5487)
|
||||
```
|
||||
|
||||
```{note}
|
||||
To install the batch invariance custom operator library, set `VLLM_BATCH_INVARIANT=1` before building vllm-ascend.
|
||||
For installation instructions, see [Set Up Using Python](https://github.com/vllm-project/vllm-ascend/blob/main/docs/source/installation.md#set-up-using-python)
|
||||
```
|
||||
|
||||
This document shows how to enable batch invariance in vLLM-Ascend. Batch invariance ensures that the output of a model is deterministic and independent of the batch size or the order of requests in a batch.
|
||||
|
||||
## Motivation
|
||||
|
||||
Batch invariance is crucial for several use cases:
|
||||
|
||||
- **Framework debugging**: Deterministic outputs make it easier to debug issues in the inference framework, as the same input will always produce the same output regardless of batching.
|
||||
- **Model debugging**: Helps identify issues in model implementations by ensuring consistent behavior across different batch configurations.
|
||||
- **Reinforcement Learning (RL)**: RL training often requires deterministic rollouts for reproducibility and stable training.
|
||||
- **Large-scale inference systems**: Systems that use vLLM as a component benefit from deterministic behavior for testing, validation, and consistency guarantees.
|
||||
|
||||
## Hardware Requirements
|
||||
|
||||
Batch invariance currently requires Ascend Atlas A2 and A3 inference products NPUs.
|
||||
We will support Ascend 950 Products and other NPUs in the future.
|
||||
|
||||
## Software Requirements
|
||||
|
||||
Batch invariance requires a custom operator library for Atlas A2 and A3 inference products, and users need to set `VLLM_BATCH_INVARIANT=1` before building vllm-ascend to install the batch invariance custom operator library during the installation process.
|
||||
|
||||
## Enabling Batch Invariance
|
||||
|
||||
Batch invariance can be enabled by setting the `VLLM_BATCH_INVARIANT` environment variable to `1`:
|
||||
|
||||
```bash
|
||||
export VLLM_BATCH_INVARIANT=1
|
||||
```
|
||||
|
||||
### Online Inference (Server Mode)
|
||||
|
||||
To start a vLLM server with batch invariance enabled:
|
||||
|
||||
```bash
|
||||
VLLM_BATCH_INVARIANT=1 vllm serve Qwen/Qwen3-8B \
|
||||
--compilation-config '{"cudagraph_mode": "PIECEWISE"}'
|
||||
```
|
||||
|
||||
Then use the OpenAI-compatible client:
|
||||
|
||||
```python
|
||||
from openai import OpenAI
|
||||
|
||||
client = OpenAI(
|
||||
api_key="EMPTY",
|
||||
base_url="http://localhost:8000/v1",
|
||||
)
|
||||
|
||||
# These requests will produce deterministic outputs
|
||||
# regardless of batch size or order
|
||||
response = client.completions.create(
|
||||
model="Qwen/Qwen3-8B",
|
||||
prompt="The future of AI is",
|
||||
max_tokens=100,
|
||||
temperature=0.7,
|
||||
seed=42,
|
||||
)
|
||||
|
||||
print(response.choices[0].text)
|
||||
```
|
||||
|
||||
### Offline Inference
|
||||
|
||||
For offline batch inference with batch invariance:
|
||||
|
||||
```python
|
||||
import os
|
||||
os.environ["VLLM_BATCH_INVARIANT"] = "1"
|
||||
|
||||
from vllm import LLM, SamplingParams
|
||||
|
||||
prompts = [
|
||||
"The future of AI is",
|
||||
"Machine learning enables",
|
||||
"Deep learning models can",
|
||||
]
|
||||
|
||||
sampling_params = SamplingParams(
|
||||
temperature=0.7,
|
||||
max_tokens=100,
|
||||
seed=42,
|
||||
)
|
||||
|
||||
llm = LLM(
|
||||
model="Qwen/Qwen3-8B",
|
||||
tensor_parallel_size=1,
|
||||
compilation_config={"cudagraph_mode": "PIECEWISE"},
|
||||
)
|
||||
|
||||
# Outputs will be deterministic regardless of batch size
|
||||
outputs = llm.generate(prompts, sampling_params)
|
||||
|
||||
for output in outputs:
|
||||
prompt = output.prompt
|
||||
generated_text = output.outputs[0].text
|
||||
print(f"Prompt: {prompt!r}")
|
||||
print(f"Generated: {generated_text!r}\n")
|
||||
```
|
||||
|
||||
## Tested Models
|
||||
|
||||
Batch invariance has been tested and verified on the following models:
|
||||
|
||||
- **Qwen3 (Dense)**: `Qwen/Qwen3-1.7B`, `Qwen/Qwen3-8B`
|
||||
- **Qwen3 (MoE)**: `Qwen/Qwen3-30B-A3B`, `Qwen/Qwen3-235B-A22B`
|
||||
|
||||
Other models may also work, but these have been explicitly validated. If you encounter issues with a specific model, please report them on the [GitHub issue tracker](https://github.com/vllm-project/vllm-ascend/issues/new/choose).
|
||||
|
||||
## Implementation Details
|
||||
|
||||
When batch invariance is enabled, vLLM:
|
||||
|
||||
1. Uses deterministic kernel implementations for attention and other operations
|
||||
2. Ensures consistent numerical behavior across different batch sizes
|
||||
3. Disables certain optimizations that may introduce non-determinism
|
||||
|
||||
```{note}
|
||||
The batch invariance attention operators currently do not support
|
||||
`FULL','FULL_DECODE_ONLY` cudagraph mode.
|
||||
```
|
||||
|
||||
```{note}
|
||||
Enabling batch invariance may impact performance compared to the default non-deterministic mode. This trade-off is intentional to guarantee reproducibility.
|
||||
```
|
||||
|
||||
## Future Improvements
|
||||
|
||||
The batch invariance feature is under active development. Planned improvements include:
|
||||
|
||||
- Support for additional NPUs series
|
||||
- Support `FULL`,`FULL_DECODE_ONLY` cudagraph mode with batch invariance attention operators
|
||||
- Expanded model coverage
|
||||
- Performance optimizations
|
||||
- Additional testing and validation
|
||||
|
||||
For the latest status and to contribute ideas, see the [tracking issue](https://github.com/vllm-project/vllm-ascend/issues/5487).
|
||||
103
docs/source/user_guide/feature_guide/context_parallel.md
Normal file
@@ -0,0 +1,103 @@
|
||||
# Context Parallel Guide
|
||||
|
||||
## Overview
|
||||
|
||||
This guide shows how to use Context Parallel, a long sequence inference optimization technique. Context Parallel includes `PCP` (Prefill Context Parallel) and `DCP` (Decode Context Parallel), which reduces NPU memory usage and improves inference speed in long sequence LLM inference.
|
||||
|
||||
## Benefits of Context Parallel
|
||||
|
||||
Context parallel mainly solves the problem of serving long context requests. As prefill and decode present quite different characteristics and have quite different SLO (service level objectives), we need to implement context parallel separately for them. The major considerations are:
|
||||
|
||||
- For long context prefill, we can use context parallel to reduce TTFT (time to first token) by amortizing the computation time of the prefill across query tokens.
|
||||
- For long context decode, we can use context parallel to reduce KV cache duplication and offer more space for KV cache to increase the batch size (and hence the throughput).
|
||||
|
||||
To learn more about the theory and implementation details of context parallel, please refer to the [context parallel developer guide](../../developer_guide/Design_Documents/context_parallel.md).
|
||||
|
||||
## Supported Scenarios
|
||||
|
||||
CP(Context Parallel) supports eager and graph execution, prefix caching, chunked prefill, speculative decoding, P/D disaggregation, and MLAPO on the model and hardware combinations documented by vLLM Ascend. The following table shows whether each feature can be combined with DCP across devices and attention backends:
|
||||
|
||||
| Device | Attention Backend | Chunked Prefill + CP | Prefix Caching + CP | Graph Mode + CP | P/D Disaggregation + CP | MLAPO + CP | Speculative Decoding + CP |
|
||||
| --- | --- | --- | --- | --- | --- | --- | --- |
|
||||
| Ascend A2/A3 | MLA/GQA | 🟢 Supported | 🟢 Supported | 🟢 Supported | 🟢 Supported | 🟢 Supported (MLA)<br>— Not applicable (GQA) | 🟢 P/D disaggregation<br>🔴 PD-mixed deployment |
|
||||
| Ascend A2/A3 | SFA | 🟢 Supported | 🟢 Supported | 🟢 Supported | 🟢 Supported | 🟢 Supported | 🟢 Supported |
|
||||
| Ascend 950 | MLA/GQA | 🔵 Experimental | 🔵 Experimental | 🔵 Experimental | 🔵 Experimental | 🔵 Experimental (MLA)<br>— Not applicable (GQA) | 🔵 P/D disaggregation<br>🔴 PD-mixed deployment |
|
||||
| Ascend 950 | SFA | 🔴 Not supported | 🔴 Not supported | 🔴 Not supported | 🔴 Not supported | 🔴 Not supported | 🔴 Not supported |
|
||||
|
||||
- 🟢 **Supported**: Combining the feature with DCP is supported.
|
||||
- 🔵 **Experimental**: Combining the feature with DCP is experimentally supported; interfaces and functionality may change.
|
||||
- 🔴 **Not supported**: Combining the feature with DCP is not supported.
|
||||
- **Not applicable**: The feature does not apply to this attention backend.
|
||||
|
||||
## How to use Context Parallel
|
||||
|
||||
You can enable `PCP` and `DCP` by `prefill_context_parallel_size` and `decode_context_parallel_size`, refer to the following example:
|
||||
|
||||
- Offline example:
|
||||
|
||||
```python
|
||||
from vllm import LLM, SamplingParams
|
||||
|
||||
prompts = [
|
||||
"The future of AI is",
|
||||
]
|
||||
sampling_params = SamplingParams(temperature=0.8, top_p=0.95)
|
||||
|
||||
llm = LLM(
|
||||
model="deepseek-ai/DeepSeek-V2-Lite",
|
||||
tensor_parallel_size=2,
|
||||
decode_context_parallel_size=2,
|
||||
prefill_context_parallel_size=2,
|
||||
)
|
||||
outputs = llm.generate(prompts, sampling_params)
|
||||
```
|
||||
|
||||
- Online example:
|
||||
|
||||
```bash
|
||||
vllm serve deepseek-ai/DeepSeek-V2-Lite \
|
||||
--tensor-parallel-size 2 \
|
||||
--decode-context-parallel-size 2 \
|
||||
--prefill-context-parallel-size 2 \
|
||||
```
|
||||
|
||||
The total world size is `tensor_parallel_size` * `prefill_context_parallel_size`, so the examples above need 4 NPUs for each.
|
||||
|
||||
## Constraints
|
||||
|
||||
- While using DCP, the following constraints must be met:
|
||||
- For MLA-based model, such as DeepSeek-R1:
|
||||
- `tensor_parallel_size >= decode_context_parallel_size`
|
||||
- `tensor_parallel_size % decode_context_parallel_size == 0`
|
||||
- For GQA-based model, such as Qwen3-235B:
|
||||
- `(tensor_parallel_size // num_key_value_heads) >= decode_context_parallel_size`
|
||||
- `(tensor_parallel_size // num_key_value_heads) % decode_context_parallel_size == 0`
|
||||
|
||||
- While using Context Parallel in KV cache transfer-needed scenario (e.g. KV pooling, PD disaggregation), to simplify KV cache transmission, `cp_kv_cache_interleave_size` must be set to the same value of KV cache `block_size`(default: 128), which specifies CP to split KV cache in a block-interleave style. For example:
|
||||
|
||||
```shell
|
||||
vllm serve deepseek-ai/DeepSeek-V2-Lite \
|
||||
--tensor-parallel-size 2 \
|
||||
--decode-context-parallel-size 2 \
|
||||
--prefill-context-parallel-size 2 \
|
||||
--cp-kv-cache-interleave-size 128 \
|
||||
--kv-transfer-config {...} \
|
||||
```
|
||||
|
||||
## Experimental Results
|
||||
|
||||
To evaluate the effectiveness of Context Parallel in long sequence LLM inference scenarios, we use **DeepSeek-R1-W8A8** and **Qwen3-235B**, deploy PD disaggregated instances in the environment of 64 cards Ascend Atlas A3 inference products*64G (A3), the configuration and performance data are as follows.
|
||||
|
||||
- DeepSeek-R1-W8A8:
|
||||
|
||||
| Configuration | Input length <br> 32k | Input length <br> 64k | Input length <br> 128k |
|
||||
| ----------------------------- | ------------------------- | ------------------------- | ------------------------- |
|
||||
| P node: (DP2 TP8 EP16) *2 <br> D node: (DP32 EP32)*1 | TTFT: 9.3s <br> TPOT: 72ms | TTFT: 22.8s <br> TPOT: 74ms | TTFT: 73.2s <br> TPOT: 82ms |
|
||||
| P node: (PCP2 TP8 DCP8 EP16) *2 <br> D node: (DP32 EP32)*1 | TTFT: 7.9s <br> TPOT: 74ms | TTFT: 15.9s <br> TPOT: 78ms | TTFT: 46.0s <br> TPOT: 83ms |
|
||||
|
||||
- Qwen3-235B:
|
||||
|
||||
| Configuration | Input length <br> 32k | Input length <br> 64k | Input length <br> 120k |
|
||||
| ----------------------------- | ------------------------- | ------------------------- | ------------------------- |
|
||||
| P node: (DP2 TP8 EP16) *2 <br> D node: (DP32 EP32)*1 | TTFT: 5.1s <br> TPOT: 65ms | TTFT: 13.1s <br> TPOT: 85ms | TTFT: 33.9s <br> TPOT: 120ms |
|
||||
| P node: (PCP2 TP8 DCP2 EP16) *2 <br> D node: (DP32 EP32)*1 | TTFT: 3.0s <br> TPOT: 66ms | TTFT: 8.9s <br> TPOT: 86ms | TTFT: 22.7s <br> TPOT: 121ms |
|
||||
149
docs/source/user_guide/feature_guide/cpu_binding.md
Normal file
@@ -0,0 +1,149 @@
|
||||
# CPU Binding
|
||||
|
||||
**Starting from vllm-ascend v0.18.0rc1, CPU binding is enabled by default on
|
||||
ARM-based Ascend servers.**
|
||||
|
||||
**You usually do not need to configure it manually.** Set `enable_cpu_binding`
|
||||
only when you want to disable it or make the default explicit.
|
||||
|
||||
## Benefits of CPU Binding
|
||||
|
||||
CPU Binding improves **host-side scheduling** for multi-socket ARM servers with
|
||||
Ascend NPUs. It is designed to solve three common host-side inference performance issues:
|
||||
|
||||
- **Lower cross-NUMA traffic.** Worker processes stay closer to the CPU and
|
||||
memory resources selected for their active NPU, reducing remote NUMA access.
|
||||
- **Lower context-switch overhead from thread preemption.** Key runtime threads
|
||||
run on stable CPU ranges, reducing scheduler movement and CPU contention on
|
||||
busy hosts.
|
||||
- **Better latency stability and multi-worker isolation.** Independent workers
|
||||
avoid sharing the same CPU/NUMA resources, which helps reduce tail-latency
|
||||
jitter and makes throughput more predictable during multi-NPU serving.
|
||||
|
||||
This feature is a host-side performance optimization. **It does not change model
|
||||
execution logic or numerical outputs.** When memory migration support is
|
||||
unavailable, CPU affinity still works, but memory locality may be worse and
|
||||
latency or throughput may degrade.
|
||||
|
||||
## Usage
|
||||
|
||||
### Online Serving
|
||||
|
||||
Default behavior:
|
||||
|
||||
```bash
|
||||
vllm serve Qwen/Qwen2.5-7B-Instruct
|
||||
```
|
||||
|
||||
Disable CPU binding:
|
||||
|
||||
```bash
|
||||
vllm serve Qwen/Qwen2.5-7B-Instruct \
|
||||
--additional-config '{"enable_cpu_binding": false}'
|
||||
```
|
||||
|
||||
### Offline Inference
|
||||
|
||||
Default behavior:
|
||||
|
||||
```python
|
||||
from vllm import LLM
|
||||
|
||||
llm = LLM(model="Qwen/Qwen2.5-7B-Instruct")
|
||||
```
|
||||
|
||||
Disable CPU binding:
|
||||
|
||||
```python
|
||||
from vllm import LLM
|
||||
|
||||
llm = LLM(
|
||||
model="Qwen/Qwen2.5-7B-Instruct",
|
||||
additional_config={"enable_cpu_binding": False},
|
||||
)
|
||||
```
|
||||
|
||||
## Requirements
|
||||
|
||||
Official vllm-ascend images have already included `util-linux` and `procps` /
|
||||
`procps-ng` in v0.18.0rc1 and earlier releases. **Starting from v0.18.0rc1, the
|
||||
official images also include `numactl`.**
|
||||
|
||||
If you are not using the official image, install the host tools manually:
|
||||
|
||||
```bash
|
||||
# Ubuntu/Debian
|
||||
sudo apt-get install -y util-linux numactl procps
|
||||
|
||||
# RHEL/CentOS/Alma/Rocky
|
||||
sudo yum install -y util-linux numactl procps-ng
|
||||
|
||||
# openEuler
|
||||
sudo dnf install -y util-linux numactl procps-ng
|
||||
```
|
||||
|
||||
**Without `numactl` / `migratepages`, vLLM Ascend skips only memory migration.**
|
||||
The worker process and runtime threads are still pinned, but pages already
|
||||
placed on remote NUMA nodes are not migrated, which **can reduce locality and
|
||||
degrade latency or throughput.**
|
||||
|
||||
For optimal locality, use a cpuset that is evenly distributed across NUMA
|
||||
nodes. Unbalanced cpusets may reduce the locality benefit of CPU binding.
|
||||
|
||||
On Ascend 950, CPU binding uses NPU-to-CPU affinity from `npu-smi info -t topo`
|
||||
to select the worker's affinity NUMA node. Each worker main process is pinned to
|
||||
one CPU cluster from that NUMA node. The cluster size is derived from `lscpu`
|
||||
`Thread(s) per core`: 8 CPUs when it is 1, and 16 CPUs when it is 2. Ascend 950
|
||||
also pins host `uvb_poll_window_thread` threads to NUMA0 CPUs except CPU0,
|
||||
constrained by the current cpuset. In Docker deployments, add `--pid=host` when
|
||||
creating the container so vLLM Ascend can discover and bind these host threads.
|
||||
Ascend 950 still can migrate memory pages when `migratepages` is available, but
|
||||
it does not separately pin ACL/release threads and does not apply IRQ binding.
|
||||
|
||||
For IRQ binding, the process also needs permission to read `/proc/interrupts`
|
||||
and write `/proc/irq/*/smp_affinity`. If `irqbalance` is running and the process
|
||||
can use `systemctl`, vLLM Ascend stops it before applying IRQ affinity. In
|
||||
containers where `systemctl` is unavailable, stop `irqbalance` on the host when
|
||||
IRQ affinity matters.
|
||||
|
||||
Ascend 950 does not apply IRQ binding. When running on Ascend 950, the log contains
|
||||
`[irq] IRQ binding skipped on Ascend 950.` and no `/proc/irq/*/smp_affinity` files are
|
||||
written by this feature.
|
||||
|
||||
Ascend 950 allocation logs use `worker=[...]` instead of `acl=[...]` or
|
||||
`release=[...]`, because ACL/release threads are not separately pinned on this
|
||||
device type. When UVB polling threads are found and bound, the log also reports
|
||||
their thread IDs and CPU pool:
|
||||
|
||||
```text
|
||||
Ascend 950 NPU0: worker=[...]
|
||||
[cpu_bind_ascend_950] uvb_poll_window_thread tids=[...] cpus=[...]
|
||||
```
|
||||
|
||||
On the host, stop `irqbalance` before starting vLLM when you need stable IRQ
|
||||
affinity:
|
||||
|
||||
```bash
|
||||
sudo systemctl stop irqbalance
|
||||
```
|
||||
|
||||
After the vLLM service exits, restart it if the host should return to the
|
||||
default IRQ balancing policy:
|
||||
|
||||
```bash
|
||||
sudo systemctl start irqbalance
|
||||
```
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
| Message | Meaning | Action |
|
||||
| --- | --- | --- |
|
||||
| `CPU binding skipped: non-ARM CPU detected.` | CPU binding only runs on ARM. | No action needed on x86_64. |
|
||||
| `Can not get running npu info.` | No running NPU was found, or `ASCEND_RT_VISIBLE_DEVICES` filtered all NPUs. | Check visible NPU IDs and `npu-smi info`. |
|
||||
| `Insufficient CPUs for binding...` | Fewer CPUs are available than the role split requires. Devices with IRQ binding need at least 5 CPUs per logical NPU. Ascend 950 needs one full cluster per worker. | Expand the cpuset or reduce visible NPUs. |
|
||||
| `NPU topo affinity not found...` | Topology affinity is unavailable. | On Ascend 950, worker CPU binding is skipped. On other topo-affinity devices, vLLM Ascend falls back to `global_slice`. Check `npu-smi info -t topo` when topology affinity is expected. |
|
||||
| `uvb_poll_window_thread not found... --pid=host` | Ascend 950 could not see host UVB polling threads. | Recreate the Docker container with `--pid=host`, then restart vLLM. |
|
||||
| `failed to bind uvb_poll_window_thread... --pid=host` | Ascend 950 found a UVB polling thread but failed to bind it. | Check permissions and recreate the Docker container with `--pid=host` if running in Docker. |
|
||||
| `The 'migratepages' command is not available...` | Memory migration is skipped, while CPU thread binding still proceeds. | Install `numactl` if NUMA locality or performance is affected. |
|
||||
| `[irq] IRQ binding skipped on Ascend 950.` | Ascend 950 does not use the IRQ binding step. | No action needed. Worker main binding and memory migration still proceed. |
|
||||
| `Bind cpus failed in rank...` | A binding step failed and CPU binding was skipped for that rank. | Check `taskset`, `lscpu`, `npu-smi`, cpuset size, and `/proc/irq` permissions. |
|
||||
54
docs/source/user_guide/feature_guide/dynamic_batch.md
Normal file
@@ -0,0 +1,54 @@
|
||||
# Dynamic Batch
|
||||
|
||||
Dynamic batch is a technique that dynamically adjusts the chunksize during each inference iteration within the chunked prefilling strategy according to the resources and SLO targets, thereby improving the effective throughput and decreasing the TBT.
|
||||
|
||||
Dynamic batch is controlled by the value of the `--SLO_limits_for_dynamic_batch`.
|
||||
Notably, only Atlas A2 inference products are supported with decode token number scales below 2048 so far.
|
||||
Especially, the improvements are quite obvious on Qwen, Llama models.
|
||||
We are working on further improvements and this feature will support more XPUs in the future.
|
||||
|
||||
## Getting started
|
||||
|
||||
### Prerequisites
|
||||
|
||||
1. Dynamic batch now depends on an offline cost model saved in a lookup table to refine the token budget. The lookup table is saved in a '.csv' file, which should be first downloaded from [A2-B3-BLK128.csv](https://vllm-ascend.obs.cn-north-4.myhuaweicloud.com/vllm-ascend/dynamic_batch_scheduler/A2-B3-BLK128.csv), renamed, and saved to the path `vllm_ascend/core/profile_table.csv`
|
||||
|
||||
2. `Pandas` is needed to load the lookup table, in case pandas is not installed.
|
||||
|
||||
```bash
|
||||
pip install pandas
|
||||
```
|
||||
|
||||
### Tuning Parameters
|
||||
|
||||
`--SLO_limits_for_dynamic_batch` is the tuning parameter (integer type) for the dynamic batch feature, larger values relax latency limitation, leading to higher effective throughput. The parameter can be selected according to the specific models or service requirements.
|
||||
|
||||
```python
|
||||
--SLO_limits_for_dynamic_batch = -1 # Default value; dynamic batching is disabled.
|
||||
--SLO_limits_for_dynamic_batch = 0 # Baseline value for dynamic batching; dynamic batching is disabled. FCFS and decode-first chunked prefilling strategy is used.
|
||||
--SLO_limits_for_dynamic_batch > 0 # User-defined positive value; dynamic batching is enabled. FCFS and decode-first chunked prefilling strategy is used.
|
||||
```
|
||||
|
||||
### Supported Models
|
||||
|
||||
So far, dynamic batch performs better on several dense models including Qwen and Llama (from 8B to 32B) with `tensor_parallel_size=8`. For different models, a proper `SLO_limits_for_dynamic_batch` parameter is needed. The empirical value of this parameter is generally `35, 50, or 75`. Therefore, some additional tests are needed to select the best parameter.
|
||||
|
||||
## Usage
|
||||
|
||||
Dynamic batch is used in the online inference. A fully executable example is as follows:
|
||||
|
||||
```shell
|
||||
SLO_LIMIT=50
|
||||
vllm serve Qwen/Qwen2.5-14B-Instruct\
|
||||
--additional_config '{"SLO_limits_for_dynamic_batch":'${SLO_LIMIT}'}' \
|
||||
--max-num-seqs 256 \
|
||||
--block-size 128 \
|
||||
--tensor_parallel_size 8 \
|
||||
--load_format dummy \
|
||||
--max_num_batched_tokens 1024 \
|
||||
--max-model-len 9000 \
|
||||
--host localhost \
|
||||
--port 12091 \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--trust-remote-code
|
||||
```
|
||||
@@ -0,0 +1,146 @@
|
||||
# Dynamic Chunked Pipeline Parallel
|
||||
|
||||
:::{note}
|
||||
For design details and mathematical models, see [Design Document](../../developer_guide/Design_Documents/dynamic_chunked_pipeline_parallel.md). For deployment tutorial, see [Dynamic Chunked Pipeline Parallel Tutorial](../../tutorials/features/dynamic_chunked_pipeline_parallel.md).
|
||||
:::
|
||||
|
||||
## Overview
|
||||
|
||||
Dynamic Chunked Pipeline Parallel (CPP) is a profiling-based dynamic chunking strategy that optimizes prefill performance for long sequences in Pipeline Parallelism (PP) scenarios.
|
||||
|
||||
### When to Use
|
||||
|
||||
- **Variable-length sequence serving**: PP does not introduce degradation on short sequences, and gains benefits through dynamic chunks on long sequences.
|
||||
- **Ultra-long sequence inference**: For sequences exceeding single-machine memory capacity (e.g., 1M tokens), dynamic chunking significantly reduces pipeline idle time.
|
||||
|
||||
## Supported Scenarios
|
||||
|
||||
Currently CPP mainly focuses on optimization during the prefill phase. CPP is recommended for PD (Prefill/Decode) disaggregation scenarios. Supported features are as follows:
|
||||
|
||||
| | Eager | Graph | Prefix <br> Cache | Chunked <br> Prefill |
|
||||
| ------- | ----- | ----- | ------ | ------ |
|
||||
| **CPP** | ✅ | ✅ | ✅ | ✅ |
|
||||
|
||||
## How to Enable
|
||||
|
||||
### Online Serving
|
||||
|
||||
```bash
|
||||
vllm serve <model_path> \
|
||||
--pipeline-parallel-size 2 \
|
||||
--enable-chunked-prefill \
|
||||
--additional-config '{"profiling_chunk_config": {"enabled": true}}'
|
||||
```
|
||||
|
||||
### Offline Inference
|
||||
|
||||
```python
|
||||
from vllm import LLM
|
||||
|
||||
llm = LLM(
|
||||
model="<model_path>",
|
||||
pipeline_parallel_size=2,
|
||||
additional_config={"profiling_chunk_config": {"enabled": True}},
|
||||
)
|
||||
```
|
||||
|
||||
## Configuration Parameters
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `enabled` | bool | False | Enable/disable Dynamic Chunked Pipeline Parallel |
|
||||
| `smooth_factor` | float | 1.0 | Smoothing factor (0 < x ≤ 1.0). Higher values trust dynamic prediction more |
|
||||
| `min_chunk` | int | 4096 | Minimum chunk size for dynamic calculation |
|
||||
| `need_timing` | bool | True | Enable/disable Online Calibration |
|
||||
| `max_fit_chunk` | int | 30 | Number of chunk-time data for Online Calibration |
|
||||
|
||||
### Parameter Tuning
|
||||
|
||||
- `smooth_factor`: Controls trust level in dynamic prediction
|
||||
- `1.0`: Strictly follow model prediction
|
||||
- `0.6~0.85`: Balance dynamic adjustment and scheduling overhead
|
||||
- `0.0`: No dynamic adjustment (degrades to fixed chunking)
|
||||
- `min_chunk`: Generally doesn't need adjustment. Should be smaller than `max-num-batched-tokens`
|
||||
|
||||
## Recommended Settings
|
||||
|
||||
### max-num-batched-tokens
|
||||
|
||||
**Notably, the TTFT of CPP is very sensitive to `max-num-batched-tokens` (considered as the initial chunk size for dynamic chunk calculation).** Because if it is too large, it will introduce significant computational waste, and if it is too small, it will lead to a decrease in operator efficiency. To leave enough room for dynamic adjustments, we recommend that the longer the sequence being processed, the larger the `max-num-batched-tokens` should be set. Recommended values:
|
||||
|
||||
| Sequence Length | `max-num-batched-tokens` |
|
||||
|-----------------|--------------------------|
|
||||
| 64k | 20480 |
|
||||
| 128k | 32768 |
|
||||
|
||||
### Online Calibration
|
||||
|
||||
For optimal performance, online calibrate with real data before production:
|
||||
|
||||
You can use ais_bench to generate fixed-length random datasets. Refer to [Using AISBench for performance evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation) for details.
|
||||
|
||||
1. Modify `<YOUR_AISBENCH_PATH>/benchmark/ais_bench/datasets/synthetic/synthetic_config.py`:
|
||||
|
||||
```python
|
||||
synthetic_config = {
|
||||
"Type": "string",
|
||||
"RequestCount": 5,
|
||||
"TrustRemoteCode": False,
|
||||
"StringConfig": {
|
||||
"Input": {
|
||||
"Method": "uniform",
|
||||
"Params": {"MinValue": 131072, "MaxValue": 131072} # Your max sequence length, max-model-len
|
||||
},
|
||||
"Output": {
|
||||
"Method": "uniform",
|
||||
"Params": {"MinValue": 1, "MaxValue": 1}
|
||||
}
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
2. Run for online calibration:
|
||||
|
||||
```bash
|
||||
ais_bench --models vllm_api_stream_chat --datasets synthetic_gen --mode perf --debug
|
||||
```
|
||||
|
||||
Configure online calibration data length to match your `max-model-len`. Use `batch_size=1` and ensure data differs to avoid cache hits if prefix caching is enabled.
|
||||
|
||||
## Performance
|
||||
|
||||
Refer to [Using AISBench for performance evaluation](../../developer_guide/evaluation/using_ais_bench.md#execute-performance-evaluation) for details.
|
||||
|
||||
To evaluate the effectiveness of Dynamic Chunked Pipeline Parallel in long sequence LLM inference scenarios, we use **DeepSeek-V3.1-W8A8** and **Qwen3-235B**, we deploy a P (Prefill) instance on Ascend Atlas A3 inference products (64 GB, A3), the configuration and performance data are as follows.
|
||||
|
||||
**Fixed-length requests, concurrency=1**:
|
||||
|
||||
- DeepSeek-V3.1-W8A8:
|
||||
|
||||
| Configuration | CPP <br> (Dynamic Chunk, <br> chunksize=32k) | PP <br>(Static Chunk, <br> chunksize=32k) |
|
||||
| ----------------------------- | ------------------------- | ------------------------- |
|
||||
| Input length 128k | TTFT: 22.5s | TTFT: 27.0s |
|
||||
|
||||
- Qwen3-235B:
|
||||
|
||||
| Configuration | CPP <br> (Dynamic Chunk, <br> chunksize=32k) | PP <br>(Static Chunk, <br> chunksize=32k) |
|
||||
| ----------------------------- | ------------------------- | ------------------------- |
|
||||
| Input length 256k | TTFT: 53.5s | TTFT: 61.4s |
|
||||
|
||||
**Variable-length requests, concurrency=4**:
|
||||
|
||||
- DeepSeek-V3.1-W8A8:
|
||||
|
||||
| Configuration | 4k~64k Input, mean=32k, std=32k <br> prefix hit rate=99% |
|
||||
| ----------------------------- | ------------------------- |
|
||||
| CPP2TP8 | Input throughput: 22424 tps/card |
|
||||
| DP2TP8 | Input throughput: 16150 tps/card |
|
||||
| PCP2TP8 | Input throughput: 18197 tps/card |
|
||||
| TP16 | Input throughput: 18875 tps/card |
|
||||
|
||||
## Constraints
|
||||
|
||||
- **Pipeline Parallelism Required**: `--pipeline-parallel-size > 1`
|
||||
- **Chunked Prefill Required**: `--enable-chunked-prefill`
|
||||
- **Incompatible with Balance Scheduling**: Cannot enable `VLLM_ASCEND_BALANCE_SCHEDULING`
|
||||
- **Startup Overhead**: Profiling adds ~64 forward passes (tens of seconds)
|
||||
85
docs/source/user_guide/feature_guide/epd_disaggregation.md
Normal file
@@ -0,0 +1,85 @@
|
||||
# Disaggregated-encoder
|
||||
|
||||
Disaggregated encoder refers to running the vision (multimodal) encoder stage of a large language model (LLM) in a separate vLLM process/instance from the language model's prefill and decode stages.
|
||||
|
||||
Similarly, disaggregated prefill isolates prompt processing (KV cache computation) from autoregressive token generation (the decode phase) across distinct vLLM instances.
|
||||
|
||||
This separation allows for targeted hardware and resource optimization for each phase, enabling precise tuning of Time-to-First-Token (TTFT) against Inter-Token Latency (ITL). Consequently, it enhances overall throughput and resource utilization during high-load serving.
|
||||
|
||||
Prefill-Decode (PD) disaggregation serves as the overarching architecture for this mechanism. In this setup, dedicated prefill instances compute KV caches and transfer them—via specialized connectors such as MooncakeLayerwise—to decode instances for token generation.
|
||||
|
||||
Within frameworks like vLLM (including the Ascend Hardware Plugin), PD disaggregation often integrates with Encoder-Prefill-Decode (EPD) architectures for multimodal models, while supporting multi-node configurations with distributed load balancing.
|
||||
|
||||
Ultimately, these architectural patterns maximize inference efficiency by addressing the contrasting computational profiles of each stage: encoding and prefilling are compute-bound and bursty, whereas decoding is memory-bound and sustained.
|
||||
|
||||
## Why disaggregated-encoder?
|
||||
|
||||
A **disaggregated encoder** runs the vision-encoder stage of a multimodal LLM in a process that is separate from the prefill / decoder stage. Deploying these two stages in independent vLLM instances brings three practical benefits:
|
||||
|
||||
1. **Independent, fine-grained scaling**
|
||||
|
||||
* Vision encoders are lightweight, while language models are orders of magnitude larger.
|
||||
* The language model can be parallelised without affecting the encoder fleet.
|
||||
* Encoder nodes can be added or removed independently.
|
||||
|
||||
2. **Lower time-to-first-token (TTFT)**
|
||||
|
||||
* Language-only requests bypass the vision encoder entirely.
|
||||
* Encoder output is injected only at required attention layers, shortening the prefill critical path.
|
||||
|
||||
3. **Cross-process reuse and caching of encoder outputs**
|
||||
|
||||
* In-process encoders confine reuse to a single worker.
|
||||
* A remote, shared cache lets any worker retrieve existing embeddings, eliminating redundant computation.
|
||||
|
||||
Design doc: <https://docs.google.com/document/d/1aed8KtC6XkXtdoV87pWT0a8OJlZ-CpnuLLzmR8l9BAE/edit>
|
||||
|
||||
---
|
||||
|
||||
## Usage
|
||||
|
||||
The current reference pathway is **ExampleConnector**.
|
||||
The ready-to-run scripts below show the workflow:
|
||||
|
||||
1 Encoder instance + 1 PD instance:
|
||||
`examples/online_serving/disaggregated_encoder/disagg_1e1pd/`
|
||||
|
||||
1 Encoder instance + 1 Prefill instance + 1 Decode instance:
|
||||
`examples/online_serving/disaggregated_encoder/disagg_1e1p1d/`
|
||||
|
||||
---
|
||||
|
||||
## Development
|
||||
|
||||

|
||||
|
||||
Disaggregated encoding is implemented by running two parts:
|
||||
|
||||
* **Encoder instance** – a vLLM instance to perform vision encoding.
|
||||
* **Prefill/Decode (PD) instance(s)** – runs language prefill and decode.
|
||||
* PD can be in either a single normal instance with (E + PD) or in disaggregated instances with (E + P + D)
|
||||
|
||||
A connector transfers encoder-cache (EC) embeddings from the encoder instance to the PD instance.
|
||||
All related code is under `vllm/distributed/ec_transfer`.
|
||||
|
||||
## Key abstractions
|
||||
|
||||
* **ECConnector** – interface for retrieving EC caches produced by the encoder.
|
||||
* *Scheduler role* – checks cache existence and schedules loads.
|
||||
* *Worker role* – loads the embeddings into memory.
|
||||
|
||||
* **EPD Load Balancing Proxy**
|
||||
* *Multi-Path Scheduling Strategy* - dynamically diverts the multimodal request or text requests to the corresponding inference path
|
||||
* *Instance-Level Dynamic Load Balancing* - dispatches multimodal requests based on a least-loaded strategy, using a priority queue to balance the active token workload across instances.
|
||||
|
||||
We create the example setup with the **MooncakeLayerwiseConnector** from `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake_layerwise_connector.py` and refer to the `examples/disaggregated_prefill_v1/load_balance_proxy_layerwise_server_example.py` to facilitate the kv transfer between P and D. For step-by-step deployment and configuration of Mooncake, refer to the following guide:
|
||||
[https://docs.vllm.ai/projects/ascend/en/latest/tutorials/features/pd_disaggregation_mooncake_multi_node.html](https://docs.vllm.ai/projects/ascend/en/latest/tutorials/features/pd_disaggregation_mooncake_multi_node.html)
|
||||
|
||||
For the PD disaggregation part, when using MooncakeLayerwiseConnector: The request first enters the Decoder instance, the Decoder triggers a remote prefill task in reverse via the Metaserver. The Prefill node then executes inference and pushes KV Cache layer-wise to the Decoder, overlapping computation with transmission. Once the transfer is complete, the Decoder seamlessly continues with the subsequent token generation.
|
||||
`docs/source/developer_guide/Design_Documents/disaggregated_prefill.md` shows the brief idea about the disaggregated prefill.
|
||||
|
||||
## Limitations
|
||||
|
||||
* Disable `--mm-processor-cache-gb 0` if you want to use cross-process caching
|
||||
|
||||
* For the PD disaggregation part, refer to the limitations of PD decomposition
|
||||
@@ -0,0 +1,220 @@
|
||||
# Expert Parallelism Load Balancer (EPLB)
|
||||
|
||||
## Overview
|
||||
|
||||
Expert balancing for MoE (Mixture of Experts) models in LLM (Large Language) serving is essential for optimal performance. Dynamically changing experts during inference can negatively impact TTFT (Time To First Token) and TPOT (Time Per Output Token) due to stop-the-world operations. Our solution aims to minimize the negative impacts caused by the operation.
|
||||
|
||||
## EPLB Effects
|
||||
|
||||
- Reduced Latency: Dynamically balances expert loads to minimize TTFT and TPOT by distributing workloads evenly across experts.
|
||||
- Adaptive Scaling: Automatically adjusts to workload fluctuations while maintaining stable performance.
|
||||
|
||||
## Support Scenarios
|
||||
|
||||
### Models
|
||||
|
||||
All MoE models supported by vLLM-Ascend.
|
||||
But we have only verified the performance on deepseek-v3.1/r1 models.
|
||||
|
||||
> [!IMPORTANT]
|
||||
> Ascend 950 Products does not support using EPLB with quant type "W4A8MXFP4", "W4A16", "W4A16MXFP4".
|
||||
|
||||
### MOE QuantType
|
||||
|
||||
| QuantType | Supported Hardware |
|
||||
| ------------------------------- | --------------------------- |
|
||||
| W8A8 / W8A8-Dynamic | A2, A3 |
|
||||
| W4A8 (with fused MC2 enabled) | A2, A3 |
|
||||
| MXFP4 | Ascend 950 Products |
|
||||
| MXFP8 | Ascend 950 Products |
|
||||
|
||||
### Usage Recommendations
|
||||
|
||||
EPLB is not recommended in the following scenarios because the load-balancing benefit may not offset its runtime overhead:
|
||||
|
||||
- P node workloads with input sequences shorter than `1024` tokens.
|
||||
- D node workloads where the number of experts per die is `<= 8` (`<= 16` on 950DT), or where the per-die load is below `128` tokens.
|
||||
|
||||
> [!WARNING]
|
||||
> Meeting the above conditions may lead to performance degradation.
|
||||
> When there are around 8 experts per die, the EPLB benefit may be comparable to its overhead. Benchmark the actual workload and enable EPLB only after confirming a performance gain.
|
||||
|
||||
## How to Use EPLB
|
||||
|
||||
EPLB has three usage modes:
|
||||
|
||||
| Mode | Config in `eplb_config` | Env Variable |
|
||||
| ---- | ----------------------- | ------------ |
|
||||
| **Dynamic EPLB** | `dynamic_eplb: true` | `DYNAMIC_EPLB=true` |
|
||||
| **Recording** (generate expert map) | `expert_map_record_path` | `DYNAMIC_EPLB=true` or `EXPERT_MAP_RECORD=true` |
|
||||
| **Static EPLB** (load pre-recorded map) | `expert_map_path` | none required |
|
||||
|
||||
> [!IMPORTANT]
|
||||
> For Dynamic EPLB and Recording modes, the env variable acts as a safety guard: setting `dynamic_eplb: true` in config alone is not enough — the assertion requires `DYNAMIC_EPLB=true` or `EXPERT_MAP_RECORD=true`. Static EPLB (loading a pre-recorded map via `expert_map_path`) does **not** require an env variable.
|
||||
|
||||
### Dynamic EPLB
|
||||
|
||||
We need to add environment variable `export DYNAMIC_EPLB="true"` to enable vLLM-Ascend EPLB. Enable dynamic balancing with auto-tuned parameters. Adjust expert_heat_collection_interval and algorithm_execution_interval based on workload patterns. In the current version, we recommend using the following: policy of SwiftBalanceEplb(2).
|
||||
|
||||
| Parameter | Description | Default |
|
||||
| --- | --- | --- |
|
||||
| dynamic_eplb | Enable dynamic EPLB. | False |
|
||||
| expert_heat_collection_interval | Interval for collecting expert heat. | 600 |
|
||||
| algorithm_execution_interval | Interval for executing the balancing algorithm. | 50 |
|
||||
| eplb_policy_type | EPLB policy type. | 2 |
|
||||
| num_redundant_experts | Number of redundant experts. | 0 |
|
||||
| eplb_heat_collection_stage | Request stage used to collect expert heat. Available values: `all`, `prefill`, and `decode`. | `all` |
|
||||
|
||||
```shell
|
||||
graph TB
|
||||
A[start] --> B(collect_heat)
|
||||
B --> C(execute_algorithm)
|
||||
C --> D(update_layer one by one)
|
||||
D --> B
|
||||
D --> F[termination upon service termination]
|
||||
```
|
||||
|
||||
```shell
|
||||
# D node or colocation
|
||||
vllm serve Qwen/Qwen3-235B-A22 \
|
||||
--tensor-parallel-size 16 \
|
||||
--enable-expert-parallel \
|
||||
--additional-config '{ "eplb_config": {
|
||||
"dynamic_eplb": true,
|
||||
"expert_heat_collection_interval": 600,
|
||||
"algorithm_execution_interval": 50,
|
||||
"eplb_policy_type": 2,
|
||||
"num_redundant_experts": 16
|
||||
}}'
|
||||
|
||||
# P node
|
||||
vllm serve Qwen/Qwen3-235B-A22 \
|
||||
--tensor-parallel-size 16 \
|
||||
--enable-expert-parallel \
|
||||
--additional-config '{ "eplb_config": {
|
||||
"dynamic_eplb": true,
|
||||
"expert_heat_collection_interval": 50,
|
||||
"algorithm_execution_interval": 5,
|
||||
"eplb_policy_type": 2,
|
||||
"num_redundant_experts": 16
|
||||
}}'
|
||||
```
|
||||
|
||||
#### EPLB Policy Types
|
||||
|
||||
The `eplb_policy_type` parameter selects the balancing algorithm used during dynamic expert redistribution:
|
||||
|
||||
| Value | Policy | Description |
|
||||
|-------|--------|-------------|
|
||||
| `0` | Random | Randomly swaps experts between ranks. Suitable for basic testing only. |
|
||||
| `1` | DefaultEplb | Open-source EPLB algorithm. Adds redundant experts to the hottest, packs via balanced assignment with local constraint exchange. |
|
||||
| `2` | SwiftBalanceEplb | Optimized for low-bandwidth environments. Supports intra-node and inter-node expert redundancy, joint optimization of expert placement. **(Recommended)** |
|
||||
| `3` | FlashLB | Statistical method using sliding-window mean/variance/covariance of expert loads. Uses FlashTree layered search for optimal replica allocation and `minimize_redeploy` for incremental adjustment. Best for high-frequency load fluctuations. |
|
||||
|
||||
#### Selective Expert Heat Collection
|
||||
|
||||
The `eplb_heat_collection_stage` option is intended for prefill-decode aggregation scenarios. Prefill requests usually process many tokens in one iteration, while decode requests usually process fewer tokens. As a result, the expert workload distribution can differ between the two stages. Collecting heat from both stages may hide the imbalance of the stage whose latency you want to optimize.
|
||||
|
||||
> [!IMPORTANT]
|
||||
> Selective heat collection is currently implemented by the Ascend model runner V1. Dynamic EPLB, including this option, is not yet supported by the Ascend model runner V2.
|
||||
|
||||
Use `eplb_heat_collection_stage` to select the stage whose expert heat contributes to EPLB:
|
||||
|
||||
| Value | Behavior | Typical use |
|
||||
| ----- | -------- | ----------- |
|
||||
| `all` | Collect expert heat from both prefill and decode iterations. | General workloads; this is the default. |
|
||||
| `prefill` | Collect expert heat only from iterations classified as prefill. | Optimize prefill workload balance and TTFT. |
|
||||
| `decode` | Collect expert heat only from iterations classified as decode. | Optimize decode workload balance and TPOT. |
|
||||
|
||||
Choose the stage according to the actual workload. The following values can be used as initial tuning guidance:
|
||||
|
||||
- For workloads whose typical input sequence length is greater than `1024` tokens, start with `prefill`.
|
||||
- For workloads whose typical input sequence length is less than `1024` tokens but concurrency is greater than `1024`, try `decode` or `all`.
|
||||
- For other or mixed workloads, benchmark `all`, `prefill`, and `decode` against the target TTFT or TPOT before choosing a setting.
|
||||
|
||||
These thresholds are empirical starting points rather than strict requirements. Production traffic distribution, concurrency, model configuration, and hardware topology can all affect the optimal stage.
|
||||
|
||||
For example, to collect only prefill heat:
|
||||
|
||||
```shell
|
||||
export DYNAMIC_EPLB="true"
|
||||
|
||||
vllm serve Qwen/Qwen3-235B-A22 \
|
||||
--tensor-parallel-size 16 \
|
||||
--enable-expert-parallel \
|
||||
--additional-config '{ "eplb_config": {
|
||||
"dynamic_eplb": true,
|
||||
"expert_heat_collection_interval": 600,
|
||||
"algorithm_execution_interval": 50,
|
||||
"eplb_policy_type": 2,
|
||||
"num_redundant_experts": 16,
|
||||
"eplb_heat_collection_stage": "prefill"
|
||||
}}'
|
||||
```
|
||||
|
||||
To collect only decode heat, set:
|
||||
|
||||
```json
|
||||
{
|
||||
"eplb_config": {
|
||||
"dynamic_eplb": true,
|
||||
"eplb_heat_collection_stage": "decode"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
> [!NOTE]
|
||||
> Stage selection applies to dynamic EPLB heat collection. Internally, vLLM-Ascend classifies each forward iteration by comparing its padded scheduled token count with the maximum expected token count of a decode iteration. An iteration above the threshold is treated as prefill; an iteration at or below the threshold is treated as decode. Classification is therefore performed per forward iteration rather than per individual request.
|
||||
|
||||
When an iteration does not match the selected stage, its expert load is not accumulated and it does not advance the heat-collection interval. Once heat collection is complete, balancing calculation and layer-by-layer expert weight updates continue normally.
|
||||
|
||||
### Static EPLB
|
||||
|
||||
> [!WARNING]
|
||||
> Static EPLB is scheduled for removal in v0.25.1.
|
||||
|
||||
#### Initial Setup (Record Expert Map)
|
||||
|
||||
We need to add environment variable `export EXPERT_MAP_RECORD="true"` to record expert map. Generate the initial expert distribution map using expert_map_record_path. This creates a baseline configuration for future deployments.
|
||||
|
||||
```shell
|
||||
vllm serve Qwen/Qwen3-235B-A22 \
|
||||
--tensor-parallel-size 16 \
|
||||
--enable-expert-parallel \
|
||||
--additional-config '{ "eplb_config": {
|
||||
"expert_map_record_path": "/path/to/eplb.json",
|
||||
"num_redundant_experts": 16,
|
||||
"expert_heat_collection_interval": 400,
|
||||
"algorithm_execution_interval": 30
|
||||
}}'
|
||||
```
|
||||
|
||||
#### Subsequent Deployments (Use Recorded Map)
|
||||
|
||||
Load the pre-recorded expert map for consistent performance. This avoids recalculating distributions at runtime.
|
||||
|
||||
```shell
|
||||
vllm serve Qwen/Qwen3-235B-A22 \
|
||||
--tensor-parallel-size 16 \
|
||||
--enable-expert-parallel \
|
||||
--additional-config '{
|
||||
"eplb_config": {"expert_map_path": "/path/to/eplb.json"}
|
||||
}'
|
||||
```
|
||||
|
||||
## Critical Considerations
|
||||
|
||||
1. Parameter Tuning:
|
||||
- expert_heat_collection_interval: Higher values (e.g., 600+) for stable workloads; lower values (e.g., 50-100) for fluctuating traffic.
|
||||
- algorithm_execution_interval: Should be ≥ 50 to avoid premature balancing during startup.
|
||||
- num_redundant_experts: (num_experts + num_redundant_experts) must be divisible by the expert-parallel size.
|
||||
|
||||
2. Hardware Requirements:
|
||||
- Ensure that all NPUs have identical memory capacity and compute capabilities.
|
||||
- Network bandwidth must support expert redistribution traffic (≥ 10 Gbps recommended).
|
||||
- shm needs to be mounted for container
|
||||
|
||||
3. Monitoring & Validation:
|
||||
- Track metrics: Search for [Expert Hotness] in the log. We will calculate the peak-to-average ratio of the load for each layer at different ranks, and then find their mean and maximum values. Current means actual peak-to-average ratio, update means estimated peak-to-average ratio after algorithm adjustment.
|
||||
- Use vLLM monitor to detect imbalances during runtime.
|
||||
- Always verify expert map JSON structure before loading (validate with jq or similar tools).
|
||||
91
docs/source/user_guide/feature_guide/external_dp.md
Normal file
@@ -0,0 +1,91 @@
|
||||
# External DP
|
||||
|
||||
For larger-scale deployments especially, it can make sense to handle the orchestration and load balancing of data parallel ranks externally.
|
||||
|
||||
In this case, it's more convenient to treat each DP rank like a separate vLLM deployment, with its own endpoint, and have an external router balance HTTP requests between them, making use of appropriate real-time telemetry from each server for routing decisions.
|
||||
|
||||
## Getting Started
|
||||
|
||||
The functionality of [external DP](https://docs.vllm.ai/en/latest/serving/data_parallel_deployment/?h=external#external-load-balancing) is already natively supported by vLLM. In vllm-ascend we provide two enhanced functionalities:
|
||||
|
||||
1. A launch script that helps to launch multiple vLLM instances in one command.
|
||||
2. A request-length-aware load-balancing proxy for external DP.
|
||||
|
||||
This tutorial will introduce the usage of them.
|
||||
|
||||
### Prerequisites
|
||||
|
||||
- Python 3.10+
|
||||
- Install dependencies needed by load-balance proxy server:
|
||||
|
||||
```shell
|
||||
pip install fastapi httpx uvicorn
|
||||
```
|
||||
|
||||
## Starting External DP Servers
|
||||
|
||||
First, you need to have at least two vLLM servers running in data parallel. These can be mock servers or actual vLLM servers. Note that this proxy also works with only one vLLM server running, but will fall back to direct request forwarding which is meaningless.
|
||||
|
||||
You can start external vLLM DP servers one-by-one manually or using the launch script in `examples/external_online_dp`. For scenarios of large DP size across multiple nodes, we recommend using our launch script for convenience.
|
||||
|
||||
### Manually Launch
|
||||
|
||||
```shell
|
||||
# This example shows how to manually launch a vLLM service with DP size 2 in one node.
|
||||
vllm serve --host 0.0.0.0 --port 8100 --data-parallel-size 2 --data-parallel-rank 0 ... # vLLM DP0
|
||||
vllm serve --host 0.0.0.0 --port 8101 --data-parallel-size 2 --data-parallel-rank 1 ... # vLLM DP1
|
||||
```
|
||||
|
||||
### Use Launch Script
|
||||
|
||||
Firstly, you need to modify the `examples/external_online_dp/run_dp_template.sh` according to your vLLM configuration. Then you can use `examples/external_online_dp/launch_online_dp.py` to launch multiple vLLM instances in one command on each node. It will internally call `examples/external_online_dp/run_dp_template.sh` for each DP rank with proper DP-related parameters.
|
||||
|
||||
An example of running external DP in one single node:
|
||||
|
||||
```shell
|
||||
cd examples/external_online_dp
|
||||
# running DP4 TP4 in a node with 16 NPUs
|
||||
python launch_online_dp.py --dp-size 4 --tp-size 4 --dp-size-local 4 --dp-rank-start 0 --dp-address x.x.x.x --dp-rpc-port 12342
|
||||
```
|
||||
|
||||
An example of running external DP in two nodes:
|
||||
|
||||
```shell
|
||||
cd examples/external_online_dp
|
||||
# running DP4 TP4 in two nodes with 8 NPUs each
|
||||
# Node 0 holds DP0 DP1 and node 1 holds DP2 DP3
|
||||
# Here x.x.x.x:12342 is served as the common data parallel RPC address
|
||||
|
||||
# On node 0:
|
||||
python launch_online_dp.py --dp-size 4 --tp-size 4 --dp-size-local 2 --dp-rank-start 0 --dp-address x.x.x.x --dp-rpc-port 12342
|
||||
|
||||
# On node 1:
|
||||
python launch_online_dp.py --dp-size 4 --tp-size 4 --dp-size-local 2 --dp-rank-start 2 --dp-address x.x.x.x --dp-rpc-port 12342
|
||||
```
|
||||
|
||||
## Starting Load-balance Proxy Server
|
||||
|
||||
After all vLLM DP instances are launched, you can now launch the load-balance proxy server, which serves as an entrypoint for coming requests and load-balances them between vLLM DP instances.
|
||||
|
||||
The proxy server has the following features:
|
||||
|
||||
- Load balances requests to multiple vLLM servers based on request length.
|
||||
- Supports OpenAI-compatible `/v1/completions` and `/v1/chat/completions` endpoints.
|
||||
- Streams responses from backend servers to clients.
|
||||
|
||||
To run the proxy server, you need to specify the host and port for each vLLM DP Instance:
|
||||
|
||||
```shell
|
||||
# For example, we have already started two DP instances in single node:
|
||||
# python launch_online_dp.py --dp-size 2 --tp-size 8 --dp-size-local 2 --dp-rank-start 0 --dp-address x.x.x.x --dp-rpc-port 12342
|
||||
# By default, launch_online_dp.py will launch vLLM instances from starting port 9000,
|
||||
# so the vLLM ports for DP0 and DP1 are 9000 and 9001 respectively.
|
||||
# Then you can start the load-balance proxy server by:
|
||||
cd examples/external_online_dp
|
||||
python dp_load_balance_proxy_server.py \
|
||||
--host 0.0.0.0 --port 8000 \
|
||||
--dp-hosts 127.0.0.1 127.0.0.1 \
|
||||
--dp-ports 9000 9001 \
|
||||
```
|
||||
|
||||
After this, you can directly send requests to the proxy server and run DP with external load balancing.
|
||||
167
docs/source/user_guide/feature_guide/flash_attention.md
Normal file
@@ -0,0 +1,167 @@
|
||||
# Flash Attention 3
|
||||
|
||||
```{note}
|
||||
Flash Attention 3 on Ascend is currently in beta. The `flash_attn_npu` package required for FA3 has been open-sourced on GitHub.
|
||||
Please refer to the [flash-attention-npu repository](https://github.com/MinghuasLab/flash-attention-npu) for more details.
|
||||
```
|
||||
|
||||
This document shows how to enable Flash Attention 3 (FA3) in vLLM-Ascend. FA3 provides a training-inference consistent attention implementation for Ascend NPUs.
|
||||
|
||||
## Motivation
|
||||
|
||||
In RL training frameworks such as veRL, the attention computation during training uses Flash Attention. When vLLM-Ascend serves as the inference backend, the default Fused Infer Attention (FIA) implementation differs from the training-side Flash Attention, which can lead to training-inference inconsistency. To address this, vLLM-Ascend introduces the FA3 attention backend to maintain consistency with the training side.
|
||||
|
||||
FA3 is crucial for the following scenarios:
|
||||
|
||||
- **Training-inference consistency**: Ensures that the attention computation during inference matches the training side, which is essential for RL workflows (e.g., veRL) where inference results are used to compute training signals.
|
||||
- **Framework debugging**: Consistent attention implementations make it easier to debug issues by eliminating discrepancies between training and inference.
|
||||
- **Reinforcement Learning (RL)**: RL training often requires deterministic and consistent rollouts for reproducibility and stable training.
|
||||
|
||||
## Feature Comparison
|
||||
|
||||
The following table compares the features of `flash_attn_with_kvcache` between GPU FA3 and Ascend NPU FA3:
|
||||
|
||||
| Feature | GPU FA3 | NPU FA3 |
|
||||
|---------|---------|---------|
|
||||
| FP16 (float16) | ✅ | ✅ |
|
||||
| BF16 (bfloat16) | ✅ | ✅ |
|
||||
| Causal Attention | ✅ | ✅ |
|
||||
| Sliding Window Attention | ✅ | - |
|
||||
| MQA/GQA | ✅ | ✅ |
|
||||
| Paged KV Cache | ✅ | ✅ |
|
||||
| Rotary Position Embedding (RoPE) | ✅ | - |
|
||||
| ALiBi | - | - |
|
||||
| Softcapping | ✅ | - |
|
||||
| FP8 Quantization | ✅ | - |
|
||||
| Variable-length Sequences | ✅ | ✅ |
|
||||
|
||||
### Differences from GPU Implementation
|
||||
|
||||
The `flash_attn_with_kvcache` interface on NPU is semantically consistent with the GPU FA3 version in terms of API parameters. The key differences are:
|
||||
|
||||
1. **Unsupported features on NPU FA3**: Sliding window attention, RoPE, ALiBi, Softcapping, and FP8 quantization are not yet supported.
|
||||
2. **Graph capture**: The tiling of `flash_attn_with_kvcache` is processed on the host side and is currently being optimized. It does not support ACL graph capture (i.e., cannot be captured into a computational graph for acceleration). Please use `compilation_config={"cudagraph_mode": "PIECEWISE"}` when enabling FA3.
|
||||
|
||||
## Hardware Requirements
|
||||
|
||||
FA3 currently requires Ascend Atlas A2 and A3 inference NPUs.
|
||||
We will support other NPUs in the future.
|
||||
|
||||
## Software Requirements
|
||||
|
||||
FA3 requires the `flash_attn_npu` package, which provides the `flash_attn_npu_v3` module with the `flash_attn_with_kvcache` operator.
|
||||
|
||||
### Installation
|
||||
|
||||
To install the `flash_attn_npu` wheel package, refer to: <https://github.com/MinghuasLab/flash-attention-npu/blob/main/README.md#installation>.
|
||||
|
||||
## Enabling Flash Attention 3
|
||||
|
||||
To enable FA3, you need to:
|
||||
|
||||
1. Set the environment variable `export VLLM_BATCH_INVARIANT=1` to enable batch invariant mode
|
||||
2. Specify the attention backend as `FLASH_ATTN` via the LLM parameter `attention_backend="FLASH_ATTN"`
|
||||
|
||||
### Online Inference (Server Mode)
|
||||
|
||||
To start a vLLM server with FA3 enabled:
|
||||
|
||||
```bash
|
||||
VLLM_BATCH_INVARIANT=1 vllm serve Qwen/Qwen3-8B \
|
||||
--attention-backend FLASH_ATTN \
|
||||
--compilation-config '{"cudagraph_mode": "PIECEWISE"}'
|
||||
```
|
||||
|
||||
Then use the OpenAI-compatible client:
|
||||
|
||||
```python
|
||||
from openai import OpenAI
|
||||
|
||||
client = OpenAI(
|
||||
api_key="EMPTY",
|
||||
base_url="http://localhost:8000/v1",
|
||||
)
|
||||
|
||||
response = client.completions.create(
|
||||
model="Qwen/Qwen3-8B",
|
||||
prompt="The future of AI is",
|
||||
max_tokens=100,
|
||||
temperature=0.7,
|
||||
seed=42,
|
||||
)
|
||||
|
||||
print(response.choices[0].text)
|
||||
```
|
||||
|
||||
### Offline Inference
|
||||
|
||||
For offline batch inference with FA3:
|
||||
|
||||
```python
|
||||
import os
|
||||
os.environ["VLLM_BATCH_INVARIANT"] = "1"
|
||||
|
||||
from vllm import LLM, SamplingParams
|
||||
|
||||
prompts = [
|
||||
"The future of AI is",
|
||||
"Machine learning enables",
|
||||
"Deep learning models can"
|
||||
]
|
||||
|
||||
sampling_params = SamplingParams(
|
||||
temperature=0.7,
|
||||
max_tokens=100,
|
||||
seed=42,
|
||||
)
|
||||
|
||||
llm = LLM(
|
||||
model="Qwen/Qwen3-8B",
|
||||
tensor_parallel_size=1,
|
||||
attention_backend="FLASH_ATTN",
|
||||
compilation_config={"cudagraph_mode": "PIECEWISE"},
|
||||
)
|
||||
|
||||
outputs = llm.generate(prompts, sampling_params)
|
||||
|
||||
for output in outputs:
|
||||
prompt = output.prompt
|
||||
generated_text = output.outputs[0].text
|
||||
print(f"Prompt: {prompt!r}")
|
||||
print(f"Generated: {generated_text!r}\n")
|
||||
```
|
||||
|
||||
## Limitations
|
||||
|
||||
- **Package not yet open-sourced**: The `flash_attn_npu` package required for FA3 has not yet been released. External users cannot use FA3 until the package is available.
|
||||
- **Sliding window not supported**: FA3 does not support sliding window attention. Models that require sliding window need to use the default FIA backend.
|
||||
- **ACL graph capture not supported**: The tiling of `flash_attn_with_kvcache` is processed on the host side and currently does not support ACL graph capture. Please use `compilation_config={"cudagraph_mode": "PIECEWISE"}` when enabling FA3.
|
||||
- **RoPE not supported**: FA3 does not support rotary position embedding within the attention kernel. vLLM-Ascend patches this by using the PyTorch native RoPE fallback instead.
|
||||
- **ALiBi not supported**: FA3 does not support ALiBi (Attention with Linear Biases).
|
||||
- **Softcapping not supported**: FA3 does not support attention logit softcapping.
|
||||
- **FP8 quantization not supported**: FA3 does not support FP8 quantized attention.
|
||||
- **MLA and SFA not supported**: FA3 does not support Multi-head Latent Attention (MLA) or Sparse Flash Attention (SFA).
|
||||
|
||||
```{note}
|
||||
Enabling FA3 may cause performance degradation compared to the default FIA backend. This trade-off is intentional to guarantee training-inference consistency.
|
||||
```
|
||||
|
||||
## Tested Models
|
||||
|
||||
FA3 has been tested and verified on the following models:
|
||||
|
||||
- **Qwen3 (Dense)**: `Qwen/Qwen3-0.6B`, `Qwen/Qwen3-1.7B`, `Qwen/Qwen3-8B`
|
||||
- **Qwen3 (MoE)**: `Qwen/Qwen3-30B-A3B`
|
||||
|
||||
Other models have not been tested yet and will be supported in the future if not supported after being tested.
|
||||
|
||||
## Future Improvements
|
||||
|
||||
The FA3 feature is under active development. Planned improvements include:
|
||||
|
||||
- Open-source the `flash_attn_npu` package
|
||||
- Support ACL graph capture (host-side tiling optimization)
|
||||
- Support for additional NPUs series
|
||||
- Expanded model coverage
|
||||
- Performance optimizations
|
||||
- Additional testing and validation
|
||||
@@ -1,78 +1,302 @@
|
||||
# Graph Mode Guide
|
||||
|
||||
```{note}
|
||||
This feature is currently experimental. In future versions, there may be behavioral changes around configuration, coverage, performance improvement.
|
||||
```
|
||||
## Overview
|
||||
|
||||
This guide provides instructions for using Ascend Graph Mode with vLLM Ascend. Please note that graph mode is only available on V1 Engine. And only Qwen, DeepSeek series models are well tested from 0.9.0rc1. We'll make it stable and generalize in the next release.
|
||||
This guide explains how graph mode is used in vLLM Ascend.
|
||||
|
||||
## Getting Started
|
||||
vLLM already provides the generic graph-mode architecture, mode definitions, and compile integration. For those upstream concepts, see:
|
||||
|
||||
From v0.9.1rc1 with V1 Engine, vLLM Ascend will run models in graph mode by default to keep the same behavior with vLLM. If you hit any issues, please feel free to open an issue on GitHub and fallback to eager mode temporarily by set `enforce_eager=True` when initializing the model.
|
||||
- [CUDA Graphs](https://docs.vllm.ai/en/latest/design/cuda_graphs/)
|
||||
- [torch.compile](https://docs.vllm.ai/en/latest/design/torch_compile/)
|
||||
|
||||
There are two kinds for graph mode supported by vLLM Ascend:
|
||||
- **ACLGraph**: This is the default graph mode supported by vLLM Ascend. In v0.9.1rc1, only Qwen series models are well tested.
|
||||
- **TorchAirGraph**: This is the GE graph mode. In v0.9.1rc1, only DeepSeek series models are supported.
|
||||
This document focuses on the Ascend-specific view: how graph mode works on Ascend, which components are involved, how to configure them, and what constraints users should keep in mind.
|
||||
|
||||
## Current Status on Ascend
|
||||
|
||||
- Graph mode is currently available only on the **V1 Engine**.
|
||||
- **ACLGraph** (capture/replay via `torch.npu.NPUGraph`) is the runtime graph execution mechanism used by the default graph path on Ascend.
|
||||
- **Npugraph_ex** is a compile-time FX graph optimization layer, enabled by default in FULL/FULL_DECODE_ONLY modes. It optimizes the graph before ACLGraph captures it.
|
||||
- **XliteGraph** is an optional graph path for selected model families and environments.
|
||||
- In context parallel scenarios, `cudagraph_mode="FULL"` is not sufficiently supported yet.
|
||||
|
||||
## Graph Paths on Ascend
|
||||
|
||||
vLLM Ascend provides two graph paths:
|
||||
|
||||
| Graph Path | Default | Description | Since |
|
||||
|---|---|---|---|
|
||||
| ACLGraph (+ Npugraph_ex) | Yes | Compile-time FX optimization (Npugraph_ex) + runtime capture/replay (ACLGraph) | v0.9.0rc1 (Npugraph_ex since v0.15.0rc1) |
|
||||
| XliteGraph | No | Preconfigured graph path for selected model families. Requires separate installation | v0.11.0 |
|
||||
|
||||
## How Graph Mode Works on Ascend
|
||||
|
||||
The default graph path on Ascend involves two stages: **compile-time optimization** and **runtime capture/replay**. ACLGraph handles the runtime capture/replay. The compile-time stage differs by `cudagraph_mode`:
|
||||
|
||||
- **FULL_AND_PIECEWISE**: Default mode, same as the upstream vLLM strategy. The compile-time path follows PIECEWISE compilation, while the runtime may still use full-graph behavior for uniform decode batches.
|
||||
- **FULL / FULL_DECODE_ONLY**: Npugraph_ex optimizes the FX graph via npugraph_ex (`force_eager=True`, compile-time only, no capture). The optimized callable is then captured and replayed by ACLGraph at runtime.
|
||||
- **PIECEWISE**: Npugraph_ex is disabled. Only basic FX fusion passes are applied at compile-time. ACLGraph captures and replays the resulting callable at runtime.
|
||||
- **NONE**: No compilation or graph capture. The model runs in eager mode.
|
||||
|
||||
| `cudagraph_mode` | Compile-time | Runtime | Npugraph_ex |
|
||||
|---|---|---|---|
|
||||
| FULL_AND_PIECEWISE | Piecewise compilation path | Mixed: PIECEWISE for mixed batches, FULL-capable for uniform decode batches | Disabled |
|
||||
| FULL / FULL_DECODE_ONLY | Npugraph_ex FX optimization | ACLGraph capture/replay | Enabled |
|
||||
| PIECEWISE | Fusion pass only | ACLGraph capture/replay | Disabled |
|
||||
| NONE | None | Eager execution | Disabled |
|
||||
|
||||
Additionally, **XliteGraph** is available as an optional alternative graph path for selected model families (see [Using XliteGraph](#using-xlitegraph)).
|
||||
|
||||
## Using ACLGraph
|
||||
ACLGraph is enabled by default. Take Qwen series models as an example, just set to use V1 Engine is enough.
|
||||
|
||||
offline example:
|
||||
ACLGraph is the runtime graph capture/replay mechanism on Ascend. It is enabled automatically when graph mode is active (i.e., `cudagraph_mode` is not `NONE`), and does not require explicit configuration.
|
||||
|
||||
### Basic usage
|
||||
|
||||
Offline example:
|
||||
|
||||
```python
|
||||
import os
|
||||
|
||||
from vllm import LLM
|
||||
|
||||
model = LLM(model="Qwen/Qwen2-7B-Instruct")
|
||||
llm = LLM(model="path/to/Qwen3-0.6B")
|
||||
outputs = llm.generate("Hello, how are you?")
|
||||
```
|
||||
|
||||
Online example:
|
||||
|
||||
```bash
|
||||
vllm serve Qwen/Qwen3-0.6B
|
||||
```
|
||||
|
||||
### Explicit `cudagraph_mode` configuration
|
||||
|
||||
The generic `cudagraph_mode` options come from upstream vLLM. On Ascend, the final effective mode may still be adjusted according to platform and backend support, so the official vLLM CUDA Graphs document remains the canonical reference for mode semantics.
|
||||
|
||||
CLI example:
|
||||
|
||||
```bash
|
||||
vllm serve Qwen/Qwen3-0.6B \
|
||||
--compilation-config '{"cudagraph_mode": "PIECEWISE"}'
|
||||
```
|
||||
|
||||
Python example:
|
||||
|
||||
```python
|
||||
from vllm import LLM
|
||||
|
||||
llm = LLM(
|
||||
model="Qwen/Qwen3-0.6B",
|
||||
compilation_config={"cudagraph_mode": "PIECEWISE"},
|
||||
)
|
||||
```
|
||||
|
||||
For the detailed meaning of `NONE`, `PIECEWISE`, `FULL`, `FULL_DECODE_ONLY`, and `FULL_AND_PIECEWISE`, as well as the generic fallback policy, see the upstream [CUDA Graphs](https://docs.vllm.ai/en/latest/design/cuda_graphs/) design doc.
|
||||
|
||||
### Attention backend compatibility
|
||||
|
||||
Not all attention backends support all graph modes. vLLM checks attention backend compatibility during compatibility checks and, when possible, automatically adjusts `cudagraph_mode` to a more compatible mode instead of failing immediately. In practice, this means a requested full-graph mode may be narrowed to a mixed or piecewise mode, and if the backend cannot support graph execution at all, graph mode may be disabled.
|
||||
|
||||
On Ascend, the current attention backend support levels are:
|
||||
|
||||
| Attention backend | Declared support | Practical meaning |
|
||||
|---|---|---|
|
||||
| `attention_v1` | `ALWAYS` | Supports graph execution for mixed prefill/decode batches |
|
||||
| `context_parallel/attention_cp` | `ALWAYS` | Supports graph execution for mixed prefill/decode batches |
|
||||
| `mla_v1` | `UNIFORM_BATCH` | Graph execution is limited to uniform batches; full graph is more restricted |
|
||||
| `context_parallel/mla_cp` | `UNIFORM_BATCH` | Graph execution is limited to uniform batches; full graph is more restricted |
|
||||
| `sfa_v1` | `UNIFORM_BATCH` | Graph execution is limited to uniform batches; full graph is more restricted |
|
||||
| `context_parallel/sfa_cp` | `UNIFORM_BATCH` | Graph execution is limited to uniform batches; full graph is more restricted |
|
||||
|
||||
This is why the effective graph mode on Ascend may differ from the mode requested in configuration.
|
||||
|
||||
### Troubleshooting capture resource exhaustion
|
||||
|
||||
If ACLGraph capture fails because the configured graph sizes exceed the runtime resources available on the current stack, vLLM Ascend now raises a dedicated error with mitigation guidance. In practice, the most useful actions are:
|
||||
|
||||
- upgrade to a newer HDK/CANN stack if one is available;
|
||||
- reduce `cudagraph_capture_sizes` or `max_cudagraph_capture_size`;
|
||||
- prefer `FULL` or `FULL_DECODE_ONLY` when the workload is mostly uniform decode;
|
||||
- temporarily disable graph mode to confirm the issue is capture-related.
|
||||
|
||||
This is most likely to appear in `PIECEWISE` or `FULL_AND_PIECEWISE` configurations because those paths tend to capture more graphs than uniform full-graph decode.
|
||||
|
||||
## Using Npugraph_ex
|
||||
|
||||
As introduced in the [RFC](https://github.com/vllm-project/vllm-ascend/issues/4715), Npugraph_ex is a compile-time FX graph optimization layer that works together with ACLGraph. It optimizes the model's FX graph before ACLGraph captures it at runtime. Its performance benefits mainly come from fusing multiple operators into single kernels (e.g., add + rms_norm → npu_add_rms_norm) to reduce kernel launch overhead.
|
||||
|
||||
```{note}
|
||||
Atlas 300I DUO and Atlas 200I Pro do not support `enable_npugraph_ex`. Set --additional-config '{"ascend_compilation_config": {"enable_npugraph_ex":false}}'.
|
||||
```
|
||||
|
||||
### Default behavior
|
||||
|
||||
Npugraph_ex is **enabled by default** when `cudagraph_mode` is `FULL` or `FULL_DECODE_ONLY`. It is automatically disabled in `PIECEWISE` or `NONE` modes.
|
||||
|
||||
This means for most users, Npugraph_ex is active without any explicit configuration:
|
||||
|
||||
```python
|
||||
from vllm import LLM
|
||||
|
||||
# Npugraph_ex is enabled by default in FULL/FULL_DECODE_ONLY mode
|
||||
llm = LLM(model="path/to/Qwen2-7B-Instruct")
|
||||
outputs = llm.generate("Hello, how are you?")
|
||||
```
|
||||
|
||||
### Explicit configuration
|
||||
|
||||
To explicitly control Npugraph_ex:
|
||||
|
||||
Offline example:
|
||||
|
||||
```python
|
||||
from vllm import LLM
|
||||
|
||||
model = LLM(
|
||||
model="path/to/Qwen2-7B-Instruct",
|
||||
additional_config={
|
||||
"ascend_compilation_config": {
|
||||
"enable_npugraph_ex": True,
|
||||
}
|
||||
}
|
||||
)
|
||||
outputs = model.generate("Hello, how are you?")
|
||||
```
|
||||
|
||||
online example:
|
||||
Online example:
|
||||
|
||||
```shell
|
||||
vllm serve Qwen/Qwen2-7B-Instruct
|
||||
```bash
|
||||
vllm serve Qwen/Qwen2-7B-Instruct \
|
||||
--additional-config '{"ascend_compilation_config":{"enable_npugraph_ex":true}}'
|
||||
```
|
||||
|
||||
## Using TorchAirGraph
|
||||
To disable Npugraph_ex explicitly:
|
||||
|
||||
If you want to run DeepSeek series models with graph mode, you should use [TorchAirGraph](https://www.hiascend.com/document/detail/zh/Pytorch/700/modthirdparty/torchairuseguide/torchair_0002.html). In this case, additional config is required.
|
||||
```bash
|
||||
vllm serve Qwen/Qwen2-7B-Instruct \
|
||||
--additional-config '{"ascend_compilation_config":{"enable_npugraph_ex":false}}'
|
||||
```
|
||||
|
||||
offline example:
|
||||
### Static kernel compilation
|
||||
|
||||
Static kernel compilation is an **optional** feature that pre-compiles operator binaries with fixed shapes at compile time, reducing runtime overhead for networks with static or near-static shapes. It is **disabled by default** and must be explicitly enabled.
|
||||
|
||||
```{note}
|
||||
Enabling static kernel triggers a compilation pass during the graph capture phase at service startup. This may add **several minutes to tens of minutes** to the startup time depending on the number of operators to compile and model complexity. Once completed, subsequent request processing is not affected.
|
||||
```
|
||||
|
||||
Offline example:
|
||||
|
||||
```python
|
||||
import os
|
||||
from vllm import LLM
|
||||
|
||||
# TorchAirGraph is only work without chunked-prefill now
|
||||
model = LLM(model="deepseek-ai/DeepSeek-R1-0528", additional_config={"torchair_graph_config": {"enabled": True},"ascend_scheduler_config": {"enabled": True,}})
|
||||
model = LLM(
|
||||
model="path/to/Qwen2-7B-Instruct",
|
||||
additional_config={
|
||||
"ascend_compilation_config": {
|
||||
"enable_npugraph_ex": True,
|
||||
"enable_static_kernel": True,
|
||||
}
|
||||
}
|
||||
)
|
||||
outputs = model.generate("Hello, how are you?")
|
||||
```
|
||||
|
||||
online example:
|
||||
Online example:
|
||||
|
||||
```shell
|
||||
vllm serve Qwen/Qwen2-7B-Instruct --additional-config='{"torchair_graph_config": {"enabled": true},"ascend_scheduler_config": {"enabled": true,}}'
|
||||
```bash
|
||||
vllm serve Qwen/Qwen2-7B-Instruct \
|
||||
--additional-config '{"ascend_compilation_config":{"enable_npugraph_ex":true, "enable_static_kernel":true}}'
|
||||
```
|
||||
|
||||
You can find more detail about additional config [here](../configuration/additional_config.md).
|
||||
#### Verifying static kernel is active
|
||||
|
||||
The recommended way to verify static kernel is in effect is through **Ascend Profiling**:
|
||||
|
||||
1. Collect a profiling trace of your running model using [Ascend PyTorch Profiler](https://www.hiascend.com/document/detail/zh/Pytorch/2600/apiref/torchnpuCustomsapi/docs/zh/custom_APIs/torch_npu-profiler/torch_npu-profiler-profile.md) (`torch_npu.profiler`).
|
||||
2. Open the generated `op_statistic.csv` file.
|
||||
3. Look for operators whose `op_type` or `name` column contains the keyword **`static_kernel`**. If such entries exist, static kernel compilation has taken effect for those operators.
|
||||
|
||||
During the compilation phase, you will see a Python warning (visible by default):
|
||||
|
||||
```text
|
||||
Starting static kernel compilation, the build directory is <path>
|
||||
```
|
||||
|
||||
This confirms that compilation has been triggered. The absence of this message means static kernel was not enabled or the cached result was reused directly.
|
||||
|
||||
For more details about Npugraph_ex, see the [npugraph_ex guide](https://www.hiascend.com/document/detail/zh/Pytorch/2600/modthirdparty/torchairuseguide/docs/zh/overview.md).
|
||||
|
||||
## Using XliteGraph
|
||||
|
||||
XliteGraph is an optional path for Llama, Qwen dense series models, Qwen MoE series models, and Qwen3-VL. It requires Xlite to be installed and configured through `xlite_graph_config`.
|
||||
|
||||
Install Xlite first:
|
||||
|
||||
```bash
|
||||
pip install xlite
|
||||
```
|
||||
|
||||
Offline example:
|
||||
|
||||
```python
|
||||
from vllm import LLM
|
||||
|
||||
# Xlite supports decode-only mode by default.
|
||||
# Full mode can be enabled with "full_mode": True.
|
||||
llm = LLM(
|
||||
model="path/to/Qwen3-32B",
|
||||
tensor_parallel_size=8,
|
||||
additional_config={
|
||||
"xlite_graph_config": {
|
||||
"enabled": True,
|
||||
"full_mode": True,
|
||||
}
|
||||
},
|
||||
)
|
||||
outputs = llm.generate("Hello, how are you?")
|
||||
```
|
||||
|
||||
Online example:
|
||||
|
||||
```bash
|
||||
vllm serve path/to/Qwen3-32B \
|
||||
--tensor-parallel-size 8 \
|
||||
--additional-config '{"xlite_graph_config": {"enabled": true, "full_mode": true}}'
|
||||
```
|
||||
|
||||
For more details about Xlite, see the [Xlite README](https://atomgit.com/openeuler/GVirt/blob/master/xlite/README.md).
|
||||
|
||||
## Common Limitations and Caveats
|
||||
|
||||
- XliteGraph should be treated as an alternative graph path, not as a drop-in replacement for ACLGraph in all scenarios.
|
||||
- Model and backend coverage is still evolving, so a configuration that works for one model family may not yet be recommended for another.
|
||||
- Encoder-decoder models currently do not keep `FULL_AND_PIECEWISE`; on Ascend they fall back to `PIECEWISE` or `NONE` depending on compilation support.
|
||||
|
||||
## Fallback to Eager Mode
|
||||
|
||||
If both `ACLGraph` and `TorchAirGraph` fail to run, you should fallback to eager mode.
|
||||
If you encounter issues with graph mode, you can temporarily fall back to eager mode by setting `enforce_eager=True`.
|
||||
|
||||
offline example:
|
||||
If ACL graph capture fails with the confirmed stream-resource signature in the error text, such as `207008` together with `Stream resources are insufficient` or `Insufficient_Stream_Resources`, vLLM Ascend will re-raise that capture failure with targeted mitigation guidance. In practice, the main levers are: upgrading to a newer HDK/CANN stack, reducing `cudagraph_capture_sizes`, lowering `max_cudagraph_capture_size`, or preferring `FULL` / `FULL_DECODE_ONLY` when the workload is mostly uniform decode.
|
||||
|
||||
**Offline example:**
|
||||
|
||||
```python
|
||||
import os
|
||||
from vllm import LLM
|
||||
|
||||
model = LLM(model="someother_model_weight", enforce_eager=True)
|
||||
outputs = model.generate("Hello, how are you?")
|
||||
llm = LLM(model="path/to/your/model", enforce_eager=True)
|
||||
outputs = llm.generate("Hello, how are you?")
|
||||
```
|
||||
|
||||
online example:
|
||||
**Online example:**
|
||||
|
||||
```shell
|
||||
vllm serve Qwen/Qwen2-7B-Instruct --enforce-eager
|
||||
```bash
|
||||
vllm serve path/to/your/model --enforce-eager
|
||||
```
|
||||
|
||||
## References
|
||||
|
||||
- [CUDA Graphs](https://docs.vllm.ai/en/latest/design/cuda_graphs/)
|
||||
- [torch.compile](https://docs.vllm.ai/en/latest/design/torch_compile/)
|
||||
- [Xlite README](https://atomgit.com/openeuler/GVirt/blob/master/xlite/README.md)
|
||||
- [Npugraph_ex guide](https://www.hiascend.com/document/detail/zh/Pytorch/2600/modthirdparty/torchairuseguide/docs/zh/overview.md)
|
||||
- [Npugraph_ex RFC](https://github.com/vllm-project/vllm-ascend/issues/4715)
|
||||
- [ACL Graph Developer Guide](../../developer_guide/Design_Documents/ACL_Graph.md)
|
||||
|
||||
BIN
docs/source/user_guide/feature_guide/images/ai_qos1.png
Normal file
|
After Width: | Height: | Size: 1007 KiB |
BIN
docs/source/user_guide/feature_guide/images/ai_qos2.png
Normal file
|
After Width: | Height: | Size: 90 KiB |
|
After Width: | Height: | Size: 572 KiB |
BIN
docs/source/user_guide/feature_guide/images/layer_sharding.png
Normal file
|
After Width: | Height: | Size: 158 KiB |
|
After Width: | Height: | Size: 39 KiB |
|
After Width: | Height: | Size: 37 KiB |
BIN
docs/source/user_guide/feature_guide/images/rfork_flowchart.jpg
Normal file
|
After Width: | Height: | Size: 160 KiB |
@@ -6,9 +6,32 @@ This section provides a detailed usage guide of vLLM Ascend features.
|
||||
:caption: Feature Guide
|
||||
:maxdepth: 1
|
||||
graph_mode
|
||||
cpu_binding
|
||||
Ai_QoS_introduction_en
|
||||
quantization
|
||||
sleep_mode
|
||||
structured_output
|
||||
lora
|
||||
eplb_swift_balancer
|
||||
expert_parallelism_load_balancer
|
||||
netloader
|
||||
rfork
|
||||
dynamic_batch
|
||||
epd_disaggregation
|
||||
kv_pool
|
||||
layerwise_kv_pool
|
||||
kv_cache_cpu_offload
|
||||
recompute_cpu_offload
|
||||
external_dp
|
||||
large_scale_ep
|
||||
ucm_deployment
|
||||
Fine_grained_TP
|
||||
layer_sharding
|
||||
speculative_decoding
|
||||
context_parallel
|
||||
weight_prefetch
|
||||
sequence_parallelism
|
||||
batch_invariance
|
||||
lmcache_ascend_deployment
|
||||
dynamic_chunk_pipeline_parallel
|
||||
flash_attention
|
||||
:::
|
||||
|
||||
108
docs/source/user_guide/feature_guide/kv_cache_cpu_offload.md
Normal file
@@ -0,0 +1,108 @@
|
||||
# KV Cache CPU Offload Guide
|
||||
|
||||
## Overview
|
||||
|
||||
KV Cache CPU Offload enables offloading inactive KV cache blocks from NPU memory to CPU memory, allowing vLLM to handle longer contexts or more concurrent requests when NPU memory is limited. When a prefix cache miss occurs on the NPU but the data exists in CPU memory, the KV cache is asynchronously loaded back to the NPU, reducing recomputation latency.
|
||||
|
||||
This feature is built on vLLM's `OffloadingConnector` framework and provides an Ascend NPU-specific implementation (`NPUOffloadingSpec`) that uses dedicated NPU streams for efficient asynchronous data transfers between NPU and CPU.
|
||||
|
||||
## Key Concepts
|
||||
|
||||
- **CPU Block Pool**: A pre-allocated pool of CPU memory blocks (optionally pinned) used to store offloaded KV cache data.
|
||||
- **Asynchronous Transfer**: NPU-to-CPU (D2H) and CPU-to-NPU (H2D) transfers are performed on separate NPU streams, overlapping with computation to minimize latency impact.
|
||||
- **LRU Eviction**: The CPU-side block pool uses an LRU (Least Recently Used) eviction policy to manage limited CPU memory efficiently.
|
||||
|
||||
## Usage
|
||||
|
||||
### Python API
|
||||
|
||||
```python
|
||||
from vllm import LLM, SamplingParams
|
||||
from vllm.config import KVTransferConfig
|
||||
|
||||
kv_transfer_config = KVTransferConfig(
|
||||
kv_connector="OffloadingConnector",
|
||||
kv_role="kv_both",
|
||||
kv_connector_extra_config={
|
||||
"num_cpu_blocks": 1000,
|
||||
"block_size": 128,
|
||||
"spec_name": "NPUOffloadingSpec",
|
||||
"spec_module_path": "vllm_ascend.kv_offload.npu",
|
||||
},
|
||||
)
|
||||
|
||||
llm = LLM(
|
||||
model="Qwen/Qwen3-0.6B",
|
||||
gpu_memory_utilization=0.5,
|
||||
kv_transfer_config=kv_transfer_config,
|
||||
)
|
||||
|
||||
sampling_params = SamplingParams(max_tokens=100, temperature=0.0)
|
||||
outputs = llm.generate(["Hello, my name is"], sampling_params)
|
||||
for output in outputs:
|
||||
print(f"Prompt: {output.prompt!r}")
|
||||
print(f"Generated: {output.outputs[0].text!r}")
|
||||
```
|
||||
|
||||
### Online Serving
|
||||
|
||||
```bash
|
||||
vllm serve Qwen/Qwen3-0.6B \
|
||||
--gpu-memory-utilization 0.5 \
|
||||
--kv-transfer-config '{
|
||||
"kv_connector": "OffloadingConnector",
|
||||
"kv_role": "kv_both",
|
||||
"kv_connector_extra_config": {
|
||||
"num_cpu_blocks": 1000,
|
||||
"block_size": 128,
|
||||
"spec_name": "NPUOffloadingSpec",
|
||||
"spec_module_path": "vllm_ascend.kv_offload.npu"
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
## Configuration Parameters
|
||||
|
||||
- `kv_connector`: Must be set to `"OffloadingConnector"`.
|
||||
- `kv_role`: Set to `"kv_both"` to enable both storing and loading of KV cache.
|
||||
- `num_cpu_blocks`: Number of blocks to allocate in CPU memory. Increase this value for longer context scenarios. Each block consumes memory proportional to `block_size × num_layers × (key_size + value_size)`.
|
||||
- `block_size`: The CPU-side block size. Should be a multiple of the NPU-side block size. Typical value: `128`.
|
||||
- `spec_name`: Must be `"NPUOffloadingSpec"` for Ascend NPU.
|
||||
- `spec_module_path`: Must be `"vllm_ascend.kv_offload.npu"`.
|
||||
|
||||
## How It Works
|
||||
|
||||
1. **Normal inference**: KV cache blocks are computed and stored on the NPU as usual.
|
||||
2. **Eviction to CPU**: When NPU memory is full and new blocks are needed, inactive KV cache blocks are asynchronously copied to CPU memory via a dedicated D2H NPU stream.
|
||||
3. **Prefix cache hit (CPU)**: When a request shares a prefix with previously computed data, and the prefix cache is not found on NPU but exists in CPU memory, the KV cache blocks are asynchronously loaded back from CPU to NPU via a dedicated H2D NPU stream.
|
||||
4. **LRU management**: The CPU block pool uses LRU eviction to discard the least recently used blocks when CPU memory is full.
|
||||
|
||||
## Optional: KV Cache Events
|
||||
|
||||
You can enable KV cache event publishing for monitoring or debugging purposes:
|
||||
|
||||
```python
|
||||
from vllm.config import KVEventsConfig
|
||||
|
||||
kv_events_config = KVEventsConfig(
|
||||
enable_kv_cache_events=True,
|
||||
publisher="zmq",
|
||||
endpoint="tcp://*:5555",
|
||||
topic="kv_events",
|
||||
)
|
||||
|
||||
llm = LLM(
|
||||
model="Qwen/Qwen3-0.6B",
|
||||
gpu_memory_utilization=0.5,
|
||||
kv_transfer_config=kv_transfer_config,
|
||||
kv_events_config=kv_events_config,
|
||||
)
|
||||
```
|
||||
|
||||
## Notes
|
||||
|
||||
- This feature requires vLLM v1 engine.
|
||||
- Adjust `num_cpu_blocks` based on available CPU memory. Using too many blocks may cause out-of-memory errors on the host.
|
||||
- Pinned (page-locked) memory is used when available for optimal transfer performance.
|
||||
- The `gpu_memory_utilization` parameter controls how much NPU memory is reserved for KV cache. Lower values leave less NPU memory for KV cache, making offloading more active.
|
||||
- For production workloads, benchmark with realistic request patterns to find the optimal `num_cpu_blocks` and `block_size` settings.
|
||||
1186
docs/source/user_guide/feature_guide/kv_pool.md
Normal file
496
docs/source/user_guide/feature_guide/large_scale_ep.md
Normal file
@@ -0,0 +1,496 @@
|
||||
# Distributed DP Server With Large-Scale Expert Parallelism
|
||||
|
||||
## Getting Started
|
||||
|
||||
vLLM-Ascend now supports prefill-decode (PD) disaggregation in the large-scale **Expert Parallelism (EP)** scenario. To achieve better performance, the distributed DP server is applied in vLLM-Ascend. In the PD separation scenario, different optimization strategies can be implemented based on the distinct characteristics of PD nodes, thereby enabling more flexible model deployment. \
|
||||
Taking the DeepSeek model as an example, using 8 Atlas 800T A3 servers to deploy the model. Assume the IP of the servers starts from 192.0.0.1 and ends by 192.0.0.8. Use the first 4 servers as prefiller nodes and the last 4 servers as decoder nodes. And the prefiller nodes are deployed as master nodes independently, while the decoder nodes use the 192.0.0.5 node as the master node.
|
||||
|
||||
## Verify Multi-Node Communication Environment
|
||||
|
||||
### Physical Layer Requirements
|
||||
|
||||
- The physical machines must be located on the same LAN, with network connectivity.
|
||||
- All NPUs must be interconnected. For the Atlas A2 generation, intra-node connectivity is via HCCS, and inter-node connectivity is via RDMA. For the Atlas A3 generation, both intra-node and inter-node connectivity are via HCCS.
|
||||
|
||||
### Verification Process
|
||||
|
||||
:::::{tab-set}
|
||||
::::{tab-item} A3
|
||||
|
||||
1. Single Node Verification:
|
||||
|
||||
Execute the following commands on each node in sequence. The results must all be `success` and the status must be `UP`:
|
||||
|
||||
```bash
|
||||
# Check the remote switch ports
|
||||
for i in {0..15}; do hccn_tool -i $i -lldp -g | grep Ifname; done
|
||||
# Get the link status of the Ethernet ports (UP or DOWN)
|
||||
for i in {0..15}; do hccn_tool -i $i -link -g ; done
|
||||
# Check the network health status
|
||||
for i in {0..15}; do hccn_tool -i $i -net_health -g ; done
|
||||
# View the network detected IP configuration
|
||||
for i in {0..15}; do hccn_tool -i $i -netdetect -g ; done
|
||||
# View gateway configuration
|
||||
for i in {0..15}; do hccn_tool -i $i -gateway -g ; done
|
||||
# View NPU network configuration
|
||||
cat /etc/hccn.conf
|
||||
```
|
||||
|
||||
2. Get NPU IP Addresses
|
||||
|
||||
```bash
|
||||
for i in {0..15}; do hccn_tool -i $i -vnic -g;done
|
||||
```
|
||||
|
||||
3. Get superpodid and SDID
|
||||
|
||||
```bash
|
||||
for i in {0..15}; do npu-smi info -t spod-info -i $i -c 0;npu-smi info -t spod-info -i $i -c 1;done
|
||||
```
|
||||
|
||||
4. Cross-Node PING Test
|
||||
|
||||
```bash
|
||||
# Execute on the target node (replace 'x.x.x.x' with actual NPU IP address)
|
||||
for i in {0..15}; do hccn_tool -i $i -hccs_ping -g address x.x.x.x;done
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} A2
|
||||
|
||||
1. Single Node Verification:
|
||||
|
||||
Execute the following commands on each node in sequence. The results must all be `success` and the status must be `UP`:
|
||||
|
||||
```bash
|
||||
# Check the remote switch ports
|
||||
for i in {0..7}; do hccn_tool -i $i -lldp -g | grep Ifname; done
|
||||
# Get the link status of the Ethernet ports (UP or DOWN)
|
||||
for i in {0..7}; do hccn_tool -i $i -link -g ; done
|
||||
# Check the network health status
|
||||
for i in {0..7}; do hccn_tool -i $i -net_health -g ; done
|
||||
# View the network detected IP configuration
|
||||
for i in {0..7}; do hccn_tool -i $i -netdetect -g ; done
|
||||
# View gateway configuration
|
||||
for i in {0..7}; do hccn_tool -i $i -gateway -g ; done
|
||||
# View NPU network configuration
|
||||
cat /etc/hccn.conf
|
||||
```
|
||||
|
||||
2. Get NPU IP Addresses
|
||||
|
||||
```bash
|
||||
for i in {0..7}; do hccn_tool -i $i -ip -g;done
|
||||
```
|
||||
|
||||
3. Cross-Node PING Test
|
||||
|
||||
```bash
|
||||
# Execute on the target node (replace 'x.x.x.x' with actual NPU IP address)
|
||||
for i in {0..7}; do hccn_tool -i $i -ping -g address x.x.x.x;done
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
:::::
|
||||
|
||||
## Large-Scale EP model deployment
|
||||
|
||||
### Generate script with configurations
|
||||
|
||||
In the PD separation scenario, we provide an optimized configuration. You can use the following shell script for configuring the prefiller and decoder nodes respectively.
|
||||
|
||||
:::::{tab-set}
|
||||
|
||||
::::{tab-item} Prefiller node
|
||||
|
||||
```shell
|
||||
# run_dp_template.sh
|
||||
#!/bin/sh
|
||||
|
||||
# this obtained through ifconfig
|
||||
# nic_name is the network interface name corresponding to local_ip
|
||||
nic_name="xxxx"
|
||||
local_ip="xxxx"
|
||||
|
||||
# basic configuration for HCCL and connection
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=10
|
||||
export HCCL_BUFFSIZE=256
|
||||
|
||||
# obtain parameters from distributed DP server
|
||||
export VLLM_DP_SIZE=$1
|
||||
export VLLM_DP_MASTER_IP=$2
|
||||
export VLLM_DP_MASTER_PORT=$3
|
||||
export VLLM_DP_RANK_LOCAL=$4
|
||||
export VLLM_DP_RANK=$5
|
||||
export VLLM_DP_SIZE_LOCAL=$7
|
||||
|
||||
#pytorch_npu settings and vllm settings
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export VLLM_USE_MODELSCOPE="True"
|
||||
|
||||
# enable the distributed DP server
|
||||
export VLLM_WORKER_MULTIPROC_METHOD="fork"
|
||||
export VLLM_ASCEND_EXTERNAL_DP_LB_ENABLED=1
|
||||
|
||||
# The w8a8 weight can be obtained from https://www.modelscope.cn/models/vllm-ascend/DeepSeek-R1-W8A8
|
||||
# "--additional-config" is used to enable characteristics from vllm-ascend
|
||||
vllm serve vllm-ascend/DeepSeek-R1-W8A8 \
|
||||
--host 0.0.0.0 \
|
||||
--port $6 \
|
||||
--tensor-parallel-size 8 \
|
||||
--enable-expert-parallel \
|
||||
--seed 1024 \
|
||||
--served-model-name deepseek_r1 \
|
||||
--max-model-len 17000 \
|
||||
--max-num-batched-tokens 16384 \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 4 \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--quantization ascend \
|
||||
--speculative-config '{"num_speculative_tokens": 1, "method":"deepseek_mtp"}' \
|
||||
--enforce-eager \
|
||||
--kv-transfer-config \
|
||||
'{"kv_connector": "MooncakeConnectorV1",
|
||||
"kv_buffer_device": "npu",
|
||||
"kv_role": "kv_producer",
|
||||
"kv_parallel_size": "1",
|
||||
"kv_port": "20001",
|
||||
}' \
|
||||
--additional-config '{"enable_weight_nz_layout":true,"enable_prefill_optimizations":true}'
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Decoder node
|
||||
|
||||
```shell
|
||||
# run_dp_template.sh
|
||||
#!/bin/sh
|
||||
|
||||
# this obtained through ifconfig
|
||||
# nic_name is the network interface name corresponding to local_ip
|
||||
nic_name="xxxx"
|
||||
local_ip="xxxx"
|
||||
|
||||
# basic configuration for HCCL and connection
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=10
|
||||
export HCCL_BUFFSIZE=1024
|
||||
|
||||
# obtain parameters from distributed DP server
|
||||
export VLLM_DP_SIZE=$1
|
||||
export VLLM_DP_MASTER_IP=$2
|
||||
export VLLM_DP_MASTER_PORT=$3
|
||||
export VLLM_DP_RANK_LOCAL=$4
|
||||
export VLLM_DP_RANK=$5
|
||||
export VLLM_DP_SIZE_LOCAL=$7
|
||||
|
||||
#pytorch_npu settings and vllm settings
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export VLLM_USE_MODELSCOPE="True"
|
||||
|
||||
# enable the distributed DP server
|
||||
export VLLM_WORKER_MULTIPROC_METHOD="fork"
|
||||
export VLLM_ASCEND_EXTERNAL_DP_LB_ENABLED=1
|
||||
|
||||
# The w8a8 weight can be obtained from https://www.modelscope.cn/models/vllm-ascend/DeepSeek-R1-W8A8
|
||||
# "--additional-config" is used to enable characteristics from vllm-ascend
|
||||
vllm serve vllm-ascend/DeepSeek-R1-W8A8 \
|
||||
--host 0.0.0.0 \
|
||||
--port $6 \
|
||||
--tensor-parallel-size 1 \
|
||||
--enable-expert-parallel \
|
||||
--seed 1024 \
|
||||
--served-model-name deepseek_r1 \
|
||||
--max-model-len 17000 \
|
||||
--max-num-batched-tokens 256 \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 28 \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--quantization ascend \
|
||||
--speculative-config '{"num_speculative_tokens": 1, "method":"deepseek_mtp"}' \
|
||||
--kv-transfer-config \
|
||||
'{"kv_connector": "MooncakeConnectorV1",
|
||||
"kv_buffer_device": "npu",
|
||||
"kv_role": "kv_consumer",
|
||||
"kv_parallel_size": "1",
|
||||
"kv_port": "20001",
|
||||
}' \
|
||||
--additional-config '{"enable_weight_nz_layout":true}'
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
:::::
|
||||
|
||||
### Start Distributed DP Server for prefill-decode disaggregation
|
||||
|
||||
Execute the following Python file on all nodes to use the distributed DP server. (We recommend using this feature on the v0.9.1 official release)
|
||||
|
||||
:::::{tab-set}
|
||||
|
||||
::::{tab-item} Prefiller node
|
||||
|
||||
```python
|
||||
import multiprocessing
|
||||
import os
|
||||
import sys
|
||||
dp_size = 2 # total number of DP engines for decode/prefill
|
||||
dp_size_local = 2 # number of DP engines on the current node
|
||||
dp_rank_start = 0 # starting DP rank for the current node
|
||||
# dp_ip is different on prefiller nodes in this example
|
||||
dp_ip = "192.0.0.1" # master node IP for DP communication
|
||||
dp_port = 13395 # port used for DP communication
|
||||
engine_port = 9000 # starting port for all DP groups on the current node
|
||||
template_path = "./run_dp_template.sh"
|
||||
if not os.path.exists(template_path):
|
||||
print(f"Template file {template_path} does not exist.")
|
||||
sys.exit(1)
|
||||
def run_command(dp_rank_local, dp_rank, engine_port_):
|
||||
command = f"bash ./run_dp_template.sh {dp_size} {dp_ip} {dp_port} {dp_rank_local} {dp_rank} {engine_port_} {dp_size_local}"
|
||||
os.system(command)
|
||||
processes = []
|
||||
for i in range(dp_size_local):
|
||||
dp_rank = dp_rank_start + i
|
||||
dp_rank_local = i
|
||||
engine_port_ = engine_port + i
|
||||
process = multiprocessing.Process(target=run_command, args=(dp_rank_local, dp_rank, engine_port_))
|
||||
processes.append(process)
|
||||
process.start()
|
||||
for process in processes:
|
||||
process.join()
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Decoder node
|
||||
|
||||
```python
|
||||
import multiprocessing
|
||||
import os
|
||||
import sys
|
||||
dp_size = 64 # total number of DP engines for decode/prefill
|
||||
dp_size_local = 16 # number of DP engines on the current node
|
||||
dp_rank_start = 0 # starting DP rank for the current node. e.g. 0/16/32/48
|
||||
# dp_ip is the same on decoder nodes in this example
|
||||
dp_ip = "192.0.0.5" # master node IP for DP communication.
|
||||
dp_port = 13395 # port used for DP communication
|
||||
engine_port = 9000 # starting port for all DP groups on the current node
|
||||
template_path = "./run_dp_template.sh"
|
||||
if not os.path.exists(template_path):
|
||||
print(f"Template file {template_path} does not exist.")
|
||||
sys.exit(1)
|
||||
def run_command(dp_rank_local, dp_rank, engine_port_):
|
||||
command = f"bash ./run_dp_template.sh {dp_size} {dp_ip} {dp_port} {dp_rank_local} {dp_rank} {engine_port_} {dp_size_local}"
|
||||
os.system(command)
|
||||
processes = []
|
||||
for i in range(dp_size_local):
|
||||
dp_rank = dp_rank_start + i
|
||||
dp_rank_local = i
|
||||
engine_port_ = engine_port + i
|
||||
process = multiprocessing.Process(target=run_command, args=(dp_rank_local, dp_rank, engine_port_))
|
||||
processes.append(process)
|
||||
process.start()
|
||||
for process in processes:
|
||||
process.join()
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
:::::
|
||||
|
||||
Note that the prefiller nodes and the decoder nodes may have different configurations. In this example, each prefiller node is deployed as a master node independently, while the decoder nodes use the 192.0.0.5 node as the master node. This leads to differences in 'dp_size_local' and 'dp_rank_start'
|
||||
|
||||
## Example proxy for Distributed DP Server
|
||||
|
||||
In the PD separation scenario, we need a proxy to distribute requests. Execute the following commands to enable the example proxy:
|
||||
|
||||
```shell
|
||||
python load_balance_proxy_server_example.py \
|
||||
--port 8000 \
|
||||
--host 0.0.0.0 \
|
||||
--prefiller-hosts \
|
||||
192.0.0.1 \
|
||||
192.0.0.2 \
|
||||
192.0.0.3 \
|
||||
192.0.0.4 \
|
||||
--prefiller-hosts-num \
|
||||
2 2 2 2 \
|
||||
--prefiller-ports \
|
||||
9000 9000 9000 9000 \
|
||||
--prefiller-ports-inc \
|
||||
2 2 2 2\
|
||||
--decoder-hosts \
|
||||
192.0.0.5 \
|
||||
192.0.0.6 \
|
||||
192.0.0.7 \
|
||||
192.0.0.8 \
|
||||
--decoder-hosts-num \
|
||||
16 16 16 16 \
|
||||
--decoder-ports \
|
||||
9000 9000 9000 9000 \
|
||||
--decoder-ports-inc \
|
||||
16 16 16 16 \
|
||||
```
|
||||
|
||||
|Parameter | meaning |
|
||||
| --- | --- |
|
||||
| --port | Proxy service Port |
|
||||
| --host | Proxy service Host IP|
|
||||
| --prefiller-hosts | Hosts of prefiller nodes |
|
||||
| --prefiller-hosts-num | Number of repetitions for prefiller node hosts |
|
||||
| --prefiller-ports | Ports of prefiller nodes |
|
||||
| --prefiller-ports-inc | Number of increments for prefiller node ports |
|
||||
| --decoder-hosts | Hosts of decoder nodes |
|
||||
| --decoder-hosts-num | Number of repetitions for decoder node hosts |
|
||||
| --decoder-ports | Ports of decoder nodes |
|
||||
| --decoder-ports-inc | Number of increments for decoder node ports |
|
||||
|
||||
You can get the proxy program in the repository's examples, [load\_balance\_proxy\_server\_example.py](https://github.com/vllm-project/vllm-ascend/blob/v0.9.1-dev/examples/disaggregate_prefill_v1/load_balance_proxy_server_example.py)
|
||||
|
||||
## Benchmark
|
||||
|
||||
We recommend using aisbench tool to assess performance. [aisbench](https://github.com/AISBench/benchmark). Execute the following commands to install aisbench
|
||||
|
||||
```shell
|
||||
git clone https://github.com/AISBench/benchmark.git
|
||||
cd benchmark/
|
||||
pip3 install -e ./
|
||||
```
|
||||
|
||||
You need to cancel the http proxy before assessing performance, as follows:
|
||||
|
||||
```shell
|
||||
# unset proxy
|
||||
unset http_proxy
|
||||
unset https_proxy
|
||||
```
|
||||
|
||||
- You can place your datasets in the directory: `benchmark/ais_bench/datasets`
|
||||
- You can change the configuration in the directory :`benchmark/ais_bench/benchmark/configs/models/vllm_api` Take `vllm_api_stream_chat.py` as an example:
|
||||
|
||||
```python
|
||||
models = [
|
||||
dict(
|
||||
attr="service",
|
||||
type=VLLMCustomAPIChatStream,
|
||||
abbr='vllm-api-stream-chat',
|
||||
path="vllm-ascend/DeepSeek-R1-W8A8",
|
||||
model="dsr1",
|
||||
request_rate = 28,
|
||||
retry = 2,
|
||||
host_ip = "192.0.0.1", # Proxy service host IP
|
||||
host_port = 8000, # Proxy service Port
|
||||
max_out_len = 10,
|
||||
batch_size=1536,
|
||||
trust_remote_code=True,
|
||||
generation_kwargs = dict(
|
||||
temperature = 0,
|
||||
seed = 1024,
|
||||
ignore_eos=False,
|
||||
)
|
||||
)
|
||||
]
|
||||
```
|
||||
|
||||
- Taking the gsm8k dataset as an example, execute the following commands to assess performance.
|
||||
|
||||
```shell
|
||||
ais_bench --models vllm_api_stream_chat --datasets gsm8k_gen_0_shot_cot_str_perf --debug --mode perf
|
||||
```
|
||||
|
||||
- For more details on commands and parameters for aisbench, refer to [aisbench](https://github.com/AISBench/benchmark)
|
||||
|
||||
## Prefill & Decode Configuration Details
|
||||
|
||||
In the PD separation scenario, we provide an optimized configuration.
|
||||
|
||||
- **prefiller node**
|
||||
|
||||
1. set HCCL_BUFFSIZE=256
|
||||
2. add '--enforce-eager' command to 'vllm serve'
|
||||
3. Take '--kv-transfer-config' as follows:
|
||||
|
||||
```shell
|
||||
--kv-transfer-config \
|
||||
'{"kv_connector": "MooncakeConnectorV1",
|
||||
"kv_buffer_device": "npu",
|
||||
"kv_role": "kv_producer",
|
||||
"kv_parallel_size": "1",
|
||||
"kv_port": "20001",
|
||||
}'
|
||||
```
|
||||
|
||||
4. Take '--additional-config' as follows:
|
||||
|
||||
```shell
|
||||
--additional-config '{"enable_weight_nz_layout":true,"enable_prefill_optimizations":true}'
|
||||
```
|
||||
|
||||
- **decoder node**
|
||||
|
||||
1. set HCCL_BUFFSIZE=1024
|
||||
2. Take '--kv-transfer-config' as follows:
|
||||
|
||||
```shell
|
||||
--kv-transfer-config
|
||||
'{"kv_connector": "MooncakeConnectorV1",
|
||||
"kv_buffer_device": "npu",
|
||||
"kv_role": "kv_consumer",
|
||||
"kv_parallel_size": "1",
|
||||
"kv_port": "20001",
|
||||
}'
|
||||
```
|
||||
|
||||
3. Take '--additional-config' as follows:
|
||||
|
||||
```shell
|
||||
--additional-config '{"enable_weight_nz_layout":true}'
|
||||
```
|
||||
|
||||
### Parameters Description
|
||||
|
||||
1. '--additional-config' Parameter Introduction:
|
||||
|
||||
- **"enable_weight_nz_layout"**: Whether to convert quantized weights to NZ format to accelerate matrix multiplication.
|
||||
- **"enable_prefill_optimizations"**: Whether to enable DeepSeek models' prefill optimizations.
|
||||
<br>
|
||||
|
||||
2. Enable MTP
|
||||
Add the following command to your configurations.
|
||||
|
||||
```shell
|
||||
--speculative-config '{"num_speculative_tokens": 1, "method":"deepseek_mtp"}'
|
||||
```
|
||||
|
||||
### Recommended Configuration Example
|
||||
|
||||
For example, if the average input length is 3.5k, and the output length is 1.1k, the context length is 16k, the max length of the input dataset is 7k. In this scenario, we give a recommended configuration for distributed DP server with high EP. Here we use 4 nodes for prefill and 4 nodes for decode.
|
||||
|
||||
| node | DP | TP | EP | max-model-len | max-num-batched-tokens | max-num-seqs | gpu-memory-utilization |
|
||||
|----------|----|----|----|---------------|------------------------|--------------|-----------|
|
||||
| prefill | 2 | 8 | 16 | 17000 | 16384 | 4 | 0.9 |
|
||||
| decode | 64 | 1 | 64 | 17000 | 256 | 28 | 0.9 |
|
||||
|
||||
:::{note}
|
||||
Note that these configurations are not related to optimization. You need to adjust these parameters based on actual scenarios.
|
||||
:::
|
||||
|
||||
## FAQ
|
||||
|
||||
### 1. Prefiller nodes need to warm up
|
||||
|
||||
Since the computation of some NPU operators requires several rounds of warm-up to achieve best performance, we recommend preheating the service with some requests before conducting performance tests to achieve the best end-to-end throughput.
|
||||
80
docs/source/user_guide/feature_guide/layer_sharding.md
Normal file
@@ -0,0 +1,80 @@
|
||||
# Layer Sharding Linear Guide
|
||||
|
||||
## Overview
|
||||
|
||||
**Layer Sharding Linear** is a memory-optimization feature designed for large language model (LLM) inference. It addresses the high memory pressure caused by **repeated linear operators across many layers** that share identical structure but have distinct weights.
|
||||
|
||||
Instead of replicating all weights on every device, **Layer Sharding Linear shards the weights of a "series" of such operators across the NPU devices in a communication group**:
|
||||
|
||||
- The **i-th layer's linear weight** is stored **only on device `i % K`**, where `K` is the number of devices in the group.
|
||||
- Other devices hold a lightweight **shared dummy tensor** during initialization and fetch the real weight **on-demand** via asynchronous broadcast during the forward pass.
|
||||
|
||||
As illustrated in the figure below, this design enables broadcast to reach weights: while the current layer (e.g., MLA or MOE) is being computed, the system **asynchronously broadcasts the next layer's weight** in the background. Because the attention computation in the MLA module is sufficiently latency-bound, the weight transfer for `o_proj` is **fully overlapped with computation**, making the communication **latency-free from the perspective of end-to-end inference**.
|
||||
|
||||
This approach **preserves exact computational semantics** while **significantly reducing NPU memory footprint**, especially critical for:
|
||||
|
||||
- Extremely deep architectures (e.g., DeepSeek-V3/R1 with 61 layers);
|
||||
- Models using **[DSA-CP](https://github.com/vllm-project/vllm-ascend/pull/4702)** or **[FlashComm2](https://github.com/vllm-project/vllm-ascend/pull/4188)**, where the full `O` (output) projection matrix must reside in memory per layer;
|
||||
- Scenarios where **attention computation latency fully overlaps** (hides) the communication cost of weight broadcasting.
|
||||
|
||||
---
|
||||
|
||||
### Flowchart
|
||||
|
||||

|
||||
|
||||
> **Figure.** Layer Sharding Linear workflow: weights are sharded by layer across devices (top), and during forward execution (bottom), asynchronous broadcast **pre-fetches** the next layer's weight while the current layer computes-enabling **zero-overhead** weight loading.
|
||||
|
||||
---
|
||||
|
||||
## Getting Started
|
||||
|
||||
To enable **Layer Sharding Linear**, specify the target linear layers using the `--additional-config` argument when launching your inference job. For example, to shard the `o_proj` and `q_b_proj` layers, use:
|
||||
|
||||
```bash
|
||||
--additional-config '{
|
||||
"layer_sharding": ["o_proj", "q_b_proj"]
|
||||
}'
|
||||
```
|
||||
|
||||
> **Restriction**
|
||||
> Layer Sharding can only be enabled in PD disaggregated's **P node**.
|
||||
> Layer Sharding is not supported by RFork weight transfer. If `--load-format rfork` is used with `layer_sharding`, RFork transfer is bypassed and the model is loaded through the default model loader.
|
||||
|
||||
---
|
||||
|
||||
## Supported Scenarios
|
||||
|
||||
This feature delivers the greatest benefit in the following cases:
|
||||
|
||||
### FlashComm2-enabled
|
||||
|
||||
When using [FlashComm2](https://github.com/vllm-project/vllm-ascend/pull/4188), the full output projection (`o_proj`) matrix must be resident in memory for each layer. Layer sharding significantly reduces memory pressure by distributing these weights across devices.
|
||||
|
||||
**Example configuration:**
|
||||
|
||||
```bash
|
||||
export VLLM_ASCEND_FLASHCOMM2_PARALLEL_SIZE=1
|
||||
vllm serve \
|
||||
--model DeepSeek-V3/R1 \
|
||||
--additional-config '{
|
||||
"layer_sharding": ["o_proj"]
|
||||
}'
|
||||
```
|
||||
|
||||
### DSA-CP-enabled
|
||||
|
||||
With [DSA-CP](https://github.com/vllm-project/vllm-ascend/pull/4702), both `q_b_proj` and `o_proj` layers require large weight matrices to be stored per layer. Sharding these layers across NPUs helps fit extremely deep models (e.g., 61-layer architectures) into limited device memory.
|
||||
|
||||
Layer Sharding can only be enabled in PD disaggregated's **P node**.
|
||||
|
||||
**Example configuration:**
|
||||
|
||||
```bash
|
||||
export VLLM_ASCEND_ENABLE_FLASHCOMM1=1
|
||||
vllm serve \
|
||||
--model DeepSeek-V3.2 \
|
||||
--additional-config '{
|
||||
"layer_sharding": ["q_b_proj", "o_proj"]
|
||||
}'
|
||||
```
|
||||
241
docs/source/user_guide/feature_guide/layerwise_kv_pool.md
Normal file
@@ -0,0 +1,241 @@
|
||||
# Layerwise KV Pool
|
||||
|
||||
Layerwise mode is an optimization for the AscendStore KV Pool that saves and
|
||||
loads KV cache **layer by layer** instead of as a single bulk copy. By pipelining
|
||||
the transfer of one layer with the attention computation of the next, it reduces
|
||||
the stall that occurs when the entire KV cache must arrive before any forward
|
||||
progress can be made.
|
||||
|
||||
Layerwise mode works in **both PD-Mixed** (`kv_role: "kv_both"`) and **PD
|
||||
disaggregation** (`kv_role: "kv_producer"` / `"kv_consumer"`) scenarios. See the
|
||||
[KV Pool guide](kv_pool.md) for the general KV Pool architecture and backend
|
||||
setup.
|
||||
|
||||
## How It Works (Brief)
|
||||
|
||||
Without layerwise mode, the KV cache for a request is saved to (or loaded from)
|
||||
the pool as one bulk operation after the full forward pass completes (or before
|
||||
it starts). For long prompts this bulk transfer introduces a serialization
|
||||
stall.
|
||||
|
||||
Layerwise mode splits the save/load at the layer granularity:
|
||||
|
||||
1. **Saving** (producer / kv_both): after computing layer *i*'s attention, the
|
||||
KV for that layer is immediately sent to the pool backend. The next layer's
|
||||
computation proceeds in parallel with the transfer.
|
||||
2. **Loading** (consumer / kv_both): before computing layer *i*'s attention, the
|
||||
system waits for layer *i*'s KV to arrive from the pool
|
||||
(`wait_for_layer_load`), then proceeds. The transfer of layer *i+1* overlaps
|
||||
with the attention computation of layer *i*.
|
||||
|
||||
The net effect: save/load latency is amortized across the forward pass rather
|
||||
than concentrated as a single blocking step.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
Layerwise mode currently requires the **memcache** backend
|
||||
(`backend: "memcache"`). Install and configure memcache_hybrid before
|
||||
proceeding — see the [KV Pool guide](kv_pool.md) for memcache installation,
|
||||
config files (`mmc-meta.conf` / `mmc-local.conf`), and MetaService startup.
|
||||
|
||||
Additional setup:
|
||||
|
||||
```bash
|
||||
# Huge pages (required by memcache device transfer)
|
||||
echo 200000 > /proc/sys/vm/nr_hugepages
|
||||
|
||||
# Source memcache environment
|
||||
source /usr/local/memcache_hybrid/set_env.sh
|
||||
source /usr/local/memfabric_hybrid/set_env.sh
|
||||
|
||||
# Uniform hashing across nodes
|
||||
export PYTHONHASHSEED=0
|
||||
```
|
||||
|
||||
## Configuration
|
||||
|
||||
Add `use_layerwise: true` to the `AscendStoreConnector` extra config:
|
||||
|
||||
```json
|
||||
{
|
||||
"kv_connector": "AscendStoreConnector",
|
||||
"kv_role": "kv_both",
|
||||
"kv_connector_extra_config": {
|
||||
"backend": "memcache",
|
||||
"mooncake_rpc_port": "0",
|
||||
"use_layerwise": true
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Change `"kv_role"` to `"kv_producer"` or `"kv_consumer"` for PD disaggregation.
|
||||
|
||||
### Key Parameters
|
||||
|
||||
| Parameter | Default | Description |
|
||||
| :--- | :--- | :--- |
|
||||
| `use_layerwise` | `false` | Enable layer-by-layer KV save/load. Requires `backend: "memcache"`. |
|
||||
| `backend` | `"mooncake"` | Storage backend. Layerwise currently supports `"memcache"` only. |
|
||||
| `mooncake_rpc_port` | `"0"` | RPC port for the scheduler↔worker lookup service. Use `"0"` to auto-assign, or a unique port per instance. |
|
||||
| `layerwise_prefetch_layers` | `1` | Number of layers to prefetch ahead of the compute frontier. Higher values improve overlap at the cost of memory. |
|
||||
| `layerwise_max_transfer_blocks` | `0` (unlimited) | Maximum number of KV blocks per transfer batch. |
|
||||
| `layerwise_max_transfer_bytes` | `0` (unlimited) | Maximum bytes per transfer batch. |
|
||||
| `h2d_stagger_us` | `0` | Stagger delay (microseconds) between H2D copies across TP ranks to avoid bus contention. |
|
||||
| `discard_partial_chunks` | `true` (non-layerwise) / `false` (layerwise) | Whether to discard KV for incomplete chunk boundaries. Layerwise defaults to `false` to preserve partial layers. |
|
||||
|
||||
## Usage Scenarios
|
||||
|
||||
### PD-Mixed (kv_both)
|
||||
|
||||
A single vLLM instance acts as both producer and consumer. The pool serves as a
|
||||
shared prefix cache: completed requests' KV is saved layer by layer, and new
|
||||
requests with overlapping prefixes load KV layer by layer. No proxy is needed.
|
||||
|
||||
```bash
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1
|
||||
|
||||
python -m vllm.entrypoints.openai.api_server \
|
||||
--model /path/to/DeepSeek-V2-Lite \
|
||||
--port 8100 \
|
||||
--trust-remote-code \
|
||||
--enforce-eager \
|
||||
--no-enable-prefix-caching \
|
||||
--tensor-parallel-size 1 \
|
||||
--max-model-len 4096 \
|
||||
--max-num-batched-tokens 4096 \
|
||||
--kv-transfer-config '{
|
||||
"kv_connector": "AscendStoreConnector",
|
||||
"kv_role": "kv_both",
|
||||
"kv_connector_extra_config": {
|
||||
"backend": "memcache",
|
||||
"mooncake_rpc_port": "0",
|
||||
"use_layerwise": true
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
Send requests directly to port 8100 — no proxy required.
|
||||
|
||||
### PD Disaggregation (kv_producer + kv_consumer)
|
||||
|
||||
Separate prefiller and decoder instances. The prefiller saves KV layer by layer;
|
||||
the decoder loads it layer by layer. A layerwise proxy coordinates request
|
||||
routing and per-layer KV placement via its `/v1/metaserver` endpoint.
|
||||
|
||||
**Prefiller:**
|
||||
|
||||
```bash
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1
|
||||
|
||||
python -m vllm.entrypoints.openai.api_server \
|
||||
--model /path/to/DeepSeek-V2-Lite \
|
||||
--port 8100 \
|
||||
--trust-remote-code \
|
||||
--enforce-eager \
|
||||
--tensor-parallel-size 1 \
|
||||
--max-model-len 4096 \
|
||||
--kv-transfer-config '{
|
||||
"kv_connector": "AscendStoreConnector",
|
||||
"kv_role": "kv_producer",
|
||||
"kv_connector_extra_config": {
|
||||
"backend": "memcache",
|
||||
"mooncake_rpc_port": "0",
|
||||
"use_layerwise": true
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
**Decoder:**
|
||||
|
||||
```bash
|
||||
export ASCEND_RT_VISIBLE_DEVICES=2,3
|
||||
|
||||
python -m vllm.entrypoints.openai.api_server \
|
||||
--model /path/to/DeepSeek-V2-Lite \
|
||||
--port 8200 \
|
||||
--trust-remote-code \
|
||||
--enforce-eager \
|
||||
--tensor-parallel-size 1 \
|
||||
--max-model-len 4096 \
|
||||
--kv-transfer-config '{
|
||||
"kv_connector": "AscendStoreConnector",
|
||||
"kv_role": "kv_consumer",
|
||||
"kv_connector_extra_config": {
|
||||
"backend": "memcache",
|
||||
"mooncake_rpc_port": "0",
|
||||
"use_layerwise": true
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
**Layerwise proxy** (different from the standard disagg proxy — serves
|
||||
`/v1/metaserver`):
|
||||
|
||||
```bash
|
||||
python examples/disaggregated_prefill_v1/load_balance_proxy_layerwise_server_example.py \
|
||||
--host 127.0.0.1 \
|
||||
--port 9000 \
|
||||
--prefiller-hosts 127.0.0.1 \
|
||||
--prefiller-ports 8100 \
|
||||
--decoder-hosts 127.0.0.1 \
|
||||
--decoder-ports 8200
|
||||
```
|
||||
|
||||
> **Note:** The proxy `--host` must **not** be `0.0.0.0` (wildcard). The
|
||||
> decoder connects back to `host:port/v1/metaserver`, so use a reachable IP.
|
||||
|
||||
Send requests to the proxy:
|
||||
|
||||
```bash
|
||||
curl -s http://127.0.0.1:9000/v1/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "/path/to/DeepSeek-V2-Lite",
|
||||
"prompt": "Hello, my name is",
|
||||
"max_tokens": 32,
|
||||
"temperature": 0.0
|
||||
}'
|
||||
```
|
||||
|
||||
## Tuning
|
||||
|
||||
### Prefetch Depth
|
||||
|
||||
Increase `layerwise_prefetch_layers` (default `1`) to prefetch more layers
|
||||
ahead of the compute frontier. This increases transfer/compute overlap but uses
|
||||
more temporary buffers. Typical values: `1–4`.
|
||||
|
||||
### Transfer Batching
|
||||
|
||||
Use `layerwise_max_transfer_blocks` or `layerwise_max_transfer_bytes` to limit
|
||||
the size of each transfer batch. This prevents a single large layer from
|
||||
monopolizing the transfer bus. Set to `0` (default) for unlimited.
|
||||
|
||||
### H2D Stagger
|
||||
|
||||
On multi-TP deployments, H2D (host-to-device) copies for all TP ranks can
|
||||
contend on the PCIe/HCCS bus. Set `h2d_stagger_us` to spread them out (e.g.
|
||||
`100` for a 100 µs stagger between ranks).
|
||||
|
||||
## Supported Models
|
||||
|
||||
Layerwise mode integrates with the **MLA** (`mla_v1`) and **SFA** (`sfa_v1`)
|
||||
attention backends. DeepSeek-V2/V3 and other MLA-based models are supported.
|
||||
|
||||
Basic full attention (`attention_v1`) and all context-parallel (CP) variants
|
||||
(`mla_cp`, `sfa_cp`, `attention_cp`) do **not** yet integrate the layerwise
|
||||
wait/save calls. Layerwise + CP is future work.
|
||||
|
||||
## Limitations
|
||||
|
||||
* **Backend**: Only `memcache` is supported for layerwise mode (`mooncake` and
|
||||
`yuanrong` do not support `use_layerwise`).
|
||||
* **Hybrid KV cache**: Not supported — layerwise raises
|
||||
`NotImplementedError` when the model has multiple KV cache group families
|
||||
(hybrid MLA + sliding-window attention).
|
||||
* **Context parallel**: Layerwise is not yet integrated with CP attention
|
||||
backends.
|
||||
* **PD disaggregation proxy**: When using `kv_producer` / `kv_consumer`, the
|
||||
dedicated layerwise proxy
|
||||
(`load_balance_proxy_layerwise_server_example.py`) is required — the standard
|
||||
disagg proxy does not serve the `/v1/metaserver` endpoint.
|
||||
@@ -0,0 +1,94 @@
|
||||
# LMCache-Ascend Deployment Guide
|
||||
|
||||
## Overview
|
||||
|
||||
LMCache-Ascend is a community maintained plugin for running LMCache on the Ascend NPU.
|
||||
|
||||
We provide a simple deployment guide here. For further info about deployment notes, please refer to [LMCache-Ascend doc](https://github.com/LMCache/LMCache-Ascend/blob/main/README.md)
|
||||
|
||||
## Getting Started
|
||||
|
||||
### Clone LMCache-Ascend Repo
|
||||
|
||||
Our repo contains a kvcache ops submodule for ease of maintenance, therefore we recommend cloning the repo with submodules.
|
||||
|
||||
```bash
|
||||
cd /workspace
|
||||
git clone --recurse-submodules https://github.com/LMCache/LMCache-Ascend.git
|
||||
```
|
||||
|
||||
### Docker
|
||||
|
||||
```bash
|
||||
cd /workspace/LMCache-Ascend
|
||||
docker build -f docker/Dockerfile.a2.openEuler -t lmcache-ascend:v0.3.12-vllm-ascend-v0.11.0-openeuler .
|
||||
```
|
||||
|
||||
Once that is built, run it with the following cmd
|
||||
|
||||
```bash
|
||||
DEVICE_LIST="0,1,2,3,4,5,6,7"
|
||||
docker run -it \
|
||||
--privileged \
|
||||
--cap-add=SYS_RESOURCE \
|
||||
--cap-add=IPC_LOCK \
|
||||
-p 8000:8000 \
|
||||
-p 8001:8001 \
|
||||
--name lmcache-ascend-dev \
|
||||
-e ASCEND_VISIBLE_DEVICES=${DEVICE_LIST} \
|
||||
-e ASCEND_RT_VISIBLE_DEVICES=${DEVICE_LIST} \
|
||||
-e ASCEND_TOTAL_MEMORY_GB=32 \
|
||||
-e VLLM_TARGET_DEVICE=npu \
|
||||
-v /usr/local/Ascend/driver:/usr/local/Ascend/driver \
|
||||
-v /etc/localtime:/etc/localtime \
|
||||
-v /var/log/npu:/var/log/npu \
|
||||
-v /dev/davinci_manager:/dev/davinci_manager \
|
||||
-v /dev/devmm_svm:/dev/devmm_svm \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /etc/hccn.conf:/etc/hccn.conf \
|
||||
lmcache-ascend:v0.3.12-vllm-ascend-v0.11.0-openeuler \
|
||||
/bin/bash
|
||||
```
|
||||
|
||||
### Manual Installation
|
||||
|
||||
Assuming your working directory is `/workspace` and vllm/vllm-ascend have already been installed.
|
||||
|
||||
1. Install LMCache Repo
|
||||
|
||||
```bash
|
||||
NO_CUDA_EXT=1 pip install lmcache==0.3.12
|
||||
```
|
||||
|
||||
2. Install LMCache-Ascend Repo
|
||||
|
||||
```bash
|
||||
cd /workspace/LMCache-Ascend
|
||||
python3 -m pip install --no-build-isolation -e .
|
||||
```
|
||||
|
||||
### Usage
|
||||
|
||||
We introduce a dynamic KVConnector via LMCacheAscendConnectorV1Dynamic, therefore LMCache-Ascend Connector can be used via the kv transfer config in the two following setting.
|
||||
|
||||
#### Online serving
|
||||
|
||||
```bash
|
||||
python \
|
||||
-m vllm.entrypoints.openai.api_server \
|
||||
--port 8100 \
|
||||
--model /data/models/Qwen/Qwen3-32B \
|
||||
--trust-remote-code \
|
||||
--disable-log-requests \
|
||||
--block-size 128 \
|
||||
--kv-transfer-config '{"kv_connector":"LMCacheAscendConnector","kv_role":"kv_both"}'
|
||||
```
|
||||
|
||||
#### Offline
|
||||
|
||||
```python
|
||||
ktc = KVTransferConfig(
|
||||
kv_connector="LMCacheAscendConnector",
|
||||
kv_role="kv_both"
|
||||
)
|
||||
```
|
||||
@@ -1,14 +1,21 @@
|
||||
# LoRA Adapters Guide
|
||||
|
||||
## Overview
|
||||
Like vLLM, vllm-ascend supports LoRA as well. The usage and more details can be found in [vLLM official document](https://docs.vllm.ai/en/latest/features/lora.html).
|
||||
|
||||
You can refer to [Supported Models](https://docs.vllm.ai/en/latest/models/supported_models.html#list-of-text-only-language-models) to find which models support LoRA in vLLM.
|
||||
Like vLLM, vllm-ascend supports LoRA as well. The usage and more details can be found in [vLLM official document](https://docs.vllm.ai/en/latest/features/lora/).
|
||||
|
||||
You can run LoRA with ACLGraph mode now. Please refer to [Graph Mode Guide](./graph_mode.md) for a better LoRA performance.
|
||||
You can refer to [Supported Models](https://docs.vllm.ai/en/latest/models/supported_models/) to find which models support LoRA in vLLM.
|
||||
|
||||
You can run LoRA with ACLGraph mode now. Please refer to [Graph Mode Guide](./graph_mode.md) for better LoRA performance.
|
||||
|
||||
Address for downloading models:
|
||||
|
||||
- base model: <https://www.modelscope.cn/models/vllm-ascend/Llama-2-7b-hf/files>
|
||||
- loRA model: <https://www.modelscope.cn/models/vllm-ascend/llama-2-7b-sql-lora-test/files>
|
||||
|
||||
## Example
|
||||
We show a simple LoRA example here, which enables the ACLGraph mode as default.
|
||||
|
||||
We provide a simple LoRA example here, which enables the ACLGraph mode by default.
|
||||
|
||||
```shell
|
||||
vllm serve meta-llama/Llama-2-7b \
|
||||
@@ -16,8 +23,8 @@ vllm serve meta-llama/Llama-2-7b \
|
||||
--lora-modules '{"name": "sql-lora", "path": "/path/to/lora", "base_model_name": "meta-llama/Llama-2-7b"}'
|
||||
```
|
||||
|
||||
## Custom LoRA Operators
|
||||
## Note
|
||||
|
||||
We have implemented LoRA-related AscendC operators, such as bgmv_shrink, bgmv_expand, sgmv_shrink and sgmv_expand. You can find them under the "csrc/kernels" directory of [vllm-ascend repo](https://github.com/vllm-project/vllm-ascend.git).
|
||||
- We have implemented LoRA-related AscendC operators, such as bgmv_shrink, bgmv_expand, sgmv_shrink and sgmv_expand. You can find them under the `csrc/kernels` directory of [vllm-ascend repo](https://github.com/vllm-project/vllm-ascend/tree/main/csrc/kernels).
|
||||
|
||||
When you install vllm and vllm-ascend, those operators mentioned above will be compiled and installed automatically. If you don't want to use AscendC operators when you run vllm-ascend, you should set `COMPILE_CUSTOM_KERNELS=0` and reinstall vllm-ascend. To require more instructions about installation and compilation, you can refer to [installation guide](../../installation.md).
|
||||
- You can enable LoRA with dense or mixture-of-experts(MoE) models now ([PR #10977](https://github.com/vllm-project/vllm-ascend/pull/10977)). However, we haven't support expert-parallel(EP) or quantification yet when you run MoE models with LoRA.
|
||||
|
||||
97
docs/source/user_guide/feature_guide/netloader.md
Normal file
@@ -0,0 +1,97 @@
|
||||
# Netloader Guide
|
||||
|
||||
This guide provides instructions for using **Netloader** as a weight-loader plugin for acceleration in **vLLM Ascend**.
|
||||
|
||||
---
|
||||
|
||||
## Overview
|
||||
|
||||
Netloader leverages high-bandwidth peer-to-peer (P2P) transfers between NPU cards to load model weights. It is implemented as a plugin (via the `register_model_loader` API added in vLLM 0.10). The workflow is:
|
||||
|
||||
1. A **server** preloads a model.
|
||||
2. A new **client** instance requests weight transfer.
|
||||
3. After validating that the model and partitioning match, the client uses HCCL collective communication (send/recv) to receive weights in the same order as stored in the model.
|
||||
|
||||
The server runs alongside normal inference tasks via sub-threads and via `stateless_init_torch_distributed_process_group` in vLLM. The client thus takes over weight initialization without needing to load from storage.
|
||||
|
||||
### Flowchart
|
||||
|
||||

|
||||
|
||||
### Timing Diagram
|
||||
|
||||

|
||||
|
||||
### Application Scenarios
|
||||
|
||||
- **Reduce startup latency**: By reusing already loaded weights and transferring them directly between NPU cards, Netloader cuts down model loading time versus conventional remote/local pull strategies.
|
||||
- **Relieve network & storage load**: Avoid repeated downloads of weight files from remote repositories, thus reducing pressure on central storage and network traffic.
|
||||
- **Improve resource utilization & lower cost**: Faster loading allows less reliance on standby compute nodes; resources can be scaled up/down more flexibly.
|
||||
- **Enhance business continuity & high availability**: In failure recovery, new instances can quickly take over without long downtime, improving system reliability and user experience.
|
||||
|
||||
---
|
||||
|
||||
## Usage
|
||||
|
||||
To enable Netloader, pass `--load-format=netloader` and provide configuration via `--model-loader-extra-config` (as a JSON string). Below are the supported configuration fields:
|
||||
|
||||
| Field Name | Type | Description | Allowed Values / Notes |
|
||||
|--------------------|---------|------------------------------------------------------------------------------------------------------|--------------------------------------------------------------------------------------------------------------|
|
||||
| **SOURCE** | List | Weight data sources. Each item is a map with `device_id` and `sources`, specifying the rank and its endpoints (IP:port). <br>Example: `{"SOURCE": [{"device_id": 0, "sources": ["10.170.22.152:19374"]}, {"device_id": 1, "sources": ["10.170.22.152:11228"]}]}` <br>If omitted or empty, fallback to default loader. The SOURCE here is second priority. | A list of objects with keys `device_id: int` and `sources: List[str]` |
|
||||
| **MODEL** | String | The model name, used to verify consistency between client and server. | Defaults to the `--model` argument if not specified. |
|
||||
| **LISTEN_PORT** | Integer | Base port for the server listener. | The actual port = `LISTEN_PORT + RANK`. If omitted, a random valid port is chosen. Valid range: 1024–65535. If out of range, that server instance won’t open a listener. |
|
||||
| **INT8_CACHE** | String | Behavior for handling int8 parameters in quantized models. | One of `["hbm", "dram", "no"]`. <br> - `hbm`: copy original int8 parameters to high-bandwidth memory (HBM) (may cost a lot of HBM). <br> - `dram`: copy to DRAM. <br> - `no`: no special handling (may lead to divergence or unpredictable behavior). Default: `"no"`. |
|
||||
| **INT8_CACHE_NAME** | List | Names of parameters to which `INT8_CACHE` is applied (i.e. filtering). | Default: `None` (means no filtering—all parameters). |
|
||||
| **OUTPUT_PREFIX** | String | Prefix for writing per-rank listener address/port files in server mode. | If set, each rank writes to `{OUTPUT_PREFIX}{RANK}.txt` (text), content = `IP:Port`. |
|
||||
| **CONFIG_FILE** | String | Path to a JSON file specifying the above configuration. | If provided, the SOURCE inside this file has **first priority** (overrides SOURCE in other configs). |
|
||||
|
||||
---
|
||||
|
||||
## Example Commands & Placeholders
|
||||
|
||||
> Replace parts in `` `<...>` `` before running.
|
||||
|
||||
### Server
|
||||
|
||||
```shell
|
||||
VLLM_SLEEP_WHEN_IDLE=1 vllm serve <model_file> \
|
||||
--tensor-parallel-size 1 \
|
||||
--served-model-name <model_name> \
|
||||
--enforce-eager \
|
||||
--port `<port>` \
|
||||
--load-format netloader
|
||||
```
|
||||
|
||||
### Client
|
||||
|
||||
```shell
|
||||
export NETLOADER_CONFIG='{"SOURCE":[{"device_id":0, "sources": ["<server_IP>:<server_Port>"]}]}'
|
||||
|
||||
VLLM_SLEEP_WHEN_IDLE=1 ASCEND_RT_VISIBLE_DEVICES=<device_id_diff_from_server> \
|
||||
vllm serve <model_file> \
|
||||
--tensor-parallel-size 1 \
|
||||
--served-model-name <model_name> \
|
||||
--enforce-eager \
|
||||
--port <client_port> \
|
||||
--load-format netloader \
|
||||
--model-loader-extra-config="${NETLOADER_CONFIG}"
|
||||
```
|
||||
|
||||
#### Placeholder Descriptions
|
||||
|
||||
- `<model_file>`: Path to the model file
|
||||
- `<model_name>`: Model name (must match between server & client)
|
||||
- `<port>`: Base listening port on server
|
||||
- `<server_IP>` + `<server_Port>`: IP and port of the Netloader server (from server log)
|
||||
- `<device_id_diff_from_server>`: Client device ID (must differ from server’s)
|
||||
- `<client_port>`: Port on which client listens
|
||||
|
||||
After startup, you can test consistency by issuing inference requests with temperature = 0 and comparing outputs.
|
||||
|
||||
---
|
||||
|
||||
## Note & Caveats
|
||||
|
||||
- If Netloader is used, **each worker process** must bind a listening port. That port may be user-specified or assigned randomly. If user-specified, ensure it is available.
|
||||
- Netloader requires extra on-chip memory to establish HCCL connections (i.e. `HCCL_BUFFERSIZE`, default ~200 MB). Users should reserve sufficient capacity (e.g. via `--gpu-memory-utilization`).
|
||||
- It is recommended to set `VLLM_SLEEP_WHEN_IDLE=1` to mitigate unstable or slow connections/transmissions. Related info: [vLLM Issue #16660](https://github.com/vllm-project/vllm/issues/16660), [vLLM PR #16226](https://github.com/vllm-project/vllm/pull/16226).
|
||||
@@ -1,71 +1,112 @@
|
||||
# Quantization Guide
|
||||
|
||||
Model quantization is a technique that reduces the size and computational requirements of a model by lowering the data precision of the weights and activation values in the model, thereby saving the memory and improving the inference speed.
|
||||
Model quantization is a technique that reduces model size and computational overhead by lowering the numerical precision of weights and activations, thereby saving memory and improving inference speed.
|
||||
|
||||
Since 0.9.0rc2 version, quantization feature is experimentally supported in vLLM Ascend. Users can enable quantization feature by specifying `--quantization ascend`. Currently, only Qwen, DeepSeek series models are well tested. We’ll support more quantization algorithm and models in the future.
|
||||
`vLLM Ascend` supports multiple quantization methods. This guide provides instructions for using different quantization tools and running quantized models on vLLM Ascend.
|
||||
|
||||
## Install modelslim
|
||||
> **Note**
|
||||
>
|
||||
> You can choose to convert the model yourself or use the quantized model we uploaded.
|
||||
> See <https://www.modelscope.cn/models/vllm-ascend/Kimi-K2-Instruct-W8A8>.
|
||||
> Before you quantize a model, ensure sufficient RAM is available.
|
||||
|
||||
To quantize a model, users should install [ModelSlim](https://gitee.com/ascend/msit/blob/master/msmodelslim/README.md) which is the Ascend compression and acceleration tool. It is an affinity-based compression tool designed for acceleration, using compression as its core technology and built upon the Ascend platform.
|
||||
## Quantization Tools
|
||||
|
||||
Install modelslim:
|
||||
vLLM Ascend supports models quantized by two main tools: `ModelSlim` and `LLM-Compressor`.
|
||||
|
||||
### 1. ModelSlim (Recommended)
|
||||
|
||||
[ModelSlim](https://gitcode.com/Ascend/msmodelslim/blob/master/README.md) is an Ascend-friendly compression tool focused on acceleration, using compression techniques, and built for Ascend hardware. It includes a series of inference optimization technologies such as quantization and compression, aiming to accelerate large language dense models, MoE models, multimodal understanding models, multimodal generation models, etc.
|
||||
|
||||
#### Installation
|
||||
|
||||
To use ModelSlim for model quantization, install it from its [Git repository](https://gitcode.com/Ascend/msmodelslim):
|
||||
|
||||
```bash
|
||||
# The branch(br_release_MindStudio_8.1.RC2_TR5_20260624) has been verified
|
||||
git clone -b br_release_MindStudio_8.1.RC2_TR5_20260624 https://gitee.com/ascend/msit
|
||||
# Install 26.0.0 version, this is currently the latest stable branch
|
||||
git clone https://gitcode.com/Ascend/msmodelslim.git -b 26.0.0
|
||||
|
||||
cd msit/msmodelslim
|
||||
cd msmodelslim
|
||||
|
||||
bash install.sh
|
||||
pip install accelerate
|
||||
```
|
||||
|
||||
## Quantize model
|
||||
#### Model Quantization
|
||||
|
||||
:::{note}
|
||||
You can choose to convert the model yourself or use the quantized model we uploaded,
|
||||
see https://www.modelscope.cn/models/vllm-ascend/Kimi-K2-Instruct-W8A8
|
||||
This conversion process will require a larger CPU memory, please ensure that the RAM size is greater than 2TB
|
||||
:::
|
||||
The following example shows how to generate W8A8 quantized weights for the [Qwen3-MoE model](https://gitcode.com/Ascend/msmodelslim/blob/master/example/Qwen3-MOE/README.md).
|
||||
|
||||
### Adapts and change
|
||||
1. Ascend does not support the `flash_attn` library. To run the model, you need to follow the [guide](https://gitee.com/ascend/msit/blob/master/msmodelslim/example/DeepSeek/README.md#deepseek-v3r1) and comment out certain parts of the code in `modeling_deepseek.py` located in the weights folder.
|
||||
2. The current version of transformers does not support loading weights in FP8 quantization format. you need to follow the [guide](https://gitee.com/ascend/msit/blob/master/msmodelslim/example/DeepSeek/README.md#deepseek-v3r1) and delete the quantization related fields from `config.json` in the weights folder
|
||||
|
||||
### Generate the w8a8 weights
|
||||
**Quantization Script:**
|
||||
|
||||
```bash
|
||||
cd example/DeepSeek
|
||||
cd example/Qwen3-MOE
|
||||
|
||||
# Support multi-card quantization
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3,4,5,6,7
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:False
|
||||
export MODEL_PATH="/root/.cache/Kimi-K2-Instruct"
|
||||
export SAVE_PATH="/root/.cache/Kimi-K2-Instruct-W8A8"
|
||||
|
||||
python3 quant_deepseek_w8a8.py --model_path $MODEL_PATH --save_path $SAVE_PATH --batch_size 4
|
||||
# Set model and save paths
|
||||
export MODEL_PATH="/path/to/your/model"
|
||||
export SAVE_PATH="/path/to/your/quantized_model"
|
||||
|
||||
# Run quantization script
|
||||
python3 quant_qwen_moe_w8a8.py --model_path $MODEL_PATH \
|
||||
--save_path $SAVE_PATH \
|
||||
--anti_dataset ../common/qwen3-moe_anti_prompt_50.json \
|
||||
--calib_dataset ../common/qwen3-moe_calib_prompt_50.json \
|
||||
--trust_remote_code True
|
||||
```
|
||||
|
||||
Here is the full converted model files except safetensors:
|
||||
After quantization completes, the output directory will contain the quantized model files.
|
||||
|
||||
For more examples, refer to the [official examples](https://gitcode.com/Ascend/msmodelslim/tree/master/example).
|
||||
|
||||
### 2. LLM-Compressor
|
||||
|
||||
[LLM-Compressor](https://github.com/vllm-project/llm-compressor) is a unified compressed model library for faster vLLM inference.
|
||||
|
||||
#### Installation
|
||||
|
||||
```bash
|
||||
.
|
||||
|-- config.json
|
||||
|-- configuration.json
|
||||
|-- configuration_deepseek.py
|
||||
|-- generation_config.json
|
||||
|-- modeling_deepseek.py
|
||||
|-- quant_model_description.json
|
||||
|-- quant_model_weight_w8a8_dynamic.safetensors.index.json
|
||||
|-- tiktoken.model
|
||||
|-- tokenization_kimi.py
|
||||
`-- tokenizer_config.json
|
||||
pip install llmcompressor
|
||||
```
|
||||
|
||||
## Run the model
|
||||
#### Model Quantization
|
||||
|
||||
Now, you can run the quantized models with vLLM Ascend. Here is the example for online and offline inference.
|
||||
`LLM-Compressor` provides various quantization scheme examples.
|
||||
|
||||
### Offline inference
|
||||
##### Dense Quantization
|
||||
|
||||
An example to generate W8A8 dynamic quantized weights for dense model:
|
||||
|
||||
```bash
|
||||
# Navigate to LLM-Compressor examples directory
|
||||
cd examples/quantization/llm-compressor
|
||||
|
||||
# Run quantization script
|
||||
python3 w8a8_int8_dynamic.py
|
||||
```
|
||||
|
||||
##### MoE Quantization
|
||||
|
||||
An example to generate W8A8 dynamic quantized weights for MoE model:
|
||||
|
||||
```bash
|
||||
# Navigate to LLM-Compressor examples directory
|
||||
cd examples/quantization/llm-compressor
|
||||
|
||||
# Run quantization script
|
||||
python3 w8a8_int8_dynamic_moe.py
|
||||
```
|
||||
|
||||
For more content, refer to the [official examples](https://github.com/vllm-project/llm-compressor/tree/main/examples).
|
||||
|
||||
The quantization types currently supported by LLM-Compressor can be viewed in the `vllm_ascend/quantization/compressed_tensors_config.py` file.
|
||||
|
||||
## Running Quantized Models
|
||||
|
||||
Once you have a quantized model which is generated by **ModelSlim**, you can use vLLM Ascend for inference by specifying the `--quantization ascend` parameter to enable quantization features, while for models quantized by **LLM-Compressor**, it is not necessary to add this parameter.
|
||||
|
||||
### Offline Inference
|
||||
|
||||
```python
|
||||
import torch
|
||||
@@ -76,12 +117,20 @@ prompts = [
|
||||
"Hello, my name is",
|
||||
"The future of AI is",
|
||||
]
|
||||
# Set sampling parameters
|
||||
sampling_params = SamplingParams(temperature=0.6, top_p=0.95, top_k=40)
|
||||
|
||||
llm = LLM(model="{quantized_model_save_path}",
|
||||
max_model_len=2048,
|
||||
llm = LLM(model="/path/to/your/quantized_model",
|
||||
max_model_len=4096,
|
||||
trust_remote_code=True,
|
||||
# Enable quantization by specifying `quantization="ascend"`
|
||||
# Set appropriate TP and DP values
|
||||
tensor_parallel_size=2,
|
||||
data_parallel_size=1,
|
||||
# Set an unused port
|
||||
port=8000,
|
||||
# Set serving model name
|
||||
served_model_name="quantized_model",
|
||||
# Specify `quantization="ascend"` to enable quantization for models quantized by ModelSlim
|
||||
quantization="ascend")
|
||||
|
||||
outputs = llm.generate(prompts, sampling_params)
|
||||
@@ -91,36 +140,22 @@ for output in outputs:
|
||||
print(f"Prompt: {prompt!r}, Generated text: {generated_text!r}")
|
||||
```
|
||||
|
||||
### Online inference
|
||||
|
||||
Enable quantization by specifying `--quantization ascend`, for more details, see DeepSeek-V3-W8A8 [tutorial](https://vllm-ascend.readthedocs.io/en/latest/tutorials/multi_node.html)
|
||||
|
||||
## FAQs
|
||||
|
||||
### 1. How to solve the KeyError: 'xxx.layers.0.self_attn.q_proj.weight' problem?
|
||||
|
||||
First, make sure you specify `ascend` quantization method. Second, check if your model is converted by this `br_release_MindStudio_8.1.RC2_TR5_20260624` modelslim version. Finally, if it still doesn't work, please
|
||||
submit a issue, maybe some new models need to be adapted.
|
||||
|
||||
### 2. How to solve the error "Could not locate the configuration_deepseek.py"?
|
||||
|
||||
Please convert DeepSeek series models using `br_release_MindStudio_8.1.RC2_TR5_20260624` modelslim, this version has fixed the missing configuration_deepseek.py error.
|
||||
|
||||
### 3. When converting deepseek series models with modelslim, what should you pay attention?
|
||||
|
||||
When the mla portion of the weights used `W8A8_DYNAMIC` quantization, if torchair graph mode is enabled, please modify the configuration file in the CANN package to prevent incorrect inference results.
|
||||
|
||||
The operation steps are as follows:
|
||||
|
||||
1. Search in the CANN package directory used, for example:
|
||||
find /usr/local/Ascend/ -name fusion_config.json
|
||||
|
||||
2. Add `"AddRmsNormDynamicQuantFusionPass":"off",` and `"MultiAddRmsNormDynamicQuantFusionPass":"off",` to the fusion_config.json you find, the location is as follows:
|
||||
### Online Inference
|
||||
|
||||
```bash
|
||||
{
|
||||
"Switch":{
|
||||
"GraphFusion":{
|
||||
"AddRmsNormDynamicQuantFusionPass":"off",
|
||||
"MultiAddRmsNormDynamicQuantFusionPass":"off",
|
||||
# Corresponding to offline inference
|
||||
python -m vllm.entrypoints.api_server \
|
||||
--model /path/to/your/quantized_model \
|
||||
--max-model-len 4096 \
|
||||
--port 8000 \
|
||||
--tensor-parallel-size 2 \
|
||||
--data-parallel-size 1 \
|
||||
--served-model-name quantized_model \
|
||||
--trust-remote-code
|
||||
```
|
||||
|
||||
## References
|
||||
|
||||
- [ModelSlim GitCode](https://gitcode.com/Ascend/msmodelslim)
|
||||
- [LLM-Compressor GitHub](https://github.com/vllm-project/llm-compressor)
|
||||
- [vLLM Quantization Guide](https://docs.vllm.ai/en/latest/features/quantization/)
|
||||
|
||||
312
docs/source/user_guide/feature_guide/recompute_cpu_offload.md
Normal file
@@ -0,0 +1,312 @@
|
||||
# Recompute CPU Offload Guide
|
||||
|
||||
## Overview
|
||||
|
||||
`RecomputeCPUOffloadConnector` preserves the KV cache of requests that are
|
||||
preempted by the Decode-side recompute scheduler. When HBM KV blocks are not
|
||||
enough, `RecomputeScheduler` may preempt a running Decode request. Without this
|
||||
connector, the request falls back to the original recompute path and may be sent
|
||||
back to the Prefill node to run prefill again. With recompute CPU offload
|
||||
enabled, the already-computed KV blocks are copied from HBM to CPU DRAM before
|
||||
the HBM blocks are reused, and copied back to HBM when the request is scheduled
|
||||
again.
|
||||
|
||||
This feature is designed for online P/D disaggregation workloads where Decode
|
||||
nodes are tuned for decode throughput and cannot efficiently recompute long
|
||||
prefills after preemption. It is opt-in and focuses on correctness first.
|
||||
|
||||
In a typical Decode deployment, `max_num_batched_tokens` is often sized around
|
||||
`max_num_seqs * (1 + num_spec_tokens)` so that the Decode node mainly schedules
|
||||
one new token, plus speculative tokens if MTP is enabled, per request. This is
|
||||
efficient for decode but too small for recomputing long prompts after
|
||||
preemption. Recompute CPU offload avoids sending such requests back through the
|
||||
full prompt recompute path when their KV state can be preserved locally.
|
||||
|
||||
## Key Concepts
|
||||
|
||||
* **Recompute preemption**: When Decode-side HBM KV cache is exhausted,
|
||||
`RecomputeScheduler` can preempt a running request and later resume it.
|
||||
* **CPU DRAM preservation**: `RecomputeCPUOffloadConnector` stores the
|
||||
preempted request's computed KV blocks in CPU memory.
|
||||
* **H2D restore**: When the preempted request is scheduled again, the connector
|
||||
restores the preserved KV blocks before model forward.
|
||||
* **Fallback behavior**: If the connector is not configured, the recompute
|
||||
scheduler is not enabled, or CPU offload capacity is insufficient, vLLM
|
||||
falls back to the original recompute behavior.
|
||||
* **`MultiConnector` integration**: In P/D disaggregation, use
|
||||
`MultiConnector` to combine the P/D connector, such as `MooncakeConnectorV1`,
|
||||
with `RecomputeCPUOffloadConnector` on Decode nodes.
|
||||
|
||||
## Configuration Parameters
|
||||
|
||||
`RecomputeCPUOffloadConnector` is configured through `kv-transfer-config`.
|
||||
|
||||
| Parameter | Description |
|
||||
| :--- | :--- |
|
||||
| `kv_connector` | Must be set to `RecomputeCPUOffloadConnector`. |
|
||||
| `kv_role` | Set to `kv_consumer` on Decode nodes. |
|
||||
| `cpu_bytes_to_use_per_rank` | Optional and recommended. CPU memory budget in bytes used by each rank/card for recompute offload. If set, it overrides `cpu_bytes_to_use / world_size`. |
|
||||
| `cpu_bytes_to_use` | Optional. Total CPU memory budget in bytes for this vLLM instance. The connector divides it by `world_size` to get the per-rank budget. This is less direct than `cpu_bytes_to_use_per_rank` and is easier to misconfigure in DP deployments. The default is 8 GiB total. |
|
||||
| `enable_offload_prefix_caching` | Optional. Enables CPU block sharing for full hashed blocks with the same prefix-cache hash. The default is `false`; keep it disabled unless explicitly testing prefix sharing. |
|
||||
|
||||
Prefer `cpu_bytes_to_use_per_rank` when you want every rank/card to use the
|
||||
same offload capacity. For example, set `cpu_bytes_to_use_per_rank` to
|
||||
`17179869184` for 16 GiB per rank/card.
|
||||
|
||||
If you use `cpu_bytes_to_use`, remember that it is divided by `world_size`. In
|
||||
typical P/D Decode deployments, this means the value is divided by the active
|
||||
Decode DP size. For example, when starting a DP2TP8 Decode service, setting
|
||||
`cpu_bytes_to_use` to 16 GiB gives each DP rank's cards about 8 GiB of recompute
|
||||
offload space. To avoid ambiguity, use `cpu_bytes_to_use_per_rank` for new
|
||||
deployments.
|
||||
|
||||
`recompute_scheduler_enable` must also be enabled in `additional-config` on
|
||||
P/D-disaggregated Decode nodes:
|
||||
|
||||
```bash
|
||||
--additional-config '{"recompute_scheduler_enable":true}'
|
||||
```
|
||||
|
||||
```{note}
|
||||
`recompute_scheduler_enable` is only valid in P/D-disaggregated mode
|
||||
(`kv_role` is `kv_producer` or `kv_consumer`). Do not enable it in PD-mixed
|
||||
mode (`kv_role` is `kv_both`).
|
||||
```
|
||||
|
||||
## Docker Shared Memory
|
||||
|
||||
Recompute CPU offload allocates pinned CPU tensors for the offloaded KV blocks.
|
||||
When running in Docker, make sure the container's shared memory is large enough.
|
||||
If `--shm-size` is too small, a large per-card offload budget, such as 16 GiB
|
||||
per card, can fail to allocate CPU tensors and may cause the service to OOM or
|
||||
hang during startup.
|
||||
|
||||
For typical A3 deployments with a large recompute-offload budget, set
|
||||
`--shm-size=1024g` when starting the container. The following snippet follows
|
||||
the Docker style used by the DeepSeek-V4-Flash tutorial:
|
||||
|
||||
```bash
|
||||
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|-a3
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1024g \
|
||||
--net=host \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
--device /dev/davinci3 \
|
||||
--device /dev/davinci4 \
|
||||
--device /dev/davinci5 \
|
||||
--device /dev/davinci6 \
|
||||
--device /dev/davinci7 \
|
||||
--device /dev/davinci8 \
|
||||
--device /dev/davinci9 \
|
||||
--device /dev/davinci10 \
|
||||
--device /dev/davinci11 \
|
||||
--device /dev/davinci12 \
|
||||
--device /dev/davinci13 \
|
||||
--device /dev/davinci14 \
|
||||
--device /dev/davinci15 \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /etc/hccn.conf:/etc/hccn.conf \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
## Usage with P/D Disaggregation
|
||||
|
||||
On Decode nodes, configure `MultiConnector` with both the P/D connector and
|
||||
`RecomputeCPUOffloadConnector`. The P/D connector handles KV transfer from
|
||||
Prefill to Decode, while `RecomputeCPUOffloadConnector` handles Decode-side
|
||||
preemption preservation and restore.
|
||||
|
||||
The following example uses `MooncakeConnectorV1` for P/D KV transfer.
|
||||
|
||||
```bash
|
||||
python3 -m vllm.entrypoints.openai.api_server \
|
||||
--model /path/to/model \
|
||||
--port 8200 \
|
||||
--trust-remote-code \
|
||||
--enforce-eager \
|
||||
--tensor-parallel-size 1 \
|
||||
--data-parallel-size 1 \
|
||||
--max-model-len 32768 \
|
||||
--block-size 128 \
|
||||
--max-num-batched-tokens 4096 \
|
||||
--additional-config '{"recompute_scheduler_enable":true}' \
|
||||
--kv-transfer-config \
|
||||
'{
|
||||
"kv_connector": "MultiConnector",
|
||||
"kv_role": "kv_consumer",
|
||||
"kv_connector_extra_config": {
|
||||
"connectors": [
|
||||
{
|
||||
"kv_connector": "MooncakeConnectorV1",
|
||||
"kv_role": "kv_consumer",
|
||||
"kv_port": "28000",
|
||||
"kv_connector_extra_config": {
|
||||
"prefill": {
|
||||
"dp_size": 1,
|
||||
"tp_size": 1
|
||||
},
|
||||
"decode": {
|
||||
"dp_size": 1,
|
||||
"tp_size": 1
|
||||
}
|
||||
}
|
||||
},
|
||||
{
|
||||
"kv_connector": "RecomputeCPUOffloadConnector",
|
||||
"kv_role": "kv_consumer",
|
||||
"kv_connector_extra_config": {
|
||||
"cpu_bytes_to_use_per_rank": 17179869184,
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
For the Prefill node, keep the normal P/D connector configuration, for example
|
||||
`MooncakeConnectorV1` with `kv_role` set to `kv_producer`.
|
||||
|
||||
When `kv_load_failure_policy` is also needed for the P/D connector, configure
|
||||
it on the top-level `MultiConnector` `kv-transfer-config`, not inside child
|
||||
connectors.
|
||||
|
||||
## Standalone Connector Example
|
||||
|
||||
`RecomputeCPUOffloadConnector` must be used together with the vLLM-Ascend
|
||||
`RecomputeScheduler`. Because `RecomputeScheduler` is only supported on
|
||||
P/D-disaggregated Decode nodes, recompute CPU offload is currently only
|
||||
available for P/D-disaggregated Decode nodes as well.
|
||||
|
||||
The following standalone connector configuration is intended only as a minimal
|
||||
configuration fragment for validating the recompute-offload path on a
|
||||
P/D-disaggregated Decode node. It is not a PD-mixed deployment mode. In
|
||||
PD-mixed or normal non-P/D deployments, do not enable recompute CPU offload; the
|
||||
engine uses the normal vLLM recompute behavior instead.
|
||||
|
||||
```python
|
||||
from vllm.config import KVTransferConfig
|
||||
|
||||
kv_transfer_config = KVTransferConfig(
|
||||
kv_connector="RecomputeCPUOffloadConnector",
|
||||
kv_role="kv_consumer",
|
||||
kv_connector_extra_config={
|
||||
"cpu_bytes_to_use_per_rank": 17179869184,
|
||||
"enable_offload_prefix_caching": False,
|
||||
},
|
||||
)
|
||||
```
|
||||
|
||||
For online serving:
|
||||
|
||||
```bash
|
||||
vllm serve /path/to/model \
|
||||
--additional-config '{"recompute_scheduler_enable":true}' \
|
||||
--kv-transfer-config '{
|
||||
"kv_connector": "RecomputeCPUOffloadConnector",
|
||||
"kv_role": "kv_consumer",
|
||||
"kv_connector_extra_config": {
|
||||
"cpu_bytes_to_use_per_rank": 17179869184,
|
||||
"enable_offload_prefix_caching": false
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
## How It Works
|
||||
|
||||
1. When `RecomputeScheduler` cannot allocate enough HBM KV blocks, it selects a
|
||||
Decode-side running request for preemption.
|
||||
2. Before the request's HBM blocks are reused, the scheduler calls the
|
||||
connector's preemption hook. If enough CPU blocks are available, the
|
||||
connector records the HBM block to CPU block mapping.
|
||||
3. The model runner calls `handle_preemptions()` before updating worker state.
|
||||
The worker copies the selected KV blocks from HBM to pinned CPU tensors.
|
||||
4. After all workers report that the store is complete, the preempted request
|
||||
becomes restorable.
|
||||
5. When the request is scheduled again, the connector reports the number of
|
||||
tokens that can be restored from CPU memory.
|
||||
6. The scheduler allocates new HBM blocks and the connector builds the CPU block
|
||||
to HBM block reload mapping.
|
||||
7. Before model forward, the worker copies the restored KV blocks back to HBM.
|
||||
Forward then continues from the restored KV state.
|
||||
|
||||
## Sliding-Window and MTP Support
|
||||
|
||||
Sliding-window models can contain logical block tables with null block IDs for
|
||||
tokens outside the attention window. The connector preserves logical block
|
||||
positions instead of compressing non-zero block IDs, so a block table such as
|
||||
`[0, 0, 20, 21]` remains aligned as `[0, 0, cpu_20, cpu_21]` on the CPU side.
|
||||
Only non-zero blocks are transferred.
|
||||
|
||||
With MTP or speculative decode enabled, the scheduler may allocate lookahead
|
||||
blocks. The reload path clips the H2D range to the GPU block table actually
|
||||
allocated for the resumed request.
|
||||
|
||||
## Notes and Limitations
|
||||
|
||||
* This feature requires the vLLM V1 engine and the vLLM-Ascend recompute
|
||||
scheduler.
|
||||
* The feature is only intended for Decode nodes in P/D disaggregation. It is not
|
||||
supported in PD-mixed or normal non-P/D deployments.
|
||||
* Do not enable `recompute_scheduler_enable` in PD-mixed deployments. Without
|
||||
P/D-disaggregated Decode-side `RecomputeScheduler`, recompute CPU offload is
|
||||
not active and vLLM follows the normal recompute path.
|
||||
* When running in Docker with a large per-rank offload budget, reserve enough
|
||||
shared memory. For typical A3 deployments with 16 GiB per-card offload, use
|
||||
`--shm-size=1024g`.
|
||||
* `enable_offload_prefix_caching` is experimental and disabled by default.
|
||||
* Current D2H and H2D transfers use basic torch copy operations. They are
|
||||
correctness-oriented and not yet optimized for transfer throughput.
|
||||
* Qwen3.5 with async scheduling is not fully supported yet. Disable async
|
||||
scheduling when using recompute CPU offload with Qwen3.5 models.
|
||||
* If CPU memory capacity is insufficient, the connector skips offload for that
|
||||
request and the scheduler falls back to the original recompute behavior.
|
||||
|
||||
## FAQ
|
||||
|
||||
### How much CPU memory should I configure?
|
||||
|
||||
Start with a budget that can hold the expected number of preempted request
|
||||
blocks. Prefer `cpu_bytes_to_use_per_rank` because it directly controls the
|
||||
offload budget for each rank/card. For example,
|
||||
`cpu_bytes_to_use_per_rank=17179869184` gives each rank/card 16 GiB.
|
||||
|
||||
`cpu_bytes_to_use` is divided by `world_size`. In a DP2TP8 Decode service,
|
||||
setting `cpu_bytes_to_use=17179869184` gives each DP rank's cards about 8 GiB
|
||||
of offload space. Increase the budget if logs show that CPU cache free blocks
|
||||
are insufficient.
|
||||
|
||||
### How do I know the offload path is active?
|
||||
|
||||
Look for logs similar to:
|
||||
|
||||
```text
|
||||
Recompute preemption offload enabled for request ...
|
||||
Created recompute offload state for request ...
|
||||
Prepared recompute offload H2D load for request ...
|
||||
```
|
||||
|
||||
If offload cannot be prepared, logs may show that CPU cache free blocks are
|
||||
insufficient, and the request will fall back to the original recompute path.
|
||||
|
||||
### Is this the same as KV Cache CPU Offload?
|
||||
|
||||
No. `KV Cache CPU Offload` is a prefix-cache offload path for inactive KV cache
|
||||
blocks. `RecomputeCPUOffloadConnector` is specifically for preserving KV blocks
|
||||
of requests preempted by the Decode-side recompute scheduler.
|
||||
|
||||
## Reference
|
||||
|
||||
For design background and implementation details, see
|
||||
[#10820: Recompute CPU Offload Connector](https://github.com/vllm-project/vllm-ascend/issues/10820).
|
||||
156
docs/source/user_guide/feature_guide/rfork.md
Normal file
@@ -0,0 +1,156 @@
|
||||
# RFork Guide
|
||||
|
||||
This guide explains how to use **RFork** as a model-loader plugin in **vLLM Ascend**.
|
||||
|
||||
---
|
||||
|
||||
## Overview
|
||||
|
||||
RFork is a warm-start weight loading path for vLLM Ascend. Instead of always reading model weights from storage, a new instance can request a compatible **seed** instance from an external planner, then pull weights directly from that seed through `YuanRong TransferEngine`.
|
||||
|
||||
The RFork loading flow in the current implementation is:
|
||||
|
||||
1. vLLM starts with `--load-format rfork`.
|
||||
2. RFork builds a **seed key** from the model identity and deployment topology.
|
||||
3. RFork asks the planner for an available seed matching that key.
|
||||
4. If a seed is returned, the new instance initializes the model structure on its local NPU, registers local weight memory, fetches the remote transfer-engine metadata from the seed, and performs batch weight transfer into local parameter buffers.
|
||||
5. If no seed is available, or any step fails, RFork cleans up and falls back to the default loader.
|
||||
6. After the instance finishes loading, it starts a local seed service and periodically reports heartbeat to the planner, so later instances can reuse it.
|
||||
|
||||
## Flowchart
|
||||
|
||||

|
||||
|
||||
## Application Scenarios
|
||||
|
||||
- **Scale-out after a first successful load**: The first instance may still load from storage, but later instances with the same deployment identity can reuse it as a seed and shorten startup time.
|
||||
- **Elastic serving clusters**: Because RFork asks a planner for available seeds, it fits clusters where instances are created and reclaimed dynamically.
|
||||
- **Topology-sensitive deployments**: RFork encodes `kv_role`, `node_rank`, optional `pp_rank`, `tp_rank`, optional `ep_rank`, and optional `draft` role into the seed key, so only topology-compatible instances are matched together.
|
||||
|
||||
---
|
||||
|
||||
## Usage
|
||||
|
||||
To enable RFork, pass `--load-format rfork` and provide RFork settings through `--model-loader-extra-config` as a JSON string.
|
||||
|
||||
### RFork Prerequisites
|
||||
|
||||
- Install the runtime dependency `YuanRong TransferEngine` on every RFork instance.
|
||||
- Run a planner service that implements the RFork seed protocol. A simple mock planner script is provided at [`rfork_planner.py`](../../../../examples/rfork/rfork_planner.py).
|
||||
|
||||
### Configuration Fields
|
||||
|
||||
| Field Name | Type | Description | Allowed Values / Notes |
|
||||
|------------|------|-------------|------------------------|
|
||||
| **model_url** | String | Logical model identifier used to build the RFork seed key. | Required for RFork transfer. Instances that should share seeds must use the same value. |
|
||||
| **model_deploy_strategy_name** | String | Deployment strategy identifier used together with `model_url` to build the seed key. | Required for RFork transfer. Instances that should share seeds must use the same value. |
|
||||
| **rfork_scheduler_url** | String | Base URL of the planner service used for seed allocation, release, and heartbeat. | Required for planner-based matching. Example: `http://127.0.0.1:1223`. |
|
||||
| **rfork_seed_timeout_sec** | Number | Timeout for waiting until the local seed HTTP service becomes healthy after startup. | Optional. Default: `5.0`. Must be greater than `0`. Invalid values fall back to the default. |
|
||||
| **rfork_seed_key_separator** | String | Separator used when building the RFork seed key string. | Optional. Default: `$`. Keep the same value across compatible instances. |
|
||||
|
||||
### How RFork Matches Seeds
|
||||
|
||||
RFork does not match instances by `model_url` alone. The local seed key is composed from:
|
||||
|
||||
- `model_url`
|
||||
- `model_deploy_strategy_name`
|
||||
- disaggregation mode derived from `kv_transfer_config.kv_role` or `kv_both`
|
||||
- `node_rank`
|
||||
- `pp_rank` when pipeline parallel size is greater than 1
|
||||
- `tp_rank`
|
||||
- `ep_rank` when expert parallelism is enabled for an MoE model
|
||||
- optional `draft` suffix when the worker runs as a draft model
|
||||
|
||||
This means two instances must agree on both model identity and deployment topology before the planner will treat them as interchangeable seeds.
|
||||
For deployments without pipeline or expert parallelism, the existing seed-key format is unchanged.
|
||||
|
||||
### Quantized Models
|
||||
|
||||
For quantized models, RFork transfers tensors after Ascend weight post-processing instead of raw checkpoint parameters. The receiver first builds the same post-load tensor layout as the seed, then RFork copies the live NPU tensors used by inference.
|
||||
|
||||
This path handles Ascend quantization changes such as weight transposition, NZ format conversion, packed weights, derived scale tensors, and MLA/SFA runtime tensors such as `W_UV` and `W_UK_T`. Empty tensors that were released during post-processing are not included in the transfer manifest.
|
||||
|
||||
When validating RFork for a quantized model:
|
||||
|
||||
- Apply the same vLLM Ascend code to both the seed instance and the receiver instance.
|
||||
- Restart the planner and all vLLM instances after changing RFork code, because existing seeds keep their old transfer metadata.
|
||||
- Use a new `model_deploy_strategy_name` after changing model arguments or RFork code, so the planner does not match a receiver with an incompatible old seed.
|
||||
- A successful RFork transfer logs `transfer weights starts` and `transfer weights time`. The fallback path logs `RFork transfer failed`.
|
||||
|
||||
## Tested Models
|
||||
|
||||
The following table records models that have been explicitly tested with RFork weight transfer. A model should be added here only after RFork transfer succeeds and the loaded instance passes basic inference validation.
|
||||
|
||||
| Model | Precision / Quantization | Hardware | Validation Status | Notes |
|
||||
|-------|--------------------------|----------|-------------------|-------|
|
||||
| Qwen2.5-7B | BF16 | A2 | Tested | RFork transfer has been validated. |
|
||||
| Qwen3-32B | BF16 | A2 | Tested | RFork transfer has been validated. |
|
||||
| Qwen3-235B-A22B | BF16 | A2 | Tested | RFork transfer has been validated. |
|
||||
| DeepSeek-V4-Flash-W8A8-MTP | W8A8 | A2 | Tested | RFork transfer with MTP draft model has been validated. |
|
||||
| GLM5-W4A8 | W4A8 | A2 | Tested | RFork transfer has been validated. |
|
||||
| Kimi2.5-W4A8 | W4A8 | A2 | Tested | RFork transfer has been validated. |
|
||||
|
||||
---
|
||||
|
||||
## Example Commands & Placeholders
|
||||
|
||||
> Replace parts in `<...>` before running.
|
||||
|
||||
### 1. Install YuanRong TransferEngine
|
||||
|
||||
```shell
|
||||
pip install openyuanrong-transfer-engine
|
||||
```
|
||||
|
||||
### 2. Start the Planner
|
||||
|
||||
A simple planner implementation is provided at [`rfork_planner.py`](../../../../examples/rfork/rfork_planner.py).
|
||||
|
||||
```shell
|
||||
python rfork_planner.py \
|
||||
--host 0.0.0.0 \
|
||||
--port <planner_port>
|
||||
```
|
||||
|
||||
### 3. Start vLLM Instances
|
||||
|
||||
Use the same RFork startup command for both the first instance and later instances in the same deployment.
|
||||
|
||||
For the first instance, the planner usually has no compatible seed yet, so RFork falls back to the default loader. After loading finishes, that instance starts its local seed service and reports itself to the planner.
|
||||
|
||||
For later instances, if the planner can allocate a compatible seed, RFork will try to transfer weights from the existing seed instance before falling back to the default loader.
|
||||
|
||||
```shell
|
||||
export RFORK_CONFIG='{
|
||||
"model_url": "<model_url>",
|
||||
"model_deploy_strategy_name": "<deploy_strategy>",
|
||||
"rfork_scheduler_url": "http://<planner_ip>:<planner_port>"
|
||||
}'
|
||||
|
||||
vllm serve <model_path> \
|
||||
--tensor-parallel-size 1 \
|
||||
--served-model-name <served_model_name> \
|
||||
--port <port> \
|
||||
--load-format rfork \
|
||||
--model-loader-extra-config "${RFORK_CONFIG}"
|
||||
```
|
||||
|
||||
### Placeholder Descriptions
|
||||
|
||||
- `<model_path>`: Model path or model identifier passed to `vllm serve`.
|
||||
- `<served_model_name>`: Service name exposed by vLLM.
|
||||
- `<planner_ip>`: IP address or hostname of the RFork planner.
|
||||
- `<planner_port>`: Listening port of the RFork planner.
|
||||
- `<model_url>`: Stable model identity string used to build the RFork seed key.
|
||||
- `<deploy_strategy>`: Stable deployment-strategy name used to build the RFork seed key.
|
||||
- `<port>`: Serving port of the vLLM instance being started.
|
||||
|
||||
---
|
||||
|
||||
## Note & Caveats
|
||||
|
||||
- RFork requires `YuanRong TransferEngine` at runtime. If the package is missing, RFork cannot initialize the transfer backend.
|
||||
- If RFORK is used, **each worker process** must bind a listening port. That port is assigned randomly.
|
||||
- RFork weight transfer does not support `additional_config.layer_sharding`. If `--load-format rfork` is used together with `layer_sharding`, RFork transfer is bypassed and the model is loaded through the default model loader.
|
||||
- RFork weight transfer does not support dynamic EPLB because expert weights and placement can change after the seed service starts. If `eplb_config.dynamic_eplb` or `eplb_config.expert_map_record_path` enables dynamic EPLB, RFork transfer is bypassed and the model is loaded through the default model loader.
|
||||
- The example [`rfork_planner.py`](../../../../examples/rfork/rfork_planner.py) is only a simple mock implementation. If you need stronger scheduling, capacity management, or production-grade availability behavior, implement your own planner based on the RFork seed protocol.
|
||||
109
docs/source/user_guide/feature_guide/sequence_parallelism.md
Normal file
@@ -0,0 +1,109 @@
|
||||
# Sequence Parallelism
|
||||
|
||||
## What is Sequence Parallelism
|
||||
|
||||
Sequence Parallelism (SP) was first introduced in [Megatron](https://arxiv.org/pdf/2205.05198), with the original intention of reducing training activation memory. The core modification was changing `Allreduce->LayerNorm` to `ReduceScatter->LayerNorm->Allgather`. This technique was later applied to inference by vllm. It should be noted that splitting Allreduce into ReduceScatter and Allgather does not inherently bring performance benefits; it reduces the computation load of LayerNorm, but this gain is minimal. The real benefits of SP come from:
|
||||
|
||||
1. LLM inference deployment often uses quantization. Taking INT8 quantization commonly used on NPUs as an example, after LayerNorm, a Quant operator quantizes the hidden states from BF16 to INT8. The communication volume of Allgather is halved, and the time consumption is almost halved.
|
||||
2. ReduceScatter and Allgather can be fused with the preceding and following Matmul operations respectively into communication-computation parallel operators, reducing latency.
|
||||
|
||||
## How to Use
|
||||
|
||||
Currently, vllm-ascend has implemented Sequence Parallelism for VL-class models based on the Inductor pass. It can be enabled in the following way:
|
||||
|
||||
```bash
|
||||
vllm serve Qwen/Qwen3-VL-2B-Instruct \
|
||||
--tensor-parallel-size 2 \
|
||||
--compilation-config '{"pass_config": {"enable_sp": true , "sp_min_token_num": 1000}}'
|
||||
```
|
||||
|
||||
- `"enable_sp"`: This is the switch for SP. Since SP relies on graph mode, it is not supported in eager mode.
|
||||
- `sp_min_token_num` (from upstream vllm's `pass_config`): Based on our experiments, when the number of tokens is small (empirical value is less than 1000), SP can actually bring negative impact. This is because when the communication volume is small, the fixed overhead of the communication operator becomes the dominant factor. SP will only take effect when `num_tokens >= sp_min_token_num`. **The default value is 1000 on Ascend, which generally does not need to be modified.** To customize, use `--compilation-config '{"pass_config": {"enable_sp": true, "sp_min_token_num": 512}}'`. The value will be appended into `compile_ranges_split_points`, which splits the graph compilation range and checks whether the pass is applicable per range.
|
||||
|
||||
Without modifying `sp_min_token_num`, the simplest way and recommended way to enable SP is:
|
||||
|
||||
```bash
|
||||
vllm serve Qwen/Qwen3-VL-2B-Instruct \
|
||||
--tensor-parallel-size 2 \
|
||||
--compilation-config '{"pass_config": {"enable_sp": true}}'
|
||||
```
|
||||
|
||||
## Difference Between SP and Flash Comm V1
|
||||
|
||||
[Flash Comm V1 (FC1)](https://gitcode.com/ascend-tribe/ascend-inference-cluster/blob/main/FlashComm/ascend-inference-cluster-flashcomm.md) is an enhanced version of Sequence Parallelism developed based on NPU. The enhancements include:
|
||||
|
||||
1. For models using the MLA structure, Allgather is postponed until after QKV projection, further reducing communication volume.
|
||||
2. For MoE models, Allgather is postponed until after Gating+DynamicQuant, also aiming to reduce communication volume.
|
||||
|
||||
FC1 is a unique optimization in vllm-ascend, currently implemented based on Custom OP, but it is difficult to support VL-class models (reasons detailed in [[RFC]: support sequence parallelism by pass](https://github.com/vllm-project/vllm-ascend/issues/5712)). Therefore, currently FC1 and SP are complementary.
|
||||
|
||||
## Support Matrix
|
||||
|
||||
### Without Quantization
|
||||
|
||||
| | VL + Dense | VL + MoE | non-VL + Dense | non-VL + MoE |
|
||||
| -------------------- | ----------- | ----------- | -------------- | ------------ |
|
||||
| Sequence Parallelism | x | x | x | x |
|
||||
| Flash Comm V1 | eager/graph | eager/graph | eager/graph | eager/graph |
|
||||
|
||||
### With Quantization
|
||||
|
||||
SP currently does not support quantization and is under adaptation.
|
||||
|
||||
| | VL + Dense | VL + MoE | non-VL + Dense | non-VL + MoE |
|
||||
| -------------------- | ----------- | ----------- | -------------- | ------------ |
|
||||
| Sequence Parallelism | x | x | x | x |
|
||||
| Flash Comm V1 | eager/graph | eager/graph | eager/graph | eager/graph |
|
||||
|
||||
## Pass Design
|
||||
|
||||
When SP is enabled, the following passes run in order: `SequenceParallelismPass` then `SequenceParallelismMoePass`.
|
||||
|
||||
### SequenceParallelismPass
|
||||
|
||||
Runs `NoOpEliminationPass` first to eliminate redundant view-like operations, then applies AllReduce-based patterns:
|
||||
|
||||
| Pattern | Match | Replacement |
|
||||
| -------------------------------------- | -------------------------------- | ------------------------------------------------------------------------------------- |
|
||||
| `MiddleAllReduceRMSNormPattern` | `all_reduce` + `layernorm` | `reduce_scatter` + `layernorm` + `all_gather` |
|
||||
| `LastAllReduceRMSNormPattern` | Same (last layer, no residual) | Same |
|
||||
| `Qwen3VLMiddleAllReduceRMSNormPattern` | `all_reduce` + add + `layernorm` | `reduce_scatter` + chunk(`deepstack_input_embeds`) + add + `layernorm` + `all_gather` |
|
||||
|
||||
**Why Qwen3 VL needs special handling by Qwen3VLMiddleAllReduceRMSNormPattern**
|
||||
|
||||
Qwen3-VL middle layers insert an extra add between `all_reduce` and `layernorm`: `hidden_states=hidden_states + deepstack_input_embeds`. Under SP, `hidden_states` (i.e., `input`) is reduced-scattered to shape `[seq_len/tp, hidden]` per rank, while `deepstack_input_embeds` comes from the vision/deepstack path and stays full-sequence `[seq_len, hidden]` (typically replicated across TP ranks). Simply doing `reduce_scatter(input) + deepstack_input_embeds` would cause a shape mismatch.
|
||||
The fix is to chunk `deepstack_input_embeds` by `tp_size` so each rank uses `add(reduce_scatter, chunk(deepstack_input_embeds)[tp_rank])`, keeping shapes consistent before `layernorm` and `all_gather`.
|
||||
|
||||
### SequenceParallelismMoePass
|
||||
|
||||
After `SequenceParallelismPass` applies, the MoE model computation graph looks like:
|
||||
|
||||

|
||||
|
||||
**Overview**
|
||||
|
||||
1. **Postponing allgather**: Under SP, `residual` is chunked by tensor parallelism. This causes a shape mismatch between hidden states and residual in the next layer's layernorm: hidden states are gathered (full sequence) while residual remains chunked. The fix is to move `all_gather` to after layernorm so that layernorm operates on consistent shapes per rank. `MiddleLayerAllgatherAddRMSNormPattern`, `LastLayerAllgatherRMSNormPattern`, and `Qwen3VLMiddleLayerAllgatherAddRMSNormPattern` are designed for this purpose, each handling different layer and structure variants (see the table below).
|
||||
|
||||
2. **AllGatherChunkNoOp cleanup**: When MoE SP is enabled, vllm introduces a `sequence_parallel_chunk` op (corresponding to `sp_chunk` in the diagram). Together with the preceding `all_gather`, the pair forms a redundant no-op (all_gather gathers, then chunk re-splits). `AllGatherChunkNoOpPattern` replaces this pair with identity to eliminate the redundant communication and computation.
|
||||
|
||||
**Pattern details:**
|
||||
|
||||
| Pattern | Match | Replacement |
|
||||
| ---------------------------------- | ---------------------------------------- | --------------------------------------- |
|
||||
| `MiddleLayerAllgatherAddRMSNormPattern` | `all_gather` + slice + `layernorm` | `layernorm` + `all_gather` |
|
||||
| `LastLayerAllgatherRMSNormPattern` | Same (last layer, no residual) | Same |
|
||||
| `Qwen3VLMiddleLayerAllgatherAddRMSNormPattern` | `all_gather` + slice + add + `layernorm` | add(chunk) + `layernorm` + `all_gather` |
|
||||
| `AllGatherChunkNoOpPattern` | `all_gather` + `sequence_parallel_chunk_impl` | identity (no-op) |
|
||||
|
||||
### FAQ
|
||||
|
||||
#### Q1: Is SP enabled by default?
|
||||
|
||||
No, SP is not enabled by default. SP is currently in the experimental stage and will be enabled by default in the future.
|
||||
|
||||
The processing flow of `enable_sp` in the code is:
|
||||
|
||||
- In `pass_config`, `enable_sp` and `sp_min_token_num` default to `None`
|
||||
- `NPUPlatform.apply_config_platform_defaults`: If `enable_sp` is `True` and `sp_min_token_num` is None, set default `sp_min_token_num` (1000 for Dense models, 1 for MoE models)
|
||||
- `VllmConfig._apply_optimization_level_defaults`: `enable_sp` is set to `True` for dense models.
|
||||
- `VllmConfig.__post_init__`: If `sp_min_token_num` is still `None`, then `enable_sp` is set to `False`
|
||||
@@ -2,15 +2,15 @@
|
||||
|
||||
## Overview
|
||||
|
||||
Sleep Mode is an API designed to offload model weights and discard KV cache from NPU memory. This functionality is essential for reinforcement learning (RL) post-training workloads, particularly in online algorithms such as PPO, GRPO, or DPO. During training, the policy model typically performs auto-regressive generation using inference engines like vLLM, followed by forward and backward passes for optimization.
|
||||
Sleep Mode is an API designed to offload model weights and discard KV cache from NPU memory. This functionality is essential for reinforcement learning (RL) post-training workloads, particularly in online algorithms such as PPO, GRPO, or DPO. During training, the policy model typically performs autoregressive generation using inference engines like vLLM, followed by forward and backward passes for optimization.
|
||||
|
||||
Since the generation and training phases may employ different model parallelism strategies, it becomes crucial to free KV cache and even offload model parameters stored within vLLM during training. This ensures efficient memory utilization and avoids resource contention on the NPU.
|
||||
|
||||
## Getting started
|
||||
|
||||
With `enable_sleep_mode=True`, the way we manage memory(malloc, free) in vllm will under a specific memory pool, during loading model and initialize kv_caches, we tag the memory as a map: `{"weight": data, "kv_cache": data}`.
|
||||
With `enable_sleep_mode=True`, the way we manage memory (malloc, free) in vLLM is under a specific memory pool. During model loading and KV cache initialization, we tag the memory as a map: `{"weight": data, "kv_cache": data}`.
|
||||
|
||||
The engine(v0/v1) supports two sleep levels to manage memory during idle periods:
|
||||
The engine (v0/v1) supports two sleep levels to manage memory during idle periods:
|
||||
|
||||
- Level 1 Sleep
|
||||
- Action: Offloads model weights and discards the KV cache.
|
||||
@@ -20,27 +20,96 @@ The engine(v0/v1) supports two sleep levels to manage memory during idle periods
|
||||
|
||||
- Level 2 Sleep
|
||||
- Action: Discards both model weights and KV cache.
|
||||
- Memory: The content of both the model weights and kv cache is forgotten.
|
||||
- Memory: The content of both the model weights and KV cache is forgotten.
|
||||
- Use Case: Ideal when switching to a different model or updating the current one.
|
||||
|
||||
Since this feature uses the low-level API [AscendCL](https://www.hiascend.com/document/detail/zh/CANNCommunityEdition/82RC1alpha002/API/appdevgapi/appdevgapi_07_0000.html), in order to use sleep mode, you should follow the [installation guide](https://vllm-ascend.readthedocs.io/en/latest/installation.html) and building from source, if you are using v0.7.3, remember to set `export COMPILE_CUSTOM_KERNELS=1`, for the latest version(v0.9.x+), the environment variable `COMPILE_CUSTOM_KERNELS` will be set 1 by default while building from source.
|
||||
Since this feature uses the low-level API [AscendCL](https://www.hiascend.com/document/detail/zh/CANNCommunityEdition/82RC1alpha002/API/appdevgapi/appdevgapi_07_0000.html), in order to use sleep mode, you should follow the [installation guide](https://docs.vllm.ai/projects/ascend/en/latest/installation.html) and build from source. If you are using < v0.12.0rc1, remember to set `export COMPILE_CUSTOM_KERNELS=1`.
|
||||
|
||||
## Optional extra cleanup
|
||||
|
||||
By default, sleep mode only releases memory managed by the sleep-mode allocator. For RL workloads that need to return more NPU memory to the trainer, vLLM Ascend also provides an optional extra cleanup path:
|
||||
|
||||
```python
|
||||
llm = LLM(
|
||||
"Qwen/Qwen2.5-0.5B-Instruct",
|
||||
enable_sleep_mode=True,
|
||||
additional_config={"enable_sleep_mode_extra_cleanup": True},
|
||||
)
|
||||
```
|
||||
|
||||
For online serving, pass the same option through `--additional-config`:
|
||||
|
||||
```bash
|
||||
vllm serve Qwen/Qwen2.5-0.5B-Instruct \
|
||||
--enable-sleep-mode \
|
||||
--additional-config '{"enable_sleep_mode_extra_cleanup": true}'
|
||||
```
|
||||
|
||||
When `enable_sleep_mode_extra_cleanup` is enabled, `sleep()` additionally:
|
||||
|
||||
- clears ACL graph attention workspaces and invalidates captured ACL graph caches when ACL graph is enabled;
|
||||
- resets the model runner graph manager so ACL graphs can be captured again after wakeup;
|
||||
- waits for pending pipeline-parallel send work, synchronizes the NPU, and destroys HCCL process groups.
|
||||
|
||||
During `wake_up()`, vLLM Ascend restores the HCCL process groups, refreshes MoE dispatcher HCCL metadata, restores sleep-mode allocator memory, and recaptures ACL graphs when needed.
|
||||
|
||||
:::{note}
|
||||
Extra cleanup trades lower sleep-time NPU memory usage for longer wakeup latency. In particular, if ACL graph is enabled, `wake_up()` must call `capture_model()` again after the model state has been restored. Keep `enable_sleep_mode_extra_cleanup` disabled when lower wakeup latency is more important than releasing HCCL and ACL graph workspace memory.
|
||||
:::
|
||||
|
||||
For level 2 sleep, wakeup can be split into two phases:
|
||||
|
||||
```python
|
||||
llm.wake_up(tags=["weights"])
|
||||
# Reload or update model weights here.
|
||||
llm.wake_up(tags=["kv_cache"])
|
||||
```
|
||||
|
||||
With extra cleanup enabled, ACL graphs are recaptured only when `tags` is `None` or contains `"kv_cache"`. This avoids recapturing graphs before externally reloaded weights and KV-cache state are ready.
|
||||
|
||||
### Expert weight layout restoration
|
||||
|
||||
For dense models, `wake_up()` simply restores the model weights to NPU memory; the tensor layout is unchanged.
|
||||
|
||||
For **unquantized MoE models** (`quant_config is None`), the fused expert weights are stored in a transposed layout for NPU matmul efficiency. This layout is produced once at model load time by `process_weights_after_loading()`: after the weights are loaded, the method transposes the second and third dimensions (`transpose(1, 2)`) of `w13_weight` and `w2_weight` to convert the standard checkpoint layout into the format required by the `torch_npu.npu_grouped_matmul` operator.
|
||||
|
||||
After the sleep-mode allocator restores the original (untransposed) memory, `wake_up()` re-applies the same transpose to the affected expert weights when the `"weights"` tag is being restored:
|
||||
|
||||
- `w13_weight` (gate/up projection): transposed back to the runtime layout when its second dimension matches `hidden_size`;
|
||||
- `w2_weight` (down projection): transposed back to the runtime layout when its third dimension matches `hidden_size`.
|
||||
|
||||
This step is skipped entirely for dense models (which have no expert weights) and for quantized models (whose weights are handled by the quantization method).
|
||||
|
||||
## Prepare Model Weights
|
||||
|
||||
Use the `Qwen2.5-0.5B-Instruct` model weights. With `VLLM_USE_MODELSCOPE=True`, the model will be downloaded automatically from ModelScope.
|
||||
|
||||
```{list-table}
|
||||
:header-rows: 1
|
||||
|
||||
* - Model
|
||||
- ModelScope Link
|
||||
* - Qwen2.5-0.5B-Instruct
|
||||
- [Qwen/Qwen2.5-0.5B-Instruct](https://www.modelscope.cn/models/Qwen/Qwen2.5-0.5B-Instruct)
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
The following is a simple example of how to use sleep mode.
|
||||
|
||||
- offline inference:
|
||||
- Offline inference:
|
||||
|
||||
```python
|
||||
import os
|
||||
|
||||
import torch
|
||||
from vllm import LLM, SamplingParams
|
||||
from vllm.utils import GiB_bytes
|
||||
from vllm.utils.mem_constants import GiB_bytes
|
||||
|
||||
|
||||
os.environ["VLLM_USE_MODELSCOPE"] = "True"
|
||||
os.environ["VLLM_WORKER_MULTIPROC_METHOD"] = "spawn"
|
||||
os.environ["VLLM_ASCEND_ENABLE_NZ"] = "0"
|
||||
|
||||
if __name__ == "__main__":
|
||||
prompt = "How are you?"
|
||||
@@ -68,19 +137,20 @@ The following is a simple example of how to use sleep mode.
|
||||
assert output[0].outputs[0].text == output2[0].outputs[0].text
|
||||
```
|
||||
|
||||
- online serving:
|
||||
- Online serving:
|
||||
:::{note}
|
||||
Considering there may be a risk of malicious access, please make sure you are under a dev-mode, and explicit specify the develop env: `VLLM_SERVER_DEV_MODE` to expose these endpoints(sleep/wake up).
|
||||
Considering there may be a risk of malicious access, please make sure you are under a dev-mode, and explicitly specify the dev environment `VLLM_SERVER_DEV_MODE` to expose these endpoints (sleep/wake up).
|
||||
:::
|
||||
|
||||
```bash
|
||||
export VLLM_SERVER_DEV_MODE="1"
|
||||
export VLLM_WORKER_MULTIPROC_METHOD="spawn"
|
||||
export VLLM_USE_MODELSCOPE="True"
|
||||
export VLLM_ASCEND_ENABLE_NZ="0"
|
||||
|
||||
vllm serve Qwen/Qwen2.5-0.5B-Instruct --enable-sleep-mode
|
||||
|
||||
# after serveing is up, post these endpoints
|
||||
# after serving is up, post to these endpoints
|
||||
|
||||
# sleep level 1
|
||||
curl -X POST http://127.0.0.1:8000/sleep \
|
||||
|
||||
344
docs/source/user_guide/feature_guide/speculative_decoding.md
Normal file
@@ -0,0 +1,344 @@
|
||||
# Speculative Decoding Guide
|
||||
|
||||
This guide shows how to use Speculative Decoding with vLLM Ascend. Speculative decoding is a technique which improves inter-token latency in memory-bound LLM inference.
|
||||
|
||||
## Overview
|
||||
|
||||
vLLM Ascend implements speculative decoding through a **proposer-verifier** architecture:
|
||||
|
||||
1. **Proposer** (`vllm_ascend/spec_decode/`): Generates draft (speculative) tokens using various methods — from simple n-gram matching to neural-network-based draft models.
|
||||
2. **Rejection Sampler** (`vllm_ascend/sample/`): Verifies draft tokens against the target model's output, accepting matches and rejecting mismatches, with optional optimizations including [Block Verify and Entropy Verify](#block-verify-and-entropy-verify).
|
||||
|
||||
The following speculative decoding methods are supported:
|
||||
|
||||
| Method | Description |
|
||||
| ------ | ----------- |
|
||||
| `ngram` | Match n-grams from the prompt |
|
||||
| `suffix` | Suffix-based pattern matching (requires Arctic Inference) |
|
||||
| `medusa` | Medusa heads embedded in the target model |
|
||||
| `eagle` | EAGLE-based draft model |
|
||||
| `eagle3` | EAGLE-3 based draft model |
|
||||
| `mtp` | Multi-Token Prediction with shared embedding head |
|
||||
| `dflash` | Draft-and-Flash with cross-attention |
|
||||
| `draft_model` | Generic external draft LLM |
|
||||
| `extract_hidden_states` | Extract hidden states for EAGLE training |
|
||||
|
||||
## Common Configuration
|
||||
|
||||
All speculative decoding methods are configured through the `speculative_config` parameter when initializing the model or starting the server:
|
||||
|
||||
- **`method`** (str, required): The speculative decoding method. Must be one of the supported method names listed in the table above.
|
||||
- **`num_speculative_tokens`** (int, required): Number of speculative tokens to generate per forward pass. Auto-filled from the draft model's `n_predict` config (e.g., MTP) or `suffix_decoding_max_tree_depth` (suffix method) when available.
|
||||
> **Note**: For PD Separation deployment, `num_speculative_tokens` should be subject to one of the following conditions:
|
||||
>
|
||||
> 1. Hybrid Mamba models (e.g., Qwen-Next and Qwen3.5 series): `num_speculative_tokens` should be equal on P nodes and D nodes.
|
||||
> 2. Other models: `num_speculative_tokens` on P nodes should be 1, and `num_speculative_tokens` on D nodes should be greater or equal to 1.
|
||||
- **`model`** (str, optional): Path or HF repo ID for the draft model. Required for `eagle`, `eagle3`, `dflash`, `medusa`, and `draft_model`. Automatically resolved for `mtp` (reuses target model), `ngram`, `suffix`, and `extract_hidden_states`.
|
||||
- **`draft_tensor_parallel_size`** (int, optional): Tensor parallelism size for the draft model. Can only be `1` or the same as the target model's tensor parallel size.
|
||||
- **`disable_padded_drafter_batch`** (bool, default: `False`): Disable input padding for speculative decoding. If set to `True`, speculative input batches can contain sequences of different lengths, which may only be supported by certain attention backends. **Note:** Only effective with `eagle`, `eagle3`, `mtp`, `dflash`, `draft_model`, and `extract_hidden_states` methods.
|
||||
|
||||
**Offline inference** — pass `speculative_config` as a Python dict to `LLM()`:
|
||||
|
||||
```python
|
||||
from vllm import LLM
|
||||
|
||||
llm = LLM(
|
||||
model="path/to/target/model",
|
||||
speculative_config={
|
||||
"method": "eagle3",
|
||||
"model": "path/to/draft/model",
|
||||
"num_speculative_tokens": 3,
|
||||
},
|
||||
)
|
||||
```
|
||||
|
||||
**Online serving** — pass `--speculative-config` (or `-sc`) as a JSON string:
|
||||
|
||||
```shell
|
||||
vllm serve path/to/target/model \
|
||||
--speculative-config '{"method": "eagle3", "model": "path/to/draft/model", "num_speculative_tokens": 3}'
|
||||
```
|
||||
|
||||
> [!NOTE]
|
||||
> On Ascend NPUs, the `npu_fused_infer_attention_score` operator supports a maximum of 16 tokens per decode round. Therefore, `(num_speculative_tokens + 1)` must be ≤ 15.
|
||||
|
||||
## Speculating by matching n-grams in the prompt
|
||||
|
||||
The following code configures vLLM Ascend to use speculative decoding where proposals are generated by matching n-grams in the prompt.
|
||||
|
||||
- Offline inference
|
||||
|
||||
```python
|
||||
from vllm import LLM, SamplingParams
|
||||
|
||||
prompts = [
|
||||
"The future of AI is",
|
||||
]
|
||||
sampling_params = SamplingParams(temperature=0.8, top_p=0.95)
|
||||
|
||||
llm = LLM(
|
||||
model="meta-llama/Meta-Llama-3.1-8B-Instruct",
|
||||
tensor_parallel_size=1,
|
||||
speculative_config={
|
||||
"method": "ngram",
|
||||
"num_speculative_tokens": 5,
|
||||
"prompt_lookup_max": 4,
|
||||
},
|
||||
)
|
||||
outputs = llm.generate(prompts, sampling_params)
|
||||
|
||||
for output in outputs:
|
||||
prompt = output.prompt
|
||||
generated_text = output.outputs[0].text
|
||||
print(f"Prompt: {prompt!r}, Generated text: {generated_text!r}")
|
||||
```
|
||||
|
||||
## Speculating using EAGLE based draft models
|
||||
|
||||
The following code configures vLLM Ascend to use speculative decoding where proposals are generated by an [EAGLE (Extrapolation Algorithm for Greater Language-model Efficiency)](https://arxiv.org/pdf/2401.15077) based draft model.
|
||||
|
||||
In v0.12.0rc1 of vLLM Ascend, the async scheduler is more stable and ready to be enabled. We have adapted it to support EAGLE, and you can use it by setting `async_scheduling=True` as follows. If you encounter any issues, please feel free to open an issue on GitHub. As a workaround, you can disable this feature by unsetting `async_scheduling=True` when initializing the model.
|
||||
|
||||
- Offline inference
|
||||
|
||||
```python
|
||||
from vllm import LLM, SamplingParams
|
||||
|
||||
prompts = [
|
||||
"The future of AI is",
|
||||
]
|
||||
sampling_params = SamplingParams(temperature=0.8, top_p=0.95)
|
||||
|
||||
llm = LLM(
|
||||
model="meta-llama/Meta-Llama-3.1-8B-Instruct",
|
||||
tensor_parallel_size=4,
|
||||
distributed_executor_backend="mp",
|
||||
enforce_eager=True,
|
||||
async_scheduling=True,
|
||||
speculative_config={
|
||||
"method": "eagle",
|
||||
"model": "yuhuili/EAGLE-LLaMA3.1-Instruct-8B",
|
||||
"draft_tensor_parallel_size": 1,
|
||||
"num_speculative_tokens": 2,
|
||||
},
|
||||
)
|
||||
|
||||
outputs = llm.generate(prompts, sampling_params)
|
||||
|
||||
for output in outputs:
|
||||
prompt = output.prompt
|
||||
generated_text = output.outputs[0].text
|
||||
print(f"Prompt: {prompt!r}, Generated text: {generated_text!r}")
|
||||
```
|
||||
|
||||
A few important things to consider when using the EAGLE based draft models:
|
||||
|
||||
1. The EAGLE draft models available in the [HF repository for EAGLE models](https://huggingface.co/yuhuili) should
|
||||
be loaded and used directly by vLLM. This functionality was added in PR [#4893](https://github.com/vllm-project/vllm-ascend/pull/4893).
|
||||
If you are using a vLLM version released before this pull request was merged, please update to a more recent version.
|
||||
|
||||
2. The EAGLE based draft models need to be run without tensor parallelism
|
||||
(i.e. draft_tensor_parallel_size is set to 1 in `speculative_config`), although
|
||||
it is possible to run the main model using tensor parallelism (see example above).
|
||||
|
||||
3. When using EAGLE-3 based draft model, option "method" must be set to "eagle3".
|
||||
That is, to specify `"method": "eagle3"` in `speculative_config`.
|
||||
|
||||
4. After enabling EAGLE, the main model needs to verify `(1 + K)` tokens generated by the main model and the draft model in one decoding process.
|
||||
And the fullgraph mode will fix the number of tokens during the verification stage,
|
||||
so `cudagraph_capture_sizes` must be a list of capture sizes, where each size is calculated as `n * (K + 1)` for each batch size `n` you want to support.
|
||||
For instance, to support batch sizes from 1 to 4 with `num_speculative_tokens = 4`, `cudagraph_capture_sizes` should be set to `[5, 10, 15, 20]`.
|
||||
|
||||
## Speculating using MTP
|
||||
|
||||
MTP (Multi-Token Prediction) boosts inference performance by parallelizing the prediction of multiple tokens, shifting from single-token to multi-token generation. This approach significantly increases generation throughput and achieves multiplicative acceleration in inference speed — all without compromising output quality.
|
||||
|
||||
- Online inference
|
||||
|
||||
```shell
|
||||
vllm serve /deepseek-ai/DeepSeek-V3.2-Exp-W8A8 \
|
||||
--port 20004 \
|
||||
--data-parallel-size 1 \
|
||||
--tensor-parallel-size 16 \
|
||||
--enable-expert-parallel \
|
||||
--seed 1024 \
|
||||
--served-model-name dsv3 \
|
||||
--max-model-len 36768 \
|
||||
--max-num-batched-tokens 5000 \
|
||||
--max-num-seqs 10 \
|
||||
--quantization ascend \
|
||||
--trust-remote-code \
|
||||
--gpu-memory-utilization 0.9 \
|
||||
--compilation-config '{"cudagraph_mode": "FULL_DECODE_ONLY"}' \
|
||||
--speculative-config '{"num_speculative_tokens": 2, "method":"mtp", "disable_padded_drafter_batch": false}'
|
||||
```
|
||||
|
||||
> [!NOTE]
|
||||
> Due to the fact that only a single layer of weights is exposed in DeepSeek's MTP, accuracy and performance are not effectively guaranteed in scenarios where `num_speculative_tokens > 1` (especially ≥ 3).
|
||||
>
|
||||
> In the fullgraph mode with `num_speculative_tokens > 1`, the capture size of each ACLGraph must be an integer multiple of `(num_speculative_tokens + 1)`.
|
||||
|
||||
## Speculating using Suffix Decoding
|
||||
|
||||
The following code configures vLLM to use speculative decoding where proposals are generated using Suffix Decoding [(SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications)](https://arxiv.org/abs/2411.04975).
|
||||
|
||||
Like n-gram, Suffix Decoding can generate draft tokens by pattern-matching using the last `n` generated tokens. Unlike n-gram, Suffix Decoding (1) can pattern-match against both the prompt and previous generations, (2) uses frequency counts to propose the most likely continuations, and (3) speculates an adaptive number of tokens for each request at each iteration to get better acceptance rates.
|
||||
|
||||
Suffix Decoding can achieve better performance for tasks with high repetition, such as code-editing, agentic loops (e.g. self-reflection, self-consistency), and RL rollouts.
|
||||
|
||||
> [!NOTE]
|
||||
> Suffix Decoding requires Arctic Inference. You can install it with `pip install arctic-inference`.
|
||||
|
||||
- Offline inference
|
||||
|
||||
```python
|
||||
from vllm import LLM, SamplingParams
|
||||
|
||||
prompts = [
|
||||
"The future of AI is",
|
||||
]
|
||||
sampling_params = SamplingParams(temperature=0.8, top_p=0.95)
|
||||
|
||||
llm = LLM(
|
||||
model="meta-llama/Meta-Llama-3.1-8B-Instruct",
|
||||
tensor_parallel_size=1,
|
||||
enforce_eager=True,
|
||||
speculative_config={
|
||||
"method": "suffix",
|
||||
"num_speculative_tokens": 15,
|
||||
},
|
||||
)
|
||||
|
||||
outputs = llm.generate(prompts, sampling_params)
|
||||
|
||||
for output in outputs:
|
||||
prompt = output.prompt
|
||||
generated_text = output.outputs[0].text
|
||||
print(f"Prompt: {prompt!r}, Generated text: {generated_text!r}")
|
||||
```
|
||||
|
||||
## Extracting Hidden States
|
||||
|
||||
The `extract_hidden_states` method is a special speculative decoding mode that does not perform actual speculation. Instead, it extracts hidden states from specified layers of the target model and saves them to disk. This is primarily used for collecting training data for EAGLE-style draft models.
|
||||
|
||||
> [!NOTE]
|
||||
> This method produces only 1 output token per request. The primary output is the hidden states saved to disk, not the generated text.
|
||||
|
||||
- Offline inference
|
||||
|
||||
```python
|
||||
import tempfile
|
||||
|
||||
from safetensors import safe_open
|
||||
from vllm import LLM, SamplingParams
|
||||
|
||||
|
||||
def main():
|
||||
with tempfile.TemporaryDirectory() as tmpdirname:
|
||||
llm = LLM(
|
||||
model="Qwen/Qwen3-8B",
|
||||
tensor_parallel_size=1,
|
||||
speculative_config={
|
||||
"method": "extract_hidden_states",
|
||||
"num_speculative_tokens": 1,
|
||||
"draft_model_config": {
|
||||
"hf_config": {
|
||||
# Layer indices to extract hidden states from
|
||||
"eagle_aux_hidden_state_layer_ids": [2, 18, 34],
|
||||
}
|
||||
},
|
||||
},
|
||||
kv_transfer_config={
|
||||
"kv_connector": "ExampleHiddenStatesConnector",
|
||||
"kv_role": "kv_producer",
|
||||
"kv_connector_extra_config": {
|
||||
"shared_storage_path": tmpdirname,
|
||||
},
|
||||
},
|
||||
)
|
||||
|
||||
prompts = ["Hello, how are you?", "What is machine learning?"]
|
||||
sampling_params = SamplingParams(max_tokens=1)
|
||||
outputs = llm.generate(prompts, sampling_params)
|
||||
|
||||
for output in outputs:
|
||||
print("Prompt:", output.prompt)
|
||||
print("Prompt token ids:", output.prompt_token_ids)
|
||||
|
||||
hidden_states_path = output.kv_transfer_params.get("hidden_states_path")
|
||||
print("Hidden states saved to:", hidden_states_path)
|
||||
|
||||
with safe_open(hidden_states_path, "pt") as f:
|
||||
token_ids = f.get_tensor("token_ids")
|
||||
hidden_states = f.get_tensor("hidden_states")
|
||||
print("Shape:", hidden_states.shape)
|
||||
# Shape: (num_tokens, num_layers, hidden_size)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
```
|
||||
|
||||
Key configuration parameters:
|
||||
|
||||
1. **`num_speculative_tokens`**: Must be set to `1`. This method does not perform actual speculation, so the value is fixed.
|
||||
|
||||
2. **`eagle_aux_hidden_state_layer_ids`**: List of layer indices from which to extract hidden states. For example, `[2, 18, 34]` extracts from layers 2, 18, and 34.
|
||||
|
||||
3. **`kv_connector`**: Must be set to `"ExampleHiddenStatesConnector"` to enable saving hidden states to disk.
|
||||
|
||||
4. **`kv_role`**: Must be set to `"kv_producer"` for the extraction mode.
|
||||
|
||||
5. **`shared_storage_path`**: Directory where hidden states will be saved as `.safetensors` files (one per request).
|
||||
|
||||
## Block Verify and Entropy Verify
|
||||
|
||||
vLLM Ascend provides two optional optimizations for the rejection sampler in speculative decoding: **Block Verify** and **Entropy Verify**. These features trade a small amount of output precision for improved inference throughput.
|
||||
|
||||
> [!WARNING]
|
||||
> Both Block Verify and Entropy Verify modify the token acceptance criteria and may cause minor precision degradation (e.g., slightly different output tokens compared to the standard rejection sampler). Evaluate the quality impact on your specific workload before enabling them in production.
|
||||
|
||||
### Block Verify
|
||||
|
||||
Block Verify evaluates all draft tokens as a block using cumulative probability products, rather than checking each token independently. This can improve the acceptance rate and reduce the overhead of rejection sampling, especially when `num_speculative_tokens >= 3`.
|
||||
|
||||
### Entropy Verify
|
||||
|
||||
Entropy Verify adjusts the acceptance threshold based on the entropy of the target distribution:
|
||||
|
||||
- **High entropy** (uncertain distribution) → lower effective threshold → more tokens accepted
|
||||
- **Low entropy** (confident distribution) → higher effective threshold → stricter rejection
|
||||
|
||||
This entropy-aware threshold is controlled by two parameters:
|
||||
|
||||
- **`posterior_threshold`** (default: `0.95`, range: `(0, 1]`): The upper bound of the modified threshold. Even when entropy is very low, the effective threshold will not exceed this value.
|
||||
- **`posterior_alpha`** (default: `0.4`, range: `>= 0`): Controls how strongly entropy influences the threshold. A higher alpha makes the threshold more sensitive to entropy changes, resulting in a higher acceptance rate for speculative tokens but also greater precision loss. You need to tune this value based on your specific model and dataset. When alpha is `0`, entropy has no effect and the threshold equals `posterior_threshold`.
|
||||
|
||||
### Usage
|
||||
|
||||
- Online inference
|
||||
|
||||
```shell
|
||||
vllm serve <model> --additional-config \
|
||||
'{"rejection_sampler_config": {"enable_block_verify": true, \
|
||||
"enable_entropy_verify": true, "posterior_threshold": 0.95, \
|
||||
"posterior_alpha": 0.4}}'
|
||||
```
|
||||
|
||||
- Offline inference
|
||||
|
||||
```python
|
||||
llm = LLM(
|
||||
model,
|
||||
additional_config={
|
||||
"rejection_sampler_config": {
|
||||
"enable_block_verify": True,
|
||||
"enable_entropy_verify": True,
|
||||
"posterior_threshold": 0.95,
|
||||
"posterior_alpha": 0.4,
|
||||
}
|
||||
},
|
||||
)
|
||||
```
|
||||
|
||||
Both features can be enabled independently or together. When used together, the cumulative acceptance from Block Verify is combined with the entropy-adjusted threshold from Entropy Verify.
|
||||
@@ -2,162 +2,16 @@
|
||||
|
||||
## Overview
|
||||
|
||||
### What is Structured Output?
|
||||
### What is structured output?
|
||||
|
||||
LLMs can be unpredictable when you need output in specific formats. Think of asking a model to generate JSON - without guidance, it might produce valid text that breaks JSON specification. **Structured Output (also called Guided Decoding)** enables LLMs to generate outputs that follow a desired structure while preserving the non-deterministic nature of the system.
|
||||
LLMs can be unpredictable when you need output in specific formats. Think of asking a model to generate JSON without guidance: it might produce valid text that breaks the JSON specification. **Structured Output (also known as Guided Decoding)** enables LLMs to generate outputs that follow a desired structure while preserving the non-deterministic nature of the system.
|
||||
|
||||
In simple terms, structured decoding gives LLMs a “template” to follow. Users provide a schema that “influences” the model’s output, ensuring compliance with the desired structure.
|
||||
In simple terms, structured decoding gives LLMs a "template" to follow. Users provide a schema that "influences" the model output, ensuring compliance with the desired structure.
|
||||
|
||||

|
||||
|
||||
### Structured Output in vllm-ascend
|
||||
## Usage in vllm-ascend
|
||||
|
||||
Currently, vllm-ascend supports **xgrammar** and **guidance** backend for structured output with vllm v1 engine.
|
||||
Currently, the usage of structured output feature in vllm-ascend is totally the same as that in vLLM.
|
||||
|
||||
XGrammar introduces a new technique that batch constrained decoding via pushdown automaton (PDA). You can think of a PDA as a “collection of FSMs, and each FSM represents a context-free grammar (CFG).” One significant advantage of PDA is its recursive nature, allowing us to execute multiple state transitions. They also include additional optimisation (for those who are interested) to reduce grammar compilation overhead. Besides, you can also find more details about guidance by yourself.
|
||||
|
||||
## How to Use Structured Output?
|
||||
|
||||
### Online Inference
|
||||
|
||||
You can also generate structured outputs using the OpenAI's Completions and Chat API. The following parameters are supported, which must be added as extra parameters:
|
||||
|
||||
- `guided_choice`: the output will be exactly one of the choices.
|
||||
- `guided_regex`: the output will follow the regex pattern.
|
||||
- `guided_json`: the output will follow the JSON schema.
|
||||
- `guided_grammar`: the output will follow the context free grammar.
|
||||
|
||||
Structured outputs are supported by default in the OpenAI-Compatible Server. You can choose to specify the backend to use by setting the `--guided-decoding-backend` flag to vllm serve. The default backend is `auto`, which will try to choose an appropriate backend based on the details of the request. You may also choose a specific backend, along with some options.
|
||||
|
||||
Now let´s see an example for each of the cases, starting with the guided_choice, as it´s the easiest one:
|
||||
|
||||
```python
|
||||
from openai import OpenAI
|
||||
client = OpenAI(
|
||||
base_url="http://localhost:8000/v1",
|
||||
api_key="-",
|
||||
)
|
||||
|
||||
completion = client.chat.completions.create(
|
||||
model="Qwen/Qwen2.5-3B-Instruct",
|
||||
messages=[
|
||||
{"role": "user", "content": "Classify this sentiment: vLLM is wonderful!"}
|
||||
],
|
||||
extra_body={"guided_choice": ["positive", "negative"]},
|
||||
)
|
||||
print(completion.choices[0].message.content)
|
||||
```
|
||||
|
||||
The next example shows how to use the guided_regex. The idea is to generate an email address, given a simple regex template:
|
||||
|
||||
```python
|
||||
completion = client.chat.completions.create(
|
||||
model="Qwen/Qwen2.5-3B-Instruct",
|
||||
messages=[
|
||||
{
|
||||
"role": "user",
|
||||
"content": "Generate an example email address for Alan Turing, who works in Enigma. End in .com and new line. Example result: alan.turing@enigma.com\n",
|
||||
}
|
||||
],
|
||||
extra_body={"guided_regex": r"\w+@\w+\.com\n", "stop": ["\n"]},
|
||||
)
|
||||
print(completion.choices[0].message.content)
|
||||
```
|
||||
|
||||
One of the most relevant features in structured text generation is the option to generate a valid JSON with pre-defined fields and formats. For this we can use the guided_json parameter in two different ways:
|
||||
|
||||
- Using a JSON Schema.
|
||||
- Defining a Pydantic model and then extracting the JSON Schema from it.
|
||||
|
||||
The next example shows how to use the guided_json parameter with a Pydantic model:
|
||||
|
||||
```python
|
||||
from pydantic import BaseModel
|
||||
from enum import Enum
|
||||
|
||||
class CarType(str, Enum):
|
||||
sedan = "sedan"
|
||||
suv = "SUV"
|
||||
truck = "Truck"
|
||||
coupe = "Coupe"
|
||||
|
||||
class CarDescription(BaseModel):
|
||||
brand: str
|
||||
model: str
|
||||
car_type: CarType
|
||||
|
||||
json_schema = CarDescription.model_json_schema()
|
||||
|
||||
completion = client.chat.completions.create(
|
||||
model="Qwen/Qwen2.5-3B-Instruct",
|
||||
messages=[
|
||||
{
|
||||
"role": "user",
|
||||
"content": "Generate a JSON with the brand, model and car_type of the most iconic car from the 90's",
|
||||
}
|
||||
],
|
||||
extra_body={"guided_json": json_schema},
|
||||
)
|
||||
print(completion.choices[0].message.content)
|
||||
```
|
||||
|
||||
Finally we have the guided_grammar option, which is probably the most difficult to use, but it´s really powerful. It allows us to define complete languages like SQL queries. It works by using a context free EBNF grammar. As an example, we can use to define a specific format of simplified SQL queries:
|
||||
|
||||
```python
|
||||
simplified_sql_grammar = """
|
||||
root ::= select_statement
|
||||
|
||||
select_statement ::= "SELECT " column " from " table " where " condition
|
||||
|
||||
column ::= "col_1 " | "col_2 "
|
||||
|
||||
table ::= "table_1 " | "table_2 "
|
||||
|
||||
condition ::= column "= " number
|
||||
|
||||
number ::= "1 " | "2 "
|
||||
"""
|
||||
|
||||
completion = client.chat.completions.create(
|
||||
model="Qwen/Qwen2.5-3B-Instruct",
|
||||
messages=[
|
||||
{
|
||||
"role": "user",
|
||||
"content": "Generate an SQL query to show the 'username' and 'email' from the 'users' table.",
|
||||
}
|
||||
],
|
||||
extra_body={"guided_grammar": simplified_sql_grammar},
|
||||
)
|
||||
print(completion.choices[0].message.content)
|
||||
```
|
||||
|
||||
Find more examples [here](https://github.com/vllm-project/vllm/blob/main/examples/offline_inference/structured_outputs.py).
|
||||
|
||||
### Offline Inference
|
||||
|
||||
To use Structured Output, we'll need to configure the guided decoding using the class `GuidedDecodingParams` inside `SamplingParams`. The main available options inside `GuidedDecodingParams` are:
|
||||
|
||||
- json
|
||||
- regex
|
||||
- choice
|
||||
- grammar
|
||||
|
||||
One example for the usage of the choice parameter is shown below:
|
||||
|
||||
```python
|
||||
from vllm import LLM, SamplingParams
|
||||
from vllm.sampling_params import GuidedDecodingParams
|
||||
|
||||
llm = LLM(model="Qwen/Qwen2.5-7B-Instruct",
|
||||
guided_decoding_backend="xgrammar")
|
||||
|
||||
guided_decoding_params = GuidedDecodingParams(choice=["Positive", "Negative"])
|
||||
sampling_params = SamplingParams(guided_decoding=guided_decoding_params)
|
||||
outputs = llm.generate(
|
||||
prompts="Classify this sentiment: vLLM is wonderful!",
|
||||
sampling_params=sampling_params,
|
||||
)
|
||||
print(outputs[0].outputs[0].text)
|
||||
```
|
||||
|
||||
Find more examples of other usages [here](https://github.com/vllm-project/vllm/blob/main/examples/offline_inference/structured_outputs.py).
|
||||
Find more examples and explanations about these usages in [vLLM official document](https://docs.vllm.ai/en/stable/features/structured_outputs/).
|
||||
|
||||
885
docs/source/user_guide/feature_guide/ucm_deployment.md
Normal file
@@ -0,0 +1,885 @@
|
||||
# UCM Store Deployment Guide
|
||||
|
||||
## Why Use UCM
|
||||
|
||||
Unified Cache Manager (UCM) provides an external KV-cache storage layer designed for prefix-caching scenarios in vLLM/vLLM-Ascend. Unlike KV Pooling, which expands prefix-cache capacity only by aggregating device memory and therefore remains limited by HBM/DRAM size and lacks persistence, UCM decouples compute from storage and adopts a tiered design. Each node uses local DRAM as a fast cache, while a shared backend—such as NFS, 3FS, or enterprise-grade storage—serves as the persistent KV store.
|
||||
|
||||
**Key benefits of using UCM:**
|
||||
|
||||
1. **Breaks Device Memory Capacity Limits**: Traditional prefix caching is constrained by HBM/DRAM size. UCM removes this ceiling by offloading KV cache to external storage, enabling cache capacity to scale with the storage system rather than with compute resources.
|
||||
|
||||
2. **Persistent and Reliable KV Cache**: UCM provides durable KV cache storage, ensuring that cached prefix blocks survive across service restarts, instance failures, or scheduling migrations. This is critical for production-grade inference systems.
|
||||
|
||||
3. **Multi-Scenario Acceleration**: UCM not only supports prefix caching but also offers training-free sparse attention methods (e.g., GSA, CacheBlend) for handling extremely long sequence inference tasks. Additionally, UCM provides PD disaggregation solutions based on storage-compute separation architecture, enabling flexible management of heterogeneous computing resources.
|
||||
|
||||
4. **Significant Performance Improvement**: When integrated with vLLM, UCM achieves **3-10x reduction** in inference latency across various scenarios, including multi-turn dialogue and long-context reasoning tasks. Benchmarks show up to **8x improvement in TTFT** for prefix caching scenarios.
|
||||
|
||||
## How UCM Works
|
||||
|
||||
### Architecture
|
||||
|
||||
UCM adopts a **centralized architecture** for KV cache management, constructing a three-tier cache hierarchy:
|
||||
|
||||
```bash
|
||||
HBM (GPU Memory) → DRAM (Local Cache) → Storage Backend (SSD/NFS/3FS)
|
||||
```
|
||||
|
||||
This three-tier design enables:
|
||||
|
||||
- **HBM (Tier 1)**: Fastest access for active inference computation
|
||||
- **DRAM (Tier 2)**: High-speed local cache for frequently accessed KV blocks
|
||||
- **Storage Backend (Tier 3)**: Persistent storage layer including local SSD, NFS-mounted storage, or dedicated systems like 3FS for unlimited capacity scaling
|
||||
|
||||
UCM chose the centralized approach (similar to DeepSeek's 3FS) over decentralized designs for several reasons:
|
||||
|
||||
1. **Simplicity**: Avoids complex affinity scheduling required in decentralized architectures
|
||||
2. **Decoupling**: Keeps inference instances independent without reporting KV cache status to schedulers
|
||||
3. **No Data Silos**: Centralized storage prevents redundant KV cache accumulation across isolated instances
|
||||
4. **Better Compatibility**: Superior compatibility with PD disaggregation and large-scale deployment
|
||||
|
||||
### Capabilities
|
||||
|
||||
UCM currently provides the following capabilities:
|
||||
|
||||
| Capability | Description |
|
||||
|------------|-------------|
|
||||
| **Prefix Cache** | Persistent KV cache storage with support for NFS Store, 3FS Store, and Pipeline Store |
|
||||
| **Sparse Attention** | Training-free sparse attention methods including GSA (Graph-based Sparse Attention) and CacheBlend for long-context acceleration |
|
||||
| **PD Disaggregation** | Prefill-Decode disaggregation with multiple modes: P2P, Centralized PD, NPGD, and xPYD |
|
||||
| **ReRoPE** | Support for Rotary Position Embedding extensions |
|
||||
|
||||
**Supported Platforms:**
|
||||
|
||||
- CUDA (NVIDIA H100, H20, L40, L20)
|
||||
- CANN (Atlas A2 inference products, Atlas A3 inference products)
|
||||
- MUSA (Mthreads S5000)
|
||||
- MACA (MetaX C500)
|
||||
|
||||
**Supported Frameworks:**
|
||||
|
||||
- vLLM (main branch)
|
||||
- vLLM-Ascend (main branch)
|
||||
- SGLang (main branch)
|
||||
|
||||
> **Note**: For the complete and latest support matrix, refer to [UCM Support Matrix](https://ucm.readthedocs.io/en/latest/user-guide/support-matrix/support_matrix.html).
|
||||
|
||||
## Deployment Guide
|
||||
|
||||
### Prerequisites
|
||||
|
||||
- OS: Linux
|
||||
- Hardware with Ascend NPUs. It is typically the Atlas 800 A2 series.
|
||||
- vLLM: main branch
|
||||
- vLLM Ascend: main branch
|
||||
|
||||
### UCM Installation
|
||||
|
||||
**Please refer to the [official UCM installation guide for Ascend NPU](https://ucm.readthedocs.io/en/latest/getting-started/quickstart_vllm_ascend.html)**
|
||||
|
||||
### PD Disaggregation Scenario
|
||||
|
||||
UCM supports two types of PD disaggregation architectures:
|
||||
|
||||
| Type | KV Transfer Method | Characteristics |
|
||||
|------|-------------------|-----------------|
|
||||
| **Centralized PD** | Via unified storage backend (NFS/3FS) | Simple architecture, complete decoupling, stateless instances |
|
||||
| **Distributed PD (P2P)** | Direct transfer via Mooncake + UCM prefix cache | Lower latency, suitable for homogeneous P/D nodes |
|
||||
|
||||
#### Centralized PD Disaggregation
|
||||
|
||||
In centralized PD disaggregation, KV cache is transmitted via a unified storage pool. The Prefill node offloads KV cache to the storage backend, and the Decode node retrieves it with high prefix cache hit rates. This approach achieves the highest degree of decoupling and simplifies scheduling logic.
|
||||
|
||||
> **Important**: For cross-node deployment, all Prefill and Decode nodes must have access to a **shared storage backend** (e.g., NFS-mounted directory or 3FS). Ensure the storage path is accessible from all nodes before proceeding.
|
||||
|
||||
**Example: 2P2D Setup**
|
||||
|
||||
Assume 2 Prefill instances on node 192.168.10.1 (ports 7800, 7801) and 2 Decode instances on node 192.168.10.2 (ports 7802, 7803), with a shared NFS storage at `/mnt/test1`.
|
||||
|
||||
**Step 1: Prepare UCM Configuration File**
|
||||
|
||||
Create a configuration file (e.g., `ucm_config_example.yaml`) with PipelineStore:
|
||||
|
||||
```yaml
|
||||
ucm_connectors:
|
||||
- ucm_connector_name: "UcmPipelineStore"
|
||||
ucm_connector_config:
|
||||
store_pipeline: "Cache|Posix"
|
||||
storage_backends: "/mnt/test1"
|
||||
cache_buffer_capacity_gb: 64
|
||||
enable_event_sync: true
|
||||
use_layerwise: false
|
||||
```
|
||||
|
||||
Key configuration parameters:
|
||||
|
||||
- **storage_backends**: The shared storage directory accessible from all nodes (e.g., NFS-mounted path or 3FS). For cross-node PD disaggregation, this must be a shared storage path.
|
||||
|
||||
> **Note**: PipelineStore is the recommended connector for UCM. It chains Cache Store (Device ↔ Host) and Posix Store (Host ↔ Storage backend) for optimal transfer performance. For more configuration options, refer to [UCM PipelineStore Documentation](https://ucm.readthedocs.io/en/latest/user-guide/prefix-cache/pipeline_store.html).
|
||||
|
||||
**Step 2: Run Prefill Servers**
|
||||
|
||||
```bash
|
||||
export PYTHONHASHSEED=123456
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3
|
||||
vllm serve /models/QwQ-32B \
|
||||
--host 0.0.0.0 \
|
||||
--port 7800 \
|
||||
--gpu-memory-utilization 0.92 \
|
||||
--data-parallel-size 1 \
|
||||
--tensor-parallel-size 4 \
|
||||
--seed 1024 \
|
||||
--max-model-len 17000 \
|
||||
--max-num-batched-tokens 8000 \
|
||||
--max-num-seqs 20 \
|
||||
--trust-remote-code \
|
||||
--enforce-eager \
|
||||
--kv-transfer-config \
|
||||
'{
|
||||
"kv_connector": "UCMConnector",
|
||||
"kv_role": "kv_both",
|
||||
"kv_connector_extra_config": {"UCM_CONFIG_FILE": "/path/to/ucm_config_example.yaml"}
|
||||
}'
|
||||
```
|
||||
|
||||
To start the second Prefill instance on the same node, modify `--port` (e.g., port 7801) and `ASCEND_RT_VISIBLE_DEVICES` accordingly.
|
||||
|
||||
**Step 3: Run Decode Servers**
|
||||
|
||||
```bash
|
||||
export PYTHONHASHSEED=123456
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3
|
||||
vllm serve /models/QwQ-32B \
|
||||
--host 0.0.0.0 \
|
||||
--port 7802 \
|
||||
--gpu-memory-utilization 0.92 \
|
||||
--data-parallel-size 1 \
|
||||
--tensor-parallel-size 4 \
|
||||
--seed 1024 \
|
||||
--max-model-len 17000 \
|
||||
--max-num-batched-tokens 8000 \
|
||||
--max-num-seqs 20 \
|
||||
--trust-remote-code \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \
|
||||
--kv-transfer-config \
|
||||
'{
|
||||
"kv_connector": "UCMConnector",
|
||||
"kv_role": "kv_both",
|
||||
"kv_connector_extra_config": {"UCM_CONFIG_FILE": "/path/to/ucm_config_example.yaml"}
|
||||
}'
|
||||
```
|
||||
|
||||
To start the second Decode instance on the same node, modify `--port` (e.g., port 7803) and `ASCEND_RT_VISIBLE_DEVICES` accordingly.
|
||||
|
||||
**Step 4: Run Load Balancing Service**
|
||||
|
||||
```bash
|
||||
python /vllm-workspace/vllm-ascend/examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py \
|
||||
--port 7805 \
|
||||
--host 0.0.0.0 \
|
||||
--prefiller-hosts 192.168.10.1 192.168.10.1 \
|
||||
--prefiller-ports 7800 7801 \
|
||||
--decoder-hosts 192.168.10.2 192.168.10.2 \
|
||||
--decoder-ports 7802 7803
|
||||
```
|
||||
|
||||
**Step 5: Performance Testing**
|
||||
|
||||
```bash
|
||||
vllm bench serve \
|
||||
--backend vllm \
|
||||
--model /models/QwQ-32B \
|
||||
--host 192.168.10.1 \
|
||||
--port 7805 \
|
||||
--seed 123456 \
|
||||
--dataset-name random \
|
||||
--num-prompts 10 \
|
||||
--random-input-len 8000 \
|
||||
--random-output-len 1000 \
|
||||
--request-rate inf \
|
||||
--ignore-eos
|
||||
```
|
||||
|
||||
#### Distributed PD Disaggregation (P2P)
|
||||
|
||||
In P2P distributed PD disaggregation, Mooncake handles direct KV cache transfer from Prefill to Decode nodes via high-speed network, while UCM provides prefix cache on Prefill nodes for KV cache reuse. This mode is suitable for scenarios with homogeneous P/D nodes and lower latency requirements.
|
||||
|
||||
> **Note**: From vLLM-Ascend 0.11.0, the official image includes pre-installed Mooncake. For installation details, refer to [kvcache-ai/Mooncake](https://github.com/kvcache-ai/Mooncake).
|
||||
|
||||
**Example: 2P2D Setup**
|
||||
|
||||
Assume 2 Prefill instances on node 192.168.10.1 (ports 9000, 9001) and 2 Decode instances on node 192.168.10.2 (ports 9000, 9001).
|
||||
|
||||
**Step 1: Run Mooncake Master Service**
|
||||
|
||||
Run Mooncake master on any node (e.g., 192.168.10.1):
|
||||
|
||||
```bash
|
||||
export LD_LIBRARY_PATH=/usr/local/lib:$LD_LIBRARY_PATH
|
||||
mooncake_master --port 50088 \
|
||||
--eviction_high_watermark_ratio 0.9 \
|
||||
--eviction_ratio 0.1 \
|
||||
--default_kv_lease_ttl 11000
|
||||
```
|
||||
|
||||
Prepare `mooncake.json` on each node:
|
||||
|
||||
```json
|
||||
{
|
||||
"metadata_server": "P2PHANDSHAKE",
|
||||
"protocol": "ascend",
|
||||
"device_name": "",
|
||||
"master_server_address": "192.168.10.1:50088",
|
||||
"global_segment_size": "1GB"
|
||||
}
|
||||
```
|
||||
|
||||
Also prepare a UCM configuration file (`ucm_config_example.yaml`) for prefix cache on Prefill nodes:
|
||||
|
||||
```yaml
|
||||
ucm_connectors:
|
||||
- ucm_connector_name: "UcmPipelineStore"
|
||||
ucm_connector_config:
|
||||
store_pipeline: "Cache|Posix"
|
||||
storage_backends: "/mnt/test1"
|
||||
cache_buffer_capacity_gb: 64
|
||||
enable_event_sync: true
|
||||
use_layerwise: true
|
||||
```
|
||||
|
||||
> **Note**: For more configuration options, refer to [UCM PipelineStore Documentation](https://ucm.readthedocs.io/en/latest/user-guide/prefix-cache/pipeline_store.html).
|
||||
|
||||
**Step 2: Run Prefill Service**
|
||||
|
||||
```bash
|
||||
export LD_LIBRARY_PATH=/usr/local/lib:/usr/local/Ascend/ascend-toolkit/latest/python/site-packages:$LD_LIBRARY_PATH
|
||||
export PYTHONHASHSEED=0
|
||||
export PYTHONPATH=$PYTHONPATH:/vllm-workspace/vllm
|
||||
export MOONCAKE_CONFIG_PATH="./mooncake.json"
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3
|
||||
|
||||
vllm serve /models/QwQ-32B \
|
||||
--host 0.0.0.0 \
|
||||
--port 9000 \
|
||||
--gpu-memory-utilization 0.92 \
|
||||
--data-parallel-size 1 \
|
||||
--tensor-parallel-size 4 \
|
||||
--seed 1024 \
|
||||
--max-model-len 17000 \
|
||||
--max-num-batched-tokens 8000 \
|
||||
--max-num-seqs 20 \
|
||||
--trust-remote-code \
|
||||
--enforce-eager \
|
||||
--kv-transfer-config \
|
||||
'{
|
||||
"kv_connector": "MultiConnector",
|
||||
"kv_role": "kv_producer",
|
||||
"kv_connector_extra_config": {
|
||||
"connectors": [
|
||||
{
|
||||
"kv_connector": "MooncakeConnectorV1",
|
||||
"kv_role": "kv_producer",
|
||||
"kv_port": 20001,
|
||||
"kv_connector_extra_config": {
|
||||
"prefill": {"dp_size": 1, "tp_size": 4},
|
||||
"decode": {"dp_size": 1, "tp_size": 4}
|
||||
}
|
||||
},
|
||||
{
|
||||
"kv_connector": "UCMConnector",
|
||||
"kv_role": "kv_both",
|
||||
"kv_connector_extra_config": {"UCM_CONFIG_FILE": "/vllm-workspace/unified-cache-management/examples/ucm_config_example.yaml"}
|
||||
}
|
||||
]
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
To start multiple Prefill instances on the same node, modify `--port`, `kv_port`, and `ASCEND_RT_VISIBLE_DEVICES` for each instance (e.g., port 9001 with kv_port 20002 for the second instance).
|
||||
|
||||
**Step 3: Run Decode Service**
|
||||
|
||||
```bash
|
||||
export LD_LIBRARY_PATH=/usr/local/lib:/usr/local/Ascend/ascend-toolkit/latest/python/site-packages:$LD_LIBRARY_PATH
|
||||
export PYTHONHASHSEED=0
|
||||
export PYTHONPATH=$PYTHONPATH:/vllm-workspace/vllm
|
||||
export MOONCAKE_CONFIG_PATH="./mooncake.json"
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3
|
||||
|
||||
vllm serve /models/QwQ-32B \
|
||||
--host 0.0.0.0 \
|
||||
--port 9000 \
|
||||
--data-parallel-size 1 \
|
||||
--tensor-parallel-size 4 \
|
||||
--seed 1024 \
|
||||
--max-model-len 17000 \
|
||||
--max-num-batched-tokens 8000 \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 4 \
|
||||
--gpu-memory-utilization 0.92 \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \
|
||||
--kv-transfer-config \
|
||||
'{
|
||||
"kv_connector": "MooncakeConnectorV1",
|
||||
"kv_role": "kv_consumer",
|
||||
"kv_port": 20001,
|
||||
"kv_connector_extra_config": {
|
||||
"prefill": {"dp_size": 1, "tp_size": 4},
|
||||
"decode": {"dp_size": 1, "tp_size": 4}
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
To start multiple Decode instances on the same node, modify `--port`, `kv_port`, and `ASCEND_RT_VISIBLE_DEVICES` for each instance (e.g., port 9001 with kv_port 20002 for the second instance).
|
||||
|
||||
**Step 4: Run Load Balancing Service**
|
||||
|
||||
```bash
|
||||
python /vllm-workspace/vllm-ascend/examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py \
|
||||
--port 7850 \
|
||||
--host 0.0.0.0 \
|
||||
--prefiller-hosts 192.168.10.1 192.168.10.1 \
|
||||
--prefiller-ports 9000 9001 \
|
||||
--decoder-hosts 192.168.10.2 192.168.10.2 \
|
||||
--decoder-ports 9000 9001
|
||||
```
|
||||
|
||||
**Step 5: Performance Testing**
|
||||
|
||||
```bash
|
||||
vllm bench serve \
|
||||
--backend vllm \
|
||||
--model /models/QwQ-32B \
|
||||
--host 192.168.10.1 \
|
||||
--port 7850 \
|
||||
--seed 123456 \
|
||||
--dataset-name random \
|
||||
--num-prompts 10 \
|
||||
--random-input-len 8000 \
|
||||
--random-output-len 1000 \
|
||||
--request-rate inf \
|
||||
--ignore-eos
|
||||
```
|
||||
|
||||
### PD-Mixed Inference
|
||||
|
||||
PD-Mixed Inference refers to the standard vLLM serving mode where Prefill and Decode phases for different requests are processed concurrently within the same instance. Unlike PD Disaggregation which physically separates Prefill and Decode into dedicated instances, PD-Mixed handles both phases in a unified scheduler, allowing interleaved execution: while one request is in Decode phase, another request can simultaneously undergo Prefill.
|
||||
|
||||
UCM enhances PD-Mixed by providing persistent KV cache storage, enabling:
|
||||
|
||||
- Prefix cache reuse across requests with shared prefixes
|
||||
- KV cache persistence across service restarts
|
||||
- Offloading KV cache to external storage to reduce GPU memory pressure
|
||||
|
||||
**Step 1: Prepare UCM Configuration File**
|
||||
|
||||
Create a configuration file (e.g., `ucm_config_example.yaml`) with PipelineStore:
|
||||
|
||||
```yaml
|
||||
ucm_connectors:
|
||||
- ucm_connector_name: "UcmPipelineStore"
|
||||
ucm_connector_config:
|
||||
store_pipeline: "Cache|Posix"
|
||||
storage_backends: "/mnt/test1"
|
||||
cache_buffer_capacity_gb: 64
|
||||
enable_event_sync: true
|
||||
use_layerwise: true
|
||||
```
|
||||
|
||||
Key configuration parameters:
|
||||
|
||||
- **storage_backends**: Directory for KV cache storage. Can be local SSD or NFS-mounted path.
|
||||
|
||||
> **Note**: For more configuration options, refer to [UCM PipelineStore Documentation](https://ucm.readthedocs.io/en/latest/user-guide/prefix-cache/pipeline_store.html).
|
||||
|
||||
**Step 2: Run PD-Mixed Service**
|
||||
|
||||
```bash
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3,4,5,6,7
|
||||
vllm serve /models/QwQ-32B \
|
||||
--host 0.0.0.0 \
|
||||
--port 7800 \
|
||||
--gpu-memory-utilization 0.92 \
|
||||
--data-parallel-size 2 \
|
||||
--tensor-parallel-size 4 \
|
||||
--seed 1024 \
|
||||
--max-model-len 17000 \
|
||||
--max-num-batched-tokens 8000 \
|
||||
--max-num-seqs 20 \
|
||||
--trust-remote-code \
|
||||
--enforce-eager \
|
||||
--block-size 128 \
|
||||
--kv-transfer-config \
|
||||
'{
|
||||
"kv_connector": "UCMConnector",
|
||||
"kv_role": "kv_both",
|
||||
"kv_connector_extra_config": {"UCM_CONFIG_FILE": "/path/to/ucm_config_example.yaml"}
|
||||
}'
|
||||
```
|
||||
|
||||
**Step 3: Performance Testing**
|
||||
|
||||
Run the benchmark twice to observe the prefix cache effect:
|
||||
|
||||
```bash
|
||||
# First run - no cache hit
|
||||
vllm bench serve \
|
||||
--backend vllm \
|
||||
--model /models/QwQ-32B \
|
||||
--host localhost \
|
||||
--port 7800 \
|
||||
--seed 123456 \
|
||||
--dataset-name random \
|
||||
--num-prompts 10 \
|
||||
--random-input-len 8000 \
|
||||
--random-output-len 1000 \
|
||||
--request-rate inf \
|
||||
--ignore-eos
|
||||
|
||||
# Second run - observe cache hit improvement
|
||||
vllm bench serve \
|
||||
--backend vllm \
|
||||
--model /models/QwQ-32B \
|
||||
--host localhost \
|
||||
--port 7800 \
|
||||
--seed 123456 \
|
||||
--dataset-name random \
|
||||
--num-prompts 10 \
|
||||
--random-input-len 8000 \
|
||||
--random-output-len 1000 \
|
||||
--request-rate inf \
|
||||
--ignore-eos
|
||||
```
|
||||
|
||||
After the second run, a significant reduction in TTFT should be observed due to UCM prefix cache hits. Review the vLLM logs for cache hit information:
|
||||
|
||||
```bash
|
||||
INFO ucm_connector.py:xxx: request_id: xxx, total_blocks_num: xxx, hit hbm: 0, hit external: xxx
|
||||
```
|
||||
|
||||
## Example: PD Disaggregation with Large Scale Expert Parallelism
|
||||
|
||||
This section demonstrates PD disaggregation for MoE models with large-scale Expert Parallelism. MoE models require enabling data parallelism to distribute expert weights across multiple nodes.
|
||||
|
||||
**Deployment Configuration:**
|
||||
|
||||
- **Prefill Instance**: 4 nodes (192.168.10.1-4), DP4TP8 (4 DP processes, each with TP8)
|
||||
- **Decode Instance**: 4 nodes (192.168.10.5-8), DP8TP4 (8 DP processes, each with TP4)
|
||||
- **Total**: 8 Atlas 800T A2 servers with 8 Ascend 910B3 NPU cards each
|
||||
- **Storage**: 8 servers connected to AI storage device A800 via CE8875 switch
|
||||
|
||||
> **Note**: For external load balancing data parallelism, refer to the vLLM official documentation: [Data Parallel Deployment of external Load Balancing](https://docs.vllm.ai/en/latest/serving/data_parallel_deployment/#external-load-balancing).
|
||||
|
||||
### Deployment Steps
|
||||
|
||||
**Step 1: Run Mooncake Master Service**
|
||||
|
||||
Run Mooncake master on any node (e.g., 192.168.10.1):
|
||||
|
||||
```bash
|
||||
export LD_LIBRARY_PATH=/usr/local/lib:$LD_LIBRARY_PATH
|
||||
mooncake_master --port 50088 \
|
||||
--eviction_high_watermark_ratio 0.9 \
|
||||
--eviction_ratio 0.1 \
|
||||
--default_kv_lease_ttl 11000
|
||||
```
|
||||
|
||||
Prepare `mooncake.json` on all 8 nodes:
|
||||
|
||||
```json
|
||||
{
|
||||
"metadata_server": "P2PHANDSHAKE",
|
||||
"protocol": "ascend",
|
||||
"device_name": "",
|
||||
"master_server_address": "192.168.10.1:50088",
|
||||
"global_segment_size": "1GB"
|
||||
}
|
||||
```
|
||||
|
||||
**Step 2: Run Prefill Service (DP4TP8)**
|
||||
|
||||
First, prepare a UCM configuration file (`ucm_config_example.yaml`) for prefix cache on Prefill nodes (192.168.10.1-4):
|
||||
|
||||
```yaml
|
||||
ucm_connectors:
|
||||
- ucm_connector_name: "UcmPipelineStore"
|
||||
ucm_connector_config:
|
||||
store_pipeline: "Cache|Posix"
|
||||
storage_backends: "/mnt/test1"
|
||||
cache_buffer_capacity_gb: 64
|
||||
enable_event_sync: true
|
||||
use_layerwise: true
|
||||
```
|
||||
|
||||
Key configuration parameters:
|
||||
|
||||
- **storage_backends**: The shared storage directory accessible from all nodes (e.g., NFS-mounted path or 3FS).
|
||||
|
||||
> **Note**: For more configuration options, refer to [UCM PipelineStore Documentation](https://ucm.readthedocs.io/en/latest/user-guide/prefix-cache/pipeline_store.html).
|
||||
|
||||
Prepare `prefill.sh` on Prefill nodes (192.168.10.1-4):
|
||||
|
||||
```bash
|
||||
#!/bin/sh
|
||||
|
||||
export LD_LIBRARY_PATH=/usr/local/lib:/usr/local/Ascend/ascend-toolkit/latest/python/site-packages:$LD_LIBRARY_PATH
|
||||
export PYTHONHASHSEED=0
|
||||
export PYTHONPATH=$PYTHONPATH:/vllm-workspace/vllm
|
||||
export MOONCAKE_CONFIG_PATH="./mooncake.json"
|
||||
|
||||
device_list=$1
|
||||
local_ip=$2
|
||||
nic_name=$3
|
||||
server_port=$4
|
||||
tp_size=$5
|
||||
dp_size=$6
|
||||
dp_rank=$7
|
||||
dp_address=$8
|
||||
dp_rpc_port=$9
|
||||
mooncake_port=${10}
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=10
|
||||
export HCCL_BUFFSIZE=256
|
||||
export ASCEND_RT_VISIBLE_DEVICES=$device_list
|
||||
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export VLLM_USE_MODELSCOPE="True"
|
||||
|
||||
vllm serve /models/GLM-5.1-w4a8 \
|
||||
--host 0.0.0.0 \
|
||||
--port $server_port \
|
||||
--data-parallel-size $dp_size \
|
||||
--data-parallel-address $dp_address \
|
||||
--data-parallel-rpc-port $dp_rpc_port \
|
||||
--data-parallel-rank $dp_rank \
|
||||
--tensor-parallel-size $tp_size \
|
||||
--enable-expert-parallel \
|
||||
--seed 1024 \
|
||||
--max-model-len 17000 \
|
||||
--max-num-batched-tokens 8000 \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 4 \
|
||||
--gpu-memory-utilization 0.92 \
|
||||
--quantization ascend \
|
||||
--enforce-eager \
|
||||
--additional-config '{"enable_weight_nz_layout":true,"enable_prefill_optimizations":true}' \
|
||||
--kv-transfer-config \
|
||||
'{
|
||||
"kv_connector": "MultiConnector",
|
||||
"kv_role": "kv_producer",
|
||||
"kv_connector_extra_config": {
|
||||
"connectors": [
|
||||
{
|
||||
"kv_connector": "MooncakeConnectorV1",
|
||||
"kv_role": "kv_producer",
|
||||
"kv_port": '$mooncake_port',
|
||||
"kv_connector_extra_config": {
|
||||
"prefill": {"dp_size": '$dp_size', "tp_size": '$tp_size'},
|
||||
"decode": {"dp_size": 8, "tp_size": 4}
|
||||
}
|
||||
},
|
||||
{
|
||||
"kv_connector": "UCMConnector",
|
||||
"kv_role": "kv_both",
|
||||
"kv_connector_extra_config": {"UCM_CONFIG_FILE": "/path/to/ucm_config_example.yaml"}
|
||||
}
|
||||
]
|
||||
}
|
||||
}' 2>&1 | tee "prefiller_dp_$dp_rank.log"
|
||||
```
|
||||
|
||||
Prepare `run_multi_dp.sh` for Prefill nodes:
|
||||
|
||||
```bash
|
||||
#!/bin/bash
|
||||
|
||||
local_ip="xxxx" # IP of current node (192.168.10.1/2/3/4)
|
||||
nic_name="xxxx" # Network interface name corresponding to local_ip
|
||||
tp_size=8
|
||||
dp_size=4 # Total DP engines for Prefill
|
||||
dp_size_local=1 # 1 DP process per node (TP8 uses all 8 cards)
|
||||
dp_rank_start=xxxx # 0 for node1, 1 for node2, 2 for node3, 3 for node4
|
||||
dp_address="192.168.10.1" # Master node for DP communication
|
||||
dp_rpc_port=13395
|
||||
server_port=9000
|
||||
mooncake_port=20001
|
||||
template_path="./prefill.sh"
|
||||
cards_per_node=8
|
||||
|
||||
cards_per_process=$((cards_per_node / dp_size_local))
|
||||
|
||||
for ((i=0; i<dp_size_local; i++)); do
|
||||
dp_rank=$((dp_rank_start + i))
|
||||
server_port=$((server_port + i))
|
||||
mooncake_port=$((mooncake_port + i * tp_size))
|
||||
|
||||
start_card=$((i * cards_per_process))
|
||||
device_list=$(seq -s, $start_card $((start_card + cards_per_process - 1)))
|
||||
|
||||
bash $template_path $device_list $local_ip $nic_name $server_port $tp_size $dp_size $dp_rank $dp_address $dp_rpc_port $mooncake_port &
|
||||
done
|
||||
|
||||
wait
|
||||
```
|
||||
|
||||
Execute `run_multi_dp.sh` on each Prefill node (192.168.10.1-4) with appropriate `local_ip` and `dp_rank_start`:
|
||||
|
||||
- 192.168.10.1: `dp_rank_start=0`
|
||||
- 192.168.10.2: `dp_rank_start=1`
|
||||
- 192.168.10.3: `dp_rank_start=2`
|
||||
- 192.168.10.4: `dp_rank_start=3`
|
||||
|
||||
**Step 3: Run Decode Service (DP8TP4)**
|
||||
|
||||
Prepare `decode.sh` on Decode nodes (192.168.10.5-8):
|
||||
|
||||
```bash
|
||||
#!/bin/sh
|
||||
|
||||
export LD_LIBRARY_PATH=/usr/local/lib:/usr/local/Ascend/ascend-toolkit/latest/python/site-packages:$LD_LIBRARY_PATH
|
||||
export PYTHONHASHSEED=0
|
||||
export PYTHONPATH=$PYTHONPATH:/vllm-workspace/vllm
|
||||
export MOONCAKE_CONFIG_PATH="./mooncake.json"
|
||||
|
||||
device_list=$1
|
||||
local_ip=$2
|
||||
nic_name=$3
|
||||
server_port=$4
|
||||
tp_size=$5
|
||||
dp_size=$6
|
||||
dp_rank=$7
|
||||
dp_address=$8
|
||||
dp_rpc_port=$9
|
||||
mooncake_port=${10}
|
||||
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
export OMP_PROC_BIND=false
|
||||
export OMP_NUM_THREADS=10
|
||||
export HCCL_BUFFSIZE=256
|
||||
export ASCEND_RT_VISIBLE_DEVICES=$device_list
|
||||
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
|
||||
export TASK_QUEUE_ENABLE=1
|
||||
export VLLM_USE_MODELSCOPE="True"
|
||||
|
||||
vllm serve /models/GLM-5.1-w4a8 \
|
||||
--host 0.0.0.0 \
|
||||
--port $server_port \
|
||||
--data-parallel-size $dp_size \
|
||||
--data-parallel-address $dp_address \
|
||||
--data-parallel-rpc-port $dp_rpc_port \
|
||||
--data-parallel-rank $dp_rank \
|
||||
--tensor-parallel-size $tp_size \
|
||||
--enable-expert-parallel \
|
||||
--seed 1024 \
|
||||
--max-model-len 17000 \
|
||||
--max-num-batched-tokens 8000 \
|
||||
--trust-remote-code \
|
||||
--max-num-seqs 4 \
|
||||
--gpu-memory-utilization 0.92 \
|
||||
--quantization ascend \
|
||||
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \
|
||||
--kv-transfer-config \
|
||||
'{
|
||||
"kv_connector": "MooncakeConnectorV1",
|
||||
"kv_role": "kv_consumer",
|
||||
"kv_port": '$mooncake_port',
|
||||
"kv_connector_extra_config": {
|
||||
"prefill": {"dp_size": 4, "tp_size": 8},
|
||||
"decode": {"dp_size": '$dp_size', "tp_size": '$tp_size'}
|
||||
}
|
||||
}' 2>&1 | tee "decoder_dp_$dp_rank.log"
|
||||
```
|
||||
|
||||
Prepare `run_multi_dp.sh` for Decode nodes:
|
||||
|
||||
```bash
|
||||
#!/bin/bash
|
||||
|
||||
local_ip="xxxx" # IP of current node (192.168.10.5/6/7/8)
|
||||
nic_name="xxxx" # Network interface name corresponding to local_ip
|
||||
tp_size=4
|
||||
dp_size=8 # Total DP engines for Decode
|
||||
dp_size_local=2 # 2 DP processes per node (TP4 uses 4 cards each)
|
||||
dp_rank_start=xxxx # 0 for node5, 2 for node6, 4 for node7, 6 for node8
|
||||
dp_address="192.168.10.5" # Master node for DP communication
|
||||
dp_rpc_port=13395
|
||||
server_port=9000
|
||||
mooncake_port=20001
|
||||
template_path="./decode.sh"
|
||||
cards_per_node=8
|
||||
|
||||
cards_per_process=$((cards_per_node / dp_size_local))
|
||||
|
||||
for ((i=0; i<dp_size_local; i++)); do
|
||||
dp_rank=$((dp_rank_start + i))
|
||||
server_port=$((server_port + i))
|
||||
mooncake_port=$((mooncake_port + i * tp_size))
|
||||
|
||||
start_card=$((i * cards_per_process))
|
||||
device_list=$(seq -s, $start_card $((start_card + cards_per_process - 1)))
|
||||
|
||||
bash $template_path $device_list $local_ip $nic_name $server_port $tp_size $dp_size $dp_rank $dp_address $dp_rpc_port $mooncake_port &
|
||||
done
|
||||
|
||||
wait
|
||||
```
|
||||
|
||||
Execute `run_multi_dp.sh` on each Decode node (192.168.10.5-8) with appropriate `local_ip` and `dp_rank_start`:
|
||||
|
||||
- 192.168.10.5: `dp_rank_start=0`
|
||||
- 192.168.10.6: `dp_rank_start=2`
|
||||
- 192.168.10.7: `dp_rank_start=4`
|
||||
- 192.168.10.8: `dp_rank_start=6`
|
||||
|
||||
**Step 4: Run Load Balancing Service**
|
||||
|
||||
```bash
|
||||
python /vllm-workspace/vllm-ascend/examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py \
|
||||
--port 7850 \
|
||||
--host 0.0.0.0 \
|
||||
--prefiller-hosts 192.168.10.1 192.168.10.2 192.168.10.3 192.168.10.4 \
|
||||
--prefiller-ports 9000 9000 9000 9000 \
|
||||
--decoder-hosts 192.168.10.5 192.168.10.5 192.168.10.6 192.168.10.6 192.168.10.7 192.168.10.7 192.168.10.8 192.168.10.8 \
|
||||
--decoder-ports 9000 9001 9000 9001 9000 9001 9000 9001
|
||||
```
|
||||
|
||||
### Benchmark Results
|
||||
|
||||
The following benchmark demonstrates UCM prefix cache effectiveness in large-scale Expert Parallelism PD disaggregation scenarios.
|
||||
|
||||
**Test Configuration:**
|
||||
|
||||
- Total requests: 128
|
||||
- Request concurrency: 128
|
||||
- Constraint: Total requests kept within Prefill instance's available HBM capacity for KV cache storage
|
||||
|
||||
**KV Cache Pre-seeding Procedure:**
|
||||
|
||||
Before each test, KV cache must be pre-seeded with a prefix ratio of **0.8**:
|
||||
|
||||
1. **Pre-seed Phase**: Send 128 requests with input length = `target_input_length × 0.8` and output length = 1 to establish the KV cache prefix
|
||||
2. **Test Phase**: Send 128 requests with full target input length and output length = 1000
|
||||
|
||||
Example for 32K input scenario:
|
||||
|
||||
- Pre-seed: 128 requests with 25600 (32K × 0.8) input tokens + 1 output token
|
||||
- Test: 128 requests with 32000 input tokens + 1000 output tokens
|
||||
|
||||
This procedure ensures the prefix portion (80% of input) is cached before measuring performance, simulating real-world prefix reuse scenarios.
|
||||
|
||||
**Test Commands:**
|
||||
|
||||
```bash
|
||||
# Step 1: Pre-seed KV cache (25600 = 32000 * 0.8)
|
||||
vllm bench serve \
|
||||
--backend vllm \
|
||||
--model /models/GLM-5.1-w4a8 \
|
||||
--host 192.168.10.1 \
|
||||
--port 7850 \
|
||||
--seed 123456 \
|
||||
--dataset-name random \
|
||||
--num-prompts 128 \
|
||||
--random-input-len 25600 \
|
||||
--random-output-len 1 \
|
||||
--request-rate inf \
|
||||
--ignore-eos
|
||||
|
||||
# Step 2: Run performance test
|
||||
vllm bench serve \
|
||||
--backend vllm \
|
||||
--model /models/GLM-5.1-w4a8 \
|
||||
--host 192.168.10.1 \
|
||||
--port 7850 \
|
||||
--seed 123456 \
|
||||
--dataset-name random \
|
||||
--num-prompts 128 \
|
||||
--random-input-len 32000 \
|
||||
--random-output-len 1000 \
|
||||
--request-rate inf \
|
||||
--ignore-eos
|
||||
```
|
||||
|
||||
**Test Scenarios:**
|
||||
|
||||
| Scenario | Description |
|
||||
|----------|-------------|
|
||||
| **Recalculation** | Baseline without UCM, HBM prefix cache disabled (full recomputation) |
|
||||
| **HBM PC** | Without UCM, HBM prefix cache enabled |
|
||||
| **UCM PC** | With UCM prefix cache enabled |
|
||||
|
||||
**Performance Results:**
|
||||
|
||||
<table>
|
||||
<thead>
|
||||
<tr>
|
||||
<th rowspan="2">Input Length</th>
|
||||
<th rowspan="2">Output Length</th>
|
||||
<th colspan="3">Recalculation</th>
|
||||
<th colspan="3">HBM PC</th>
|
||||
<th colspan="3">UCM PC</th>
|
||||
</tr>
|
||||
<tr>
|
||||
<th>TTFT (ms)</th>
|
||||
<th>TPOT (ms)</th>
|
||||
<th>E2EL (ms)</th>
|
||||
<th>TTFT (ms)</th>
|
||||
<th>TPOT (ms)</th>
|
||||
<th>E2EL (ms)</th>
|
||||
<th>TTFT (ms)</th>
|
||||
<th>TPOT (ms)</th>
|
||||
<th>E2EL (ms)</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td><strong>32K</strong></td>
|
||||
<td><strong>1K</strong></td>
|
||||
<td>140730</td>
|
||||
<td>64</td>
|
||||
<td>173820</td>
|
||||
<td>108879</td>
|
||||
<td>65</td>
|
||||
<td>142228</td>
|
||||
<td>51861</td>
|
||||
<td>66</td>
|
||||
<td>85615</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><strong>64K</strong></td>
|
||||
<td><strong>1K</strong></td>
|
||||
<td>181864</td>
|
||||
<td>64</td>
|
||||
<td>214988</td>
|
||||
<td>144444</td>
|
||||
<td>65</td>
|
||||
<td>177561</td>
|
||||
<td>69718</td>
|
||||
<td>66</td>
|
||||
<td>103752</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><strong>128K</strong></td>
|
||||
<td><strong>1K</strong></td>
|
||||
<td>268016</td>
|
||||
<td>65</td>
|
||||
<td>301648</td>
|
||||
<td>267680</td>
|
||||
<td>65</td>
|
||||
<td>301135</td>
|
||||
<td>105083</td>
|
||||
<td>66</td>
|
||||
<td>138946</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
> **Note**: Due to data parallelism, requests during the test phase may not be routed to the same DP process that was used for KV cache pre-seeding. As a result, HBM PC achieves an actual cache hit rate lower than the intended 0.8. UCM addresses this limitation by storing all KV cache in shared external storage, ensuring that requests can hit cached data regardless of which DP process handles them. This guarantees a true cache hit rate equal to the pre-seeding ratio of 0.8, significantly reducing TTFT compared to HBM PC. The improved TTFT effectively increases Prefill instance throughput, thereby boosting the overall system throughput.
|
||||
73
docs/source/user_guide/feature_guide/weight_prefetch.md
Normal file
@@ -0,0 +1,73 @@
|
||||
# Weight Prefetch Guide
|
||||
|
||||
Weight prefetching optimizes memory usage by preloading weights into the cache before they are needed, minimizing delays caused by memory access during model execution. Linear layers sometimes exhibit relatively high MTE utilization. To address this, we create a separate pipeline specifically for weight prefetching, which runs in parallel with the original vector computation pipeline, such as quantize, MoE gating top_k, RMSNorm and SwiGlu. This approach allows the weights to be preloaded to L2 cache ahead of time, reducing MTE utilization during the linear layer computations and indirectly improving Cube computation efficiency by minimizing resource contention and optimizing data flow.
|
||||
|
||||
Since we use vector computations to hide the weight prefetching pipeline, this has an effect on computation. If you prioritize low latency over high throughput, it is best not to enable prefetching.
|
||||
|
||||
## Quick Start
|
||||
|
||||
Use `--additional-config '{"weight_prefetch_config": {"enabled": true}}'` to enable weight prefetch.
|
||||
|
||||
## Fine-tune Prefetch Ratio
|
||||
|
||||
Since weight prefetch uses vector computations to hide the weight prefetching pipeline, the setting of the prefetch size is crucial. If the size is too small, the optimization benefits will not be fully realized, while a larger size may lead to resource contention, resulting in performance degradation. To accommodate different scenarios, we have added `prefetch_ratio` to allow for flexible size configuration based on the specific workload, details as follows:
|
||||
|
||||
With `prefetch_ratio` in `"weight_prefetch_config"` to customize the weight prefetch ratio for specific linear layers.
|
||||
|
||||
The “attn” and “moe” configuration options are used for MoE model, details as follows:
|
||||
|
||||
`"attn": { "qkv": 1.0, "o": 1.0}, "moe": {"gate_up": 0.8}`
|
||||
|
||||
The “mlp” configuration option is used to optimize the performance of the Dense model, details as follows:
|
||||
|
||||
`"mlp": {"gate_up": 1.0, "down": 1.0}`
|
||||
|
||||
Above values are the default config, the default value has a good performance for Qwen3-235B-A22B-W8A8 when `--max-num-seqs` is 144, for Qwen3-32B-W8A8 when `--max-num-seqs` is 72.
|
||||
|
||||
However, this may not be the optimal configuration for your scenario. For higher concurrency, you can try increasing the prefetch size. For lower concurrency, prefetching may not offer any advantages, so you can decrease the size or disable prefetching. Determine if the prefetch size is appropriate by collecting profiling data. Specifically, check if the time required for the prefetch operation (e.g., MLP Down Proj weight prefetching) overlaps with the time required for parallel vector computation operators (e.g., SwiGlu computation), and whether the prefetch operation is no later than the completion time of the vector computation operator. In the profiling timeline, a prefetch operation appears as a CMO operation on a single stream; this CMO operation is the prefetch operation.
|
||||
|
||||
Notes:
|
||||
|
||||
1) Weight prefetch of MLP `down` project prefetch depends on sequence parallel, if you want to open for mlp `down` please also enable sequence parallel.
|
||||
2) Due to the current size of the L2 cache, the maximum prefetch cannot exceed 18MB. If `prefetch_ratio * linear_layer_weight_size >= 18 * 1024 * 1024` bytes, the backend will only prefetch 18 MB.
|
||||
|
||||
## Example
|
||||
|
||||
1) For MoE model:
|
||||
|
||||
```shell
|
||||
--additional-config \
|
||||
'{
|
||||
"weight_prefetch_config": {
|
||||
"enabled": true,
|
||||
"prefetch_ratio": {
|
||||
"attn": {
|
||||
"qkv": 1.0,
|
||||
"o": 1.0
|
||||
},
|
||||
"moe": {
|
||||
"gate_up": 0.8
|
||||
}
|
||||
}
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
2) For dense model:
|
||||
|
||||
Following is the default configuration that can get a good performance for `--max-num-seqs` is 72 for Qwen3-32B-W8A8
|
||||
|
||||
```shell
|
||||
--additional-config \
|
||||
'{
|
||||
"weight_prefetch_config": {
|
||||
"enabled": true,
|
||||
"prefetch_ratio": {
|
||||
"mlp": {
|
||||
"gate_up": 1.0,
|
||||
"down": 1.0
|
||||
}
|
||||
}
|
||||
}
|
||||
}'
|
||||
```
|
||||
0
docs/source/user_guide/image-1.png
Normal file
0
docs/source/user_guide/image.png
Normal file
44
docs/source/user_guide/support_matrix/feature_matrix.md
Normal file
@@ -0,0 +1,44 @@
|
||||
# Feature Matrix
|
||||
|
||||
The table below shows mutually exclusive features and the support on Ascend hardware, extended from the [vLLM table](https://docs.vllm.ai/en/latest/features/#feature-x-feature).
|
||||
|
||||
The symbols used have the following meanings:
|
||||
|
||||
- ✅ = Full compatibility
|
||||
- 🟠 = Partial compatibility
|
||||
- ❌ = No compatibility
|
||||
- ❔ = Unknown or TBD
|
||||
|
||||
| Feature | [ACLGraph Full_Decode_Only](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/Design_Documents/ACL_Graph.html) | [ACLGraph Piecewise](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/Design_Documents/ACL_Graph.html) | Async Scheduling | [<abbr title="Automatic Prefix Caching">APC</abbr>](https://docs.vllm.ai/en/latest/features/automatic_prefix_caching/) | [Chunked Prefill](https://docs.vllm.ai/en/stable/configuration/optimization/#chunked-prefill) | [Context Parallel](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/Design_Documents/context_parallel.html) | [Cpu Binding](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/Design_Documents/cpu_binding.html) | [<abbr title="Data Parallel">DP</abbr>](https://docs.vllm.ai/en/latest/serving/data_parallel_deployment/) | [Disaggregated Prefill](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/Design_Documents/disaggregated_prefill.html) | [Eagle3](https://docs.vllm.ai/en/latest/features/speculative_decoding/eagle/) | [Eplb](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/Design_Documents/eplb_swift_balancer.html) | [<abbr title="Expert-Parallel">EP</abbr>](https://docs.vllm.ai/en/latest/serving/expert_parallel_deployment/) | Flashcomm1 | [KV Cache Pool](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/Design_Documents/KV_Cache_Pool_Guide.html) | Layer Sharding | Lmhead TP | Mlapo | [<abbr title="Multimodal Inputs">mm</abbr>](https://docs.vllm.ai/en/latest/features/multimodal_inputs/) | Multistream Moe | Shared Expert DP | [Quantization W4A4](https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/quantization.html) | [Quantization W4A8](https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/quantization.html) | [Quantization W8A8](https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/quantization.html) | <abbr title="Tensor Parallel">TP</abbr> | Weight nz |
|
||||
| - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - |
|
||||
| [ACLGraph Full_Decode_Only](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/Design_Documents/ACL_Graph.html) | ✅ | | | | | | | | | | | | | | | | | | | | | | | | |
|
||||
| [ACLGraph Piecewise](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/Design_Documents/ACL_Graph.html) | ❌ | ✅ | | | | | | | | | | | | | | | | | | | | | | | |
|
||||
| Async Scheduling | ✅ | ✅ | ✅ | | | | | | | | | | | | | | | | | | | | | | |
|
||||
| [<abbr title="Automatic Prefix Caching">APC</abbr>](https://docs.vllm.ai/en/latest/features/automatic_prefix_caching/) | ✅ | ✅ | ✅ | ✅ | | | | | | | | | | | | | | | | | | | | | |
|
||||
| [Chunked Prefill](https://docs.vllm.ai/en/stable/configuration/optimization/#chunked-prefill) | ✅ | ✅ | ✅ | ✅ | ✅ | | | | | | | | | | | | | | | | | | | | |
|
||||
| [Context Parallel](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/Design_Documents/context_parallel.html) | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | | | | | | | | | | | | | | | | | | | |
|
||||
| [Cpu Binding](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/Design_Documents/cpu_binding.html) | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | | | | | | | | | | | | | | | | | | |
|
||||
| [<abbr title="Data Parallel">DP</abbr>](https://docs.vllm.ai/en/latest/serving/data_parallel_deployment/) | ✅ | ✅ | ✅ | ✅ | ✅ | 🟠<sup>1</sup> | ✅ | ✅ | | | | | | | | | | | | | | | | | |
|
||||
| [Disaggregated Prefill](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/Design_Documents/disaggregated_prefill.html) | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | | | | | | | | | | | | | | | | |
|
||||
| [Eagle3](https://docs.vllm.ai/en/latest/features/speculative_decoding/eagle/) | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | | | | | | | | | | | | | | | |
|
||||
| [Eplb](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/Design_Documents/eplb_swift_balancer.html) | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | | | | | | | | | | | | | | |
|
||||
| [<abbr title="Expert-Parallel">EP</abbr>](https://docs.vllm.ai/en/latest/serving/expert_parallel_deployment/) | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | | | | | | | | | | | | | |
|
||||
| Flashcomm1 | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 🟠<sup>2</sup> | ✅ | ✅ | ✅ | ✅ | | | | | | | | | | | | |
|
||||
| [KV Cache Pool](https://docs.vllm.ai/projects/ascend/en/latest/developer_guide/Design_Documents/KV_Cache_Pool_Guide.html) | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | | | | | | | | | | | |
|
||||
| Layer Sharding | ✅ | ✅ | ✅ | ✅ | ✅ | 🟠 | ✅ | ✅ | 🟠<sup>3</sup> | ✅ | ✅ | ✅ | ✅ | ❔ | ✅ | | | | | | | | | | |
|
||||
| Lmhead TP | ✅ | ✅ | ✅ | ✅ | ✅ | ❔ | ✅ | 🟠<sup>4</sup> | ✅ | ✅ | ✅ | ✅ | ❌ | ❔ | ✅ | ✅ | | | | | | | | | |
|
||||
| Mlapo | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 🟠<sup>5</sup> | ✅ | ✅ | ✅ | ❌ | ❔ | ❌ | ✅ | ✅ | | | | | | | | |
|
||||
| [<abbr title="Multimodal Inputs">mm</abbr>](https://docs.vllm.ai/en/latest/features/multimodal_inputs/) | ✅ | ✅ | ✅ | ✅ | ✅ | 🟠 | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | ✅ | ✅ | | | | | | | |
|
||||
| Multistream Moe | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ❔ | ✅ | ✅ | ✅ | ✅ | | | | | | |
|
||||
| Shared Expert DP | ✅ | ✅ | ✅ | ✅ | ✅ | 🟠<sup>1</sup> | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ❔ | ✅ | ✅ | ❔ | ✅ | | | | | |
|
||||
| [Quantization W4A4](https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/quantization.html) | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | ❔ | ❔ | ✅ | ❔ | ✅ | ❔ | ❌ | ❔ | ❔ | ✅ | | | | |
|
||||
| [Quantization W4A8](https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/quantization.html) | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | ✅ | ❔ | ✅ | ❔ | ❌ | ✅ | ✅ | ❔ | ✅ | | | |
|
||||
| [Quantization W8A8](https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/quantization.html) | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ❔ | ✅ | ✅ | | |
|
||||
| <abbr title="Tensor Parallel">TP</abbr> | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | |
|
||||
| Weight nz | ✅ | ✅ | ✅ | ✅ | ✅ | ❔ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | 🟠 | ✅ | ✅ | ✅ |
|
||||
|
||||
- <sup>1</sup> Only dcp supports dp while pcp does not support dp.
|
||||
- <sup>2</sup> Flashcomm is only enabled on the prefill stage.
|
||||
- <sup>3</sup> Layer sharding is only enabled on the prefill stage.
|
||||
- <sup>4</sup> Lmhead TP is only enabled in the pure dp scenarios.
|
||||
- <sup>5</sup> MLAPO is only supported on the decode stage.
|
||||
@@ -1,10 +1,11 @@
|
||||
# Features and models
|
||||
# Features and Models
|
||||
|
||||
This section provides a detailed supported matrix by vLLM Ascend.
|
||||
This section provides a detailed matrix supported by vLLM Ascend.
|
||||
|
||||
:::{toctree}
|
||||
:caption: Support Matrix
|
||||
:maxdepth: 1
|
||||
supported_models
|
||||
supported_features
|
||||
feature_matrix
|
||||
:::
|
||||
|
||||
@@ -1,45 +1,48 @@
|
||||
# Feature Support
|
||||
# Supported Features
|
||||
|
||||
The feature support principle of vLLM Ascend is: **aligned with the vLLM**. We are also actively collaborating with the community to accelerate support.
|
||||
The feature support principle of vLLM Ascend is: **aligned with vLLM**. We are also actively collaborating with the community to accelerate support.
|
||||
|
||||
Functional call: <https://docs.vllm.ai/en/latest/features/tool_calling/>
|
||||
|
||||
You can check the [support status of vLLM V1 Engine][v1_user_guide]. Below is the feature support status of vLLM Ascend:
|
||||
|
||||
| Feature | Status | Next Step |
|
||||
|-------------------------------|----------------|------------------------------------------------------------------------|
|
||||
| Chunked Prefill | 🟢 Functional | Functional, see detail note: [Chunked Prefill][cp] |
|
||||
| Automatic Prefix Caching | 🟢 Functional | Functional, see detail note: [vllm-ascend#732][apc] |
|
||||
| LoRA | 🟢 Functional | [vllm-ascend#396][multilora], [vllm-ascend#893][v1 multilora] |
|
||||
| Speculative decoding | 🟢 Functional | Basic support |
|
||||
| Pooling | 🟢 Functional | CI needed and adapting more models; V1 support rely on vLLM support. |
|
||||
| Enc-dec | 🟡 Planned | vLLM should support this feature first. |
|
||||
| Multi Modality | 🟢 Functional | [Tutorial][multimodal], optimizing and adapting more models |
|
||||
| LogProbs | 🟢 Functional | CI needed |
|
||||
| Prompt logProbs | 🟢 Functional | CI needed |
|
||||
| Async output | 🟢 Functional | CI needed |
|
||||
| Beam search | 🟢 Functional | CI needed |
|
||||
| Guided Decoding | 🟢 Functional | [vllm-ascend#177][guided_decoding] |
|
||||
| Tensor Parallel | 🟢 Functional | Make TP >4 work with graph mode |
|
||||
| Pipeline Parallel | 🟢 Functional | Write official guide and tutorial. |
|
||||
| Expert Parallel | 🟢 Functional | Dynamic EPLB support. |
|
||||
| Data Parallel | 🟢 Functional | Data Parallel support for Qwen3 MoE. |
|
||||
| Prefill Decode Disaggregation | 🟢 Functional | Functional, xPyD is supported. |
|
||||
| Quantization | 🟢 Functional | W8A8 available; working on more quantization method support(W4A8, etc) |
|
||||
| Graph Mode | 🔵 Experimental| Experimental, see detail note: [vllm-ascend#767][graph_mode] |
|
||||
| Sleep Mode | 🟢 Functional | |
|
||||
| Chunked Prefill | 🟢 Functional | Functional, see detailed note: [Chunked Prefill][cp] |
|
||||
| Automatic Prefix Caching | 🟢 Functional | Functional, see detailed note: [vllm-ascend#732][apc] |
|
||||
| LoRA | 🔵 Experimental | Functional, see detailed note: [LoRA][LoRA] |
|
||||
| Speculative decoding | 🟢 Functional | Basic support |
|
||||
| Pooling | 🔵 Experimental | CI needed to adapt to more models; V1 support relies on vLLM support. |
|
||||
| Enc-dec | 🟡 Planned | vLLM should support this feature first. |
|
||||
| Multi Modality | 🟢 Functional | [Multi Modality][multimodal], optimizing and adapting more models |
|
||||
| LogProbs | 🟢 Functional | CI needed |
|
||||
| Prompt logProbs | 🟢 Functional | CI needed |
|
||||
| Async output | 🟢 Functional | CI needed |
|
||||
| Beam search | 🔵 Experimental | CI needed |
|
||||
| Guided Decoding | 🟢 Functional | [vllm-ascend#177][guided_decoding] |
|
||||
| Tensor Parallel | 🟢 Functional | Make TP >4 work with graph mode. |
|
||||
| Pipeline Parallel | 🟢 Functional | Write official guide and tutorial. |
|
||||
| Expert Parallel | 🟢 Functional | Support dynamic EPLB. |
|
||||
| Data Parallel | 🟢 Functional | Data Parallel support for Qwen3 MoE. |
|
||||
| Prefill Decode Disaggregation | 🟢 Functional | Functional, xPyD is supported. |
|
||||
| Quantization | 🟢 Functional | W8A8 available; working on more quantization method support (W4A8, etc) |
|
||||
| Graph Mode | 🟢 Functional | Functional, see detailed note: [Graph Mode][graph_mode] |
|
||||
| Sleep Mode | 🟢 Functional | Functional, see detailed note: [Sleep Mode][sleep_mode] |
|
||||
| Context Parallel | 🟢 Functional | Functional, see detailed note: [Context Parallel][context_parallel] |
|
||||
|
||||
- 🟢 Functional: Fully operational, with ongoing optimizations.
|
||||
- 🔵 Experimental: Experimental support, interfaces and functions may change.
|
||||
- 🚧 WIP: Under active development, will be supported soon.
|
||||
- 🟡 Planned: Scheduled for future implementation (some may have open PRs/RFCs).
|
||||
- 🔴 NO plan / Deprecated: No plan or deprecated by vLLM.
|
||||
- 🔴 NO plan/Deprecated: No plan or deprecated by vLLM.
|
||||
|
||||
[v1_user_guide]: https://docs.vllm.ai/en/latest/getting_started/v1_user_guide.html
|
||||
[multimodal]: https://vllm-ascend.readthedocs.io/en/latest/tutorials/single_npu_multimodal.html
|
||||
[v1_user_guide]: https://docs.vllm.ai/en/latest/usage/v1_guide/
|
||||
[multimodal]: https://docs.vllm.ai/projects/ascend/en/latest/tutorials/models/Qwen-VL-Dense.html
|
||||
[guided_decoding]: https://github.com/vllm-project/vllm-ascend/issues/177
|
||||
[multilora]: https://github.com/vllm-project/vllm-ascend/issues/396
|
||||
[v1 multilora]: https://github.com/vllm-project/vllm-ascend/pull/893
|
||||
[graph_mode]: https://github.com/vllm-project/vllm-ascend/issues/767
|
||||
[LoRA]: https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/lora.html
|
||||
[graph_mode]: https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/graph_mode.html
|
||||
[apc]: https://github.com/vllm-project/vllm-ascend/issues/732
|
||||
[cp]: https://docs.vllm.ai/en/stable/performance/optimization.html#chunked-prefill
|
||||
[cp]: https://docs.vllm.ai/en/stable/configuration/optimization/
|
||||
[1P1D]: https://github.com/vllm-project/vllm-ascend/pull/950
|
||||
[ray]: https://github.com/vllm-project/vllm-ascend/issues/1751
|
||||
[context_parallel]: https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/context_parallel.html
|
||||
[sleep_mode]: https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/sleep_mode.html
|
||||
|
||||
@@ -1,79 +1,194 @@
|
||||
# Model Support
|
||||
# Supported Models
|
||||
|
||||
Get the newest info here: https://github.com/vllm-project/vllm-ascend/issues/1608
|
||||
Get the latest info here: <https://github.com/vllm-project/vllm-ascend/issues/1608>
|
||||
|
||||
## Text-only Language Models
|
||||
**Legend Description**:
|
||||
|
||||
- ✅ = Supported model/feature
|
||||
- 🔵 = Experimental supported model/feature
|
||||
- ❌ = Not supported model/feature
|
||||
- 🟡 = Not tested or verified
|
||||
|
||||
## Text-Only Language Models
|
||||
|
||||
### Generative Models
|
||||
|
||||
| Model | Supported | Note |
|
||||
|-------------------------------|-----------|----------------------------------------------------------------------|
|
||||
| DeepSeek v3 | ✅ | |
|
||||
| DeepSeek R1 | ✅ | |
|
||||
| DeepSeek Distill (Qwen/LLama) | ✅ | |
|
||||
| Qwen3 | ✅ | |
|
||||
| Qwen3-based | ✅ | |
|
||||
| Qwen3-Coder | ✅ | |
|
||||
| Qwen3-Moe | ✅ | |
|
||||
| Qwen2.5 | ✅ | |
|
||||
| Qwen2 | ✅ | |
|
||||
| Qwen2-based | ✅ | |
|
||||
| QwQ-32B | ✅ | |
|
||||
| LLama2/3/3.1 | ✅ | |
|
||||
| Internlm | ✅ | [#1962](https://github.com/vllm-project/vllm-ascend/issues/1962) |
|
||||
| Baichuan | ✅ | |
|
||||
| Baichuan2 | ✅ | |
|
||||
| Phi-4-mini | ✅ | |
|
||||
| MiniCPM | ✅ | |
|
||||
| MiniCPM3 | ✅ | |
|
||||
| Ernie4.5 | ✅ | |
|
||||
| Ernie4.5-Moe | ✅ | |
|
||||
| Gemma-2 | ✅ | |
|
||||
| Gemma-3 | ✅ | |
|
||||
| Phi-3/4 | ✅ | |
|
||||
| Mistral/Mistral-Instruct | ✅ | |
|
||||
| GLM-4.5 | ✅ | |
|
||||
| GLM-4 | ❌ | [#2255](https://github.com/vllm-project/vllm-ascend/issues/2255) |
|
||||
| GLM-4-0414 | ❌ | [#2258](https://github.com/vllm-project/vllm-ascend/issues/2258) |
|
||||
| ChatGLM | ❌ | [#554](https://github.com/vllm-project/vllm-ascend/issues/554) |
|
||||
| DeepSeek v2.5 | 🟡 | Need test |
|
||||
| Mllama | 🟡 | Need test |
|
||||
| MiniMax-Text | 🟡 | Need test |
|
||||
#### Core Supported Models
|
||||
|
||||
:::::{tab-set}
|
||||
|
||||
::::{tab-item} Ascend 950 Products
|
||||
|
||||
| Model | Support | Note | BF16 | Supported Hardware | W8A8 | Chunked Prefill | Automatic Prefix Cache | LoRA | Speculative Decoding | Async Scheduling | Tensor Parallel | Pipeline Parallel | Expert Parallel | Data Parallel | Prefill-decode Disaggregation | Piecewise AclGraph | Fullgraph AclGraph | max-model-len | MLP Weight Prefetch | Doc |
|
||||
|-------|--------|--------|------|------|------|---------|-------|------|------|--------|-------|--------|--------|-------|-------|--------|----------|---------|----------|-----|
|
||||
|DeepSeek V4-Flash|✅|Native mixed MXFP8/MXFP4 weights||Ascend 950 Products|✅|✅|✅||✅|✅||✅|✅|✅|✅||✅|1M||[DeepSeek V4-Flash](../../tutorials/models/DeepSeek-V4-Flash.md)|
|
||||
|DeepSeek V4-Pro|✅|Native mixed MXFP8/MXFP4 weights||Ascend 950 Products|✅|✅|✅||✅|✅||✅|✅|✅|✅||✅|1M||[DeepSeek V4-Pro](../../tutorials/models/DeepSeek-V4-Pro.md)|
|
||||
|DeepSeek-V3.1|✅| |✅| Ascend 950 Products |✅|✅|✅||✅|✅|✅||✅|✅|✅|✅|✅|240k|| [DeepSeek-V3.1](../../tutorials/models/DeepSeek-V3.1.md) |
|
||||
|GLM-5.1|✅| |✅| Ascend 950 Products |✅|✅|✅||✅|✅|✅||✅|✅|✅||✅|200k||[GLM-5.1](../../tutorials/models/GLM5.md) |
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} A2/A3
|
||||
|
||||
| Model | Support | Note | BF16 | Supported Hardware | W8A8 | Chunked Prefill | Automatic Prefix Cache | LoRA | Speculative Decoding | Async Scheduling | Tensor Parallel | Pipeline Parallel | Expert Parallel | Data Parallel | Prefill-decode Disaggregation | Piecewise AclGraph | Fullgraph AclGraph | max-model-len | MLP Weight Prefetch | Doc |
|
||||
|-------------------------------|---------|-----------------------------------------------------|------|--------------------|------|-----------------|------------------------|------|----------------------|------------------|-----------------|-------------------|-----------------|---------------|-------------------------------|--------------------|--------------------|---------------|---------------------|-----|
|
||||
| DeepSeek V4-Flash | ✅ | | ✅ | A2/A3 | ✅ | ✅ |✅|| ✅ |✅| ✅ || ✅ | ✅ | ✅ || ✅ | 1M || [DeepSeek V4-Flash](../../tutorials/models/DeepSeek-V4-Flash.md) |
|
||||
| DeepSeek V4-Pro | ✅ | | ✅ | A2/A3 | ✅ | ✅ |✅|| ✅ |✅| ✅ || ✅ | ✅ | ✅ || ✅ | 1M || [DeepSeek-V4-Pro](../../tutorials/models/DeepSeek-V4-Pro.md) |
|
||||
| DeepSeek V3/3.1 | ✅ | | ✅ | A2/A3 | ✅ | ✅ | ✅ || ✅ || ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 240k || [DeepSeek-V3.1](../../tutorials/models/DeepSeek-V3.1.md) |
|
||||
| DeepSeek V3.2 | ✅ | | ✅ | A2/A3 | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 160k | ✅ | [DeepSeek-V3.2](../../tutorials/models/DeepSeek-V3.2.md) |
|
||||
| DeepSeek R1 | ✅ | | ✅ | A2/A3 | ✅ | ✅ | ✅ || ✅ || ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 128k || [DeepSeek-R1](../../tutorials/models/DeepSeek-R1.md) |
|
||||
| Qwen3-Dense | ✅ | | ✅ | A2/A3 | ✅ | ✅ | ✅ ||| ✅ | ✅ ||| ✅ || ✅ | ✅ | 128k | ✅ | [Qwen3-Dense](../../tutorials/models/Qwen3-Dense.md) |
|
||||
| Qwen3-30B-A3B | ✅ | | ✅ | A2/A3 | ✅ | ✅ | ✅ || ✅ | ✅ | ✅ || ✅ | ✅ || ✅ | ✅ ||| [Qwen3-30B-A3B](../../tutorials/models/Qwen3-30B-A3B.md) |
|
||||
| Qwen3-Coder-30B-A3B | 🔵 | | ✅ | A2/A3 | ✅ | ✅ | ✅ || ✅ | ✅ | ✅ || ✅ | ✅ || ✅ | ✅ ||| [Qwen3-Coder-30B-A3B](../../tutorials/models/Qwen3-Coder-30B-A3B.md) |
|
||||
| Qwen3-235B-A22B | ✅ | | ✅ | A2/A3 | ✅ | ✅ | ✅ ||| ✅ | ✅ || ✅ | ✅ | ✅ | ✅ | ✅ | 256k || [Qwen3-235B-A22B](../../tutorials/models/Qwen3-235B-A22B.md) |
|
||||
| Qwen3-Next | 🔵 | | ✅ | A2/A3 | ✅ |||||| ✅ ||| ✅ || ✅ | ✅ ||| [Qwen3-Next](../../tutorials/models/Qwen3-Next.md) |
|
||||
| GLM-4.x | ✅ | | | A2/A3 |✅|✅|✅||✅|✅|✅||✅|✅|✅|✅|✅|198k||[GLM-4.x](../../tutorials/models/GLM4.x.md)|
|
||||
| GLM-5/5.1 | ✅ | | ✅ | A2/A3 | ✅ | ✅ | ✅ || ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 200k || [GLM-5](../../tutorials/models/GLM5.md) |
|
||||
| GLM-5.2 | 🔵 | | ✅ | A2/A3 | ✅ | ✅ | ✅ || ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 200k || [GLM-5](../../tutorials/models/GLM5.2.md) |
|
||||
| Kimi-K2-Thinking | 🔵 | | | A2/A3 |||||||||||||||| [Kimi-K2-Thinking](../../tutorials/models/Kimi-K2-Thinking.md) |
|
||||
| DeepSeekOCR2 | ✅ | | ✅ | A2/A3 ||✅||||✅|||||||||| [DeepSeekOCR2](../../tutorials/models/DeepSeekOCR2.md) |
|
||||
| MiniMax-M2.5/2.7 | ✅ | | ✅ | A2/A3/Ascend950 (Ascend950 experimental) |✅|✅|✅|❌|✅|✅|✅|🟡|✅|✅|✅|🟡|✅|200k|🟡| [MiniMax-M2](../../tutorials/models/MiniMax-M2.md) |
|
||||
| Qwen2.5-Math-RM-72B | 🔵 | vllm-rm, tensor_parallel_size=4, max_model_len=4096 | ✅ | A2 | ✅ | 🟡 | 🟡 | ❌ | 🟡 | ✅ | ✅ | 🟡 | 🟡 | 🟡 | 🟡 | 🟡 | 🟡 | 4096 | 🟡 | [Qwen2.5-Math-RM-72B](../../tutorials/models/Qwen2.5-Math-RM-72B.md) |
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
|
||||
| Model | Support | Note | BF16 | Supported Hardware | W8A8 | Chunked Prefill | Automatic Prefix Cache | LoRA | Speculative Decoding | Async Scheduling | Tensor Parallel | Prefill-decode Disaggregation | Piecewise AclGraph | Fullgraph AclGraph | max-model-len | Doc |
|
||||
|---------------|---------|------|------|--------------------|------|-----------------|------------------------|------|----------------------|------------------|-----------------|-------------------------------|--------------------|--------------------|---------------|-----|
|
||||
| Qwen3-Dense | ✅ | | ❌ | Atlas 300I DUO | ✅ | ✅ | ✅ | ❌ | 🟡 | ✅ | ✅ | ❌ | ✅ | ✅ | 20k | [Qwen3-Dense](../../tutorials/models/Qwen3-Dense.md) |
|
||||
| Qwen3-30B-A3B | ✅ | | ❌ | Atlas 300I DUO | ✅ | ✅ | ✅ | ❌ | 🟡 | ✅ | ✅ | ❌ | ✅ | ✅ | 16k | [Qwen3-30B-A3B](../../tutorials/models/Qwen3-30B-A3B.md) |
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
#### Extended Compatible Models
|
||||
|
||||
| Model | Support | Note | Supported Hardware |
|
||||
|-------------------------------|-----------|----------------------------------------------------------------------|--------------------|
|
||||
| DeepSeek Distill (Qwen/Llama) | 🔵 | | A2/A3 |
|
||||
| Qwen3-based | 🔵 | | A2/A3 |
|
||||
| Qwen2 | 🔵 | | A2/A3 |
|
||||
| Qwen2.5 | ✅ | | A2/A3 |
|
||||
| Qwen2-based | 🔵 | | A2/A3 |
|
||||
| QwQ-32B | 🔵 | | A2/A3 |
|
||||
| Llama2/3/3.1/3.2 | 🔵 | | A2/A3 |
|
||||
| Internlm | 🔵 | [#1962](https://github.com/vllm-project/vllm-ascend/issues/1962) | A2/A3 |
|
||||
| Baichuan | 🔵 | | A2/A3 |
|
||||
| Baichuan2 | 🔵 | | A2/A3 |
|
||||
| Phi-4-mini | 🔵 | | A2/A3 |
|
||||
| MiniCPM | 🔵 | | A2/A3 |
|
||||
| MiniCPM3 | 🔵 | | A2/A3 |
|
||||
| Ernie4.5 | 🔵 | | A2/A3 |
|
||||
| Ernie4.5-Moe | 🔵 | | A2/A3 |
|
||||
| Gemma-2 | 🔵 | | A2/A3 |
|
||||
| Gemma-3 | 🔵 | | A2/A3 |
|
||||
| Phi-3/4 | 🔵 | | A2/A3 |
|
||||
| Mistral/Mistral-Instruct | 🔵 | | A2/A3 |
|
||||
| Hy3-preview | 🔵 | | A3 |
|
||||
| DeepSeek V2.5 | 🟡 | Need test | |
|
||||
| Mllama | 🟡 | Need test | |
|
||||
| MiniMax-Text | 🟡 | Need test | |
|
||||
|
||||
### Pooling Models
|
||||
|
||||
| Model | Supported | Note |
|
||||
|-------------------------------|-----------|----------------------------------------------------------------------|
|
||||
| Qwen3-Embedding | ✅ | |
|
||||
| Molmo | ✅ | [1942](https://github.com/vllm-project/vllm-ascend/issues/1942) |
|
||||
| XLM-RoBERTa-based | ❌ | [1960](https://github.com/vllm-project/vllm-ascend/issues/1960) |
|
||||
:::::{tab-set}
|
||||
::::{tab-item} A2/A3
|
||||
|
||||
| Model | Support | Note | Supported Hardware | W8A8 | Doc |
|
||||
|-------------------------------|-----------|----------------------------------------------------------------------|------------------------------|------|------|
|
||||
| Qwen3-Embedding | 🔵 | | A2/A3 |🟡| [Qwen3-Embedding](../../tutorials/models/Qwen3-Embedding.md)|
|
||||
| Qwen3-VL-Embedding | 🔵 | | A2/A3 |🔵| [Qwen3-VL-Embedding](../../tutorials/models/Qwen3-VL-Embedding.md)|
|
||||
| Qwen3-Reranker | 🔵 | | A2/A3 |🟡| [Qwen3-Reranker](../../tutorials/models/Qwen3-Reranker.md)|
|
||||
| Qwen3-VL-Reranker | 🔵 | | A2/A3 |🔵| [Qwen3-VL-Reranker](../../tutorials/models/Qwen3-VL-Reranker.md)|
|
||||
| Molmo | 🔵 | [1942](https://github.com/vllm-project/vllm-ascend/issues/1942) | A2/A3 |🟡| |
|
||||
| XLM-RoBERTa-based | 🔵 | | A2/A3 |🟡| |
|
||||
| Bert | 🔵 | | A2/A3 |🟡| |
|
||||
| Qwen2.5-Math-RM-72B | 🔵 | Reward Model, gsm8k_correctness accuracy=0.80 | A2 |🟡| [Qwen2.5-Math-RM-72B](../../tutorials/models/Qwen2.5-Math-RM-72B.md) |
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
|
||||
| Model | Support | Note | Supported Hardware | W8A8| Doc |
|
||||
|-------------------|---------|------|--------------------|-----|--------------------------------------------------------------------|
|
||||
| Qwen3-Embedding | 🔵 | FP16 | Atlas 300I DUO |🟡| [Qwen3-Embedding](../../tutorials/models/Qwen3-Embedding.md) |
|
||||
| Qwen3-VL-Embedding| 🔵 | FP16 | Atlas 300I DUO |🔵| [Qwen3-VL-Embedding](../../tutorials/models/Qwen3-VL-Embedding.md) |
|
||||
| Qwen3-Reranker | 🔵 | FP16 | Atlas 300I DUO |🟡| [Qwen3-Reranker](../../tutorials/models/Qwen3-Reranker.md) |
|
||||
| Qwen3-VL-Reranker | 🔵 | FP16 | Atlas 300I DUO |🔵| [Qwen3-VL-Reranker](../../tutorials/models/Qwen3-VL-Reranker.md) |
|
||||
| XLM-RoBERTa-based | 🔵 | FP16; embedding and scoring | Atlas 300I DUO |🟡| |
|
||||
| Qwen2.5-based | 🔵 | FP16 classification | Atlas 300I DUO |🟡| |
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
## Multimodal Language Models
|
||||
|
||||
### Generative Models
|
||||
|
||||
| Model | Supported | Note |
|
||||
|--------------------------------|---------------|----------------------------------------------------------------------|
|
||||
| Qwen2-VL | ✅ | |
|
||||
| Qwen2.5-VL | ✅ | |
|
||||
| Qwen2.5-Omni | ✅ | [1760](https://github.com/vllm-project/vllm-ascend/issues/1760) |
|
||||
| QVQ | ✅ | |
|
||||
| LLaVA 1.5/1.6 | ✅ | [1962](https://github.com/vllm-project/vllm-ascend/issues/1962) |
|
||||
| InternVL2 | ✅ | |
|
||||
| InternVL2.5 | ✅ | |
|
||||
| Qwen2-Audio | ✅ | |
|
||||
| Aria | ✅ | |
|
||||
| LLaVA-Next | ✅ | |
|
||||
| LLaVA-Next-Video | ✅ | |
|
||||
| MiniCPM-V | ✅ | |
|
||||
| Mistral3 | ✅ | |
|
||||
| Phi-3-Vison/Phi-3.5-Vison | ✅ | |
|
||||
| Gemma3 | ✅ | |
|
||||
| LLama4 | ❌ | [1972](https://github.com/vllm-project/vllm-ascend/issues/1972) |
|
||||
| LLama3.2 | ❌ | [1972](https://github.com/vllm-project/vllm-ascend/issues/1972) |
|
||||
| Keye-VL-8B-Preview | ❌ | [1963](https://github.com/vllm-project/vllm-ascend/issues/1963) |
|
||||
| Florence-2 | ❌ | [2259](https://github.com/vllm-project/vllm-ascend/issues/2259) |
|
||||
| GLM-4V | ❌ | [2260](https://github.com/vllm-project/vllm-ascend/issues/2260) |
|
||||
| InternVL2.0/2.5/3.0<br>InternVideo2.5/Mono-InternVL | ❌ | [2064](https://github.com/vllm-project/vllm-ascend/issues/2064) |
|
||||
| Whisper | ❌ | [2262](https://github.com/vllm-project/vllm-ascend/issues/2262) |
|
||||
| Ultravox | 🟡 Need test | |
|
||||
#### Core Supported Models
|
||||
|
||||
:::::{tab-set}
|
||||
|
||||
::::{tab-item} Ascend 950 Products
|
||||
|
||||
| Model | Support | Note | BF16 | Supported Hardware | W8A8 | Chunked Prefill | Automatic Prefix Cache | LoRA | Speculative Decoding | Async Scheduling | Tensor Parallel | Pipeline Parallel | Expert Parallel | Data Parallel | Prefill-decode Disaggregation | Piecewise AclGraph | Fullgraph AclGraph | max-model-len | MLP Weight Prefetch | Doc |
|
||||
|-----------------|----------|--------|------|------|------|---------|-------|------|------|--------|-------|--------|--------|-------|-------|--------|----------|---------|----------|-----|
|
||||
|Qwen3.5-397B-A17B|✅ | |✅ | Ascend 950DT |✅|✅|✅||✅|✅|✅||✅|✅|✅|✅|✅|1010000|| [Qwen3.5-397B-A17B](../../tutorials/models/Qwen3.5-397B-A17B.md) |
|
||||
|Qwen3.6-27B |✅ | |✅ | Ascend 950 Products |✅|✅|✅||✅|✅|✅||✅|✅|✅|✅|✅|262144|| [Qwen3.5-27B / Qwen3.6-27B](../../tutorials/models/Qwen3.5-27B-Qwen3.6-27B.md) |
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} A2/A3
|
||||
|
||||
| Model | Support | Note | BF16 | Supported Hardware | W8A8 | Chunked Prefill | Automatic Prefix Cache | LoRA | Speculative Decoding | Async Scheduling | Tensor Parallel | Pipeline Parallel | Expert Parallel | Data Parallel | Prefill-decode Disaggregation | Piecewise AclGraph | Fullgraph AclGraph | max-model-len | MLP Weight Prefetch | Doc |
|
||||
|-------------------------------------|---------------|------|------|--------------------|------|-----------------|------------------------|------|----------------------|------------------|-----------------|-------------------|-----------------|---------------|-------------------------------|--------------------|--------------------|---------------|---------------------|-----|
|
||||
| Qwen3-VL | ✅ | | |A2/A3|||||||✅|||||✅|✅||| [Qwen-VL-Dense](../../tutorials/models/Qwen-VL-Dense.md) |
|
||||
| Qwen3-VL-30B-A3B/Qwen3-VL-235B-A22B | ✅ | | ✅ | A2/A3 | ✅ | ✅ | ✅ | | | ✅ | ✅ | | ✅ | ✅ | ✅ | ✅ | ✅ | 262144 || [Qwen3-VL-30B-A3B](../../tutorials/models/Qwen3-VL-30B-A3B-Instruct.md)/[Qwen3-VL-235B-A22B](../../tutorials/models/Qwen3-VL-235B-A22B-Instruct.md) |
|
||||
| Qwen3.5-397B-A17B | ✅ | | ✅ | A2/A3 |✅|✅|✅||✅|✅|✅||✅|✅|✅|✅|✅|1010000|| [Qwen3.5-397B-A17B](../../tutorials/models/Qwen3.5-397B-A17B.md) |
|
||||
| Qwen3.5-27B / Qwen3.6-27B | ✅ | | ✅ | A2/A3 |✅|✅|✅||✅|✅|✅||✅|✅|✅|✅|✅|262144|| [Qwen3.5-27B / Qwen3.6-27B](../../tutorials/models/Qwen3.5-27B-Qwen3.6-27B.md) |
|
||||
| Qwen3.6-35B-A3B | ✅ | | ✅ | A2/A3 |✅|✅|✅||🔵|✅|✅||✅|✅|❌|✅|✅|262144|| [Qwen3.6-35B-A3B](../../tutorials/models/Qwen3.6-35B-A3B.md) |
|
||||
| Qwen3-Omni-30B-A3B-Thinking | ✅ | | |A2/A3|||||||✅||✅|||||||[Qwen3-Omni-30B-A3B-Thinking](../../tutorials/models/Qwen3-Omni-30B-A3B-Thinking.md)|
|
||||
| Kimi-K2.5/Kimi-K2.6 | ✅ | | |A2/A3||✅|✅||✅|✅|✅||✅|✅|✅|✅|✅|262144||[Kimi-K2.5](../../tutorials/models/Kimi-K2.5.md)/[Kimi-K2.6](../../tutorials/models/Kimi-K2.6.md)|
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Atlas 300I DUO
|
||||
|
||||
| Model | Support | Note | BF16 | Supported Hardware | W8A8 | Chunked Prefill | Automatic Prefix Cache | LoRA | Speculative Decoding | Async Scheduling | Tensor Parallel | Prefill-decode Disaggregation | Piecewise AclGraph | Fullgraph AclGraph | max-model-len | Doc |
|
||||
|-----------------|---------|------|------|--------------------|------|-----------------|------------------------|------|----------------------|------------------|-----------------|-------------------------------|--------------------|--------------------|---------------|-----|
|
||||
| Qwen3-VL | ✅ | | ❌ | Atlas 300I DUO | ✅ | ✅ | ✅ | ❌ | 🟡 | ✅ | ✅ | ❌ | ✅ | ✅ | 16k | [Qwen-VL-Dense](../../tutorials/models/Qwen-VL-Dense.md) |
|
||||
| Qwen3.5-Dense | ✅ | | ❌ | Atlas 300I DUO | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | 256k | [Qwen3.5-Dense](../../tutorials/models/Qwen3.5-Dense.md) |
|
||||
| Qwen3.5-35B-A3B | ✅ | | ❌ | Atlas 300I DUO | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | 256k | [Qwen3.5-35B-A3B](../../tutorials/models/Qwen3.6-35B-A3B.md) |
|
||||
| Qwen3.6-27B | ✅ | | ❌ | Atlas 300I DUO | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | 256k | [Qwen3.6-27B](../../tutorials/models/Qwen3.5-27B-Qwen3.6-27B.md) |
|
||||
| Qwen3.6-35B-A3B | ✅ | | ❌ | Atlas 300I DUO | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | 256k | [Qwen3.6-35B-A3B](../../tutorials/models/Qwen3.6-35B-A3B.md) |
|
||||
| PaddleOCR-VL | 🔵 | | ❌ | Atlas 300I DUO | ❌ | ✅ | ✅ | ❌ | ❌ | ✅ | ❌ | ❌ | ✅ | ✅ | 16k | [PaddleOCR-VL](../../tutorials/models/PaddleOCR-VL.md) |
|
||||
| Qwen3-ASR | 🔵 | | ❌ | Atlas 300I DUO | ❌ | ✅ | ✅ | ❌ | ❌ | ✅ | 🟡 | ❌ | ✅ | ✅ | 4096 | [Qwen3-ASR-1.7B](../../tutorials/models/Qwen3-ASR-1.7B.md) |
|
||||
|
||||
::::
|
||||
:::::
|
||||
|
||||
#### Extended Compatible Models
|
||||
|
||||
| Model | Support | Note | Supported Hardware |
|
||||
|--------------------------------|---------------|----------------------------------------------------------------------|--------------------|
|
||||
| Qwen2-VL | 🔵 | | A2/A3 |
|
||||
| Qwen3-Omni | 🔵 | | A2/A3 |
|
||||
| QVQ | 🔵 | | A2/A3 |
|
||||
| Qwen2-Audio | 🔵 | | A2/A3 |
|
||||
| Aria | 🔵 | | A2/A3 |
|
||||
| LLaVA-Next | 🔵 | | A2/A3 |
|
||||
| LLaVA-Next-Video | 🔵 | | A2/A3 |
|
||||
| MiniCPM-V | 🔵 | | A2/A3 |
|
||||
| Mistral3 | 🔵 | | A2/A3 |
|
||||
| Phi-3-Vision/Phi-3.5-Vision | 🔵 | | A2/A3 |
|
||||
| Gemma3 | 🔵 | | A2/A3 |
|
||||
| Llama3.2 | 🔵 | | A2/A3 |
|
||||
| PaddleOCR-VL | 🔵 | | A2/A3 |
|
||||
| Llama4 | ❌ | [1972](https://github.com/vllm-project/vllm-ascend/issues/1972) | |
|
||||
| Keye-VL-8B-Preview | ❌ | [1961](https://github.com/vllm-project/vllm-ascend/issues/1961) | |
|
||||
| Florence-2 | ❌ | [2259](https://github.com/vllm-project/vllm-ascend/issues/2259) | |
|
||||
| GLM-4V | ❌ | [2260](https://github.com/vllm-project/vllm-ascend/issues/2260) | |
|
||||
| InternVL2.0/2.5/3.0<br>InternVideo2.5/Mono-InternVL | ❌ | [2064](https://github.com/vllm-project/vllm-ascend/issues/2064) | |
|
||||
| Whisper | ❌ | [2262](https://github.com/vllm-project/vllm-ascend/issues/2262) | |
|
||||
| Ultravox | 🟡 | Need test | |
|
||||
|
||||