133
docs/source/developer_guide/Design_Documents/ACL_Graph.md
Normal file
133
docs/source/developer_guide/Design_Documents/ACL_Graph.md
Normal file
@@ -0,0 +1,133 @@
|
||||
# ACL Graph
|
||||
|
||||
## Overview
|
||||
|
||||
ACL Graph is the Ascend realization of vLLM static graph execution. Upstream vLLM and PyTorch documents already describe the generic graph model, including `CUDAGraphMode`, runtime dispatch, batch descriptors, bucketing and padding, and the definitions of full graph and piecewise graph. This document focuses on what is specific to Ascend in `vllm-ascend`: the platform integration points, the extra constraints introduced by ACL graph capture, and the mechanisms used to keep attention parameters correct during replay.
|
||||
|
||||
On Ascend, the design goal is the same as upstream static graph execution: reduce host launch overhead for small and medium runtime shapes. The implementation boundary is different. vLLM provides the generic dispatch path, while `vllm-ascend` supplies the platform wrapper, capture-size trimming, and attention-specific update logic needed by ACL graph replay.
|
||||
|
||||
## Prerequisites and References
|
||||
|
||||
- Upstream vLLM design doc for generic graph concepts: [CUDA Graphs](https://docs.vllm.ai/en/latest/design/cuda_graphs/).
|
||||
- PyTorch graph documentation for generic capture and replay semantics: [Accelerating PyTorch with CUDA Graphs](https://pytorch.org/blog/accelerating-pytorch-with-cuda-graphs/).
|
||||
- Ascend user guide for operational enablement: [Graph Mode Guide](https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/graph_mode.html).
|
||||
- Existing repo design note: [ACL Graph](https://docs.vllm.ai/projects/ascend/zh-cn/latest/developer_guide/Design_Documents/ACL_Graph.html)
|
||||
|
||||
This document intentionally does not re-explain upstream topics such as graph mode selection, dispatcher behavior, batch descriptor construction, capture bucketing, padding policy, or the generic meaning of full versus piecewise execution.
|
||||
|
||||
## How ACL Graph Fits into vLLM
|
||||
|
||||
vLLM owns the generic static graph flow. On Ascend, `NPUPlatform.get_static_graph_wrapper_cls()` returns `vllm_ascend.compilation.acl_graph.ACLGraphWrapper`, which is the platform-specific wrapper used when vLLM enables static graph mode.
|
||||
|
||||
`ACLGraphWrapper` is responsible for:
|
||||
|
||||
- reading the runtime mode and `batch_descriptor` from the forward context,
|
||||
- deciding whether to run eagerly, capture a new ACL graph, or replay a cached ACL graph,
|
||||
- caching graph entries per batch descriptor,
|
||||
- preserving the graph pool and replay bookkeeping needed by the Ascend backend.
|
||||
|
||||
The wrapper does not define the upstream dispatch policy. It assumes the runtime mode and batch descriptor have already been chosen correctly by vLLM, then applies Ascend capture or replay to that concrete runtime shape.
|
||||
|
||||
## Capture Sizes and Bucketing
|
||||
|
||||
vLLM graph replay requires stable runtime shapes, so vLLM does not try to capture every possible batch shape. Instead, it prepares a finite set of capture sizes and dispatches a runtime batch to the nearest supported size. If the runtime batch is larger than the largest configured capture size, graph mode is skipped and execution falls back to eager mode.
|
||||
|
||||
By default, vLLM builds capture sizes as:
|
||||
|
||||
- `1`, `2`, `4`
|
||||
- multiples of `8` from `8` up to `255`
|
||||
- multiples of `16` from `256` up to `max_cudagraph_capture_size`
|
||||
|
||||
Conceptually, the default list looks like:
|
||||
|
||||
```text
|
||||
[1, 2, 4, 8, 16, 24, 32, ..., 248, 256, 272, 288, ...]
|
||||
```
|
||||
|
||||
The smaller step at small batch sizes reduces padding overhead where latency is most sensitive, while the larger step at bigger sizes keeps the number of captured graphs under control.
|
||||
|
||||
On Ascend, this generic upstream bucketing strategy is still the starting point, but the final capture sizes may be reduced further by platform-specific constraints:
|
||||
|
||||
- sequence-parallel filtering may remove unsupported sizes,
|
||||
- runtime resource limits may still prevent some configured sizes from being captured,
|
||||
- some runtime modes may be normalized before capture begins.
|
||||
|
||||
## Ascend-Specific Design Constraints
|
||||
|
||||
### Capture breadth is still constrained by runtime resources
|
||||
|
||||
Unlike CUDA Graph on CUDA devices, ACL graph capture on Ascend can still fail when the selected graph sizes consume more runtime resources than the current backend can supply. Piecewise mode is the most sensitive case because it captures many subgraphs and the total capture cost scales with model depth and configured size coverage.
|
||||
|
||||
Older versions of vLLM Ascend applied a local `update_aclgraph_sizes()` heuristic to shrink the PIECEWISE capture-size set before final capture. That heuristic has been removed. The current implementation keeps upstream sizing and dispatch behavior intact, then intercepts the confirmed capture-time stream-resource signature in `vllm_ascend/compilation/acl_graph.py` and re-raises it with clearer mitigation guidance.
|
||||
|
||||
In practice, this means users should treat `cudagraph_capture_sizes` and `max_cudagraph_capture_size` as the primary tuning levers when capture fails. Newer HDK/CANN combinations can materially improve ACL graph capacity, while communication-heavy configurations may still require a smaller configured size set.
|
||||
|
||||
### Platform mode normalization is stricter than generic upstream behavior
|
||||
|
||||
Ascend currently narrows some generic upstream modes in `vllm_ascend.platform.NPUPlatform.check_and_update_config()`.
|
||||
|
||||
- Encoder-decoder models are forced to `PIECEWISE`.
|
||||
- `use_inductor` is disabled for ACL graph paths.
|
||||
- `ASCEND_LAUNCH_BLOCKING=1` is rejected when ACL graph is enabled.
|
||||
- Xlite graph mode can disable ACL graph full mode or fall back to `FULL_DECODE_ONLY`, depending on configuration.
|
||||
|
||||
These checks document the subset of upstream graph behavior that the current Ascend backend can execute safely. Some of them are long-term platform constraints, while others are clearly transitional in the current implementation.
|
||||
|
||||
## Key Ascend-Specific Mechanisms
|
||||
|
||||
### Host-side attention parameter update for full graph replay
|
||||
|
||||
Full graph replay on Ascend has an extra problem that upstream generic documentation does not cover in detail: some attention operators need runtime metadata updates even when the overall graph is static. The Ascend implementation handles this by separating graph capture from host-side task parameter updates.
|
||||
|
||||
The flow is:
|
||||
|
||||
1. During capture, attention backends record per-graph task handles, events, workspaces, and weak references to the tensors or metadata that must be refreshed.
|
||||
2. Before replay, `update_full_graph_params()` calls the backend specific `update_graph_params()` implementation.
|
||||
3. That backend runs parameter refresh on an update stream with `torch.npu.graph_task_update_begin(...)` and `torch.npu.graph_task_update_end(...)` around the underlying attention operator launch.
|
||||
4. `torch.npu.ExternalEvent` objects are used to enforce ordering between the host-side update stream and the replay stream.
|
||||
|
||||
This mechanism is implemented in attention backends such as:
|
||||
|
||||
- `vllm_ascend/attention/attention_v1.py`
|
||||
- `vllm_ascend/attention/mla_v1.py`
|
||||
- `vllm_ascend/attention/context_parallel/attention_cp.py`
|
||||
- `vllm_ascend/attention/context_parallel/mla_cp.py`
|
||||
|
||||
The important design point is that Ascend full graph support depends on backend-provided `update_graph_params()` hooks. Without that hook, capture alone is not enough to replay the correct attention state.
|
||||
|
||||
### Replay ordering and synchronization
|
||||
|
||||
`ACLGraphWrapper` synchronizes the current stream before replay in the common path to ensure that host-side parameter updates stay aligned with the graph execution that will consume them. This is especially relevant in asynchronous scheduling or multi-threaded execution.
|
||||
|
||||
If ordering is not preserved, the parameter update for iteration *i* can be observed by the replay of iteration *i-1*, or the replay of iteration *i* can start before its own parameter update has completed. In practice, this means the attention operator may run with mismatched runtime metadata, which can cause incorrect results, precision issues, or even hangs. The code keeps a narrower path for the main full-graph eagle case, but the general design assumption is the same: replay must not overtake pending parameter update work.
|
||||
|
||||
## Full vs Piecewise on Ascend
|
||||
|
||||
Upstream docs already define full graph and piecewise graph semantically. On Ascend, the practical difference is driven by backend support and resource cost.
|
||||
|
||||
### Piecewise mode
|
||||
|
||||
Piecewise mode is the conservative path. It relies on the generic vLLM split execution strategy, then applies ACL graph capture to the non-attention segments selected by the compilation path. On Ascend, this mode is currently the more widely supported option, but it is also the most sensitive to stream pressure because the number of captured graphs scales with model depth.
|
||||
|
||||
### Full graph mode
|
||||
|
||||
Full graph mode is the more performance-oriented path when the attention backend can support runtime parameter patching through `update_graph_params()`. On Ascend, full graph support is tied to those attention-specific update hooks, workspace caching, and replay ordering guarantees.
|
||||
|
||||
## Diagnostics and Operational Notes
|
||||
|
||||
- The simplest way to confirm that graph mode is active is to enable cudagraph metrics and keep log stats enabled. In CLI usage, use `--cudagraph-metrics` and do not pass `--disable-log-stats`. In Python usage, set `cudagraph_metrics=True` and `disable_log_stats=False`. Then inspect the emitted metrics and logs.
|
||||
- Profiling can also confirm whether replay is happening, and developers can add temporary prints before replay when debugging locally, but those are secondary methods and are not expanded here.
|
||||
- Capture-size selection primarily follows upstream configuration and dispatch behavior; only the confirmed stream-resource capture failure is rewritten with user-facing guidance at runtime.
|
||||
- In debug mode, `ACLGraphWrapper` asserts that replay uses the same tensor addresses recorded during capture.
|
||||
- `ASCEND_LAUNCH_BLOCKING=1` is incompatible with ACL graph enablement in the current implementation.
|
||||
- For debugging inside graph execution, the repo also provides graph-aware print helpers in `vllm_ascend.utils`, but those are developer diagnostics rather than part of the execution design.
|
||||
|
||||
## Related Files
|
||||
|
||||
- `vllm_ascend/platform.py`, mode normalization, platform hooks, and static graph wrapper selection.
|
||||
- `vllm_ascend/compilation/acl_graph.py`, ACL graph wrapper, capture and replay cache, graph parameter containers, and full graph update dispatch.
|
||||
- `vllm_ascend/compilation/acl_graph.py`, runtime ACL graph capture, replay, and capture-failure guidance.
|
||||
- `vllm_ascend/attention/attention_v1.py`, full graph attention parameter capture and update logic.
|
||||
- `vllm_ascend/attention/mla_v1.py`, MLA (Multi-Head Latent Attention) specific full graph parameter capture and update logic.
|
||||
- `vllm_ascend/attention/context_parallel/attention_cp.py`, context parallel attention update path.
|
||||
- `vllm_ascend/attention/context_parallel/mla_cp.py`, context parallel MLA update path.
|
||||
@@ -0,0 +1,91 @@
|
||||
# KV Cache Pool
|
||||
|
||||
## Why KV Cache Pool?
|
||||
|
||||
Prefix caching is an important feature in LLM inference that can reduce prefill computation time drastically.
|
||||
|
||||
However, the performance gain from prefix caching is highly dependent on the cache hit rate, while the cache hit rate can be limited if one only uses on-chip memory for KV cache storage.
|
||||
|
||||
Hence, KV Cache Pool is proposed to utilize various types of storage including on-chip memory, DRAM, and SSD, making a pool for KV Cache storage while making the prefix of requests visible across all nodes, increasing the cache hit rate for all requests.
|
||||
|
||||
vLLM Ascend currently supports [MooncakeStore](https://github.com/kvcache-ai/Mooncake), one of the most recognized KV Cache storage engines.
|
||||
|
||||
While one can utilize MooncakeStore in vLLM V1 engine by setting it as a remote backend of LMCache with GPU (see [Tutorial](https://github.com/LMCache/LMCache/blob/dev/examples/kv_cache_reuse/remote_backends/mooncakestore/README.md)), we find it would be better to integrate a connector that directly supports MooncakeStore and can utilize the data transfer strategy that best fits Huawei NPU hardware.
|
||||
|
||||
Hence, we propose to integrate MooncakeStore with a brand new **MooncakeStoreConnectorV1**, which is indeed largely inspired by **LMCacheConnectorV1** (see the [How is MooncakeStoreConnectorV1 Implemented?](#how-is-mooncakestoreconnectorv1-implemented) section).
|
||||
|
||||
## Usage
|
||||
|
||||
vLLM Ascend currently supports MooncakeStore for KV Cache Pool. To enable MooncakeStore, one needs to configure `kv-transfer-config` and choose `MooncakeStoreConnector` as the KV Connector.
|
||||
|
||||
For step-by-step deployment and configuration, please refer to the [KV Pool User Guide](https://docs.vllm.ai/projects/ascend/en/latest/user_guide/feature_guide/kv_pool.html).
|
||||
|
||||
## How it works?
|
||||
|
||||
The KV Cache Pool integrates multiple memory tiers (on-chip memory, DRAM, SSD, etc.) through a connector-based architecture.
|
||||
|
||||
Each connector implements a unified interface for storing, retrieving, and transferring KV blocks between tiers, depending on access frequency and hardware bandwidth.
|
||||
|
||||
When combined with vLLM's Prefix Caching mechanism, the pool enables efficient caching both locally (in on-chip memory) and globally (via Mooncake), ensuring that frequently used prefixes remain hot while less frequently accessed KV data can spill over to lower-cost memory.
|
||||
|
||||
### 1. Combining KV Cache Pool with on-chip memory Prefix Caching
|
||||
|
||||
Prefix Caching with on-chip memory is already supported by the vLLM V1 Engine.
|
||||
By introducing KV Connector V1, users can seamlessly combine on-chip memory-based Prefix Caching with Mooncake-backed KV Pool.
|
||||
|
||||
The user can enable both features simply by enabling Prefix Caching, which is enabled by default in vLLM V1 unless the `--no-enable-prefix-caching` flag is set, and setting up the KV Connector for KV Pool (e.g., the MooncakeStoreConnector).
|
||||
|
||||
**Workflow**:
|
||||
|
||||
1. The engine first checks for prefix hits in the on-chip memory cache.
|
||||
|
||||
2. After getting the number of hit tokens on on-chip memory, it queries the KV Pool via the connector. If there are additional hits in the KV Pool, we get the **additional blocks only** from the KV Pool, and get the rest of the blocks directly from on-chip memory to minimize the data transfer latency.
|
||||
|
||||
3. After the KV Caches in the KV Pool are loaded into on-chip memory, the remaining process is the same as Prefix Caching in on-chip memory.
|
||||
|
||||
### 2. Combining KV Cache Pool with Mooncake PD Disaggregation
|
||||
|
||||
When used together with Mooncake PD (Prefill-Decode) Disaggregation, the KV Cache Pool can further decouple prefill and decode stages across devices or nodes.
|
||||
|
||||
Currently, we only perform put and get operations of KV Pool for **Prefill Nodes**, and Decode Nodes get their KV Cache from Mooncake P2P KV Connector, i.e., MooncakeConnector.
|
||||
|
||||
The key benefit of doing this is that we can keep the gain in performance by computing less with Prefix Caching from on-chip memory and KV Pool for Prefill Nodes, while not sacrificing the data transfer efficiency between Prefill and Decode nodes with P2P KV Connector that transfers KV Caches between NPU devices directly.
|
||||
|
||||
To enable this feature, we need to set up both Mooncake Connector and MooncakeStore Connector with a Multi Connector, which is a KV Connector class provided by vLLM that can call multiple KV Connectors in a specific order.
|
||||
|
||||
For details, please also refer to the [Mooncake connector deployment guide](https://github.com/vllm-project/vllm-ascend/blob/main/examples/disaggregated_prefill_v1/mooncake_connector_deployment_guide.md).
|
||||
|
||||
## How is MooncakeStoreConnectorV1 Implemented?
|
||||
|
||||
**MooncakeStoreConnectorV1** inherits the KV Connector V1 class in vLLM V1: through implementing the required methods defined in the KV connector V1 base class, one can integrate a third-party KV cache transfer/storage backend into the vLLM framework.
|
||||
|
||||
MooncakeStoreConnectorV1 is also largely inspired by LMCacheConnectorV1 in terms of the `Lookup Engine`/`Lookup Client` design for looking up KV cache keys, and the `ChunkedTokenDatabase` class for processing tokens into prefix-aware hashes as well as other hashing related designs. On top of this, we have also added our own design including `KVTransferThread` that allows async `get` and `put` of KV caches with multi-threading, and NPU-related data transfer optimization such as removing the `LocalBuffer` in LMCache to remove redundant data transfer.
|
||||
|
||||
The KV Connector methods that need to be implemented can be categorized into scheduler-side methods that are called in V1 scheduler and worker-side methods that are called in V1 worker, namely:
|
||||
|
||||
### KV Connector Scheduler-Side Methods
|
||||
|
||||
`get_num_new_matched_tokens`: Get prefix cache hit in number of tokens through looking up into the KV pool.
|
||||
`update_states_after_alloc`: Update KVConnector state after temporary buffer alloc.
|
||||
`build_connector_meta`: Attach the connector metadata to the request object.
|
||||
`request_finished`: Once a request is finished, determine whether request blocks should be freed now or will be sent asynchronously and freed later.
|
||||
|
||||
### Connector Worker-Side Methods
|
||||
|
||||
`register_kv_caches`: Register KV cache buffers needed for KV cache transfer.
|
||||
`start_load_kv`: Perform KV cache load operation that transfers KV cache from storage to device.
|
||||
`wait_for_layer_load`: Optional; Wait for layer load in layerwise + async KV load scenario.
|
||||
`save_kv_layer`: Optional; Do layerwise KV cache put into KV Pool.
|
||||
`wait_for_save`: Wait for KV Save to finish if async KV cache save/put.
|
||||
`get_finished`: Get request that finished KV transfer, `done_sending` if `put` finished, `done_receiving` if `get` finished.
|
||||
|
||||
## DFX
|
||||
|
||||
1. When looking up a key in KV Pool, if we cannot find the key, there is no Cache Hit for this specific block; we return no hit for this block and do not look up further blocks for the current request.
|
||||
2. Similarly, when we are trying to put a block into KV Pool and it fails, we do not put further blocks (subject to change).
|
||||
|
||||
## Limitations
|
||||
|
||||
1. Currently, MooncakeStore for vLLM Ascend only supports DRAM as the storage for KV Cache Pool.
|
||||
|
||||
2. For now, if we successfully looked up a key and found it exists, but failed to get it when calling KV Pool's get function, we just output a log indicating the get operation failed and keep going; hence, the accuracy of that specific request may be affected. We will handle this situation by falling back the request and re-compute everything assuming there's no prefix cache hit (or even better, revert only one block and keep using the Prefix Caches before that).
|
||||
@@ -0,0 +1,286 @@
|
||||
# Prepare inputs for model forwarding
|
||||
|
||||
## Purpose
|
||||
|
||||
Information required to perform model forward pass:
|
||||
|
||||
- the inputs
|
||||
- the corresponding attention metadata of the inputs
|
||||
|
||||
The following diagram shows what we should prepare for model inference.
|
||||
|
||||
```text
|
||||
+---------------+
|
||||
inputs --> | |
|
||||
| model | --> output
|
||||
attn_meta --> | |
|
||||
+---------------+
|
||||
```
|
||||
|
||||
Therefore, as long as we have these two pieces of information mentioned above, we can perform the model's forward propagation.
|
||||
|
||||
This document will explain **how we obtain the inputs and their corresponding attention metadata**.
|
||||
|
||||
## Overview
|
||||
|
||||
### 1. Obtain inputs
|
||||
|
||||
The workflow of obtaining inputs:
|
||||
|
||||
1. Get `token positions`: relative position of each token within its request sequence.
|
||||
|
||||
2. Get `token indices`: index of each scheduled token in the token table.
|
||||
|
||||
3. Get `Token IDs`: using token indices to retrieve the Token IDs from **token id table**.
|
||||
|
||||
At last, these `Token IDs` are required to be fed into a model, and `positions` should also be sent into the model to create `RoPE` (Rotary positional embedding). Both of them are the inputs of the model.
|
||||
|
||||
**Note**: The `Token IDs` are the inputs of a model, so we also call them `Input IDs`.
|
||||
|
||||
### 2. Build inputs attention metadata
|
||||
|
||||
A model requires these attention metadata during the forward pass:
|
||||
|
||||
- `query start location`: start and end location of each request corresponding to the scheduled tokens.
|
||||
- `sequence length`: length of each request including both computed tokens and newly scheduled tokens.
|
||||
- `number of computed tokens`: number of computed tokens for each request.
|
||||
- `number of requests`: number of requests in this batch.
|
||||
- `number of tokens`: total number of scheduled tokens in this batch.
|
||||
- **`block table`**: translates the logical address (within its sequence) of each block to its global physical address in the device's memory.
|
||||
- `max query len`: the longest scheduled tokens length in this request batch.
|
||||
- `slot mapping`: indices of each token that input token will be stored into.
|
||||
- `attention mask`: mask matrix applied to attention scores before softmax to control which tokens can attend to each other (usually a causal attention).
|
||||
|
||||
## Before start
|
||||
|
||||
There are mainly three types of variables.
|
||||
|
||||
- token level: represents one attribute corresponding to each scheduled token, so the length of this variable is the number of scheduled tokens.
|
||||
- request level: represents one attribute of each scheduled request, whose length usually is the number of scheduled requests. (`query start location` is a special case, which has one more element.)
|
||||
- system level:
|
||||
1. **Token IDs table**: stores the token IDs (i.e. the inputs of a model) of each request. The shape of this table is `(max num request, max model len)`. Here, `max num request` is the maximum count of concurrent requests allowed in a forward batch and `max model len` is the maximum token count that can be handled at one request sequence in this model.
|
||||
2. **Block table**: translates the logical address (within its sequence) of each block to its global physical address in the device's memory. The shape of this table is `(max num request, max model len / block size)`
|
||||
|
||||
**Note**: Both of these two tables come from the `_update_states` method before **preparing inputs**. You can take a look if you need more inspiration.
|
||||
|
||||
### Tips
|
||||
|
||||
Simply put, a `token ID` is an **integer** (usually `int32`), which represents a token.
|
||||
Example of `Token ID`:
|
||||
|
||||
```shell
|
||||
| Token ID | Token |
|
||||
|--------------|---------------|
|
||||
| 0 | [PAD] |
|
||||
| 1 | <|endoftext|> |
|
||||
| 2 | <|start|> |
|
||||
| 3 | [SEP] |
|
||||
| 4 | I |
|
||||
| 5 | the |
|
||||
| 6 | be |
|
||||
| 7 | of |
|
||||
| 8 | and |
|
||||
| ... | ... |
|
||||
| ... | ... |
|
||||
| vocab_size-1 | <|im_end|> |
|
||||
```
|
||||
|
||||
## Go through details
|
||||
|
||||
Assumptions:
|
||||
|
||||
- maximum number of tokens that can be scheduled at once: 10
|
||||
- `block size`: 2
|
||||
- Totally schedule 3 requests. Their prompt lengths are 3, 2, and 8 respectively.
|
||||
- `max model length`: 12 (the maximum token count that can be handled at one request sequence in a model).
|
||||
|
||||
These assumptions are configured at the beginning when starting vLLM. They are not fixed, so you can manually set them.
|
||||
|
||||
### Step 1: All requests in the prefill phase
|
||||
|
||||
#### Obtain inputs
|
||||
|
||||
As the maximum number of tokens that can be scheduled is 10, the scheduled tokens of each request can be represented as `{'0': 3, '1': 2, '2': 5}`. Note that `request_2` uses chunked prefill, leaving 3 prompt tokens unscheduled.
|
||||
|
||||
##### 1. Get token positions
|
||||
|
||||
First, determine which request each token belongs to: tokens 0–2 are assigned to **request_0**, tokens 3–4 to **request_1**, and tokens 5–9 to **request_2**. To represent this mapping, we use `request indices`, for example, `request indices`: `[0, 0, 0, 1, 1, 2, 2, 2, 2, 2]`.
|
||||
|
||||
For each request, use **the number of computed tokens** + **the relative position of current scheduled tokens** (`request_0: [0 + 0, 0 + 1, 0 + 2]`, `request_1: [0 + 0, 0 + 1]`, `request_2: [0 + 0, 0 + 1,..., 0 + 4]`) and then concatenate them together (`[0, 1, 2, 0, 1, 0, 1, 2, 3, 4]`).
|
||||
|
||||
Note: there is a more efficient way (using `request indices`) to create positions in actual code.
|
||||
|
||||
Finally, `token positions` can be obtained as `[0, 1, 2, 0, 1, 0, 1, 2, 3, 4]`. This variable is **token level**.
|
||||
|
||||
##### 2. Get token indices
|
||||
|
||||
The shape of the current **Token IDs table** is `(max num request, max model len)`.
|
||||
|
||||
Why are these `T_3_5`, `T_3_6`, `T_3_7` in this table without being scheduled?
|
||||
|
||||
- We fill all Token IDs in one request sequence to this table at once, but we only retrieve the tokens we scheduled this time. Then we retrieve the remaining Token IDs next time.
|
||||
|
||||
```shell
|
||||
| T_0_0 | T_0_1 | T_0_2 | ? | ? | ? | ? | ? | ? | ? | ? | ? |
|
||||
| T_1_0 | T_1_1 | ? | ? | ? | ? | ? | ? | ? | ? | ? | ? |
|
||||
| T_2_0 | T_2_1 | T_3_2 | T_3_3 | T_3_4 | T_3_5 | T_3_6 | T_3_7 | ? | ? | ? | ? |
|
||||
| ? | ? | ? | ? | ? | ? | ? | ? | ? | ? | ? | ? |
|
||||
......
|
||||
......
|
||||
......
|
||||
```
|
||||
|
||||
Note that `T_x_x` is an `int32`.
|
||||
|
||||
Let's say `M = max model len`. Then we can use `token positions` together with `request indices` of each token to construct `token indices`.
|
||||
|
||||
So `token indices` = `[0 + 0 * M, 1 + 0 * M, 2 + 0 * M, 0 + 1 * M, 1 + 1 * M, 0 + 2 * M, 1 + 2 * M, 2 + 2 * M, 3 + 2 * M, 4 + 2 * M]` = `[0, 1, 2, 12, 13, 24, 25, 26, 27, 28]`
|
||||
|
||||
##### 3. Retrieve the Token IDs
|
||||
|
||||
We use `token indices` to select out the corresponding `Input IDs` from the token table. The pseudocode is as follows:
|
||||
|
||||
```shell
|
||||
input_ids = token_table[token_indices]
|
||||
```
|
||||
|
||||
As mentioned before, we refer to these `Token IDs` as `Input IDs`.
|
||||
|
||||
- `Input IDs` = `[T_0_0, T_0_1, T_0_2, T_1_0, T_1_1, T_2_0, T_2_1, T_3_2, T_3_3, T_3_4]`
|
||||
|
||||
#### Build inputs attention metadata
|
||||
|
||||
In the current **Block Table**, we use the first block (i.e. block_0) to mark the unused block. The shape of the block is `(max num request, max model len / block size)`, where `max model len / block size = 12 / 2 = 6`.
|
||||
|
||||
```shell
|
||||
| 1 | 2 | 0 | 0 | 0 | 0 |
|
||||
| 3 | 0 | 0 | 0 | 0 | 0 |
|
||||
| 4 | 5 | 6 | 0 | 0 | 0 |
|
||||
| 0 | 0 | 0 | 0 | 0 | 0 |
|
||||
......
|
||||
......
|
||||
......
|
||||
```
|
||||
|
||||
The KV cache block in the device memory is like:
|
||||
|
||||
```shell
|
||||
| 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | ......
|
||||
```
|
||||
|
||||
Let's say `K = max model len / block size = 6`, and we can get token `device block number`.
|
||||
|
||||
The workflow of achieving slot mapping:
|
||||
|
||||
1. Get `block table indices` using `K`, `positions` and `request indices`.
|
||||
|
||||
Purpose: For each token, it could be used to select `device block number` from `block table`.
|
||||
|
||||
2. Get `device block number` using `block table indices`.
|
||||
|
||||
Purpose: `device block number` indicates which device block each token belongs to.
|
||||
|
||||
3. Get `block offsets` using `positions` and `block size`.
|
||||
|
||||
Purpose: `block offsets` indicates the offsets of each token within a block.
|
||||
|
||||
4. construct `slot mapping` using `device block number` and `block offsets`.
|
||||
|
||||
Purpose: we can use `slot mapping` to store Token IDs into token slots.
|
||||
|
||||
Details:
|
||||
|
||||
1. (**Token level**) Use a simple formula to calculate `block table indices`: `request indices * K + positions / block size`. So it equals `[0 * 6 + 0 / 2, 0 * 6 + 1 / 2, 0 * 6 + 2 / 2, 1 * 6 + 0 / 2, 1 * 6 + 1 / 2, 2 * 6 + 0 / 2, 2 * 6 + 1 / 2, 2 * 6 + 2 / 2, 2 * 6 + 3 / 2, 2 * 6 + 4 / 2] = [0, 0, 1, 6, 6, 12, 12, 13, 13, 14]`. This could be used to select `device block number` from `block table`.
|
||||
2. (**Token level**) Use `block table indices` to select out `device block number` for each scheduled token. The pseudocode is `block_numbers = block_table[block_table_indices]`. So `device block number=[1, 1, 2, 3, 3, 4, 4, 5, 5, 6]`
|
||||
3. (**Token level**) `block offsets` could be computed by `block offsets = positions % block size = [0, 1, 0, 0, 1, 0, 1, 0, 1, 0]`.
|
||||
4. Finally, use `block offsets` and `device block number` to create `slot mapping`: `device block number * block size + block_offsets = [2, 3, 4, 6, 7, 8, 9, 10, 11, 12]`
|
||||
|
||||
(**Request level**) As we know the scheduled token count is `[3, 2, 5]`:
|
||||
|
||||
- (**Request level**) Use prefix sum to calculate `query start location`: `[0, 3, 5, 10]`.
|
||||
- (**Request level**) All tokens in step 1 are in the prefill stage, and the computed tokens count is 0; then `sequence length` = `[3, 2, 5]`.
|
||||
- (**Request level**) As mentioned above, `number of computed tokens` are all 0s: `[0, 0, 0]`.
|
||||
- `number of requests`: `3`
|
||||
- (**Request level**) `number of tokens`: `[3, 2, 5]`
|
||||
- `max query len`: `5`
|
||||
- (**Token level**) `slot mapping`: `[2, 3, 4, 6, 7, 8, 9, 10, 11, 12]`
|
||||
- `attention mask`: For all requests that initiate a prefill process, we simply create only one mask matrix for reuse across different requests. The shape of this mask matrix is `5 * 5`:
|
||||
|
||||
### Step 2: Chunked prefill
|
||||
|
||||
In Step 2, we no longer provide explanations or perform calculations; instead, we directly present the final result.
|
||||
|
||||
#### Obtain inputs
|
||||
|
||||
Scheduled token of each request: `{'0': 1, '1': 1, '2': 3}`
|
||||
|
||||
1. `request indices`: `[0, 1, 2, 2, 2]`
|
||||
2. `token positions`: `[3, 2, 5, 6, 7]`
|
||||
|
||||
Current **Token IDs table**:
|
||||
|
||||
```shell
|
||||
| T_0_0 | T_0_1 | T_0_2 | T_0_3 | ? | ? | ? | ? | ? | ? | ? | ? |
|
||||
| T_1_0 | T_1_1 | T_1_2 | ? | ? | ? | ? | ? | ? | ? | ? | ? |
|
||||
| T_2_0 | T_2_1 | T_3_2 | T_3_3 | T_3_4 | T_3_5 | T_3_6 | T_3_7 | ? | ? | ? | ? |
|
||||
| ? | ? | ? | ? | ? | ? | ? | ? | ? | ? | ? | ? |
|
||||
......
|
||||
......
|
||||
......
|
||||
```
|
||||
|
||||
**Note**: **T_0_3**, **T_1_2** are new Token IDs of **request_0** and **request_1** respectively. They are sampled from the output of the model.
|
||||
|
||||
3. `token indices`: `[3, 14, 29, 30, 31]`
|
||||
4. `Input IDs`: `[T_0_3, T_1_2, T_3_5, T_3_6, T_3_7]`
|
||||
|
||||
#### Build inputs attention metadata
|
||||
|
||||
We allocate the blocks `7` and `8` to `request_1` and `request_2` respectively, as they need more space in device to store KV cache following token generation or chunked prefill.
|
||||
|
||||
Current **Block Table**:
|
||||
|
||||
```shell
|
||||
| 1 | 2 | 0 | 0 | 0 | 0 |
|
||||
| 3 | 7 | 0 | 0 | 0 | 0 |
|
||||
| 4 | 5 | 6 | 8 | 0 | 0 |
|
||||
| 0 | 0 | 0 | 0 | 0 | 0 |
|
||||
......
|
||||
......
|
||||
......
|
||||
```
|
||||
|
||||
KV cache block in the device memory:
|
||||
|
||||
```shell
|
||||
| 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | ......
|
||||
```
|
||||
|
||||
1. (**Token level**) `block table indices`: `[1, 7, 14, 15, 15]`
|
||||
2. (**Token level**) `device block number`: `[2, 7, 6, 8, 8]`
|
||||
3. (**Token level**) `block offsets`: `[1, 0, 1, 0, 1]`
|
||||
4. (**Token level**) `slot mapping`: `[5, 14, 13, 16, 17]`
|
||||
|
||||
Scheduled token count: `[1, 1, 3]`
|
||||
|
||||
- `query start location`: `[0, 1, 2, 5]`
|
||||
|
||||
- `sequence length`: `[4, 3, 8]`
|
||||
|
||||
- `number of computed tokens`: `[3, 2, 5]`
|
||||
|
||||
- `number of requests`: `3`
|
||||
|
||||
- `max query len`: `3`
|
||||
|
||||
- `slot mapping`: `[5, 14, 13, 16, 17]`
|
||||
|
||||
- `attention mask`: `5 * 8`
|
||||
|
||||
Each token has a `1 * 8` vector, and there are 5 scheduled tokens.
|
||||
|
||||
## At last
|
||||
|
||||
If you understand step 1 and step 2, you will know all the following steps.
|
||||
|
||||
Hope this document helps you better understand how vLLM prepares inputs for model forwarding. If you have any good ideas, you are welcome to contribute to us.
|
||||
@@ -0,0 +1,25 @@
|
||||
# Adding a custom aclnn operation
|
||||
|
||||
This document describes how to add a custom aclnn operation to vllm-ascend.
|
||||
|
||||
## How custom aclnn operation works in vllm-ascend?
|
||||
|
||||
Custom aclnn operations are built and installed into `vllm_ascend/cann_ops_custom` directory during the build process of vllm-ascend. Then the aclnn operators are bound to `torch.ops._C_ascend` module, enabling users to invoke them in vllm-ascend python code.
|
||||
|
||||
To enable custom operations, use the following code:
|
||||
|
||||
```python
|
||||
from vllm_ascend.utils import enable_custom_op
|
||||
|
||||
enable_custom_op()
|
||||
```
|
||||
|
||||
## How to add a custom aclnn operation?
|
||||
|
||||
1. Create a new operation folder under `csrc` directory.
|
||||
2. Create `op_host` and `op_kernel` directories for host and kernel source code.
|
||||
3. Add build options in `csrc/build_aclnn.sh` for supported SOC. Note that multiple ops should be separated with `;`, i.e. `CUSTOM_OPS="op1;op2;op3"`.
|
||||
4. Bind aclnn operators to torch.ops._C_ascend module in `csrc/torch_binding.cpp`.
|
||||
5. Write a meta implementation in `csrc/torch_binding_meta.cpp` for the op to be captured into the aclgraph.
|
||||
|
||||
After a successful build of vllm-ascend, the custom aclnn operation can be invoked in python code.
|
||||
197
docs/source/developer_guide/Design_Documents/context_parallel.md
Normal file
197
docs/source/developer_guide/Design_Documents/context_parallel.md
Normal file
@@ -0,0 +1,197 @@
|
||||
# Context Parallel (CP)
|
||||
|
||||
TL;DR: PCP accelerates prefill via sequence splitting. DCP eliminates KV cache redundancy.
|
||||
|
||||

|
||||
|
||||
For the main discussions during the development process, please refer to the [RFC](https://github.com/vllm-project/vllm/issues/25749) and the relevant links referenced by or referencing this RFC.
|
||||
|
||||
## What is CP?
|
||||
|
||||
**Context Parallel (CP)** is a strategy for parallelizing computation along the sequence dimension across multiple devices.
|
||||
|
||||
**Prefill Context Parallel (PCP)** expands the world size of devices and uses dedicated communication domains.
|
||||
Its primary goal is to partition the sequence dimension during the prefill phase, enabling different devices to compute distinct chunks of the sequence simultaneously.
|
||||
The KV cache is sharded along the sequence dimension across devices.
|
||||
This approach impacts the computational logic of both the Prefill and Decode stages to varying degrees.
|
||||
|
||||
**Decode Context Parallel (DCP)** reuses the communication domain of Tensor Parallelism (TP) and does not require additional devices.
|
||||
Its main objective is to eliminate duplicated storage of the KV cache by sharding it along the sequence dimension across devices within the TP domain that would otherwise hold redundant copies.
|
||||
DCP primarily influences the Decode logic, as well as the logic for chunked prefill and cached prefill.
|
||||
|
||||
## How to Use CP?
|
||||
|
||||
Please refer to the [context parallel user guide](../../user_guide/feature_guide/context_parallel.md) for detailed information.
|
||||
|
||||
## How It Works?
|
||||
|
||||
### Device Distribution
|
||||
|
||||
We introduce new communication domains for PCP and reuse TP for DCP, and this is the new layout of devices for PCP2, DCP2, and TP4.
|
||||

|
||||
|
||||
### Block Table
|
||||
|
||||
CP performs sequence sharding on the KV cache storage. To facilitate efficient storage and access, tokens are stored in an interleaved manner across devices, with the interleaving granularity determined by `cp_kv_cache_interleave_size`, whose default value is `cp_kv_cache_interleave_size=1`, a.k.a. 'token interleave'.
|
||||
|
||||
Given that PCP and DCP behave similarly for KV cache sharding, we refer to them collectively as CP. Specifically, `cp_size = pcp_size * dcp_size`, and `cp_rank = pcp_rank * dcp_size + dcp_rank`.
|
||||
|
||||
As illustrated, a virtual block is defined in the block table, where blocks within the same CP device group form a virtual block. The virtual block size is `virtual_block_size = block_size * cp_size`.
|
||||
|
||||
For any token `x`, referencing the following figure, its (virtual) block index is `x // virtual_block_size`, and the offset within the virtual block is `offset_within_virtual_block = x % virtual_block_size`.
|
||||
The local block index is `local_block_index = offset_within_virtual_block // cp_kv_cache_interleave_size`, and the device number is `target_rank = local_block_index % cp_size`.
|
||||
The offset within the local block is `(local_block_index // cp_size) * cp_kv_cache_interleave_size + offset_within_virtual_block % cp_kv_cache_interleave_size`.
|
||||
|
||||

|
||||
|
||||
Based on the logic above, the `slot_mapping` calculation process is adjusted, and the `slot_mapping` values on each device are modified to ensure the KV cache is sharded along the sequence dimension and stored across different devices as expected.
|
||||
|
||||
The current implementation requires that `block_size % cp_kv_cache_interleave_size == 0`.
|
||||
|
||||
### Decode Context Parallel (DCP)
|
||||
|
||||
As mentioned above, the primary function of DCP is to shard the KV cache along the sequence dimension for storage. Its impact lies in the logic of the decode and chunked prefill phases.
|
||||
|
||||
**Prefill Phase:**
|
||||
As illustrated, during the Chunked Prefill computation, two distinct logic implementations are employed for MLA and GQA backends.
|
||||
|
||||
- In the **MLA backend**, a Context KV Cache `all_gather` operation is performed to aggregate the full KV values.
|
||||
These are then used for attention computation with the Q values of the current chunk.
|
||||
Note that in multi-request scenarios, the directly gathered KV results are interleaved across requests.
|
||||
The `reorg_kvcache` function is used to reorganize the KV cache, ensuring that the KV cache of the same request is stored contiguously.
|
||||
|
||||
- In the **GQA backend**, an `all_gather` is performed along the head dimension for Q.
|
||||
This is because DCP overlaps with the TP communication domain, and the Q heads within a DCP group differ.
|
||||
However, they need to exchange results with the locally computed KV cache for online Softmax updates.
|
||||
To ensure correctness during result updates, the Q values are synchronized across the DCP group via head-dimension `all_gather`.
|
||||
During the result update process, `cp_lse_ag_out_rs` is invoked to aggregate `attn_output` and `attn_lse`, update the results, and perform a reduce-scatter operation on the outputs.
|
||||
Alternatively, we can use an all-to-all communication to exchange the output and LSE results, followed by direct local updates. This approach aligns with the logic adapted for PCP compatibility.
|
||||
|
||||

|
||||
|
||||
**Decode Phase:**
|
||||
The logic during the decode phase is consistent with that of GQA's chunked prefill: an all-gather operation is first performed along the Q head dimension to ensure consistency within the DCP group.
|
||||
After computing the results with the local KV cache, the results are updated via the `cp_lse_ag_out_rs` function.
|
||||
|
||||

|
||||
|
||||
### GLM-5.2 SFA DCP Replicated Indexer
|
||||
|
||||
GLM-5.2 uses Sparse Flash Attention (SFA) with a LightningIndexer. For DCP,
|
||||
the indexer needs a full-sequence view to select the same sparse top-k blocks
|
||||
as non-DCP SFA, while the much larger SFA KV cache should remain sharded to
|
||||
retain DCP's memory benefit. The replicated-indexer path provides this split
|
||||
layout:
|
||||
|
||||
- The LightningIndexer cache is replicated on every DCP rank. Index selection
|
||||
therefore uses the complete sequence and produces globally consistent sparse
|
||||
top-k indices.
|
||||
- The SFA KV cache remains DCP-local. The global indices from the replicated
|
||||
indexer view are remapped to local KV indices before SFA runs.
|
||||
- During prefill or a mixed batch, only the KV blocks referenced by the sparse
|
||||
block table are compacted and all-gathered after the current layer has
|
||||
written its KV cache. The gathered KV uses a remapped block table for SFA,
|
||||
so this path does not all-gather Q and does not need LSE or output
|
||||
post-processing.
|
||||
- Decode-only batches retain the DCP SFA Q-gather and result-merge path.
|
||||
|
||||
This mode is selected automatically for SFA sparse models when
|
||||
`prefill_context_parallel_size=1` and `decode_context_parallel_size>1`. It
|
||||
requires `decode_context_parallel_size == tensor_parallel_size`; PCP combined
|
||||
with this replicated-indexer path is not supported.
|
||||
|
||||
For a GLM-5.2 DSA-CP deployment, enable FlashComm1 and DSA CP and keep the CP
|
||||
interleave size equal to the KV-cache block size:
|
||||
|
||||
```bash
|
||||
export VLLM_ASCEND_ENABLE_FLASHCOMM1=1
|
||||
|
||||
vllm serve <glm-5.2-model> \
|
||||
--tensor-parallel-size <N> \
|
||||
--prefill-context-parallel-size 1 \
|
||||
--decode-context-parallel-size <N> \
|
||||
--block-size <B> \
|
||||
--cp-kv-cache-interleave-size <B> \
|
||||
--additional-config '{"enable_dsa_cp": true}'
|
||||
```
|
||||
|
||||
The replicated indexer increases indexer-cache memory in proportion to the
|
||||
DCP world size; the SFA KV cache itself remains sharded. For this SFA CP path,
|
||||
`cp_kv_cache_interleave_size` must equal `block_size`. A mismatched setting is
|
||||
overridden during configuration validation, but deployments should set both
|
||||
values explicitly to avoid relying on that fallback.
|
||||
|
||||
### Prefill Context Parallel (PCP)
|
||||
|
||||
**Tokens Partition in Head-Tail Style**
|
||||
|
||||
PCP requires splitting the input sequence and ensuring balanced computational load across devices during the prefill phase.
|
||||
We employ a head-tail style for splitting and concatenation: specifically, the sequence is first padded to a length of `2*pcp_size`, then divided into `2*pcp_size` equal parts.
|
||||
The first part is merged with the last part, the second part with the second last part, and so on, thereby assigning computationally balanced chunks to each device.
|
||||
Additionally, since allgather aggregation of KV or Q results in interleaved chunks from different requests, we compute `pcp_allgather_restore_idx` to quickly restore the original order.
|
||||
|
||||
These logics are implemented in the function `_update_tokens_for_pcp`.
|
||||
|
||||

|
||||
|
||||
**Prefill Phase:**
|
||||
|
||||
During the Prefill phase (excluding chunked prefill), we employ an all-gather KV approach to address the issue of incomplete sequences on individual GPUs.
|
||||
It is important to note that we only aggregate the KV values for the current layer at a time, and these are discarded immediately after use, avoiding excessive peak memory usage.
|
||||
This method can also be directly applied to KV cache storage (since the KV cache partitioning method differs from PCP sequence partitioning, it is inevitable that each GPU requires a complete copy of the KV values).
|
||||
All attention backends maintain consistency in this logic.
|
||||
|
||||
Note: While a Ring Attention approach could also facilitate information exchange with lower peak memory and enable computation-communication overlap, we prioritized the all-gather KV implementation after evaluating that the development complexity was high and the benefits of overlap were limited.
|
||||
|
||||

|
||||
|
||||
**Decode Phase:**
|
||||
|
||||
During the decode phase, we only need to add an allgather within the PCP group after the DCP all-to-all communication exchanges the output and LSE, before proceeding with the output update.
|
||||
|
||||

|
||||
|
||||
**Chunked Prefill:**
|
||||
|
||||
Currently, there are three viable approaches for Chunked Prefill compatibility: **AllGatherQ**, **AllGatherKV**, and **Ring-Attn**.
|
||||
Since PCP performs sequence sharding on both the query sequence and the KV cache, we need to ensure that one side has complete information or employ a method like Ring-Attn to perform computations sequentially.
|
||||
The advantages and disadvantages of Ring-Attn will not be elaborated here.
|
||||
|
||||
We have implemented the **AllGatherQ** approach in the GQA attention backend and the **AllGatherKV** approach in the MLA attention backend.
|
||||
The workflow after **AllGatherQ** is identical to the decode phase, while the workflow after **AllGatherKV** is the same as the standard prefill phase.
|
||||
For details, please refer to the diagram below; specific steps will not be repeated.
|
||||
|
||||
One important note: **AllGatherKV** may lead to significant peak memory usage when the context length becomes excessively long.
|
||||
To mitigate this, we adopt a segmented processing strategy.
|
||||
By predefining the maximum amount of KV cache processed per round, we sequentially complete the attention computation and online softmax updates for each segment.
|
||||
|
||||

|
||||
|
||||
### SFA DSA-CP Mixed `o_proj` Path
|
||||
|
||||
SFA DSA-CP mixed execution intentionally reuses the normal TP-sharded `o_proj`.
|
||||
This is part of the DSA-CP mixed data path, not a standalone user-facing `o_proj` TP switch.
|
||||
The mixed path is used when one instance may handle both decode-only and prefill/mixed batches, so `o_proj` must support two layouts at runtime:
|
||||
|
||||
- **Decode-only batches** keep the decode TP path.
|
||||
SFA outputs are exchanged with an all-to-all in the TP group, then the original TP-sharded `o_proj` runs normally.
|
||||
- **Prefill or mixed batches** produce SFA outputs that are not directly compatible with the TP-sharded `o_proj` input layout.
|
||||
Before `o_proj` forward, each rank all-gathers the TP-sharded `o_proj` weight and all input-sharded quantization parameters into temporary full-weight buffers.
|
||||
The full-weight `o_proj` forward runs once for that batch, and the module is then restored to the TP parameter aliases.
|
||||
|
||||
The storage invariant is that the original TP-sharded `o_proj` parameter remains the only persistent source of truth.
|
||||
`o_proj_tp_*` tensors are aliases of the original parameter storage.
|
||||
`o_proj_full_*` tensors are reusable communication buffers for prefill/mixed full-gather execution only.
|
||||
They must not become a second persistent copy of the TP weight.
|
||||
|
||||
This coupling preserves the existing decode TP behavior, supports prefill/mixed DSA-CP batches, and avoids adding an extra configuration path whose state can drift from DSA-CP mixed execution.
|
||||
|
||||
### Related Files
|
||||
|
||||
- slot_mapping computation: `vllm_ascend/worker/block_table.py`
|
||||
- sequences splitting and metadata prepare: `vllm_ascend/worker/model_runner_v1.py`
|
||||
- PCP token splitting and metadata generation: `vllm_ascend/worker/pcp_utils.py`
|
||||
- GQA backend: `vllm_ascend/attention/context_parallel/attention_cp.py`
|
||||
- MLA backend: `vllm_ascend/attention/context_parallel/mla_cp.py`
|
||||
- DSA backend: `vllm_ascend/attention/context_parallel/dsa_cp.py`
|
||||
- SFA backend: `vllm_ascend/attention/context_parallel/sfa_cp.py`
|
||||
281
docs/source/developer_guide/Design_Documents/cpu_binding.md
Normal file
281
docs/source/developer_guide/Design_Documents/cpu_binding.md
Normal file
@@ -0,0 +1,281 @@
|
||||
# CPU Binding
|
||||
|
||||
## Overview
|
||||
|
||||
CPU binding is an **Ascend-native host-side optimization** for vLLM workers on
|
||||
ARM servers. **Starting from vllm-ascend v0.18.0rc1, it is enabled by default
|
||||
through `enable_cpu_binding=True`.**
|
||||
|
||||
The feature does not change model execution logic or numerical results. It only
|
||||
controls CPU placement for the worker process, key runtime threads, memory
|
||||
pages, and NPU IRQs when the host environment allows it. By keeping the main
|
||||
worker, ACL, and release threads on dedicated CPU ranges, it **helps reduce
|
||||
context-switch overhead from scheduler preemption on busy hosts.**
|
||||
|
||||
## Why CPU Binding?
|
||||
|
||||
On multi-socket ARM systems, the Linux scheduler may place worker threads on
|
||||
CPUs far from the NPU that the worker drives. This can increase cross-NUMA
|
||||
traffic, increase thread preemption, and introduce latency jitter. The Ascend
|
||||
backend therefore owns a CPU allocation policy to **reduce cross-NUMA traffic,
|
||||
reduce thread preemption, and improve latency stability** instead of relying on
|
||||
upstream GPU NUMA binding flags.
|
||||
|
||||
This is also why upstream NUMA flags are adapted on Ascend:
|
||||
|
||||
- `--numa-bind` is converted to `additional_config={"enable_cpu_binding": true}`.
|
||||
- `--numa-bind-nodes` and `--numa-bind-cpus` are ignored because Ascend computes CPU pools from NPU topology or global logical NPU IDs.
|
||||
|
||||
## How It Works?
|
||||
|
||||
The allocator derives its plan from runtime host state:
|
||||
|
||||
| Input | Source | Purpose |
|
||||
| --- | --- | --- |
|
||||
| Allowed CPUs | `/proc/self/status` `Cpus_allowed_list` | The only CPUs eligible for binding. Container cpusets are respected. |
|
||||
| Logical NPU map | `npu-smi info -m` | Maps card/chip IDs to global logical NPU IDs and gives `total_logic_npus`. On Ascend 950, `Chip Logic ID` is not reported, so `NPU ID` is used as the logical ID. |
|
||||
| Running NPUs | `npu-smi info` process table, filtered by `ASCEND_RT_VISIBLE_DEVICES` | Identifies the logical NPUs used by this worker process. A2/A3 process rows use `NPU Chip`; Ascend 950 process rows use `NPU ID`. |
|
||||
| Topology affinity | `npu-smi info -t topo` | Provides NPU-to-CPU affinity for `topo_affinity` mode. |
|
||||
| CPU NUMA map | `lscpu -e=CPU,NODE` | Used to extend single-NUMA affinity pools to the next NUMA node. |
|
||||
| Thread topology | `lscpu` `Thread(s) per core` | Determines Ascend 950 cluster size: 8 CPUs for 1 thread per core, 16 CPUs for 2 threads per core. |
|
||||
| UVB polling threads | `ps -Te` | Finds host `uvb_poll_window_thread` threads for Ascend 950 UVB CPU binding. Docker containers must use `--pid=host` to see these host threads. |
|
||||
|
||||
### Strategy Selection
|
||||
|
||||
The binding strategy is selected by Ascend device type:
|
||||
|
||||
| Device type | Strategy | Reason |
|
||||
| --- | --- | --- |
|
||||
| A3 | `global_slice` | A3 uses HCCS card-to-card interconnect. Each NPU is nearly equidistant from all NUMA nodes, so there is no strong NPU-to-NUMA affinity signal. Global logical NPU ID based slicing gives deterministic, non-overlapping CPU pools and CPU/NUMA isolation between workers. |
|
||||
| Ascend 950 | `topo_affinity` | Ascend 950 uses NPU-to-CPU affinity from `npu-smi info -t topo` to choose an affinity NUMA node, then assigns one CPU cluster from that NUMA node to each worker. It also reports process rows by `NPU ID` instead of `NPU Chip`, skips IRQ binding, and binds host UVB polling threads. |
|
||||
| A2 and Atlas 300 inference products | `topo_affinity` | A2 and Atlas 300 inference products provide NPU-to-CPU affinity information through `npu-smi info -t topo`, so they use this topology signal when available. |
|
||||
|
||||
If `topo_affinity` is selected but topo affinity is unavailable, the allocator falls back to `global_slice`.
|
||||
|
||||
### CPU Pool Construction
|
||||
|
||||
#### global_slice
|
||||
|
||||
`global_slice` is designed for devices without a useful NPU-to-CPU affinity
|
||||
signal, including A3. Because A3's **HCCS interconnect makes the distance
|
||||
from each NPU to each NUMA node nearly the same**, topology affinity is not a
|
||||
useful placement signal. The allocator therefore partitions the sorted
|
||||
`allowed_cpus` list by global logical NPU ID.
|
||||
|
||||
1. Determine `total_npus` in this order:
|
||||
- `total_logic_npus` from `npu-smi info -m`
|
||||
- number of topo affinity entries
|
||||
- number of running NPUs
|
||||
2. Compute:
|
||||
- `base = len(allowed_cpus) // total_npus`
|
||||
- `extra = len(allowed_cpus) % total_npus`
|
||||
3. Each logical NPU gets a deterministic slice:
|
||||
- NPU IDs `< extra` receive `base + 1` CPUs.
|
||||
- Remaining NPU IDs receive `base` CPUs.
|
||||
4. Only running NPUs are materialized into `npu_cpu_pool`.
|
||||
|
||||
This is the key property: two independent worker processes with the same cpuset
|
||||
but different visible NPU IDs still get **non-overlapping CPU pools** because
|
||||
both processes slice against the same global NPU ID space. With a NUMA-aligned
|
||||
cpuset, this also provides **CPU/NUMA isolation between workers**, so one worker
|
||||
does not share the same CPU or NUMA slice with another worker.
|
||||
|
||||
`global_slice` requires enough CPUs for the selected device's role split:
|
||||
|
||||
- Devices with IRQ binding require `base >= 5`:
|
||||
2 CPUs for SQ/CQ IRQ binding, at least 1 CPU for the main worker, 1 CPU for
|
||||
ACL thread, and 1 CPU for release thread.
|
||||
|
||||
#### topo_affinity
|
||||
|
||||
`topo_affinity` is designed for A2, Atlas 300 inference products, Ascend 950,
|
||||
and other non-A3 device types. A2 and Atlas 300 inference products expose
|
||||
**meaningful NPU-to-CPU affinity information**, so the allocator starts from NPU
|
||||
topology affinity when it is available and then avoids overlap for shared
|
||||
affinity groups.
|
||||
|
||||
1. Build candidate NPUs from all logical NPUs:
|
||||
- always include running NPUs
|
||||
- include non-running NPUs only when their affinity overlaps this process's allowed cpuset
|
||||
2. For each candidate NPU, intersect topo affinity with `allowed_cpus`.
|
||||
3. If the intersection is empty for a candidate, binding fails for this rank.
|
||||
4. If the affinity CPUs are all on one NUMA node, extend the pool with CPUs from the next NUMA node, constrained by `allowed_cpus`.
|
||||
5. Group NPUs with identical extended pools and split each shared pool evenly across that group.
|
||||
6. Keep only running NPUs in the final `npu_cpu_pool`.
|
||||
|
||||
The non-running candidate step is intentional. It prevents two independent
|
||||
single-card workers from selecting the same CPU range when their visible NPUs
|
||||
share the same topology affinity.
|
||||
|
||||
For Ascend 950, topology affinity is used differently:
|
||||
|
||||
1. Bind all visible host `uvb_poll_window_thread` threads to NUMA0 CPUs except CPU0, constrained by `allowed_cpus`. Docker containers must use `--pid=host` to make these host threads visible.
|
||||
2. Use topo affinity to identify each NPU's single affinity NUMA node.
|
||||
3. Parse `Thread(s) per core` from `lscpu` and set cluster size to 8 CPUs when it is 1, or 16 CPUs when it is 2.
|
||||
4. Split each affinity NUMA's sorted allowed CPU list into contiguous clusters.
|
||||
5. Assign clusters by sorted logical NPU ID, including hidden NPUs that share the same affinity NUMA.
|
||||
6. Keep only running NPUs in the final `npu_cpu_pool`.
|
||||
|
||||
If Ascend 950 topo affinity is missing, spans multiple NUMA nodes, has too few
|
||||
clusters, or reports an unsupported `Thread(s) per core`, worker CPU binding is
|
||||
skipped without raising to the worker process.
|
||||
|
||||
### Role Split
|
||||
|
||||
After a CPU pool is built, the allocator splits it by role:
|
||||
|
||||
For devices with IRQ binding:
|
||||
|
||||
| Role | CPUs |
|
||||
| --- | --- |
|
||||
| SQ/CQ IRQ | `pool[0]`, `pool[1]` |
|
||||
| Main worker process and subthreads | `pool[2:-2]` |
|
||||
| ACL thread | `pool[-2]` |
|
||||
| Release thread | `pool[-1]` |
|
||||
|
||||
For Ascend 950:
|
||||
|
||||
| Role | CPUs |
|
||||
| --- | --- |
|
||||
| Main worker process and subthreads | the whole assigned cluster |
|
||||
| ACL thread | not separately pinned |
|
||||
| Release thread | not separately pinned |
|
||||
|
||||
If a final pool has fewer CPUs than the selected role split requires, binding
|
||||
fails for this rank and the worker logs a warning from the caller. The minimum
|
||||
is 5 CPUs per NPU for devices with IRQ binding. Ascend 950 requires one full
|
||||
cluster per worker.
|
||||
|
||||
## Conditional Host Tuning
|
||||
|
||||
After CPU affinity is applied, CPU binding can also apply two host-side tuning
|
||||
steps when the environment supports them:
|
||||
|
||||
- Memory migration uses `migratepages` to move the worker process's existing
|
||||
pages to the selected NUMA node. This keeps the worker closer to the memory it
|
||||
reads and reduces remote-NUMA memory read latency.
|
||||
- IRQ binding places NPU IRQ handling on the CPUs reserved for the corresponding
|
||||
NPU when `/proc/irq` is writable and IRQ files can be resolved.
|
||||
Ascend 950 skips this step.
|
||||
|
||||
These are conditional parts of CPU binding, not separate feature switches. If a
|
||||
host prerequisite is missing, that step is skipped while CPU thread binding
|
||||
still proceeds. Missing `migratepages` can still leave pages on remote NUMA
|
||||
nodes, so **latency or throughput may regress compared with a full CPU binding
|
||||
setup.**
|
||||
|
||||
## Examples
|
||||
|
||||
### A3 inference server with 640 CPUs and 16 NPUs
|
||||
|
||||
Inputs:
|
||||
|
||||
- `allowed_cpus = [0..639]`
|
||||
- `total_logic_npus = 16`
|
||||
- `running_npu_list = [0..15]`
|
||||
|
||||
Computation:
|
||||
|
||||
- `base = 640 // 16 = 40`
|
||||
- `extra = 0`
|
||||
- Worker `i` driving logical NPU `i` receives CPU slice
|
||||
`[i * 40 .. i * 40 + 39]`.
|
||||
|
||||
Global slice view:
|
||||
|
||||
```text
|
||||
CPU range: 0 639
|
||||
|-- worker0/NPU0 --|-- worker1/NPU1 --| ... |-- worker15/NPU15 --|
|
||||
| 0-39 | 40-79 | ... | 600-639 |
|
||||
```
|
||||
|
||||
Role split inside each worker slice:
|
||||
|
||||
```text
|
||||
40-CPU worker slice
|
||||
| IRQ CPUs | main worker process and subthreads | ACL thread | release thread |
|
||||
| c0-c1 | c2-c37 | c38 | c39 |
|
||||
```
|
||||
|
||||
Concrete examples:
|
||||
|
||||
| Worker | Logical NPU | CPU pool | IRQ CPUs | Main CPUs | ACL CPU | Release CPU |
|
||||
| --- | --- | --- | --- | --- | --- | --- |
|
||||
| 0 | 0 | 0-39 | 0-1 | 2-37 | 38 | 39 |
|
||||
| 1 | 1 | 40-79 | 40-41 | 42-77 | 78 | 79 |
|
||||
| ... | ... | ... | ... | ... | ... | ... |
|
||||
| 15 | 15 | 600-639 | 600-601 | 602-637 | 638 | 639 |
|
||||
|
||||
This layout remains deterministic even when different worker processes share
|
||||
the same cpuset, because slicing is based on the global logical NPU ID.
|
||||
|
||||
### A2 topo_affinity with hidden same-affinity NPUs
|
||||
|
||||
Inputs from an A2 topology:
|
||||
|
||||
- NPU0 affinity: 144-167
|
||||
- NPU2 affinity: 144-167
|
||||
- Process A sees only NPU0
|
||||
- Process B sees only NPU2
|
||||
- Both processes have `allowed_cpus = [144..191]`
|
||||
|
||||
The allocator includes the hidden same-affinity NPU as a candidate in each
|
||||
process, splits the shared extended pool, and then keeps only the visible NPU in
|
||||
the final pool.
|
||||
|
||||
Final pools:
|
||||
|
||||
| Process | Visible NPU | Final CPU pool |
|
||||
| --- | --- | --- |
|
||||
| A | 0 | 144-167 |
|
||||
| B | 2 | 168-191 |
|
||||
|
||||
This avoids overlapping CPU pools even when the two workers are launched as independent single-card services.
|
||||
|
||||
## Logs
|
||||
|
||||
The allocator logs the selected mode and allocation plan:
|
||||
|
||||
```text
|
||||
[cpu_bind_mode] mode=topo_affinity rank=0 visible_npus=[0]
|
||||
The CPU allocation plan is as follows:
|
||||
NPU0: main=[...] acl=[...] release=[...]
|
||||
```
|
||||
|
||||
Ascend 950 uses a different role split, so its plan log does not include ACL or
|
||||
release fields. UVB polling thread binding is reported separately when matching
|
||||
threads are found:
|
||||
|
||||
```text
|
||||
[cpu_bind_mode] mode=topo_affinity rank=0 visible_npus=[0]
|
||||
The CPU allocation plan is as follows:
|
||||
Ascend 950 NPU0: worker=[...]
|
||||
[cpu_bind_ascend_950] uvb_poll_window_thread tids=[...] cpus=[...]
|
||||
```
|
||||
|
||||
## Limitations
|
||||
|
||||
- CPU binding runs only on ARM. It is skipped on x86_64.
|
||||
- Each final NPU pool must have enough CPUs for its role split: at least 5 CPUs
|
||||
for devices with IRQ binding. Ascend 950 requires one complete CPU cluster per worker.
|
||||
- `global_slice` is deterministic and provides CPU/NUMA isolation when the
|
||||
cpuset is NUMA-aligned, but it cannot guarantee NUMA-local pools when CPU
|
||||
numbering or cpuset layout crosses NUMA boundaries.
|
||||
- `topo_affinity` depends on usable output from `npu-smi info -t topo`.
|
||||
- IRQ binding requires writable `/proc/irq` and resolvable PCI/IRQ information.
|
||||
Ascend 950 skips IRQ binding even when `/proc/irq` is writable.
|
||||
- Ascend 950 UVB polling thread binding requires visibility into the host PID
|
||||
namespace. Docker containers must be created with `--pid=host`; otherwise
|
||||
`uvb_poll_window_thread` may not be found.
|
||||
- Memory migration requires `migratepages`; otherwise only memory migration is
|
||||
skipped. CPU affinity still applies, but performance may degrade because
|
||||
existing pages are not moved to the target NUMA node and may be read through
|
||||
higher-latency remote NUMA access.
|
||||
- If an exception escapes the binding flow, `NPUWorker` logs a warning and skips CPU binding for that rank.
|
||||
|
||||
## References
|
||||
|
||||
- Implementation: `vllm_ascend/cpu_binding.py`
|
||||
- Worker integration: `vllm_ascend/worker/worker.py`
|
||||
- Config: `vllm_ascend/ascend_config.py` and `docs/source/user_guide/configuration/additional_config.md`
|
||||
- Tests: `tests/ut/device_allocator/test_cpu_binding.py`
|
||||
@@ -0,0 +1,105 @@
|
||||
# Disaggregated-prefill
|
||||
|
||||
## Why disaggregated-prefill?
|
||||
|
||||
This feature addresses the need to optimize the **Time Per Output Token (TPOT)** and **Time To First Token (TTFT)** in large-scale inference tasks. The motivation is two-fold:
|
||||
|
||||
1. **Adjusting Parallel Strategy and Instance Count for P and D Nodes**
|
||||
Using the disaggregated-prefill strategy, this feature allows the system to flexibly adjust the parallelization strategy (e.g., data parallelism (dp), tensor parallelism (tp), and expert parallelism (ep)) and the instance count for both P (Prefiller) and D (Decoder) nodes. This leads to better system performance tuning, particularly for **TTFT** and **TPOT**.
|
||||
|
||||
2. **Optimizing TPOT**
|
||||
Without the disaggregated-prefill strategy, prefill tasks are inserted during decoding, which results in inefficiencies and delays. Disaggregated-prefill solves this by allowing for better control over the system's **TPOT**. By managing chunked prefill tasks effectively, the system avoids the challenge of determining the optimal chunk size and provides more reliable control over the time taken for generating output tokens.
|
||||
|
||||
---
|
||||
|
||||
## Usage
|
||||
|
||||
vLLM Ascend currently supports two types of connectors for handling KV cache management:
|
||||
|
||||
- **MooncakeConnector**: D nodes pull KV cache from P nodes.
|
||||
- **MooncakeLayerwiseConnector**: P nodes push KV cache to D nodes in a layered manner.
|
||||
|
||||
For step-by-step deployment and configuration, refer to the following guide:
|
||||
[https://docs.vllm.ai/projects/ascend/en/latest/tutorials/features/pd_disaggregation_mooncake_multi_node.html](https://docs.vllm.ai/projects/ascend/en/latest/tutorials/features/pd_disaggregation_mooncake_multi_node.html)
|
||||
|
||||
---
|
||||
|
||||
## How It Works
|
||||
|
||||
### 1. Design Approach
|
||||
|
||||
Under the disaggregated-prefill, a global proxy receives external requests, forwarding prefill to P nodes and decode to D nodes; the KV cache (key-value cache) is exchanged between P and D nodes via peer-to-peer (P2P) communication.
|
||||
|
||||
### 2. Implementation Design
|
||||
|
||||
Our design diagram is shown below, illustrating the pull and push schemes respectively.
|
||||

|
||||

|
||||
|
||||
#### Mooncake Connector
|
||||
|
||||
1. The request is sent to the Proxy's `_handle_completions` endpoint.
|
||||
2. The Proxy calls `select_prefiller` to choose a P node and forwards the request, configuring `kv_transfer_params` with `do_remote_decode=True`, `max_completion_tokens=1`, and `min_tokens=1`.
|
||||
3. After the P node's scheduler finishes prefill, `update_from_output` invokes the schedule connector's `request_finished` to defer KV cache release, constructs `kv_transfer_params` with `do_remote_prefill=True`, and returns to the Proxy.
|
||||
4. The Proxy calls `select_decoder` to choose a D node and forwards the request.
|
||||
5. On the D node, the scheduler marks the request as `RequestStatus.WAITING_FOR_REMOTE_KVS`, pre-allocates KV cache, calls `kv_connector_no_forward` to pull the remote KV cache, then notifies the P node to release KV cache and proceeds with decoding to return the result.
|
||||
|
||||
#### Mooncake Layerwise Connector
|
||||
|
||||
1. The request is sent to the Proxy's `_handle_completions` endpoint.
|
||||
2. The Proxy calls `select_decoder` to choose a D node and forwards the request, configuring `kv_transfer_params` with `do_remote_prefill=True` and setting the `metaserver` endpoint.
|
||||
3. On the D node, the scheduler uses `kv_transfer_params` to mark the request as `RequestStatus.WAITING_FOR_REMOTE_KVS`, pre-allocates KV cache, then calls `kv_connector_no_forward` to send a request to the metaserver and waits for the KV cache transfer to complete.
|
||||
4. The Proxy's `metaserver` endpoint receives the request, calls `select_prefiller` to choose a P node, and forwards it with `kv_transfer_params` set to `do_remote_decode=True`, `max_completion_tokens=1`, and `min_tokens=1`.
|
||||
5. During processing, the P node's scheduler pushes KV cache layer-wise; once all layers pushing is complete, it releases the request and notifies the D node to begin decoding.
|
||||
6. The D node performs decoding and returns the result.
|
||||
|
||||
### 3. Interface Design
|
||||
|
||||
Taking MooncakeConnector as an example, the system is organized into three primary classes:
|
||||
|
||||
- **MooncakeConnector**: Base class that provides core interfaces.
|
||||
- **MooncakeConnectorScheduler**: Interface for scheduling the connectors within the engine core, responsible for managing KV cache transfer requirements and completion.
|
||||
- **MooncakeConnectorWorker**: Interface for managing KV cache registration and transfer in worker processes.
|
||||
|
||||
### 4. Specifications Design
|
||||
|
||||
This feature is flexible and supports various configurations, including setups with MLA and GQA models. It is compatible with A2 and A3 hardware configurations and facilitates scenarios involving equal TP setups and certain unequal TP setups across multiple P and D nodes.
|
||||
|
||||
| Feature | Status |
|
||||
|-------------------------------|----------------|
|
||||
| A2 | 🟢 Functional |
|
||||
| A3 | 🟢 Functional |
|
||||
| equal TP configuration | 🟢 Functional |
|
||||
| unequal TP configuration | 🟢 Functional |
|
||||
| MLA | 🟢 Functional |
|
||||
| GQA | 🟢 Functional |
|
||||
|
||||
- 🟢 Functional: Fully operational, with ongoing optimizations.
|
||||
- 🔵 Experimental: Experimental support, interfaces and functions may change.
|
||||
- 🚧 WIP: Under active development, will be supported soon.
|
||||
- 🟡 Planned: Scheduled for future implementation (some may have open PRs/RFCs).
|
||||
- 🔴 NO plan/Deprecated: No plan or deprecated by vLLM.
|
||||
|
||||
---
|
||||
|
||||
## DFX Analysis
|
||||
|
||||
### 1. Config Parameter Validation
|
||||
|
||||
Validate KV transfer config by checking whether the kv_connector type is supported. On transfer failures, emit clear error logs for diagnostics.
|
||||
|
||||
### 2. Port Conflict Detection
|
||||
|
||||
Before startup, perform a port-usage check on configured ports (e.g., rpc_port, metrics_port, http_port/metaserver) by attempting to bind. If a port is already in use, fail fast and log an error.
|
||||
|
||||
### 3. PD Ratio Validation
|
||||
|
||||
Under non-symmetric PD scenarios, validate the P-to-D tp ratio against expected and scheduling constraints to ensure correct and reliable operation.
|
||||
|
||||
---
|
||||
|
||||
## Limitations
|
||||
|
||||
- Heterogeneous P and D nodes are not supported, for example, running P nodes on A2 and D nodes on A3.
|
||||
|
||||
- In non-symmetric TP configurations, only cases where the P nodes have a higher TP degree than the D nodes and the P TP count is an integer multiple of the D TP count are supported (i.e., P_tp > D_tp and P_tp % D_tp = 0).
|
||||
@@ -0,0 +1,158 @@
|
||||
# Dynamic Chunked Pipeline Parallel (CPP)
|
||||
|
||||
TL;DR CPP uses profiling-based dynamic chunking to equalize per-chunk latency and eliminate pipeline bubbles in PP scenarios.
|
||||
|
||||
## Background
|
||||
|
||||
### Problem Statement
|
||||
|
||||
In Pipeline Parallelism (PP) + Chunked Prefill scenarios, long sequences are split into fixed-size chunks that pass through the pipeline sequentially. Due to the O(n²) computational complexity of Self-Attention, **chunks of the same size take increasingly longer to process as the prefix sequence grows**:
|
||||
|
||||
```text
|
||||
Chunk 1 (history=0): ██████ → Time T1
|
||||
Chunk 2 (history=4K): ████████ → Time T2 > T1
|
||||
Chunk 3 (history=8K): ██████████ → Time T3 > T2
|
||||
Chunk 4 (history=12K): ████████████ → Time T4 > T3
|
||||
```
|
||||
|
||||
This time variance propagates across pipeline stages, causing increased idle waiting (Pipeline Bubble) and significantly reducing GPU utilization.
|
||||
|
||||
### Solution Overview
|
||||
|
||||
Dynamic Chunked Pipeline Parallel uses a **profile-first, then predict** strategy:
|
||||
|
||||
```text
|
||||
Fixed Chunking (equal chunk size, unequal time):
|
||||
|
||||
Stage 0 |■■■■|■■■■■■|■■■■■■■■|■■■■■■■■■■|
|
||||
Stage 1 | |■■■■ |■■■■■■ |■■■■■■■■ |■■■■■■■■■■|
|
||||
↑ bubble ↑ bubble ↑ bubble
|
||||
|
||||
Dynamic Chunking (unequal chunk size, equal time):
|
||||
|
||||
Stage 0 |■■■■■■|■■■■■■|■■■■■■|■■■■■■|
|
||||
Stage 1 | |■■■■■■|■■■■■■|■■■■■■|■■■■■■|
|
||||
↑ no bubble — stages stay in sync
|
||||
```
|
||||
|
||||
The core idea is borrowed from [SGLang's dynamic chunking mechanism](https://lmsys.org/blog/2026-01-15-chunked-pipeline/), with additional enhancements such as online calibration.
|
||||
|
||||
## Design
|
||||
|
||||
### Quadratic Latency Model
|
||||
|
||||
Transformer prefill latency grows quadratically with sequence length due to the O(n²) Self-Attention mechanism:
|
||||
|
||||
$$f(l) = a \cdot l^2 + b \cdot l + c$$
|
||||
|
||||
Where:
|
||||
|
||||
- $a \cdot l^2$: Attention overhead (quadratic)
|
||||
- $b \cdot l$: Linear operations (FFN, projection)
|
||||
- $c$: Fixed overhead (kernel launch)
|
||||
|
||||
### Startup Phase: Profiling
|
||||
|
||||
During engine initialization, the system profiles actual model performance:
|
||||
|
||||
1. **Sampling**: Uniformly sample 64 different chunk sizes from `base_chunk_size` down to near 0
|
||||
2. **Execution**: Perform real model forward passes for each chunk size and precisely measure latency (milliseconds)
|
||||
3. **Fitting**: Fit the quadratic model using least squares
|
||||
4. **Target Setting**: Calculate target per-chunk latency based on `base_chunk_size`
|
||||
|
||||
In PP mode, all workers execute forward passes to stay synchronized, but only the first PP rank's timing results are used for scheduling decisions.
|
||||
|
||||
### Runtime Phase: Dynamic Prediction
|
||||
|
||||
Given current prefix length $L$ and target latency $T = f(\text{base\_chunk\_size}) - f(0)$, the system solves for the next chunk size $x$:
|
||||
|
||||
$$f(L + x) - f(L) = T$$
|
||||
|
||||
Expanding to:
|
||||
|
||||
$$a \cdot x^2 + (2aL + b) \cdot x - T = 0$$
|
||||
|
||||
Solved using the quadratic formula:
|
||||
|
||||
$$x = \frac{-(2aL + b) + \sqrt{(2aL + b)^2 + 4aT}}{2a}$$
|
||||
|
||||
The result goes through post-processing:
|
||||
|
||||
1. **Smoothing**: Blend predicted chunk size with `base_chunk_size` using `smooth_factor`
|
||||
2. **Alignment**: Round down to multiple of `page_size` (minimum 64)
|
||||
3. **Constraints**: Not exceeding `max_model_len - history_len` and `max_num_scheduled_tokens`
|
||||
|
||||
### Online Calibration
|
||||
|
||||
Since profiling only covers sequences up to `max_num_batched_tokens` (typically shorter than real workloads), the system continuously refines the model at runtime.
|
||||
|
||||
**Extended Model (two variables):**
|
||||
|
||||
$$f(C, H) = a \cdot C(C+H) + b \cdot (C+H) + c$$
|
||||
|
||||
Where $C$ is chunk size and $H$ is prefix history length.
|
||||
|
||||
After each batch, feature vectors `[Σ(C+H)·C, Σ(C+H), N]` and actual execution time are recorded. Once enough data points accumulate (5-30), model parameters are updated using least squares.
|
||||
|
||||
## Architecture
|
||||
|
||||
### Key Components
|
||||
|
||||
| Component | Location | Responsibility |
|
||||
|-----------|----------|---------------|
|
||||
| **ChunkSizePredictor** | `vllm_ascend/core/profiling_chunk_predictor.py` | Quadratic model fitting and prediction |
|
||||
| **ProfilingChunkManager** | `vllm_ascend/core/profiling_chunk_predictor.py` | Manage profiling workflow and predictor |
|
||||
| **Scheduler** | `vllm_ascend/core/scheduler_profiling_chunk.py` | Integrate CPP scheduling |
|
||||
| **EngineCore** | `vllm_ascend/patch/platform/patch_profiling_chunk.py` | Startup profiling, record execution time |
|
||||
| **NPUWorker** | `vllm_ascend/worker/worker.py` | Execute real forward pass profiling |
|
||||
| **NPUModelRunner** | `vllm_ascend/worker/model_runner_v1.py` | `profile_cpp=True` mode |
|
||||
|
||||
### Workflow
|
||||
|
||||
```text
|
||||
┌─────────────────────────────────────────────────────────────┐
|
||||
│ Startup Phase │
|
||||
├─────────────────────────────────────────────────────────────┤
|
||||
│ 1. EngineCore.init() triggers profiling │
|
||||
│ 2. ProfilingChunkManager samples 64 chunk sizes │
|
||||
│ 3. NPUWorker executes forward passes │
|
||||
│ 4. ChunkSizePredictor fits quadratic model │
|
||||
│ 5. Target latency = f(base_chunk_size) - f(0) │
|
||||
└─────────────────────────────────────────────────────────────┘
|
||||
↓
|
||||
┌─────────────────────────────────────────────────────────────┐
|
||||
│ Runtime Phase │
|
||||
├─────────────────────────────────────────────────────────────┤
|
||||
│ For each prefill chunk: │
|
||||
│ 1. Scheduler queries ChunkSizePredictor │
|
||||
│ 2. Given history length L, solve for optimal chunk size │
|
||||
│ 3. Apply smoothing and alignment │
|
||||
│ 4. Execute chunk │
|
||||
│ 5. Record actual timing for online calibration │
|
||||
│ 6. Update model if enough samples collected │
|
||||
└─────────────────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
## Comparison with SGLang
|
||||
|
||||
| Feature | SGLang Dynamic Chunking | Dynamic Chunked Pipeline Parallel |
|
||||
|---------|------------------------|-----------------------------------|
|
||||
| Profiling method | Preset quadratic function | Real forward pass profiling at startup |
|
||||
| Model fitting | $f(l) = a \cdot l^2 + b \cdot l + c$ | Same + online calibration $f(C,H)$ |
|
||||
| Online updates | None | History-based fitting |
|
||||
| Accuracy | May deviate on different hardware | Adapts to actual hardware performance |
|
||||
| Startup cost | None | ~64 forward passes (tens of seconds) |
|
||||
|
||||
## Constraints
|
||||
|
||||
- **Pipeline Parallelism Required**: Must set `--pipeline-parallel-size > 1`
|
||||
- **Chunked Prefill Required**: Must enable `--enable-chunked-prefill`
|
||||
- **Incompatible with Balance Scheduling**: Cannot enable `VLLM_ASCEND_BALANCE_SCHEDULING`
|
||||
- **Startup Overhead**: Profiling phase adds tens of seconds to initialization
|
||||
- **Memory**: No additional runtime memory overhead; profiling reuses existing dummy_run mechanism
|
||||
|
||||
## References
|
||||
|
||||
- [SGLang Dynamic Chunking Blog](https://lmsys.org/blog/2026-01-15-chunked-pipeline/)
|
||||
- [User Guide](../../user_guide/feature_guide/dynamic_chunk_pipeline_parallel.md)
|
||||
- [Tutorial](../../tutorials/features/dynamic_chunked_pipeline_parallel.md)
|
||||
@@ -0,0 +1,245 @@
|
||||
# Expert Parallelism Load Balancer (EPLB)
|
||||
|
||||
## Why We Need EPLB?
|
||||
|
||||
When using Expert Parallelism (EP), different experts are assigned to different NPUs. Given that the load of various experts may vary depending on the current workload, it is crucial to maintain balanced loads across different NPUs. We adopt a redundant experts strategy by duplicating heavily-loaded experts. Then, we heuristically pack these duplicated experts onto NPUs to ensure load balancing across them. Moreover, thanks to the group-limited expert routing used in MoE models, we also attempt to place experts of the same group on the same node to reduce inter-node data traffic, whenever possible.
|
||||
|
||||
To facilitate reproduction and deployment, vLLM Ascend supports the deployed EP load balancing algorithm in `vllm_ascend/eplb/core/policy`. The algorithm computes a balanced expert replication and placement plan based on the estimated expert loads. Note that the exact method for predicting expert loads is outside the scope of this repository. A common method is to use a moving average of historical statistics.
|
||||
|
||||

|
||||
|
||||
## How to Use EPLB?
|
||||
|
||||
Please refer to the EPLB section of the user guide for detailed information: [How to Use EPLB](../../user_guide/feature_guide/expert_parallelism_load_balancer.md)
|
||||
|
||||
## How It Works?
|
||||
|
||||
**EPLB Module Architecture**
|
||||
|
||||
```shell
|
||||
vllm_ascend
|
||||
├── eplb
|
||||
│ ├── adaptor
|
||||
│ │ └── vllm_adaptor.py
|
||||
│ ├── core
|
||||
│ │ ├── policy
|
||||
│ │ │ ├── policy_abstract.py
|
||||
│ │ │ ├── policy_default_eplb.py
|
||||
│ │ │ ├── policy_factory.py
|
||||
│ │ │ ├── policy_flashlb.py
|
||||
│ │ │ ├── policy_random.py
|
||||
│ │ │ └── policy_swift_balancer.py
|
||||
│ │ ├── eplb_device_transfer_loader.py
|
||||
│ │ ├── eplb_utils.py
|
||||
│ │ └── eplb_worker.py
|
||||
│ ├── eplb_updator.py
|
||||
│ └── utils.py
|
||||
└───────────
|
||||
```
|
||||
|
||||
**1. Adaptor Module**
|
||||
*Handles registration and adaptation for different MoE model types*
|
||||
|
||||
- `vllm_adaptor.py`
|
||||
Implementation supporting Qwen3-MoE and DeepSeek models, standardizing parameter handling for policy algorithms
|
||||
|
||||
**2. Core Module**
|
||||
*Implements core algorithms, updates, and asynchronous processing*
|
||||
|
||||
- **Policy Submodule**
|
||||
*Load balancing algorithms with factory pattern instantiation*
|
||||
- `policy_abstract.py`
|
||||
Abstract class for load balancing strategy interfaces
|
||||
- `policy_default_eplb.py`
|
||||
Default implementation of open-source EPLB paper algorithm
|
||||
- `policy_swift_balancer.py`
|
||||
Enhanced version optimizing expert swaps for low-bandwidth devices (e.g., A2)
|
||||
- `policy_flashlb.py`
|
||||
Threshold-based adjustment reducing operational costs through layer-wise fluctuation detection
|
||||
- `policy_random.py`
|
||||
Random policy for basic testing
|
||||
- `policy_factory.py`
|
||||
Strategy factory for automatic algorithm instantiation
|
||||
|
||||
- `eplb_device_transfer_loader.py`
|
||||
Manages expert table/weight transmission and updates
|
||||
- `eplb_utils.py`
|
||||
Utilities for expert table initialization and mapping
|
||||
- `eplb_worker.py`
|
||||
Asynchronous algorithm orchestration and result processing
|
||||
|
||||
**3. System Components**
|
||||
|
||||
- `eplb_updator.py`
|
||||
Central coordinator for load balancing during inference workflows
|
||||
- `utils.py`
|
||||
General utilities for EPLB interface registration
|
||||
|
||||
*Key Optimizations:*
|
||||
|
||||
1. Maintained original structure while improving technical clarity
|
||||
2. Standardized terminology
|
||||
3. Enhanced algorithm differentiation through concise descriptors
|
||||
4. Improved scoping through hierarchical presentation
|
||||
5. Preserved file/class relationships while optimizing readability
|
||||
|
||||
### Default Algorithm
|
||||
|
||||
#### Hierarchical Load Balancing
|
||||
|
||||
When the number of server nodes evenly divides the number of expert groups, we use the hierarchical load balancing policy to leverage group-limited expert routing. We first pack the expert groups onto nodes evenly, ensuring balanced loads across different nodes. Then, we replicate the experts within each node. Finally, we pack the replicated experts onto individual NPUs to ensure load balancing across them. The hierarchical load balancing policy can be used in the prefilling stage with a smaller expert-parallel size.
|
||||
|
||||
#### Global Load Balancing
|
||||
|
||||
In other cases, we use the global load balancing policy, which replicates experts globally regardless of expert groups, and packs the replicated experts onto individual NPUs. This policy can be adopted in the decoding stage with a larger expert-parallel size.
|
||||
|
||||
### Add a New EPLB Policy
|
||||
|
||||
If you want to add a new eplb policy to vllm_ascend, you must follow these steps:
|
||||
|
||||
1. Inherit the `EplbPolicy` abstract class of `policy_abstract.py` and override the `rebalance_experts` interface, ensuring consistent input parameters `current_expert_table`, `expert_workload` and return types `newplacement`.
|
||||
For example:
|
||||
|
||||
```python
|
||||
class RandomLoadBalance(EplbPolicy):
|
||||
def rebalance_experts(self, current_expert_table, expert_workload):
|
||||
new_table = copy.deepcopy(current_expert_table)
|
||||
num_layers = len(current_expert_table)
|
||||
|
||||
for i in range(num_layers):
|
||||
# randomly choose two card
|
||||
# indices = random.sample(range(num_card), 2)
|
||||
indices = [3, 1]
|
||||
|
||||
# swap redundant experts
|
||||
expert_id_to_exchange = new_table[i][indices[0]][-1].clone()
|
||||
new_table[i][indices[0]][-1] = new_table[i][indices[1]][-1]
|
||||
new_table[i][indices[1]][-1] = expert_id_to_exchange
|
||||
|
||||
return 1, [-i for i in range(num_layers)], new_table
|
||||
```
|
||||
|
||||
2. To add a new EPLB algorithm, include the policy type and its corresponding implementation class in the `PolicyFactory` of `policy_factory.py`.
|
||||
|
||||
### Add a New MoE Model
|
||||
|
||||
**Implementation Guide for Model Integration**
|
||||
|
||||
1. **Adapter File Modification**
|
||||
- Inherit or modify `vllm_ascend/eplb/adaptor/vllm_adaptor.py`
|
||||
- Add processing logic for key parameters:
|
||||
- `num_dense_layers`
|
||||
- `global_expert_num`
|
||||
- `num_roe_layers`
|
||||
- Ensure parameter synchronization in the `model_register` function.
|
||||
|
||||
For example:
|
||||
|
||||
Modify `__init__` of `vllm_adaptor.py` to add a new moe model eplb params:
|
||||
|
||||
```python
|
||||
if self.model.config.model_type == "qwen3_moe":
|
||||
self.num_dense_layers = 0
|
||||
self.global_expert_num = self.model.config.num_experts
|
||||
```
|
||||
|
||||
Modify `model_register` of `vllm_adaptor.py` to register eplb params for new moe model:
|
||||
|
||||
```python
|
||||
if config.model_type == "qwen3_moe":
|
||||
model.num_moe_layers = config.num_hidden_layers
|
||||
```
|
||||
|
||||
2. **MoE Feature Integration**
|
||||
- Extend `vllm_ascend/eplb/utils.py` with MoE-specific methods
|
||||
- Implement required functionality for expert routing or weight management
|
||||
|
||||
3. **Registration Logic Update**
|
||||
- Add patch logic within the `model_register` function
|
||||
- Maintain backward compatibility with existing model types
|
||||
|
||||
4. **Validation & Testing**
|
||||
- Verify parameter consistency across layers
|
||||
- Test cross-device communication for expert tables
|
||||
- Benchmark against baseline implementations (e.g., Qwen3-MoE)
|
||||
|
||||
*Key Implementation Notes:*
|
||||
|
||||
- Preserve existing interface contracts in abstract classes
|
||||
- Use decorators for non-intrusive patch integration
|
||||
- Leverage `eplb_utils.py` for shared expert mapping operations
|
||||
|
||||
## DFX
|
||||
|
||||
### Parameter Validation
|
||||
|
||||
#### Integer Parameters
|
||||
|
||||
All integer input parameters must explicitly specify their maximum and minimum values and be subject to valid value validation. For example, `expert_heat_collection_interval` must be greater than 0:
|
||||
|
||||
```python
|
||||
@staticmethod
|
||||
def check_iterations(iterations):
|
||||
if not isinstance(iterations, int):
|
||||
raise TypeError(f"The {iterations} is not int.")
|
||||
if iterations <= 0:
|
||||
raise ValueError(
|
||||
f"The {iterations} can not be less than or equal to 0.")
|
||||
if iterations > sys.maxsize:
|
||||
raise ValueError(
|
||||
f"The {iterations} can not be larger than {sys.maxsize}")
|
||||
```
|
||||
|
||||
#### File Path
|
||||
|
||||
The file path for EPLB must be checked for legality, such as whether the file path is valid and whether it has appropriate read and write permissions. For example:
|
||||
|
||||
```python
|
||||
@staticmethod
|
||||
def check_expert_map_path(expert_map):
|
||||
if expert_map is None:
|
||||
return
|
||||
if not isinstance(expert_map, str):
|
||||
raise TypeError("The expert_map is not str.")
|
||||
if not expert_map.strip():
|
||||
raise ValueError("The expert_map is not empty.")
|
||||
_, ext = os.path.splitext(expert_map)
|
||||
if ext.lower() != ".json":
|
||||
raise TypeError("The expert_map is not json.")
|
||||
if not os.path.exists(expert_map):
|
||||
raise ValueError("The expert_map does not exist.")
|
||||
try:
|
||||
with open(expert_map, "w", encoding='utf-8') as f:
|
||||
f.read()
|
||||
except Exception as e:
|
||||
raise IOError(
|
||||
f"Fail read expert info from {expert_map}, please check the reading permission of {expert_map} : {e}"
|
||||
)
|
||||
|
||||
```
|
||||
|
||||
### Function Specifications
|
||||
|
||||
#### Initialization Function
|
||||
|
||||
All EPLB parameters must be initialized by default during initialization, with specified parameter types and default values for proper handling.
|
||||
|
||||
#### General Functions
|
||||
|
||||
All method arguments must specify parameter types and default values, and functions must include default return value handling for default arguments. It is recommended to use `try-except` blocks to handle the function body, specifying the type of exception captured and the failure handling (e.g., logging exceptions or returning a failure status).
|
||||
|
||||
### Consistency
|
||||
|
||||
#### Expert Map
|
||||
|
||||
The expert map must be globally unique during initialization and update. In a multi-node scenario during initialization, distributed communication should be used to verify the consistency of expert maps across each rank. If they are inconsistent, the user should be notified of which ranks have inconsistent maps.
|
||||
During the update process, if only a few layers or the expert table of a certain rank has been changed, the updated expert table must be synchronized with the EPLB's context to ensure global consistency.
|
||||
|
||||
#### Expert Weight
|
||||
|
||||
When updating expert weights, ensure that the memory allocated for the expert weights has been released, or that the expert (referring to the old version) is no longer in use.
|
||||
|
||||
## Limitations
|
||||
|
||||
Before using EPLB, start the script and add `export DYNAMIC_EPLB="true"`.
|
||||
Before performing load data collection (or performance data collection), start the script and add `export EXPERT_MAP_RECORD="true"`.
|
||||
20
docs/source/developer_guide/Design_Documents/index.md
Normal file
20
docs/source/developer_guide/Design_Documents/index.md
Normal file
@@ -0,0 +1,20 @@
|
||||
# Design Documents
|
||||
|
||||
This section provides an overview of the features implemented in vLLM Ascend. Developers can refer to this guide to understand how vLLM Ascend works.
|
||||
|
||||
:::{toctree}
|
||||
:caption: Design Documents
|
||||
:maxdepth: 1
|
||||
patch
|
||||
cpu_binding
|
||||
ModelRunner_prepare_inputs
|
||||
disaggregated_prefill
|
||||
eplb_swift_balancer
|
||||
ACL_Graph
|
||||
KV_Cache_Pool_Guide
|
||||
add_custom_aclnn_op
|
||||
context_parallel
|
||||
dynamic_chunked_pipeline_parallel
|
||||
quantization
|
||||
npugraph_ex
|
||||
:::
|
||||
105
docs/source/developer_guide/Design_Documents/npugraph_ex.md
Normal file
105
docs/source/developer_guide/Design_Documents/npugraph_ex.md
Normal file
@@ -0,0 +1,105 @@
|
||||
# Npugraph_ex
|
||||
|
||||
## How Does It Work?
|
||||
|
||||
This is an optimization based on FX graphs, which can be considered an acceleration solution for the aclgraph mode.
|
||||
|
||||
You can get its code [code](https://gitcode.com/Ascend/torchair)
|
||||
|
||||
```{note}
|
||||
Atlas 300I DUO and Atlas 200I Pro do not support `enable_npugraph_ex`. Set --additional-config '{"ascend_compilation_config": {"enable_npugraph_ex":false}}'.
|
||||
```
|
||||
|
||||
## Default FX Graph Optimization
|
||||
|
||||
### FX Graph pass
|
||||
|
||||
- For the intermediate nodes of the model, replace the non-in-place operators contained in the nodes with in-place operators to reduce memory movement during computation and improve performance.
|
||||
- For the original input parameters of the model, if they include in-place operators, Dynamo's Functionalize process will replace the in-place operators with a form of non-in-place operators + copy operators. npugraph_ex will reverse this process, restoring the in-place operators and reducing memory movement.
|
||||
|
||||
### FX fusion pass
|
||||
|
||||
npugraph_ex now provides some operator fusion passes, and more will be added in the future.
|
||||
|
||||
Operator combinations that meet the replacement rules can be replaced with the corresponding fused operators.
|
||||
|
||||
You can get the default [fusion pass list](https://www.hiascend.com/document/detail/zh/Pytorch/2600/modthirdparty/torchairuseguide/docs/zh/npugraph_ex/basic/pattern_fusion_pass.md#功能简介)
|
||||
|
||||
## Custom fusion pass
|
||||
|
||||
Users can register a custom graph fusion pass in npugraph_ex to modify PyTorch FX graphs. The registration relies on the register_replacement API.
|
||||
|
||||
Below is the declaration of this API and a demo of its usage.
|
||||
|
||||
```python
|
||||
register_replacement(search_fn, replace_fn, example_inputs, trace_fn=fwd_only, extra_check=_return_true, search_fn_pattern=None)
|
||||
```
|
||||
|
||||
|Parameter Name| Input/Output |Explanation|Is necessary|
|
||||
|--|--------------|---|-------|
|
||||
|search_fn|Input|This function is the operator combination or calculation logic that you want to recognize in the FX graph, such as the operator combination that needs to be fused|Yes|
|
||||
|replace_fn|Input|When the combination corresponding to search_fn is found in the target graph, this function's computation logic will replace the original subgraph to achieve operator fusion or optimization.|Yes|
|
||||
|example_inputs|Input|Example input tensors used to track search_fn and replace_fn. The shape and dtype of the input should match the actual scenario.|Yes|
|
||||
|trace_fn|Input|By default, only the forward computation graph is tracked, which is suitable for optimization during the inference phase; if training scenarios need to be supported, a function that supports backward tracking can be provided.|No|
|
||||
|extra_check|Input|Find the extra verification function after operator fusion. The function's input parameter must be a Match object from torch._inductor.pattern_matcher, and it is used for further custom checks on the matching result, such as checking whether the fused operators are on the same stream, checking the device type, checking the input shapes, and so on.|No|
|
||||
|search_fn_pattern|Input|A custom pattern object is generally unnecessary to provide. Its definition follows the rules of the native PyTorch MultiOutputPattern object. After passing this parameter, search_fn will no longer be used to match operator combinations; instead, this parameter will be used directly as the matching rule.|No|
|
||||
|
||||
### Usage Example
|
||||
|
||||
```python
|
||||
import functools
|
||||
import torch, torch_npu, npugraph_ex
|
||||
|
||||
from torch._inductor.pattern_matcher import Match
|
||||
from torch._subclasses.fake_tensor import FakeTensorMode
|
||||
from npugraph_ex.core.utils import logger
|
||||
|
||||
# Assume fusing the add operator and the npu_rms_norm operator into the npu_add_rms_norm operator
|
||||
# Define a search_fn to find the operator combinations in the original FX graph before fusion.
|
||||
def search_fn(x1, x2, gamma):
|
||||
xOut = torch.add(x1, x2)
|
||||
y, _ = torch_npu.npu_rms_norm(xOut, gamma)
|
||||
return y, xOut
|
||||
|
||||
# Define a replace_fn, that is, a fusion operator, used to replace operator combinations in the FX graph
|
||||
def replace_fn(x1, x2, gamma):
|
||||
y, _, xOut = torch_npu.npu_add_rms_norm(
|
||||
x1, x2, gamma
|
||||
)
|
||||
return y, xOut
|
||||
|
||||
# extra_check can pass in additional validation logic. Here, it is used to check whether the last dimension of the first input parameter x1 is a specific value; if it is not the specific value, fusion is not allowed.
|
||||
def extra_check(match: Match):
|
||||
x1 = match.kwargs.get("x1")
|
||||
|
||||
if x1 is None:
|
||||
return False
|
||||
if not hasattr(x1, "meta") or "val" not in x1.meta:
|
||||
return False
|
||||
|
||||
a_shape = x1.meta["val"].shape
|
||||
return a_shape[-1] == 7168
|
||||
|
||||
|
||||
# Define some sample inputs to trace search_fn and replace_fn into an FX graph
|
||||
fake_mode = FakeTensorMode()
|
||||
with fake_mode:
|
||||
# sizes/values don't actually matter for initial trace
|
||||
# once we get a possible match we re-trace with the actual values and verify the match still holds
|
||||
input_tensor = functools.partial(torch.empty, (1, 1, 2), device="npu", dtype=torch.float16)
|
||||
kwargs_tensor = functools.partial(torch.empty, 2, device="npu", dtype=torch.float16)
|
||||
|
||||
# Call the npugraph_ex.register_replacement API with search_fn, replace_fn, and example_inputs. If there are additional validations, you can pass them in as extra_check.
|
||||
npugraph_ex.register_replacement(
|
||||
search_fn=search_fn,
|
||||
replace_fn=replace_fn,
|
||||
example_inputs=(input_tensor(), input_tensor(), kwargs_tensor()),
|
||||
extra_check=extra_check
|
||||
)
|
||||
```
|
||||
|
||||
The default fusion pass in npugraph_ex is also implemented based on this API. You can see more examples of using this API in the vllm-ascend and npugraph_ex code repositories.
|
||||
|
||||
### DFX
|
||||
|
||||
By reusing the TORCH_COMPILE_DEBUG environment variable from the PyTorch community, when TORCH_COMPILE_DEBUG=1 is set, it will output the FX graphs throughout the entire process.
|
||||
75
docs/source/developer_guide/Design_Documents/patch.md
Normal file
75
docs/source/developer_guide/Design_Documents/patch.md
Normal file
@@ -0,0 +1,75 @@
|
||||
# Patch in vLLM Ascend
|
||||
|
||||
vLLM Ascend is a platform plugin for vLLM. Due to the different release cycle of vLLM and vLLM Ascend and their hardware limitations, we need to patch some code in vLLM to make it compatible with vLLM Ascend.
|
||||
|
||||
In vLLM Ascend code, we provide a patch module `vllm_ascend/patch` to adapt to changes in vLLM.
|
||||
|
||||
## Principle
|
||||
|
||||
We should keep in mind that Patch is not the best way to make vLLM Ascend compatible. It's just a temporary solution. The best way is to contribute the change to vLLM to make it compatible with vLLM Ascend initially. In vLLM Ascend, we have the basic principle for Patch strategy:
|
||||
|
||||
1. Less is more. Please do not patch unless it's the only way currently.
|
||||
2. Once a patch is added, it's required to describe the future plan for removing the patch.
|
||||
3. Anytime, cleaning the patch code is welcome.
|
||||
|
||||
## How it works
|
||||
|
||||
In `vllm_ascend/patch`, you can see the code structure as follows:
|
||||
|
||||
```shell
|
||||
vllm_ascend/
|
||||
└── patch/
|
||||
├── platform/
|
||||
│ └── patch_xxx.py
|
||||
└── worker/
|
||||
└── patch_yyy.py
|
||||
```
|
||||
|
||||
- **platform**: The patch code in this directory is for patching the code in vLLM Main process. It's called by `vllm_ascend/platform::NPUPlatform::pre_register_and_update` very early when vLLM is initialized.
|
||||
- For online mode, vLLM process calls the platform patch in `vllm/vllm/engine/arg_utils.py::AsyncEngineArgs.add_cli_args` when parsing the CLI args.
|
||||
- For offline mode, vLLM process calls the platform patch in `vllm/vllm/engine/arg_utils.py::EngineArgs.create_engine_config` when parsing the input parameters.
|
||||
- **worker**: The patch code in this directory is for patching the code in vLLM worker process. It's called by `vllm_ascend/worker/worker::NPUWorker::__init__` when the vLLM Worker process is initialized.
|
||||
- For both online and offline mode, vLLM EngineCore process calls the worker patch in `vllm/vllm/worker/worker_base.py::WorkerWrapperBase.init_worker` when initializing the worker process.
|
||||
|
||||
## How to write a patch
|
||||
|
||||
Before writing a patch, following the principle above, we should patch the least code. If it's necessary, we can patch the code in either **platform** or **worker** folder. Here is an example to patch `distributed` module in vLLM.
|
||||
|
||||
1. Decide which version of vLLM we should patch. For example, after analysis, here we want to patch both `0.10.0` and `main` of vLLM.
|
||||
2. Decide which process we should patch. For example, here `distributed` belongs to the vLLM main process, so we should patch `platform`.
|
||||
3. Create the patch file in the right folder. The file should be named as `patch_{module_name}.py`. The example here is `vllm_ascend/patch/platform/patch_distributed.py`.
|
||||
4. Write your patch code in the new file. Here is an example:
|
||||
|
||||
```python
|
||||
import vllm
|
||||
|
||||
def patch_destroy_model_parallel():
|
||||
# your patch code
|
||||
...
|
||||
|
||||
vllm.distributed.parallel_state.destroy_model_parallel = patch_destroy_model_parallel
|
||||
```
|
||||
|
||||
5. Import the patch file in `__init__.py`. In this example, add `import vllm_ascend.patch.platform.patch_distributed` into `vllm_ascend/patch/platform/__init__.py`.
|
||||
6. Add the description of the patch in `vllm_ascend/patch/__init__.py`. The description format is as follows:
|
||||
|
||||
```python
|
||||
# ** File: <The patch file name> **
|
||||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||
# 1. `<The target patch module in vLLM>`
|
||||
# Why:
|
||||
# <Describe the reason why we need to patch>
|
||||
# How:
|
||||
# <Describe the way to patch>
|
||||
# Related PR (if no, explain why):
|
||||
# <Add a link to the related PR in vLLM. If there is no related PR, explain why>
|
||||
# Future Plan:
|
||||
# <Describe the future plan to remove the patch>
|
||||
```
|
||||
|
||||
7. Add the Unit Test and E2E Test. Any newly added code in vLLM Ascend should contain the Unit Test and E2E Test as well. You can find more details in [test guide](../contribution/testing.md)
|
||||
|
||||
## Limitations
|
||||
|
||||
1. In V1 Engine, vLLM starts three kinds of processes: Main process, EngineCore process and Worker process. Now vLLM Ascend can only patch the code in Main process and Worker process by default. If you want to patch the code running in EngineCore process, you should patch EngineCore process entirely during setup. Find the entire code in `vllm.v1.engine.core`. Please override `EngineCoreProc` and `DPEngineCoreProc` entirely.
|
||||
2. If you are running edited vLLM code, the version of vLLM may be changed automatically. For example, if you run the edited vLLM based on v0.9.n, the version of vLLM may be changed to v0.9.nxxx. In this case, the patch for v0.9.n in vLLM Ascend would not work as expected, because vLLM Ascend can't distinguish the version of the vLLM you're using. In this case, you can set the environment variable `VLLM_VERSION` to specify the version of the vLLM you're using, and then the patch for that version (e.g., v0.9.n) should work.
|
||||
114
docs/source/developer_guide/Design_Documents/quantization.md
Normal file
114
docs/source/developer_guide/Design_Documents/quantization.md
Normal file
@@ -0,0 +1,114 @@
|
||||
# Quantization Adaptation Guide
|
||||
|
||||
This document provides guidance for adapting quantization algorithms and models related to **ModelSlim**.
|
||||
|
||||
## Quantization Feature Introduction
|
||||
|
||||
### Quantization Inference Process
|
||||
|
||||
The current process for registering and obtaining quantization methods in vLLM Ascend is as follows:
|
||||
|
||||

|
||||
|
||||
vLLM Ascend registers a custom Ascend quantization method. By configuring the `--quantization ascend` parameter (or `quantization="ascend"` for offline), the quantization feature is enabled. When constructing the `quant_config`, the registered `AscendModelSlimConfig` is initialized and `get_quant_method` is called to obtain the quantization method corresponding to each weight part, stored in the `quant_method` attribute.
|
||||
|
||||
Currently supported quantization methods include `AscendLinearMethod`, `AscendFusedMoEMethod`, `AscendEmbeddingMethod`, and their corresponding non-quantized methods:
|
||||
|
||||

|
||||
|
||||
The quantization method base class defined by vLLM and the overall call flow of quantization methods are as follows:
|
||||
|
||||

|
||||
|
||||
The `embedding` method is generally not implemented for quantization, focusing only on the other three methods.
|
||||
|
||||
The `create_weights` method is used for weight initialization; the `process_weights_after_loading` method is used for weight post-processing, such as transposition, format conversion, data type conversion, etc.; the `apply` method is used to perform activation quantization and quantized matrix multiplication calculations during the forward process.
|
||||
|
||||
We need to implement the `create_weights`, `process_weights_after_loading`, and `apply` methods for different **layers** (**attention**, **mlp**, **MoE (Mixture of Experts)**).
|
||||
|
||||
**Supplement**: When loading the model, the quantized model's description file **quant_model_description.json** needs to be read. This file describes the quantization configuration and parameters for each part of the model weights, for example:
|
||||
|
||||
```json
|
||||
{
|
||||
"model.layers.0.linear_attn.dt_bias": "FLOAT",
|
||||
"model.layers.0.linear_attn.A_log": "FLOAT",
|
||||
"model.layers.0.linear_attn.conv1d.weight": "FLOAT",
|
||||
"model.layers.0.linear_attn.in_proj_qkvz.weight": "W8A8_DYNAMIC",
|
||||
"model.layers.0.linear_attn.in_proj_qkvz.weight_scale": "W8A8_DYNAMIC",
|
||||
"model.layers.0.linear_attn.in_proj_qkvz.weight_offset": "W8A8_DYNAMIC",
|
||||
"model.layers.0.linear_attn.in_proj_ba.weight": "FLOAT",
|
||||
"model.layers.0.linear_attn.norm.weight": "FLOAT",
|
||||
"model.layers.0.linear_attn.out_proj.weight": "FLOAT",
|
||||
"model.layers.0.mlp.gate.weight": "FLOAT",
|
||||
"model.layers.0.mlp.experts.0.gate_proj.weight": "W8A8_DYNAMIC",
|
||||
"model.layers.0.mlp.experts.0.gate_proj.weight_scale": "W8A8_DYNAMIC",
|
||||
"model.layers.0.mlp.experts.0.gate_proj.weight_offset": "W8A8_DYNAMIC"
|
||||
}
|
||||
```
|
||||
|
||||
Based on the above content, we present a brief description of the adaptation process for quantization algorithms and quantized models.
|
||||
|
||||
### Quantization Algorithm Adaptation
|
||||
|
||||
- **Step 1: Algorithm Design**. Define the algorithm ID (e.g., `W4A8_DYNAMIC`), determine supported layers (linear, moe, attention), and design the quantization scheme (static/dynamic, pertensor/perchannel/pergroup).
|
||||
- **Step 2: Registration**. Use the `@register_scheme` decorator in `vllm_ascend/quantization/methods/registry.py` to register your quantization scheme class.
|
||||
|
||||
```python
|
||||
from vllm_ascend.quantization.methods import register_scheme, AscendLinearScheme, AscendMoEScheme
|
||||
|
||||
@register_scheme("W4A8_DYNAMIC", "linear")
|
||||
class AscendW4A8DynamicLinearMethod(AscendLinearScheme):
|
||||
...
|
||||
|
||||
@register_scheme("W4A8_DYNAMIC", "moe")
|
||||
class AscendW4A8DynamicFusedMoEMethod(AscendMoEScheme):
|
||||
...
|
||||
```
|
||||
|
||||
- **Step 3: Implementation**. Create an algorithm implementation file, such as `vllm_ascend/quantization/methods/w4a8.py`, and implement the method class and logic.
|
||||
- **Step 4: Testing**. Use your algorithm to generate quantization configurations and verify correctness and performance on target models and hardware.
|
||||
|
||||
### Quantized Model Adaptation
|
||||
|
||||
Adapting a new quantized model requires ensuring the following three points:
|
||||
|
||||
- The original model has been successfully adapted in `vLLM Ascend`.
|
||||
- **Fused Module Mapping**: Add the model's `model_type` to `packed_modules_model_mapping` in `vllm_ascend/quantization/modelslim_config.py` (e.g., `qkv_proj`, `gate_up_proj`, `experts`) to ensure sharding consistency and correct loading.
|
||||
|
||||
```python
|
||||
packed_modules_model_mapping = {
|
||||
"qwen3_moe": {
|
||||
"qkv_proj": [
|
||||
"q_proj",
|
||||
"k_proj",
|
||||
"v_proj",
|
||||
],
|
||||
"gate_up_proj": [
|
||||
"gate_proj",
|
||||
"up_proj",
|
||||
],
|
||||
"experts":
|
||||
["experts.0.gate_proj", "experts.0.up_proj", "experts.0.down_proj"],
|
||||
},
|
||||
}
|
||||
```
|
||||
|
||||
- All quantization algorithms used by the quantized model have been integrated into the `quantization` module.
|
||||
|
||||
## Currently Supported Quantization Algorithms
|
||||
|
||||
vLLM Ascend supports multiple quantization algorithms. The following table provides an overview of each quantization algorithm based on the implementation in the `vllm_ascend.quantization` module:
|
||||
|
||||
| Algorithm | Weight | Activation | Weight Granularity | Activation Granularity | Type | Description |
|
||||
| ------------------------ | ------ | ---------- | ------------------ | ---------------------- | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
||||
| `W4A16` | INT4 | FP16/BF16 | Per-Group | Per-Tensor | Static | 4-bit weight quantization with 16-bit activation precision, specifically designed for MoE model expert layers, supporting int32 format weight packing |
|
||||
| `W8A16` | INT8 | FP16/BF16 | Per-Channel | Per-Tensor | Static | 8-bit weight quantization with 16-bit activation precision, balancing accuracy and performance, suitable for linear layers |
|
||||
| `W8A8` | INT8 | INT8 | Per-Channel | Per-Tensor | Static | Static activation quantization, suitable for scenarios requiring high precision |
|
||||
| `W8A8_DYNAMIC` | INT8 | INT8 | Per-Channel | Per-Token | Dynamic | Dynamic activation quantization with per-token scaling factor calculation |
|
||||
| `W4A8_DYNAMIC` | INT4 | INT8 | Per-Group | Per-Token | Dynamic | Supports both direct per-channel quantization to 4-bit and two-step quantization (per-channel to 8-bit then per-group to 4-bit) |
|
||||
| `W4A4_FLATQUANT_DYNAMIC` | INT4 | INT4 | Per-Channel | Per-Token | Dynamic | Uses FlatQuant for activation distribution smoothing before 4-bit dynamic quantization, with additional matrix multiplications for precision preservation |
|
||||
| `W8A8_MIX` | INT8 | INT8 | Per-Channel | Per-Tensor/Token | Mixed | We support two deployment modes: PD Colocation (dynamic quantization for both P and D) and PD Disaggregation (dynamic-quant P and static-quant D) |
|
||||
|
||||
**Static vs Dynamic:** Static quantization uses pre-computed scaling factors with better performance, while dynamic quantization computes scaling factors on-the-fly for each token/activation tensor with higher precision.
|
||||
|
||||
**Granularity:** Refers to the scope of scaling factor computation (e.g., per-tensor, per-channel, per-group).
|
||||
469
docs/source/developer_guide/contribution/doc_writing.md
Normal file
469
docs/source/developer_guide/contribution/doc_writing.md
Normal file
@@ -0,0 +1,469 @@
|
||||
# Doc writing guide
|
||||
|
||||
## Guide to Writing Model Tutorial Doc
|
||||
|
||||
`docs/source/_templates/Model-Deployment-Tutorial-Template.md` is a template for writing model deployment tutorials. You can copy and modify it to create new docs.
|
||||
|
||||
## Testable doc code block generation (``model-code``)
|
||||
|
||||
- For **documentation authors**: how to insert testable command blocks into docs
|
||||
- For **developers**: how to add a new converter
|
||||
|
||||
Built-in supported `converter_tag` values:
|
||||
|
||||
| converter_tag | Renders | YAML source |
|
||||
| --- | --- | --- |
|
||||
| `single_node` | A single node's env exports + `vllm serve` script | `test_cases[case_index]` |
|
||||
| `multi_node` | One host's env exports + `vllm serve` script | `deployment[host_index]` |
|
||||
| `external_dp_template` | One external-DP node's env exports + `vllm serve` command | `templates[host_index]` |
|
||||
| `external_dp_launch` | One `launch_online_dp.py` line per node | `config[]` |
|
||||
| `external_dp_proxy` | The load-balance proxy launch command | `config[]` + `routing` |
|
||||
|
||||
### For authors: add a block
|
||||
|
||||
:::{important}
|
||||
By default, the generator scans only `.md` files under `docs/source/tutorials/models/` and produces artifacts.
|
||||
If you put ``model-code`` blocks in other directories, Sphinx builds will not automatically generate the corresponding scripts.
|
||||
:::
|
||||
|
||||
All ``model-code`` blocks need:
|
||||
|
||||
| Option | Required | Description |
|
||||
| --- | --- | --- |
|
||||
| `block_name` | Yes | Block name; must be unique within the current document |
|
||||
| `converter_tag` | Yes | Selects one of the built-in converters |
|
||||
| `test_case_path` | Yes | Repository-relative YAML path that stays within the repo; the file must exist |
|
||||
|
||||
Use the body of the block to add shell wrapper lines such as `set -eux`. Always
|
||||
place the `{{ generated }}` placeholder where the converter output should be
|
||||
inserted.
|
||||
|
||||
#### converter_tag: `single_node`
|
||||
|
||||
`single_node` reads one item from `test_cases`. The optional `case_index`
|
||||
metadata selects the item; when omitted, it defaults to `0`.
|
||||
|
||||
Only the fields read by this converter are expanded below. Other test metadata
|
||||
can be left in the YAML and is ignored by this converter.
|
||||
|
||||
```yaml
|
||||
test_cases:
|
||||
- name: qwen3-8b-single
|
||||
model: Qwen/Qwen3-8B
|
||||
envs:
|
||||
HCCL_BUFFSIZE: "1024"
|
||||
SERVER_PORT: DEFAULT_PORT
|
||||
server_cmd:
|
||||
- --tensor-parallel-size
|
||||
- "1"
|
||||
- --port
|
||||
- $SERVER_PORT
|
||||
- --trust-remote-code
|
||||
server_cmd_extra:
|
||||
- --enable-expert-parallel
|
||||
benchmarks: ...
|
||||
```
|
||||
|
||||
`envs` is rendered as `export` lines. `SERVER_PORT: DEFAULT_PORT` is resolved
|
||||
to the default single-node port `8000`. `model` becomes `vllm serve <model>`,
|
||||
and `server_cmd` plus optional `server_cmd_extra` become command arguments.
|
||||
Both command fields can be either a shell string or a flat token list.
|
||||
|
||||
Write the doc block like this:
|
||||
|
||||
````md
|
||||
```{model-code}
|
||||
:block_name: qwen3_8b_single_node
|
||||
:converter_tag: single_node
|
||||
:test_case_path: tests/e2e/nightly/single_node/models/configs/your_model.yaml
|
||||
:case_index: 0
|
||||
|
||||
set -eux
|
||||
{{ generated }}
|
||||
```
|
||||
````
|
||||
|
||||
Generated shell script:
|
||||
|
||||
```bash
|
||||
set -eux
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export SERVER_PORT=8000
|
||||
|
||||
vllm serve Qwen/Qwen3-8B \
|
||||
--tensor-parallel-size 1 \
|
||||
--port $SERVER_PORT \
|
||||
--trust-remote-code \
|
||||
--enable-expert-parallel
|
||||
```
|
||||
|
||||
#### converter_tag: `multi_node`
|
||||
|
||||
`multi_node` reads one item from `deployment`. The required `host_index`
|
||||
metadata selects which host to render.
|
||||
|
||||
```yaml
|
||||
deployment:
|
||||
- envs:
|
||||
SERVER_PORT: "8000"
|
||||
server_cmd: >
|
||||
vllm serve Qwen/Qwen3-235B-A22B
|
||||
--host 0.0.0.0
|
||||
--port $SERVER_PORT
|
||||
--data-parallel-size 2
|
||||
--tensor-parallel-size 8
|
||||
--data-parallel-address $LOCAL_IP
|
||||
- envs:
|
||||
SERVER_PORT: "8000"
|
||||
server_cmd: >
|
||||
vllm serve Qwen/Qwen3-235B-A22B
|
||||
--headless
|
||||
--port $SERVER_PORT
|
||||
--data-parallel-size 2
|
||||
--tensor-parallel-size 8
|
||||
--data-parallel-start-rank 1
|
||||
--data-parallel-address $MASTER_IP
|
||||
benchmarks: ...
|
||||
```
|
||||
|
||||
`server_cmd` must be a complete command starting with `vllm serve <model>`.
|
||||
It can be written as a shell string or a flat token list.
|
||||
|
||||
Write the doc block like this:
|
||||
|
||||
````md
|
||||
```{model-code}
|
||||
:block_name: qwen3_235b_worker_1
|
||||
:converter_tag: multi_node
|
||||
:test_case_path: tests/e2e/nightly/multi_node/internal_dp/config/your_model.yaml
|
||||
:host_index: 1
|
||||
|
||||
set -eux
|
||||
{{ generated }}
|
||||
```
|
||||
````
|
||||
|
||||
Generated shell script for `host_index: 1`:
|
||||
|
||||
```bash
|
||||
set -eux
|
||||
export MASTER_IP=192.168.1.10
|
||||
export SERVER_PORT=8000
|
||||
|
||||
vllm serve Qwen/Qwen3-235B-A22B \
|
||||
--headless \
|
||||
--port $SERVER_PORT \
|
||||
--data-parallel-size 2 \
|
||||
--tensor-parallel-size 8 \
|
||||
--data-parallel-start-rank 1 \
|
||||
--data-parallel-address $MASTER_IP
|
||||
```
|
||||
|
||||
#### converter_tag: `external_dp_template`
|
||||
|
||||
`external_dp_template` reads one item from `templates`. The required
|
||||
`host_index` metadata selects which template to render. The top-level `model`
|
||||
field is also required because the converter builds `vllm serve <model>`.
|
||||
|
||||
```yaml
|
||||
model: Eco-Tech/GLM-Test
|
||||
templates:
|
||||
- node_index: 0
|
||||
envs:
|
||||
HCCL_BUFFSIZE: "1024"
|
||||
ASCEND_RT_VISIBLE_DEVICES: "${VISIBLE_DEVICES}"
|
||||
server_cmd_template:
|
||||
- --host
|
||||
- 0.0.0.0
|
||||
- --port
|
||||
- ${PORT}
|
||||
- --data-parallel-size
|
||||
- ${DP_SIZE}
|
||||
- --data-parallel-rank
|
||||
- ${DP_RANK}
|
||||
- --data-parallel-address
|
||||
- ${DP_ADDRESS}
|
||||
- --data-parallel-rpc-port
|
||||
- ${DP_RPC_PORT}
|
||||
- --tensor-parallel-size
|
||||
- ${TP_SIZE}
|
||||
- --trust-remote-code
|
||||
config: ...
|
||||
routing: ...
|
||||
```
|
||||
|
||||
Known braced template variables are rewritten to the positional shell arguments
|
||||
that `run_dp_template.sh` receives from `launch_online_dp.py`:
|
||||
|
||||
| Template variable | Rendered positional |
|
||||
| --- | --- |
|
||||
| `${VISIBLE_DEVICES}` | `$1` |
|
||||
| `${PORT}` | `$2` |
|
||||
| `${DP_SIZE}` | `$3` |
|
||||
| `${DP_RANK}` | `$4` |
|
||||
| `${DP_ADDRESS}` | `$5` |
|
||||
| `${DP_RPC_PORT}` | `$6` |
|
||||
| `${TP_SIZE}` | `$7` |
|
||||
|
||||
Unknown braced variables and unbraced shell references such as `$SERVER_PORT`
|
||||
are left unchanged.
|
||||
|
||||
Write the doc block like this:
|
||||
|
||||
````md
|
||||
```{model-code}
|
||||
:block_name: glm_external_dp_template_node0
|
||||
:converter_tag: external_dp_template
|
||||
:test_case_path: tests/e2e/nightly/multi_node/external_dp/config/your_model.yaml
|
||||
:host_index: 0
|
||||
|
||||
set -eux
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
|
||||
{{ generated }}
|
||||
```
|
||||
````
|
||||
|
||||
Generated shell script for `host_index: 0`:
|
||||
|
||||
```bash
|
||||
set -eux
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export ASCEND_RT_VISIBLE_DEVICES=$1
|
||||
|
||||
vllm serve Eco-Tech/GLM-Test \
|
||||
--host 0.0.0.0 \
|
||||
--port $2 \
|
||||
--data-parallel-size $3 \
|
||||
--data-parallel-rank $4 \
|
||||
--data-parallel-address $5 \
|
||||
--data-parallel-rpc-port $6 \
|
||||
--tensor-parallel-size $7 \
|
||||
--trust-remote-code
|
||||
```
|
||||
|
||||
#### converter_tag: `external_dp_launch`
|
||||
|
||||
`external_dp_launch` reads the full `config` list and renders one
|
||||
`launch_online_dp.py` command per node. It does not take an index option.
|
||||
|
||||
```yaml
|
||||
config:
|
||||
- node_index: 0
|
||||
port_start: 7100
|
||||
dp_rpc_port: 12321
|
||||
dp_size: 2
|
||||
dp_size_local: 2
|
||||
dp_rank_start: 0
|
||||
tp_size: 8
|
||||
dp_address: "${NODE_0_IP}"
|
||||
- node_index: 1
|
||||
port_start: 7200
|
||||
dp_rpc_port: 12321
|
||||
dp_size: 4
|
||||
dp_size_local: 4
|
||||
dp_rank_start: 0
|
||||
tp_size: 4
|
||||
dp_address: "${NODE_1_IP}"
|
||||
templates: ...
|
||||
routing: ...
|
||||
```
|
||||
|
||||
Write the doc block like this:
|
||||
|
||||
````md
|
||||
```{model-code}
|
||||
:block_name: glm_external_dp_launch
|
||||
:converter_tag: external_dp_launch
|
||||
:test_case_path: tests/e2e/nightly/multi_node/external_dp/config/your_model.yaml
|
||||
|
||||
set -eux
|
||||
{{ generated }}
|
||||
```
|
||||
````
|
||||
|
||||
Generated shell script:
|
||||
|
||||
```bash
|
||||
set -eux
|
||||
python launch_online_dp.py --dp-size 2 --tp-size 8 --dp-size-local 2 --dp-rank-start 0 --dp-address ${NODE_0_IP} --dp-rpc-port 12321 --vllm-start-port 7100
|
||||
|
||||
python launch_online_dp.py --dp-size 4 --tp-size 4 --dp-size-local 4 --dp-rank-start 0 --dp-address ${NODE_1_IP} --dp-rpc-port 12321 --vllm-start-port 7200
|
||||
```
|
||||
|
||||
#### converter_tag: `external_dp_proxy`
|
||||
|
||||
`external_dp_proxy` reads `config` and `routing`. It renders the
|
||||
`load_balance_proxy_server_example.py` command for `routing.type:
|
||||
disaggregated_prefill`. It does not take an index option.
|
||||
|
||||
```yaml
|
||||
routing:
|
||||
type: disaggregated_prefill
|
||||
groups:
|
||||
prefiller: [0]
|
||||
decoder: [1]
|
||||
config:
|
||||
- node_index: 0
|
||||
port_start: 7100
|
||||
dp_size_local: 2
|
||||
dp_rpc_port: 12321
|
||||
dp_size: 2
|
||||
dp_rank_start: 0
|
||||
tp_size: 8
|
||||
dp_address: "${NODE_0_IP}"
|
||||
- node_index: 1
|
||||
port_start: 7200
|
||||
dp_size_local: 4
|
||||
dp_rpc_port: 12321
|
||||
dp_size: 4
|
||||
dp_rank_start: 0
|
||||
tp_size: 4
|
||||
dp_address: "${NODE_1_IP}"
|
||||
templates: ...
|
||||
```
|
||||
|
||||
`routing.groups.prefiller` and `routing.groups.decoder` contain indices into
|
||||
`config`. Each referenced node expands to `dp_size_local` host and port entries.
|
||||
The proxy itself is rendered on `${NODE_0_IP}:1999`.
|
||||
|
||||
Write the doc block like this:
|
||||
|
||||
````md
|
||||
```{model-code}
|
||||
:block_name: glm_external_dp_proxy
|
||||
:converter_tag: external_dp_proxy
|
||||
:test_case_path: tests/e2e/nightly/multi_node/external_dp/config/your_model.yaml
|
||||
|
||||
set -eux
|
||||
{{ generated }}
|
||||
```
|
||||
````
|
||||
|
||||
Generated shell script:
|
||||
|
||||
```bash
|
||||
set -eux
|
||||
python load_balance_proxy_server_example.py \
|
||||
--host ${NODE_0_IP} \
|
||||
--port 1999 \
|
||||
--prefiller-hosts \
|
||||
${NODE_0_IP} \
|
||||
${NODE_0_IP} \
|
||||
--prefiller-ports \
|
||||
7100 \
|
||||
7101 \
|
||||
--decoder-hosts \
|
||||
${NODE_1_IP} \
|
||||
${NODE_1_IP} \
|
||||
${NODE_1_IP} \
|
||||
${NODE_1_IP} \
|
||||
--decoder-ports \
|
||||
7200 \
|
||||
7201 \
|
||||
7202 \
|
||||
7203
|
||||
```
|
||||
|
||||
### Local debugging and generation
|
||||
|
||||
#### Generate only (without building the full site)
|
||||
|
||||
```bash
|
||||
# Generate all model-code artifacts under docs/source/tutorials/models/
|
||||
python3 tools/docs_codegen/cli.py
|
||||
|
||||
# Generate artifacts for a single document
|
||||
python3 tools/docs_codegen/cli.py --doc docs/source/tutorials/models/Kimi-K2-Thinking.md
|
||||
|
||||
# Generate a single block and print it (no files written)
|
||||
python3 tools/docs_codegen/cli.py \
|
||||
--block docs/source/tutorials/models/Kimi-K2-Thinking.md::kimi_k2_thinking_single_node \
|
||||
--dry-run --stdout
|
||||
```
|
||||
|
||||
By default, artifacts are written to: `docs/_build/doc_codegen/<doc_stem>/<block_name>.sh`.
|
||||
|
||||
:::{note}
|
||||
After the script is generated, please make sure to check whether the generated content is runnable, especially key parts such as environment variables and command-line parameters.
|
||||
:::
|
||||
|
||||
#### Build the site & preview locally
|
||||
|
||||
```bash
|
||||
# Install documentation build dependencies
|
||||
python3 -m pip install -r docs/requirements-docs.txt
|
||||
|
||||
# (Optional) Clean previous builds
|
||||
make -C docs clean
|
||||
|
||||
# Build the English site
|
||||
make -C docs html
|
||||
|
||||
# (Optional) Build the Chinese site
|
||||
make -C docs intl
|
||||
|
||||
# Preview locally
|
||||
python3 -m http.server -d docs/_build/html 8000
|
||||
|
||||
# Then open in a browser:
|
||||
# http://localhost:8000
|
||||
```
|
||||
|
||||
### For developers: add a new converter
|
||||
|
||||
A converter turns one loaded YAML file plus one parsed `ModelCodeBlock` into a
|
||||
`GeneratedScript`. The current pipeline is:
|
||||
|
||||
1. `BlockScanner` parses ``model-code`` fences and accepts only options listed
|
||||
in `MODEL_CODE_OPTION_NAMES`.
|
||||
2. `YamlLoader` loads `test_case_path`.
|
||||
3. `get_converter()` looks up `block.converter_tag` from
|
||||
`build_default_converters()`.
|
||||
4. The selected converter returns `GeneratedScript(content=..., language="shell")`.
|
||||
5. `GeneratorService` replaces `{{ generated }}` in the block body, validates
|
||||
that the final script is non-empty, and writes
|
||||
`docs/_build/doc_codegen/<doc_stem>/<block_name>.sh`.
|
||||
|
||||
To add a converter:
|
||||
|
||||
1. In `tools/docs_codegen/converters.py`, add a `BaseConverter` subclass with a
|
||||
unique `name`. That name is the value authors put in `:converter_tag:`.
|
||||
2. Implement `convert(self, loaded_yaml, *, block) -> GeneratedScript`. Use
|
||||
`make_docs_codegen_error(..., block=block)` for user-facing validation
|
||||
errors so the CLI and Sphinx output include document context.
|
||||
3. Reuse helpers from `tools/docs_codegen/utils.py`, such as
|
||||
`require_mapping`, `require_mapping_list`, `require_scalar_mapping`,
|
||||
`require_indexed_mapping`, `require_node_field`, `parse_command_tokens`,
|
||||
`substitute_template_positionals`, and `render_cli_command`.
|
||||
4. Register the converter in `build_default_converters()`. If it is not
|
||||
registered, `get_converter()` will reject the new `converter_tag`.
|
||||
5. If the converter needs new directive metadata, add the option name to
|
||||
`MODEL_CODE_OPTION_NAMES` in `tools/docs_codegen/scanner.py` and to
|
||||
`ModelCodeDirective.option_spec` in
|
||||
`tools/docs_codegen/sphinx_extension.py`. Read the option with
|
||||
`block.get_option("<option_name>")`.
|
||||
6. Add or update tests in `tests/ut/tools/test_docs_codegen.py`. Cover the
|
||||
successful render path, required option validation, YAML shape validation,
|
||||
and any CLI/Sphinx scanner behavior affected by new metadata.
|
||||
7. Add a real ``model-code`` example in a model tutorial, preferably under
|
||||
`docs/source/tutorials/models/`, and point it to an existing YAML file under
|
||||
`tests/`.
|
||||
8. Validate with the CLI:
|
||||
|
||||
```bash
|
||||
python3 tools/docs_codegen/cli.py --doc <your_doc> --dry-run
|
||||
python3 tools/docs_codegen/cli.py --block <your_doc>::<block_name> --dry-run --stdout
|
||||
```
|
||||
|
||||
If a converter should render something other than shell, set
|
||||
`GeneratedScript.language` accordingly so Sphinx can highlight the generated
|
||||
literal block correctly.
|
||||
154
docs/source/developer_guide/contribution/e2e_ci_test.md
Normal file
154
docs/source/developer_guide/contribution/e2e_ci_test.md
Normal file
@@ -0,0 +1,154 @@
|
||||
# E2E CI Test
|
||||
|
||||
This document explains how to trigger specific E2E tests against your PR code via a
|
||||
comment command, without running the full E2E test suite.
|
||||
|
||||
## Background
|
||||
|
||||
The `E2E-Full` workflow ([`pr_test.yaml`](https://github.com/vllm-project/vllm-ascend/blob/main/.github/workflows/pr_test.yaml)) normally runs the complete E2E test suite
|
||||
when a PR has `ready` label. This is expensive in CI resources
|
||||
and time.
|
||||
|
||||
Authorized users can trigger only the specific test files they care about by posting a
|
||||
`/e2e` comment on the PR, then adding the `ready` label.
|
||||
|
||||
## How to Trigger
|
||||
|
||||
### 1. Post a comment
|
||||
|
||||
First, post a comment on the PR specifying which test paths to run:
|
||||
|
||||
```text
|
||||
/e2e [test-path-1] [test-path-2] ...
|
||||
```
|
||||
|
||||
- Each path must be a valid pytest path relative to the repository root.
|
||||
- Multiple paths can be listed in a single comment, separated by spaces.
|
||||
- A specific test case can be targeted using `::` notation.
|
||||
|
||||
| Comment format | Effect |
|
||||
|---|---|
|
||||
| `/e2e tests/e2e/pull_request/one_card/test_foo.py` | Run one test file on one_card |
|
||||
| `/e2e tests/e2e/pull_request/two_card/test_bar.py` | Run one test file on two_card |
|
||||
| `/e2e path1 path2 path3` | Run multiple files, routed by path pattern |
|
||||
| `/e2e tests/e2e/pull_request/one_card/test_foo.py::test_case` | Run a specific test case |
|
||||
|
||||
### 2. Add the label
|
||||
|
||||
After posting the comment, add the **`ready`** label to your PR.
|
||||
Adding the label is what actually **triggers** the workflow — at that point the workflow
|
||||
reads the existing comments to find the `/e2e` command.
|
||||
|
||||
:::{note}
|
||||
Only repository **Contributors** (Triage role) and **Maintainers** (Write role) can add
|
||||
labels. If you do not have this permission, ask a maintainer to add the label for you.
|
||||
You can find the list of maintainers and contributors by checking the
|
||||
[CODEOWNERS](https://github.com/vllm-project/vllm-ascend/blob/main/.github/CODEOWNERS)
|
||||
file.
|
||||
:::
|
||||
|
||||
:::{important}
|
||||
The comment must be posted **before** the label is added. If you add the label first,
|
||||
the workflow will find no `/e2e` comment and will not trigger any per-test runs.
|
||||
:::
|
||||
|
||||
:::{note}
|
||||
Additionally, only the **PR author** or collaborators with **write or admin** repository
|
||||
access can trigger tests via comment. The workflow validates the commenter's permission
|
||||
before proceeding.
|
||||
:::
|
||||
|
||||
### 3. Wait for results
|
||||
|
||||
GitHub Actions will trigger the `E2E-Full` workflow. Only the hardware jobs matching
|
||||
the provided test paths will run, which saves CI resources.
|
||||
|
||||
## Path Routing Rules
|
||||
|
||||
The workflow automatically routes each test path to the correct hardware runner based
|
||||
on path patterns:
|
||||
|
||||
| Path pattern | Hardware | Runner |
|
||||
|---|---|---|
|
||||
| `two_card` in path | two_card A3 NPU | `linux-aarch64-a3-2` |
|
||||
| `four_card` in path | four_card A3 NPU | `linux-aarch64-a3-4` |
|
||||
| `_310p` in filename under one/two_card | Ascend 310P x1 | `linux-aarch64-310p-*` |
|
||||
| `_310p` in filename under four_card | Ascend 310P x4 | `linux-aarch64-310p-*` |
|
||||
| All other paths | one_card A2 NPU | `linux-aarch64-a2b3-1` |
|
||||
|
||||
When paths from multiple categories are listed in a single comment, each category's
|
||||
tests run on its respective hardware in parallel.
|
||||
|
||||
## Test Path Reference
|
||||
|
||||
The `tests/e2e/pull_request/` directory is organized by hardware category:
|
||||
|
||||
```text
|
||||
tests/e2e/pull_request/
|
||||
├── one_card/ # Single card tests → A2 NPU x1 runner
|
||||
├── two_card/ # Two card tests → A3 NPU x2 runner
|
||||
├── four_card/ # Four card tests → A3 NPU x4 runner
|
||||
```
|
||||
|
||||
310P tests use `_310p` subdirectories or `_310p.py` filename suffix under the
|
||||
corresponding card directory:
|
||||
|
||||
```text
|
||||
tests/e2e/pull_request/one_card/_310p/ # 310P single card
|
||||
tests/e2e/pull_request/four_card/_310p/ # 310P four card
|
||||
```
|
||||
|
||||
## Comparison with Full E2E Suite
|
||||
|
||||
| Aspect | Full E2E suite | Per-test comment trigger |
|
||||
|---|---|---|
|
||||
| Trigger | `ready` labels | `/e2e` comment + `ready` label |
|
||||
| Scope | All E2E tests | Only specified test paths |
|
||||
| Who can trigger | Anyone who can add labels | PR author or write/admin collaborator |
|
||||
| Use case | Pre-merge validation | Iterative debugging of specific tests |
|
||||
|
||||
## Examples
|
||||
|
||||
Run a single one_card test:
|
||||
|
||||
```text
|
||||
/e2e tests/e2e/pull_request/one_card/test_offline_inference.py
|
||||
```
|
||||
|
||||
Run a two_card test:
|
||||
|
||||
```text
|
||||
/e2e tests/e2e/pull_request/two_card/test_data_parallel.py
|
||||
```
|
||||
|
||||
Run tests across multiple hardware categories in one comment:
|
||||
|
||||
```text
|
||||
/e2e tests/e2e/pull_request/one_card/test_offline_inference.py tests/e2e/pull_request/two_card/test_data_parallel.py
|
||||
```
|
||||
|
||||
Re-trigger after fixing an issue: just push a new commit. The `synchronize` event
|
||||
re-runs the workflow and picks up the existing `/e2e` comment automatically — no need
|
||||
to post a new comment.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
**The workflow did not start after I added the label.**
|
||||
|
||||
- Make sure the `/e2e` comment was posted **before** the label was added.
|
||||
If the label was added first, remove it and re-add it after posting the comment.
|
||||
- Check that the comment starts exactly with `/e2e` followed by at least one path,
|
||||
with no leading spaces or extra characters before the slash.
|
||||
- To re-trigger after fixing an issue, simply push a new commit — the workflow will
|
||||
reuse the existing `/e2e` comment automatically.
|
||||
|
||||
**Tests ran on the wrong hardware.**
|
||||
|
||||
- Check that the path includes the expected directory segment (`one_card`, `two_card`,
|
||||
`four_card`, or `_310p`). Paths that do not match any of these patterns are routed to
|
||||
the one_card runner by default.
|
||||
|
||||
**The `parse-comment` job skipped with a permission error.**
|
||||
|
||||
- Only the PR author or write/admin collaborators can use the comment trigger.
|
||||
Ask a maintainer to post the `/e2e` comment instead.
|
||||
@@ -1,16 +1,17 @@
|
||||
# Contributing
|
||||
|
||||
## Building and testing
|
||||
It's recommended to set up a local development environment to build and test
|
||||
## Building and Testing
|
||||
|
||||
It's recommended to set up a local development environment to build vllm-ascend and run tests
|
||||
before you submit a PR.
|
||||
|
||||
### Setup development environment
|
||||
### Set up a development environment
|
||||
|
||||
Theoretically, the vllm-ascend build is only supported on Linux because
|
||||
`vllm-ascend` dependency `torch_npu` only supports Linux.
|
||||
|
||||
But you can still set up dev env on Linux/Windows/macOS for linting and basic
|
||||
test as following commands:
|
||||
But you can still set up a development environment on Linux/Windows/macOS for linting and running basic
|
||||
tests.
|
||||
|
||||
#### Run lint locally
|
||||
|
||||
@@ -27,20 +28,19 @@ cd vllm-ascend
|
||||
# Install lint requirement and enable pre-commit hook
|
||||
pip install -r requirements-lint.txt
|
||||
|
||||
# Run lint (You need install pre-commits deps via proxy network at first time)
|
||||
# Run lint (You need to install pre-commits deps via proxy network at first time)
|
||||
bash format.sh
|
||||
```
|
||||
|
||||
#### Run CI locally
|
||||
|
||||
After complete "Run lint" setup, you can run CI locally:
|
||||
After completing "Run lint" setup, you can run CI (Continuous integration) locally:
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
cd ~/vllm-project/
|
||||
|
||||
# Run CI need vLLM installed
|
||||
# Run CI needs vLLM installed
|
||||
git clone --branch |vllm_version| https://github.com/vllm-project/vllm.git
|
||||
cd vllm
|
||||
pip install -r requirements/build.txt
|
||||
@@ -51,7 +51,7 @@ cd ..
|
||||
cd vllm-ascend
|
||||
# For Linux:
|
||||
pip install -r requirements-dev.txt
|
||||
# For non Linux:
|
||||
# For non-Linux:
|
||||
cat requirements-dev.txt | grep -Ev '^#|^--|^$|^-r' | while read PACKAGE; do pip install "$PACKAGE"; done
|
||||
cat requirements.txt | grep -Ev '^#|^--|^$|^-r' | while read PACKAGE; do pip install "$PACKAGE"; done
|
||||
|
||||
@@ -68,13 +68,13 @@ git commit -sm "your commit info"
|
||||
|
||||
🎉 Congratulations! You have completed the development environment setup.
|
||||
|
||||
### Test locally
|
||||
### Testing locally
|
||||
|
||||
You can refer to [Testing](./testing.md) doc to help you setup testing environment and running tests locally.
|
||||
You can refer to [Testing](./testing.md) to set up a testing environment and running tests locally.
|
||||
|
||||
## DCO and Signed-off-by
|
||||
|
||||
When contributing changes to this project, you must agree to the DCO. Commits must include a `Signed-off-by:` header which certifies agreement with the terms of the DCO.
|
||||
When contributing changes to this project, you must agree to the DCO. Commits must include a `Signed-off-by:` header which certifies agreement with the terms of the DCO (Developer Certificate of Origin).
|
||||
|
||||
Using `-s` with `git commit` will automatically add this header.
|
||||
|
||||
@@ -88,8 +88,8 @@ Only specific types of PRs will be reviewed. The PR title is prefixed appropriat
|
||||
- `[Platform]` for new features or optimization in platform.
|
||||
- `[Worker]` for new features or optimization in worker.
|
||||
- `[Core]` for new features or optimization in the core vllm-ascend logic (such as platform, attention, communicators, model runner)
|
||||
- `[Kernel]` changes affecting compute kernels and ops.
|
||||
- `[Bugfix]` for bug fixes.
|
||||
- `[Kernel]` for changes affecting compute kernels and ops.
|
||||
- `[BugFix]` for bug fixes.
|
||||
- `[Doc]` for documentation fixes and improvements.
|
||||
- `[Test]` for tests (such as unit tests).
|
||||
- `[CI]` for build or continuous integration improvements.
|
||||
@@ -101,11 +101,15 @@ If the PR spans more than one category, please include all relevant prefixes.
|
||||
|
||||
## Others
|
||||
|
||||
You may find more information about contributing to vLLM Ascend backend plugin on [<u>docs.vllm.ai</u>](https://docs.vllm.ai/en/latest/contributing/overview.html).
|
||||
If you find any problem when contributing, you can feel free to submit a PR to improve the doc to help other developers.
|
||||
You may find more information about contributing to vLLM Ascend backend plugin on [<u>docs.vllm.ai</u>](https://docs.vllm.ai/en/latest/contributing).
|
||||
If you encounter any problems while contributing, feel free to submit a PR to improve the documentation to help other developers.
|
||||
|
||||
:::{toctree}
|
||||
:caption: Index
|
||||
:maxdepth: 1
|
||||
testing
|
||||
doc_writing
|
||||
multi_node_test
|
||||
nightly_ci_test
|
||||
e2e_ci_test
|
||||
:::
|
||||
|
||||
553
docs/source/developer_guide/contribution/multi_node_test.md
Normal file
553
docs/source/developer_guide/contribution/multi_node_test.md
Normal file
@@ -0,0 +1,553 @@
|
||||
# Multi Node Test
|
||||
|
||||
Multi-Node CI is designed to test distributed scenarios of very large models, for example, disaggregated_prefill multi DP across multi nodes and so on.
|
||||
|
||||
## How it works
|
||||
|
||||
The following picture shows the basic deployment view of the multi-node CI mechanism. It shows how the GitHub action interacts with [lws](https://lws.sigs.k8s.io/docs/overview/) (a kind of kubernetes crd resource).
|
||||
|
||||

|
||||
|
||||
From the workflow perspective, we can see how the final test script is executed. The key point is that the shared files `tests/e2e/nightly/multi_node/scripts/lws.yaml.jinja2` and `tests/e2e/nightly/multi_node/scripts/run.sh` define the cluster template and pod entry script. Each node executes different logic according to the [LWS_WORKER_INDEX](https://lws.sigs.k8s.io/docs/reference/labels-annotations-and-environment-variables/) environment variable, so that multiple nodes can form a distributed cluster to perform tasks. `run.sh` selects the pytest entrypoint from the config path: internal DP configs use `internal_dp/scripts/test_multi_node.py`, while external DP configs use `external_dp/scripts/test_external_dp.py`.
|
||||
|
||||

|
||||
|
||||
## How to contribute
|
||||
|
||||
1. Upload custom weights
|
||||
|
||||
If you need customized weights, for example, you quantized a w8a8 weight for DeepSeek-V3 and you want your weight to run on CI, uploading weights to ModelScope's [vllm-ascend](https://www.modelscope.cn/organization/vllm-ascend) organization is welcome. If you do not have permission to upload, please contact @Potabk
|
||||
|
||||
2. Add config yaml
|
||||
|
||||
For the normal internal DP multi-node flow, add the config yaml to `tests/e2e/nightly/multi_node/internal_dp/config/`, like `DeepSeek-V3.yaml`. External DP cases use the separate `tests/e2e/nightly/multi_node/external_dp/config/` directory and should pass that directory through `config_base_path` in workflow or `CONFIG_BASE_PATH` locally.
|
||||
|
||||
Suppose you have **2 nodes** running a 1P1D setup (1 Prefillers + 1 Decoder):
|
||||
|
||||
you may add a config file looks like:
|
||||
|
||||
```yaml
|
||||
test_name: "test DeepSeek-V3 disaggregated_prefill"
|
||||
# the model being tested
|
||||
model: "vllm-ascend/DeepSeek-V3-W8A8"
|
||||
# how large the cluster is
|
||||
num_nodes: 2
|
||||
npu_per_node: 16
|
||||
# All env vars you need should add it here
|
||||
env_common: &env_common
|
||||
VLLM_USE_MODELSCOPE: true
|
||||
OMP_PROC_BIND: false
|
||||
OMP_NUM_THREADS: 100
|
||||
HCCL_BUFFSIZE: 1024
|
||||
SERVER_PORT: 8080
|
||||
disaggregated_prefill:
|
||||
enabled: true
|
||||
# node index(a list) which meet all the conditions:
|
||||
# - prefiller
|
||||
# - no headless(have api server)
|
||||
prefiller_host_index: [0]
|
||||
# node index(a list) which meet all the conditions:
|
||||
# - decoder
|
||||
decoder_host_index: [1]
|
||||
|
||||
# Add each node's vllm serve cli command just like you run locally
|
||||
# Add each node's individual envs like follow
|
||||
deployment:
|
||||
- name: prefiller node # optional: just for description, not used in code
|
||||
envs:
|
||||
<<: *env_common
|
||||
VLLM_ASCEND_ENABLE_FLASHCOMM1: 1
|
||||
# Continue to add other envs if needed
|
||||
server_cmd: >
|
||||
vllm serve ...
|
||||
- name: decoder node # optional: just for description, not used in code
|
||||
envs:
|
||||
<<: *env_common
|
||||
VLLM_ASCEND_ENABLE_FLASHCOMM1: 1
|
||||
# Continue to add other envs if needed
|
||||
server_cmd: >
|
||||
vllm serve ...
|
||||
benchmarks:
|
||||
perf:
|
||||
# fill with performance test kwargs
|
||||
acc:
|
||||
# fill with accuracy test kwargs
|
||||
```
|
||||
|
||||
3. Add the case to nightly workflow
|
||||
|
||||
Currently, the multi-node test workflow is defined in `.github/workflows/schedule_nightly_test_a3.yaml`.
|
||||
|
||||
```yaml
|
||||
multi-node-tests:
|
||||
name: multi-node
|
||||
if: always() && (github.event_name == 'schedule' || github.event_name == 'workflow_dispatch')
|
||||
strategy:
|
||||
fail-fast: false
|
||||
max-parallel: 1
|
||||
matrix:
|
||||
test_config:
|
||||
- name: multi-node-deepseek-pd
|
||||
config_file_path: DeepSeek-V3.yaml
|
||||
size: 2
|
||||
- name: multi-node-qwen3-dp
|
||||
config_file_path: Qwen3-235B-A22B.yaml
|
||||
size: 2
|
||||
- name: GLM5_1-W8A8-EP-external
|
||||
config_file_path: GLM5_1-W8A8-EP-external.yaml
|
||||
config_base_path: tests/e2e/nightly/multi_node/external_dp/config/
|
||||
size: 4
|
||||
uses: ./.github/workflows/_e2e_nightly_multi_node.yaml
|
||||
with:
|
||||
soc_version: a3
|
||||
runner: linux-aarch64-a3-0
|
||||
image: 'swr.cn-southwest-2.myhuaweicloud.com/base_image/ascend-ci/vllm-ascend:nightly-a3'
|
||||
replicas: 1
|
||||
size: ${{ matrix.test_config.size }}
|
||||
config_file_path: ${{ matrix.test_config.config_file_path }}
|
||||
config_base_path: ${{ matrix.test_config.config_base_path || '' }}
|
||||
name: ${{ matrix.test_config.name }}
|
||||
secrets:
|
||||
KUBECONFIG_B64: ${{ secrets.KUBECONFIG_B64 }}
|
||||
```
|
||||
|
||||
The matrix above defines all the parameters required to add a multi-machine use
|
||||
case. The parameters worth noting are `size`, `config_file_path`, and
|
||||
`config_base_path`. `size` defines the number of nodes required for your use
|
||||
case. `config_file_path` is the yaml file name, and `config_base_path` tells the
|
||||
loader which config directory to use. For internal DP cases, use an empty
|
||||
`config_base_path` so the loader uses its default internal DP config directory.
|
||||
For external DP cases, set it to
|
||||
`tests/e2e/nightly/multi_node/external_dp/config/`.
|
||||
|
||||
## Run Multi-Node tests locally
|
||||
|
||||
### 1. Use kubernetes
|
||||
|
||||
This section assumes that you already have a [Kubernetes](https://kubernetes.io/docs/setup/) NPU cluster environment locally. Then you can easily start our test with one click.
|
||||
|
||||
- Step 1. Install LWS CRD resources
|
||||
|
||||
See <https://lws.sigs.k8s.io/docs/installation/> Which can be used as a reference
|
||||
|
||||
- Step 2. Deploy the following yaml file `lws.yaml` as needed
|
||||
|
||||
```yaml
|
||||
apiVersion: leaderworkerset.x-k8s.io/v1
|
||||
kind: LeaderWorkerSet
|
||||
metadata:
|
||||
name: test-server
|
||||
namespace: vllm-project
|
||||
spec:
|
||||
replicas: 1
|
||||
leaderWorkerTemplate:
|
||||
size: 2
|
||||
restartPolicy: None
|
||||
leaderTemplate:
|
||||
metadata:
|
||||
labels:
|
||||
role: leader
|
||||
spec:
|
||||
containers:
|
||||
- name: vllm-leader
|
||||
imagePullPolicy: Always
|
||||
image: swr.cn-southwest-2.myhuaweicloud.com/base_image/ascend-ci/vllm-ascend:nightly-a3
|
||||
env:
|
||||
- name: CONFIG_YAML_PATH
|
||||
value: DeepSeek-V3.yaml
|
||||
- name: CONFIG_BASE_PATH
|
||||
value: tests/e2e/nightly/multi_node/internal_dp/config/
|
||||
- name: WORKSPACE
|
||||
value: "/vllm-workspace"
|
||||
- name: FAIL_TAG
|
||||
value: FAIL_TAG
|
||||
command:
|
||||
- sh
|
||||
- -c
|
||||
- |
|
||||
bash /vllm-workspace/vllm-ascend/tests/e2e/nightly/multi_node/scripts/run.sh
|
||||
resources:
|
||||
limits:
|
||||
huawei.com/ascend-1980: 16
|
||||
memory: 512Gi
|
||||
ephemeral-storage: 100Gi
|
||||
requests:
|
||||
huawei.com/ascend-1980: 16
|
||||
memory: 512Gi
|
||||
ephemeral-storage: 100Gi
|
||||
cpu: 125
|
||||
ports:
|
||||
- containerPort: 8080
|
||||
# readinessProbe:
|
||||
# tcpSocket:
|
||||
# port: 8080
|
||||
# initialDelaySeconds: 15
|
||||
# periodSeconds: 10
|
||||
volumeMounts:
|
||||
- mountPath: /root/.cache
|
||||
name: shared-volume
|
||||
- mountPath: /usr/local/Ascend/driver/tools
|
||||
name: driver-tools
|
||||
- mountPath: /dev/shm
|
||||
name: dshm
|
||||
volumes:
|
||||
- name: dshm
|
||||
emptyDir:
|
||||
medium: Memory
|
||||
sizeLimit: 15Gi
|
||||
- name: shared-volume
|
||||
persistentVolumeClaim:
|
||||
claimName: nv-action-vllm-benchmarks-v2
|
||||
- name: driver-tools
|
||||
hostPath:
|
||||
path: /usr/local/Ascend/driver/tools
|
||||
workerTemplate:
|
||||
spec:
|
||||
containers:
|
||||
- name: vllm-worker
|
||||
imagePullPolicy: Always
|
||||
image: swr.cn-southwest-2.myhuaweicloud.com/base_image/ascend-ci/vllm-ascend:nightly-a3
|
||||
env:
|
||||
- name: CONFIG_YAML_PATH
|
||||
value: DeepSeek-V3.yaml
|
||||
- name: CONFIG_BASE_PATH
|
||||
value: tests/e2e/nightly/multi_node/internal_dp/config/
|
||||
- name: WORKSPACE
|
||||
value: "/vllm-workspace"
|
||||
- name: FAIL_TAG
|
||||
value: FAIL_TAG
|
||||
command:
|
||||
- sh
|
||||
- -c
|
||||
- |
|
||||
bash /vllm-workspace/vllm-ascend/tests/e2e/nightly/multi_node/scripts/run.sh
|
||||
resources:
|
||||
limits:
|
||||
huawei.com/ascend-1980: 16
|
||||
memory: 512Gi
|
||||
ephemeral-storage: 100Gi
|
||||
requests:
|
||||
huawei.com/ascend-1980: 16
|
||||
ephemeral-storage: 100Gi
|
||||
cpu: 125
|
||||
volumeMounts:
|
||||
- mountPath: /root/.cache
|
||||
name: shared-volume
|
||||
- mountPath: /usr/local/Ascend/driver/tools
|
||||
name: driver-tools
|
||||
- mountPath: /dev/shm
|
||||
name: dshm
|
||||
volumes:
|
||||
- name: dshm
|
||||
emptyDir:
|
||||
medium: Memory
|
||||
sizeLimit: 15Gi
|
||||
- name: shared-volume
|
||||
persistentVolumeClaim:
|
||||
claimName: nv-action-vllm-benchmarks-v2
|
||||
- name: driver-tools
|
||||
hostPath:
|
||||
path: /usr/local/Ascend/driver/tools
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: Service
|
||||
metadata:
|
||||
name: vllm-leader
|
||||
namespace: vllm-project
|
||||
spec:
|
||||
ports:
|
||||
- name: http
|
||||
port: 8080
|
||||
protocol: TCP
|
||||
targetPort: 8080
|
||||
selector:
|
||||
leaderworkerset.sigs.k8s.io/name: vllm
|
||||
role: leader
|
||||
type: ClusterIP
|
||||
```
|
||||
|
||||
```bash
|
||||
kubectl apply -f lws.yaml
|
||||
```
|
||||
|
||||
Verify the status of the pods:
|
||||
|
||||
```bash
|
||||
kubectl get pods -n vllm-project
|
||||
```
|
||||
|
||||
Should get an output similar to this:
|
||||
|
||||
```bash
|
||||
NAME READY STATUS RESTARTS AGE
|
||||
vllm-0 1/1 Running 0 2s
|
||||
vllm-0-1 1/1 Running 0 2s
|
||||
```
|
||||
|
||||
Verify that the distributed inference works:
|
||||
|
||||
```bash
|
||||
kubectl logs -f vllm-0 -n vllm-project
|
||||
```
|
||||
|
||||
Should get something similar to this:
|
||||
|
||||
```shell
|
||||
INFO 12-30 11:00:57 [__init__.py:43] Available plugins for group vllm.platform_plugins:
|
||||
INFO 12-30 11:00:57 [__init__.py:45] - ascend -> vllm_ascend:register
|
||||
INFO 12-30 11:00:57 [__init__.py:48] All plugins in this group will be loaded. Set `VLLM_PLUGINS` to control which plugins to load.
|
||||
INFO 12-30 11:00:57 [__init__.py:217] Platform plugin ascend is activated
|
||||
INFO 12-30 11:00:57 [importing.py:68] Triton not installed or not compatible; certain GPU-related functions will not be available.
|
||||
================================================================================================== test session starts ===================================================================================================
|
||||
platform linux -- Python 3.12.13, pytest-8.4.2, pluggy-1.6.0 -- /usr/local/python3.12.13/bin/python3
|
||||
cachedir: .pytest_cache
|
||||
rootdir: /vllm-workspace/vllm-ascend
|
||||
configfile: pyproject.toml
|
||||
plugins: cov-7.0.0, asyncio-1.3.0, mock-3.15.1, anyio-4.12.0
|
||||
asyncio: mode=Mode.STRICT, debug=False, asyncio_default_fixture_loop_scope=None, asyncio_default_test_loop_scope=function
|
||||
collected 1 item
|
||||
|
||||
tests/e2e/nightly/multi_node/internal_dp/scripts/test_multi_node.py::test_multi_node [2025-12-30 11:01:01] INFO multi_node_config.py:294: Loading config yaml: tests/e2e/nightly/multi_node/internal_dp/config/DeepSeek-V3.yaml
|
||||
[2025-12-30 11:01:01] INFO multi_node_config.py:348: Resolving cluster IPs via DNS...
|
||||
[2025-12-30 11:01:01] INFO multi_node_config.py:212: Node 0 envs: {'VLLM_USE_MODELSCOPE': 'True', 'OMP_PROC_BIND': 'False', 'OMP_NUM_THREADS': '100', 'HCCL_BUFFSIZE': '1024', 'SERVER_PORT': '8080', 'NUMEXPR_MAX_THREADS': '128', 'DISAGGREGATED_PREFILL_PROXY_SCRIPT': 'examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py', 'HCCL_IF_IP': '10.0.0.102', 'HCCL_SOCKET_IFNAME': 'eth0', 'GLOO_SOCKET_IFNAME': 'eth0', 'TP_SOCKET_IFNAME': 'eth0', 'LOCAL_IP': '10.0.0.102', 'NIC_NAME': 'eth0', 'MASTER_IP': '10.0.0.102'}
|
||||
[2025-12-30 11:01:01] INFO multi_node_config.py:159: Launching proxy: python examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py --host 10.0.0.102 --port 6000 --prefiller-hosts 10.0.0.102 --prefiller-ports 8080 --decoder-hosts 10.0.0.138 --decoder-ports 8080
|
||||
[2025-12-30 11:01:01] INFO conftest.py:107: Starting server with command: vllm serve vllm-ascend/DeepSeek-V3-W8A8 --host 0.0.0.0 --port 8080 --data-parallel-size 2 --data-parallel-size-local 2 --tensor-parallel-size 8 --seed 1024 --enforce-eager --enable-expert-parallel --max-num-seqs 16 --max-model-len 8192 --max-num-batched-tokens 8192 --quantization ascend --trust-remote-code --no-enable-prefix-caching --gpu-memory-utilization 0.9 --kv-transfer-config {"kv_connector": "MooncakeConnectorV1", "kv_role": "kv_producer", "kv_port": "30000",
|
||||
"kv_connector_extra_config": {
|
||||
"prefill": {
|
||||
"dp_size": 2,
|
||||
"tp_size": 8
|
||||
},
|
||||
"decode": {
|
||||
"dp_size": 2,
|
||||
"tp_size": 8
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### 2. Test without Kubernetes
|
||||
|
||||
The same `tests/e2e/nightly/multi_node/scripts/run.sh` entrypoint can be used
|
||||
on prepared bare-metal or container hosts. Without LWS, set the values that
|
||||
Kubernetes normally injects yourself:
|
||||
|
||||
- `cluster_hosts` in the config yaml, using IPs reachable from every node.
|
||||
- `LWS_WORKER_INDEX` on each node, starting from `0`.
|
||||
- `CONFIG_YAML_PATH` as the config file name and `CONFIG_BASE_PATH` as the
|
||||
config directory.
|
||||
|
||||
Use the host NIC IPs that can reach each other, for example addresses shown by
|
||||
`ip addr` or `ifconfig` on the active network interface. Do not use per-host
|
||||
Docker bridge addresses such as `172.17.0.1`, because each host has its own
|
||||
local bridge.
|
||||
|
||||
Local `cluster_hosts` edits should be removed before submitting a PR unless the
|
||||
hosts are part of a committed test environment.
|
||||
|
||||
#### 2.1 Internal DP local run
|
||||
|
||||
##### 2.1.1 Add cluster hosts
|
||||
|
||||
Edit the internal DP config you want to run, for example:
|
||||
|
||||
```text
|
||||
tests/e2e/nightly/multi_node/internal_dp/config/DeepSeek-V3.yaml
|
||||
```
|
||||
|
||||
Add `cluster_hosts` as a top-level field, for example near `num_nodes` and
|
||||
`npu_per_node`:
|
||||
|
||||
```yaml
|
||||
cluster_hosts:
|
||||
- "172.22.0.xxx"
|
||||
- "172.22.0.xxx"
|
||||
```
|
||||
|
||||
##### 2.1.2 Prepare the environment
|
||||
|
||||
Install vllm-ascend development dependencies on every cluster host:
|
||||
|
||||
```bash
|
||||
cd /vllm-workspace/vllm-ascend
|
||||
python3 -m pip install -r requirements-dev.txt
|
||||
```
|
||||
|
||||
Install AISBench on the first host, which is the node with
|
||||
`LWS_WORKER_INDEX=0`:
|
||||
|
||||
```bash
|
||||
export AIS_BENCH_TAG="v3.1-20260330-master"
|
||||
export AIS_BENCH_URL="https://github.com/AISBench/benchmark.git"
|
||||
export BENCHMARK_HOME=/vllm-workspace/vllm-ascend/benchmark
|
||||
|
||||
git clone -b ${AIS_BENCH_TAG} --depth 1 ${AIS_BENCH_URL} $BENCHMARK_HOME
|
||||
cd $BENCHMARK_HOME
|
||||
pip install -e . -r requirements/api.txt -r requirements/extra.txt
|
||||
```
|
||||
|
||||
If your local image already contains the model, benchmark data, Ascend runtime,
|
||||
and AISBench, you only need the run-time exports in the next step.
|
||||
|
||||
##### 2.1.3 Start each node
|
||||
|
||||
Run the script on each node separately. Start worker nodes first, then start
|
||||
node 0.
|
||||
|
||||
On node 1:
|
||||
|
||||
```bash
|
||||
export WORKSPACE=/vllm-workspace
|
||||
export IS_PR_TEST=false
|
||||
export CONFIG_YAML_PATH=DeepSeek-V3.yaml
|
||||
export CONFIG_BASE_PATH=tests/e2e/nightly/multi_node/internal_dp/config/
|
||||
export LWS_WORKER_INDEX=1
|
||||
|
||||
cd $WORKSPACE/vllm-ascend
|
||||
bash tests/e2e/nightly/multi_node/scripts/run.sh
|
||||
```
|
||||
|
||||
On node 0:
|
||||
|
||||
```bash
|
||||
export WORKSPACE=/vllm-workspace
|
||||
export IS_PR_TEST=false
|
||||
export CONFIG_YAML_PATH=DeepSeek-V3.yaml
|
||||
export CONFIG_BASE_PATH=tests/e2e/nightly/multi_node/internal_dp/config/
|
||||
export LWS_WORKER_INDEX=0
|
||||
|
||||
cd $WORKSPACE/vllm-ascend
|
||||
bash tests/e2e/nightly/multi_node/scripts/run.sh
|
||||
```
|
||||
|
||||
Internal DP logs are mainly printed to the terminal running `run.sh`. When
|
||||
`LOG_PREFIX` is set, the shared script also backs up Ascend logs to:
|
||||
|
||||
```text
|
||||
$LOG_PREFIX/node_<LWS_WORKER_INDEX>_plogs/
|
||||
```
|
||||
|
||||
#### 2.2 External DP local run
|
||||
|
||||
##### 2.2.1 Add cluster hosts
|
||||
|
||||
Edit the external DP config you want to run. For example:
|
||||
|
||||
```text
|
||||
tests/e2e/nightly/multi_node/external_dp/config/GLM5_1-W8A8-EP-external.yaml
|
||||
```
|
||||
|
||||
Add `cluster_hosts` as a top-level field, for example near `num_nodes` and
|
||||
`npu_per_node`:
|
||||
|
||||
```yaml
|
||||
cluster_hosts:
|
||||
- "172.22.0.xxx"
|
||||
- "172.22.0.xxx"
|
||||
- "172.22.0.xxx"
|
||||
- "172.22.0.xxx"
|
||||
```
|
||||
|
||||
##### 2.2.2 Prepare the environment
|
||||
|
||||
Install vllm-ascend development dependencies on every cluster host:
|
||||
|
||||
```bash
|
||||
cd /vllm-workspace/vllm-ascend
|
||||
python3 -m pip install -r requirements-dev.txt
|
||||
```
|
||||
|
||||
Install AISBench on node 0:
|
||||
|
||||
```bash
|
||||
export AIS_BENCH_TAG="v3.1-20260330-master"
|
||||
export AIS_BENCH_URL="https://github.com/AISBench/benchmark.git"
|
||||
export BENCHMARK_HOME=/vllm-workspace/vllm-ascend/benchmark
|
||||
|
||||
git clone -b ${AIS_BENCH_TAG} --depth 1 ${AIS_BENCH_URL} $BENCHMARK_HOME
|
||||
cd $BENCHMARK_HOME
|
||||
pip install -e . -r requirements/api.txt -r requirements/extra.txt
|
||||
```
|
||||
|
||||
If your local image already contains the model, benchmark data, Ascend runtime,
|
||||
and AISBench, you only need the run-time exports in the next step.
|
||||
|
||||
##### 2.2.3 Start each node
|
||||
|
||||
External DP uses the same shared `run.sh`. Set `CONFIG_BASE_PATH` to the
|
||||
external DP config directory so the script chooses
|
||||
`external_dp/scripts/test_external_dp.py`.
|
||||
|
||||
Then start non-master nodes first, and start node 0 last. The following example
|
||||
uses `GLM5_1-W8A8-EP-external.yaml`, which is a 4-node disaggregated prefill
|
||||
case.
|
||||
|
||||
On node 1, node 2, and node 3, set the matching `LWS_WORKER_INDEX`:
|
||||
|
||||
```bash
|
||||
export WORKSPACE=/vllm-workspace
|
||||
export IS_PR_TEST=false
|
||||
export CONFIG_BASE_PATH=tests/e2e/nightly/multi_node/external_dp/config/
|
||||
export CONFIG_YAML_PATH=GLM5_1-W8A8-EP-external.yaml
|
||||
export LWS_WORKER_INDEX=1 # Use 2 on node 2, and 3 on node 3.
|
||||
|
||||
cd $WORKSPACE/vllm-ascend
|
||||
bash tests/e2e/nightly/multi_node/scripts/run.sh
|
||||
```
|
||||
|
||||
On node 0:
|
||||
|
||||
```bash
|
||||
export WORKSPACE=/vllm-workspace
|
||||
export IS_PR_TEST=false
|
||||
export CONFIG_BASE_PATH=tests/e2e/nightly/multi_node/external_dp/config/
|
||||
export CONFIG_YAML_PATH=GLM5_1-W8A8-EP-external.yaml
|
||||
export LWS_WORKER_INDEX=0
|
||||
|
||||
cd $WORKSPACE/vllm-ascend
|
||||
bash tests/e2e/nightly/multi_node/scripts/run.sh
|
||||
```
|
||||
|
||||
For `GLM5_1-W8A8-EP-external.yaml`, node 0 and node 1 start prefiller ranks,
|
||||
node 2 and node 3 start decoder ranks, and node 0 also starts the proxy and
|
||||
benchmark.
|
||||
|
||||
##### 2.2.4 Read logs while the test is running
|
||||
|
||||
The terminal running `run.sh` prints pytest orchestration logs. For external DP,
|
||||
AISBench output is also printed on node 0, while rank and proxy stdout/stderr
|
||||
are written to `EXTERNAL_DP_LOG_DIR`. The default layout is:
|
||||
|
||||
```text
|
||||
/tmp/external_dp_logs/
|
||||
node-0/
|
||||
rank-0.log
|
||||
rank-1.log
|
||||
proxy.log
|
||||
node-1/
|
||||
rank-0.log
|
||||
rank-1.log
|
||||
```
|
||||
|
||||
The first line of each rank log records the exact command and environment used
|
||||
to start that rank. `proxy.log` exists only on the configured proxy node,
|
||||
usually node 0.
|
||||
|
||||
Use a separate log directory when running multiple local experiments:
|
||||
|
||||
```bash
|
||||
export EXTERNAL_DP_LOG_DIR=/tmp/external_dp_logs_pd_local
|
||||
```
|
||||
|
||||
To watch logs in real time, run these commands in another terminal on the
|
||||
corresponding node:
|
||||
|
||||
```bash
|
||||
# node 0: ranks and proxy
|
||||
tail -F /tmp/external_dp_logs/node-0/rank-0.log \
|
||||
/tmp/external_dp_logs/node-0/rank-1.log \
|
||||
/tmp/external_dp_logs/node-0/proxy.log
|
||||
|
||||
# node 1: ranks
|
||||
tail -F /tmp/external_dp_logs/node-1/rank-0.log \
|
||||
/tmp/external_dp_logs/node-1/rank-1.log
|
||||
```
|
||||
317
docs/source/developer_guide/contribution/nightly_ci_test.md
Normal file
317
docs/source/developer_guide/contribution/nightly_ci_test.md
Normal file
@@ -0,0 +1,317 @@
|
||||
# Nightly CI Test
|
||||
|
||||
This document explains how to trigger nightly hardware CI tests against your own PR code
|
||||
on Ascend NPU hardware (A2/A3), without waiting for the scheduled nightly run.
|
||||
|
||||
## Background
|
||||
|
||||
By default, nightly CI tests run on a fixed schedule using pre-built nightly images.
|
||||
Contributors can self-service trigger these tests directly against their PR changes
|
||||
by combining a GitHub label with a comment command.
|
||||
|
||||
## How to Trigger
|
||||
|
||||
### 1. Post a comment
|
||||
|
||||
Post one of the following comments in the PR to specify which tests to run.
|
||||
The comment itself triggers the workflow — no label is required.
|
||||
|
||||
| Comment | Effect |
|
||||
|---------|--------|
|
||||
| `/nightly` | Run **all** nightly tests |
|
||||
| `/nightly all` | Run **all** nightly tests (same as above) |
|
||||
| `/nightly test1 test2 ...` | Run only the **named** tests |
|
||||
|
||||
:::{note}
|
||||
Only repository **Contributors** (Triage role) and **Maintainers** (Write role) can
|
||||
trigger the `/nightly` command. If you do not have this permission, ask a maintainer
|
||||
to post the comment for you. You can find the list of maintainers and contributors in
|
||||
the project's [Governance](../../community/governance.md) page or by checking the
|
||||
[CODEOWNERS](https://github.com/vllm-project/vllm-ascend/blob/main/.github/CODEOWNERS)
|
||||
file.
|
||||
:::
|
||||
|
||||
### 2. Wait for results
|
||||
|
||||
GitHub Actions will trigger the `Nightly-A2` or `Nightly-A3` workflow. Only tests
|
||||
matching the filter will be dispatched, which saves hardware resources.
|
||||
|
||||
## Differences Between PR and Scheduled Runs
|
||||
|
||||
| | Scheduled / Manual Dispatch | PR-triggered |
|
||||
|---|----------------------------|---|
|
||||
| Trigger | Cron (daily) or `workflow_dispatch` | `/nightly` comment |
|
||||
| Code tested | Pre-built nightly image | Your PR's HEAD commit (source installed fresh) |
|
||||
| Test scope | All tests | Configurable via `/nightly <names>` |
|
||||
| vLLM + vllm-ascend | From image | Checked out and installed from source |
|
||||
| Test matrix | From main branch's matrix YAML | From PR branch's matrix YAML |
|
||||
|
||||
When a PR run is detected (`is_pr_test: true`), the workflow additionally:
|
||||
|
||||
1. Uninstalls any existing vllm packages in the container.
|
||||
2. Checks out the specific vllm version and your PR's vllm-ascend commit from source.
|
||||
3. Installs all dependencies from source.
|
||||
4. Installs the `aisbench` benchmark suite.
|
||||
|
||||
## Test Matrix Data Source
|
||||
|
||||
The set of nightly test cases (their names, runners, test paths, model configs) is
|
||||
declared in a single data file:
|
||||
|
||||
```text
|
||||
.github/workflows/configs/nightly_config.yaml
|
||||
```
|
||||
|
||||
The file is organized as `a2:` and `a3:` top-level keys (one per SoC). Under each
|
||||
SoC, tests are grouped by execution shape (single-node, multi-node, double-node,
|
||||
multi-card, accuracy) and each group holds a `test_config` (or `nightly` / `pr_only`
|
||||
for accuracy) list whose entries carry a `name` plus the fields consumed by the
|
||||
downstream reusable workflows (`os`, `tests`, `config_file_path`, `size`, etc.).
|
||||
|
||||
Both the `Nightly-A2` and `Nightly-A3` workflows dynamically read this file at run
|
||||
time — there is no hardcoded test matrix in the workflow YAMLs. The
|
||||
`/nightly <name>` slash command resolves names by walking the same file from the
|
||||
PR branch, so newly added entries can be exercised on a PR before they land on
|
||||
main.
|
||||
|
||||
## Adding a New Nightly Test Case
|
||||
|
||||
To add a new test case (no need to touch the workflow YAMLs):
|
||||
|
||||
1. Append an entry under the appropriate section in
|
||||
`.github/workflows/configs/nightly_config.yaml`. Each entry needs at least:
|
||||
- `name`: unique identifier used in `/nightly <name>` filters
|
||||
- `os` (for single-node / multi-card pytest+yaml tests) or `runner` is inferred
|
||||
- one of `tests:` (pytest directory) or `config_file_path:` (YAML-driven model config)
|
||||
- `size` (multi-node / double-node only)
|
||||
2. Add the actual test files (pytest modules under `tests/e2e/nightly/...` or
|
||||
YAML model configs in `tests/e2e/nightly/.../configs/`).
|
||||
3. Open a PR. Once CI is green, you can validate the new entry against real NPU
|
||||
hardware **without** merging the PR — see *Examples* below.
|
||||
|
||||
## Available Test Names
|
||||
|
||||
The test names you can pass to `/nightly` correspond to the `name` fields under
|
||||
the matching section in `.github/workflows/configs/nightly_config.yaml`. The
|
||||
tables below mirror the current contents of that file.
|
||||
|
||||
### A2 workflow (`.github/workflows/schedule_nightly_test_a2.yaml`)
|
||||
|
||||
**Single-node tests** (`a2.single_node.test_config`):
|
||||
|
||||
| Test name | Description |
|
||||
|-----------|-------------|
|
||||
| `test_custom_op_multi_card` | Custom operator tests (multi card) |
|
||||
| `qwen3-vl-32b-instruct-w8a8` | Qwen3-VL-32B-Instruct W8A8 |
|
||||
| `qwen3-32b-int8` | Qwen3-32B INT8 quantization |
|
||||
| `Qwen3.5-27B-w8a8-A2` | Qwen3.5-27B W8A8 |
|
||||
| `Qwen3.5-397B-A17B-w4a8-mtp` | Qwen3.5-397B-A17B W4A8 + MTP |
|
||||
|
||||
**Multi-node tests** (`a2.multi_node.test_config`):
|
||||
|
||||
| Test name | Description |
|
||||
|-----------|-------------|
|
||||
| `multi-node-qwen3-235b-dp` | Qwen3-235B-A22B, 2-node DP |
|
||||
| `multi-node-GLM-5.1-w8a8-A2` | GLM-5.1 W8A8, 2 nodes |
|
||||
| `multi-node-Kimi-K2.5-W4A8-A2` | Kimi-K2.5 W4A8, 2 nodes |
|
||||
|
||||
**Accuracy tests** (`a2.accuracy.nightly` and `a2.accuracy.pr_only`):
|
||||
|
||||
| Test name | Description | Scope |
|
||||
|-----------|-------------|-------|
|
||||
| `accuracy-group-1` | Qwen3-VL-8B, Qwen3-8B, Qwen2-Audio-7B, etc. | nightly |
|
||||
| `accuracy-group-2` | ERNIE-4.5, Molmo-7B, Llama-3.2-3B, etc. | nightly |
|
||||
| `accuracy-group-3` | Qwen3-30B-A3B, Qwen3-VL-30B-A3B, etc. | nightly |
|
||||
| `accuracy-group-4` | Qwen3-Next-80B-A3B, Qwen3-Omni-30B-A3B, etc. | nightly |
|
||||
| `pr-accuracy-group-1` | gemma-3-4b-it, internlm3-8b-instruct, etc. | pr_only |
|
||||
| `pr-accuracy-group-2` | Qwen2.5-Math-RM-72B, Hunyuan-A13B-Instruct | pr_only |
|
||||
|
||||
The `pr-accuracy-group-*` entries only run on `/nightly` (PR-triggered) runs;
|
||||
`/nightly all` on the schedule skips them.
|
||||
|
||||
### A3 workflow (`.github/workflows/schedule_nightly_test_a3.yaml`)
|
||||
|
||||
**Multi-node tests** (`a3.multi_node.test_config`, 4-node):
|
||||
|
||||
| Test name | Description |
|
||||
|-----------|-------------|
|
||||
| `multi-node-deepseek-v3.2-W8A8-EP` | DeepSeek-V3.2-W8A8 with EP, 4-node |
|
||||
|
||||
**Double-node tests** (`a3.double_node.test_config`, 2-node, run after multi-node):
|
||||
|
||||
| Test name | Description |
|
||||
|-----------|-------------|
|
||||
| `multi-node-deepseek-r1-w8a8-longseq` | DeepSeek-R1-W8A8 long sequence, 2-node |
|
||||
| `multi-node-qwen3-dp` | Qwen3-235B-A22B, 2-node DP |
|
||||
| `multi-node-qwenw8a8-2node-eplb` | Qwen3-235B-W8A8 with EPLB, 2-node |
|
||||
| `multi-node-dpsk3.2-2node` | DeepSeek-V3.2-W8A8, 2-node |
|
||||
| `multi-node-qwenw8a8-2node-longseq` | Qwen3-235B-W8A8 long sequence, 2-node |
|
||||
| `multi-node-qwen-disagg-pd` | Qwen3-235B disaggregated PD, 2-node |
|
||||
| `multi-node-qwen-vl-disagg-pd` | Qwen3-VL-235B disaggregated PD, 2-node |
|
||||
| `multi-node-deepseek-v3.1` | DeepSeek-V3.1-BF16, 2-node |
|
||||
| `multi-node-deepseek-v3.2-W8A8-EP` | DeepSeek-V3.2-W8A8 with EP, 4-node |
|
||||
| `multi-node-glm-5.2` | GLM-5.1-W8A8, 2-node |
|
||||
|
||||
**Single-node tests** (`a3.single_node.test_config`):
|
||||
|
||||
| Test name | Description |
|
||||
|-----------|-------------|
|
||||
| `mtpx-deepseek-r1-0528-w8a8` | MTP-X + DeepSeek-R1-0528-W8A8 |
|
||||
| `deepseek-r1-0528-w8a8` | DeepSeek-R1-0528-W8A8 |
|
||||
| `kimi-k2-thinking` | Kimi-K2-Thinking |
|
||||
| `qwen3-vl-235b-a22b-instruct-w8a8` | Qwen3-VL-235B-A22B-Instruct-W8A8 |
|
||||
| `deepseek-r1-0528-w8a8-prefix-cache` | DeepSeek-R1-0528-W8A8 prefix cache |
|
||||
| `deepseek-v3-2-w8a8` | DeepSeek-V3.2-W8A8 |
|
||||
| `glm-4.7-w8a8` | GLM-4.7 W8A8 |
|
||||
| `kimi-k2.5` | Kimi-K2.5 |
|
||||
| `qwen3-235b-a22b-w8a8` | Qwen3-235B-A22B-W8A8 |
|
||||
| `Qwen3.5-397B-A17B-w8a8-mtp` | Qwen3.5-397B-A17B W8A8 + MTP |
|
||||
| `MiniMax-M2.5-w8a8-QuaRot-A3` | MiniMax-M2.5 W8A8 + QuaRot |
|
||||
| `Qwen3.5-27B-w8a8-A3` | Qwen3.5-27B W8A8 |
|
||||
| `Qwen3.5-122B-A10B-W8A8-A3` | Qwen3.5-122B-A10B W8A8 |
|
||||
| `DeepSeek-V4-Flash-W8A8-A3` | DeepSeek-V4-Flash W8A8 |
|
||||
|
||||
**Multi-card tests** (`a3.multi_card.test_config`):
|
||||
|
||||
| Test name | Description |
|
||||
|-----------|-------------|
|
||||
| `qwen3-30b-acc` | Qwen3-30B accuracy test |
|
||||
| `qwen3-30b-a3b-w8a8` | Qwen3-30B-A3B-W8A8 |
|
||||
| `qwen3-32b-int8` | Qwen3-32B-Int8 |
|
||||
| `qwen3-32b-int8-prefix-cache` | Qwen3-32B-Int8 prefix cache |
|
||||
| `Qwen3-30B-A3B-W4A8-llm-compressor` | Qwen3-30B-A3B W4A8 via llm-compressor |
|
||||
| `Qwen3-30B-QuaRot` | Qwen3-30B QuaRot + eagle3 |
|
||||
| `Qwen3-32B-QuaRot` | Qwen3-32B QuaRot + eagle3 |
|
||||
|
||||
:::{warning}
|
||||
The A3 resource pool has a maximum concurrency of **5×16 NPUs**. Multi-node tests
|
||||
run with `max-parallel: 2` to avoid resource exhaustion. Running `/nightly all` on
|
||||
A3 will queue a large number of jobs — prefer targeting specific test names when
|
||||
possible.
|
||||
:::
|
||||
|
||||
## Examples
|
||||
|
||||
Run all available nightly tests against your PR:
|
||||
|
||||
```text
|
||||
/nightly
|
||||
```
|
||||
|
||||
Run only the custom operator multi-card test:
|
||||
|
||||
```text
|
||||
/nightly test_custom_op_multi_card
|
||||
```
|
||||
|
||||
Run two specific tests at once (one per SoC):
|
||||
|
||||
```text
|
||||
/nightly test_custom_op_multi_card mtpx-deepseek-r1-0528-w8a8
|
||||
```
|
||||
|
||||
Run a single accuracy group (with all of its models):
|
||||
|
||||
```text
|
||||
/nightly accuracy-group-1
|
||||
```
|
||||
|
||||
Run a single accuracy model (only that model from a group):
|
||||
|
||||
```text
|
||||
/nightly accuracy-group-1/Qwen3-8B
|
||||
```
|
||||
|
||||
Re-trigger after fixing an issue: just push a new commit. The `synchronize` event
|
||||
re-runs the workflow and picks up the existing `/nightly` comment automatically — no
|
||||
need to post a new comment.
|
||||
|
||||
## Adding a New Test Case — Worked Example
|
||||
|
||||
To add `my-new-test` to the A2 single-node section:
|
||||
|
||||
1. Edit `.github/workflows/configs/nightly_config.yaml`, append under
|
||||
`a2.single_node.test_config`:
|
||||
|
||||
```yaml
|
||||
- name: my-new-test
|
||||
os: linux-aarch64-a2b3-4
|
||||
tests: tests/e2e/nightly/single_node/ops/multicard_ops_a2/test_my_new.py
|
||||
```
|
||||
|
||||
2. Commit the new pytest file (`test_my_new.py`) in the same PR.
|
||||
|
||||
3. Trigger from the PR:
|
||||
|
||||
```text
|
||||
/nightly my-new-test
|
||||
```
|
||||
|
||||
The workflow will:
|
||||
|
||||
- `pr_nightly_command.yml` reads your PR's `nightly_config.yaml` and resolves
|
||||
`my-new-test` → dispatch A2 only.
|
||||
- `Nightly-A2` is dispatched at `main`, but `generate-a2-matrix` checks out your
|
||||
PR commit and reads the new entry from the matrix.
|
||||
- `single-node-tests` runs one matrix job for `my-new-test`, with
|
||||
`should_run=true`. The reusable workflow checks out your PR code (via
|
||||
`vllm_ascend_ref`) and runs your pytest.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
**The workflow didn't start after I posted the comment.**
|
||||
|
||||
- Check that the comment starts exactly with `/nightly` with no leading spaces or
|
||||
extra characters before the slash.
|
||||
- Confirm you have at least Triage permission on the repository; unauthorized
|
||||
users' comments are ignored.
|
||||
- To re-trigger after fixing an issue, simply push a new commit — the workflow will
|
||||
reuse the existing `/nightly` comment automatically.
|
||||
|
||||
**Only some tests ran, not the ones I expected.**
|
||||
|
||||
- Test names are case-sensitive and must match the `name` field in
|
||||
`.github/workflows/configs/nightly_config.yaml` exactly (see the tables above).
|
||||
- For a PR-triggered run, the matrix is loaded from your PR's
|
||||
`nightly_config.yaml`, not main. If a name isn't in your PR's file, it won't
|
||||
be recognized and the dispatch will be skipped.
|
||||
- Check the `parse-trigger` job output in GitHub Actions for the resolved
|
||||
`test_filter` value.
|
||||
|
||||
**The workflow ran with the scheduled image, not my PR code.**
|
||||
|
||||
- Confirm the workflow was triggered by `repository_dispatch` (slash command),
|
||||
not bare `workflow_dispatch`. The `pr_nightly_command.yml` workflow is what
|
||||
actually dispatches `schedule_nightly_test_a2.yaml` / `_a3.yaml` with
|
||||
`vllm_ascend_ref` pointing at your PR SHA.
|
||||
|
||||
**A new test I added isn't being recognized.**
|
||||
|
||||
- Confirm the entry is well-formed YAML under
|
||||
`.github/workflows/configs/nightly_config.yaml`. The `name` field is required
|
||||
and must be unique within the SoC's section.
|
||||
- The matrix is loaded from your PR branch, so make sure the file is committed
|
||||
to the same branch the `/nightly` comment was posted on.
|
||||
|
||||
**How to obtain more detailed logs to pinpoint problems for multi-node tests**
|
||||
|
||||
- For most issues, the stdout pop-up logs from GitHub actions are sufficient (this log always represents the logs from the first node).
|
||||
- If the logs from a first node are no longer sufficient to provide effective logging information, see the summary of your jobs to download log archive for the corresponding test, which includes the framework-side logs and plog information for each node, structured as follows:
|
||||
|
||||
```shell
|
||||
.
|
||||
├── node0
|
||||
│ ├── root
|
||||
│ │ └── ascend
|
||||
│ │ └── log
|
||||
│ └── var
|
||||
│ └── log
|
||||
│ └── vllm-deepseek-v3-0f233d-0_logs.txt
|
||||
└── node1
|
||||
├── root
|
||||
│ └── ascend
|
||||
│ └── log
|
||||
└── var
|
||||
└── log
|
||||
└── vllm-deepseek-v3-0f233d-0-1_logs.txt
|
||||
```
|
||||
@@ -1,10 +1,10 @@
|
||||
# Testing
|
||||
|
||||
This secition explains how to write e2e tests and unit tests to verify the implementation of your feature.
|
||||
This document explains how to write unit tests, E2E tests, and nightly tests to verify your feature implementation.
|
||||
|
||||
## Setup test environment
|
||||
## Set up a test environment
|
||||
|
||||
The fastest way to setup test environment is to use the main branch container image:
|
||||
The fastest way to set up a test environment is to use the main branch's container image:
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: e2e
|
||||
@@ -13,7 +13,7 @@ The fastest way to setup test environment is to use the main branch container im
|
||||
:selected:
|
||||
:sync: cpu
|
||||
|
||||
You can run the unit tests on CPU with the following steps:
|
||||
You can run the unit tests on CPUs with the following steps:
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
@@ -22,39 +22,50 @@ cd ~/vllm-project/
|
||||
# ls
|
||||
# vllm vllm-ascend
|
||||
|
||||
# Use mirror to speedup download
|
||||
# docker pull quay.nju.edu.cn/ascend/cann:|cann_image_tag|
|
||||
# Use mirror to speed up download
|
||||
# docker pull m.daocloud.io/quay.io/ascend/cann:|cann_image_tag|
|
||||
export IMAGE=quay.io/ascend/cann:|cann_image_tag|
|
||||
docker run --rm --name vllm-ascend-ut \
|
||||
-v $(pwd):/vllm-project \
|
||||
-v ~/.cache:/root/.cache \
|
||||
-ti $IMAGE bash
|
||||
|
||||
# (Optional) Configure mirror to speedup download
|
||||
# (Optional) Configure mirror to speed up download
|
||||
sed -i 's|ports.ubuntu.com|mirrors.huaweicloud.com|g' /etc/apt/sources.list
|
||||
pip config set global.index-url https://mirrors.huaweicloud.com/repository/pypi/simple/
|
||||
|
||||
# For torch-npu dev version or x86 machine
|
||||
# For TorchNPU dev version or x86 machine
|
||||
export PIP_EXTRA_INDEX_URL="https://download.pytorch.org/whl/cpu/ https://mirrors.huaweicloud.com/ascend/repos/pypi"
|
||||
|
||||
# src path
|
||||
export SRC_WORKSPACE=/vllm-workspace
|
||||
mkdir -p $SRC_WORKSPACE
|
||||
cd $SRC_WORKSPACE
|
||||
|
||||
apt-get update -y
|
||||
apt-get install -y python3-pip git vim wget net-tools gcc g++ cmake libnuma-dev curl gnupg2
|
||||
|
||||
# Install vllm
|
||||
cd /vllm-project/vllm
|
||||
VLLM_TARGET_DEVICE=empty python3 -m pip -v install .
|
||||
git clone -b |vllm_ascend_version| --depth 1 https://github.com/vllm-project/vllm-ascend.git
|
||||
git clone --depth 1 https://github.com/vllm-project/vllm.git
|
||||
|
||||
# Install vllm-ascend
|
||||
cd /vllm-project/vllm-ascend
|
||||
# [IMPORTANT] Import LD_LIBRARY_PATH to enumerate the CANN environment under CPU
|
||||
# vllm
|
||||
cd $SRC_WORKSPACE/vllm
|
||||
VLLM_TARGET_DEVICE=empty python3 -m pip install .
|
||||
python3 -m pip uninstall -y triton
|
||||
|
||||
# vllm-ascend
|
||||
cd $SRC_WORKSPACE/vllm-ascend
|
||||
export LD_LIBRARY_PATH=$LD_LIBRARY_PATH:/usr/local/Ascend/ascend-toolkit/latest/$(uname -m)-linux/devlib
|
||||
# For cpu environment, set SOC_VERSION for different chips.
|
||||
# See https://github.com/vllm-project/vllm-ascend/blob/3cb0af0bcf3299089ca7e72159fa36e825a470f8/setup.py#L132 for detail.
|
||||
export SOC_VERSION="ascend910b1"
|
||||
python3 -m pip install .
|
||||
python3 -m pip install -r requirements-dev.txt
|
||||
python3 -m pip install -v .
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Single card
|
||||
::::{tab-item} Single-card
|
||||
:sync: single
|
||||
|
||||
```{code-block} bash
|
||||
@@ -66,6 +77,7 @@ export DEVICE=/dev/davinci0
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:main
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--device $DEVICE \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
@@ -86,13 +98,16 @@ After starting the container, you should install the required packages:
|
||||
# Prepare
|
||||
pip config set global.index-url https://mirrors.tuna.tsinghua.edu.cn/pypi/web/simple
|
||||
|
||||
# Switch to the /vllm-workspace/vllm-ascend directory
|
||||
cd /vllm-workspace/vllm-ascend/
|
||||
|
||||
# Install required packages
|
||||
pip install -r requirements-dev.txt
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Multi cards
|
||||
::::{tab-item} Multi-cards
|
||||
:sync: multi
|
||||
|
||||
```{code-block} bash
|
||||
@@ -101,6 +116,7 @@ pip install -r requirements-dev.txt
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:main
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
@@ -136,13 +152,13 @@ pip install -r requirements-dev.txt
|
||||
|
||||
## Running tests
|
||||
|
||||
### Unit test
|
||||
### Unit tests
|
||||
|
||||
There are several principles to follow when writing unit tests:
|
||||
|
||||
- The test file path should be consistent with source file and start with `test_` prefix, such as: `vllm_ascend/worker/worker_v1.py` --> `tests/ut/worker/test_worker_v1.py`
|
||||
- The vLLM Ascend test are using unittest framework, see [here](https://docs.python.org/3/library/unittest.html#module-unittest) to understand how to write unit tests.
|
||||
- All unit tests can be run on CPU, so you must mock the device-related function to host.
|
||||
- The test file path should be consistent with the source file and start with the `test_` prefix, such as: `vllm_ascend/worker/worker.py` --> `tests/ut/worker/test_worker.py`
|
||||
- The vLLM Ascend test uses unittest framework. See [the Python unittest documentation](https://docs.python.org/3/library/unittest.html#module-unittest) to understand how to write unit tests.
|
||||
- All unit tests can be run on CPUs, so you must mock the device-related functions on the host.
|
||||
- Example: [tests/ut/test_ascend_config.py](https://github.com/vllm-project/vllm-ascend/blob/main/tests/ut/test_ascend_config.py).
|
||||
- You can run the unit tests using `pytest`:
|
||||
|
||||
@@ -161,12 +177,12 @@ TORCH_DEVICE_BACKEND_AUTOLOAD=0 pytest -sv tests/ut
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Single card
|
||||
::::{tab-item} Single-card
|
||||
:sync: single
|
||||
|
||||
```bash
|
||||
cd /vllm-workspace/vllm-ascend/
|
||||
# Run all single card the tests
|
||||
# Run all single-card tests
|
||||
pytest -sv tests/ut
|
||||
|
||||
# Run single test
|
||||
@@ -175,12 +191,12 @@ pytest -sv tests/ut/test_ascend_config.py
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Multi cards test
|
||||
::::{tab-item} Multi-card
|
||||
:sync: multi
|
||||
|
||||
```bash
|
||||
cd /vllm-workspace/vllm-ascend/
|
||||
# Run all single card the tests
|
||||
# Run all multi-card tests
|
||||
pytest -sv tests/ut
|
||||
|
||||
# Run single test
|
||||
@@ -193,8 +209,63 @@ pytest -sv tests/ut/test_ascend_config.py
|
||||
|
||||
### E2E test
|
||||
|
||||
Although vllm-ascend CI provide [e2e test](https://github.com/vllm-project/vllm-ascend/blob/main/.github/workflows/vllm_ascend_test.yaml) on Ascend CI, you can run it
|
||||
locally.
|
||||
Although vllm-ascend CI provides E2E tests on Ascend CI (for example,
|
||||
[schedule_nightly_test_a2.yaml](https://github.com/vllm-project/vllm-ascend/blob/main/.github/workflows/schedule_nightly_test_a2.yaml), [schedule_nightly_test_a3.yaml](https://github.com/vllm-project/vllm-ascend/blob/main/.github/workflows/schedule_nightly_test_a3.yaml), [pr_test.yaml](https://github.com/vllm-project/vllm-ascend/blob/main/.github/workflows/pr_test.yaml)), you can run them locally.
|
||||
|
||||
#### PR-triggered E2E test
|
||||
|
||||
You can run tests with `pytest` as well. Typical examples:
|
||||
:::::{tab-set}
|
||||
:sync-group: e2e
|
||||
|
||||
::::{tab-item} Local (CPU)
|
||||
:sync: cpu
|
||||
|
||||
You can't run the E2E test on CPUs.
|
||||
::::
|
||||
|
||||
::::{tab-item} Single-card
|
||||
:selected:
|
||||
:sync: single
|
||||
|
||||
```bash
|
||||
cd /vllm-workspace/vllm-ascend/
|
||||
# Run all single-card tests
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/pull_request/one_card/
|
||||
|
||||
# Run a certain test script
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/pull_request/one_card/test_camem.py
|
||||
|
||||
# Run a certain case in test script
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/pull_request/one_card/test_camem.py::test_end_to_end
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Multi-card
|
||||
:sync: multi
|
||||
|
||||
```bash
|
||||
cd /vllm-workspace/vllm-ascend/
|
||||
# Run all multi-card tests
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/pull_request/two_card/
|
||||
|
||||
# Run a certain test script
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/pull_request/two_card/test_qwen3_moe_eplb.py
|
||||
|
||||
# Run a certain case in test script
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/pull_request/two_card/test_qwen3_moe_eplb.py::test_qwen3_moe_w8a8_distributed_tp2_ep_dynamic_eplb
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
:::::
|
||||
|
||||
This will reproduce the E2E test behavior.
|
||||
|
||||
#### Nightly-triggered E2E test
|
||||
|
||||
You can run tests with `pytest` as well. Typical examples:
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: e2e
|
||||
@@ -202,84 +273,102 @@ locally.
|
||||
::::{tab-item} Local (CPU)
|
||||
:sync: cpu
|
||||
|
||||
You can't run e2e test on CPU.
|
||||
You can't run the E2E test on CPUs.
|
||||
::::
|
||||
|
||||
::::{tab-item} Single card
|
||||
::::{tab-item} Single-card
|
||||
:selected:
|
||||
:sync: single
|
||||
|
||||
```bash
|
||||
cd /vllm-workspace/vllm-ascend/
|
||||
# Run all single card the tests
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/singlecard/
|
||||
|
||||
# Run a certain test script
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/singlecard/test_offline_inference.py
|
||||
|
||||
# Run a certain case in test script
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/singlecard/test_offline_inference.py::test_models
|
||||
# run all single-card op tests
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/nightly/single_node/ops/singlecard_ops/
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Multi cards test
|
||||
::::{tab-item} Multi-card
|
||||
:sync: multi
|
||||
|
||||
```bash
|
||||
cd /vllm-workspace/vllm-ascend/
|
||||
# Run all single card the tests
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/multicard/
|
||||
# run all multi-card op tests on A2
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/nightly/single_node/ops/multicard_ops_a2/
|
||||
|
||||
# Run a certain test script
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/multicard/test_dynamic_npugraph_batchsize.py
|
||||
|
||||
# Run a certain case in test script
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/multicard/test_offline_inference.py::test_models
|
||||
# run all multi-card op tests on A3
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/nightly/single_node/ops/multicard_ops_a3/
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
:::::
|
||||
|
||||
This will reproduce e2e test: [vllm_ascend_test.yaml](https://github.com/vllm-project/vllm-ascend/blob/main/.github/workflows/vllm_ascend_test.yaml).
|
||||
For running nightly single-node model test cases locally, refer to the following example.
|
||||
|
||||
#### E2E test example:
|
||||
```bash
|
||||
export CONFIG_YAML_PATH=Qwen3-32B.yaml
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/nightly/single_node/models/scripts/test_single_node.py
|
||||
```
|
||||
|
||||
- Offline test example: [`tests/e2e/singlecard/test_offline_inference.py`](https://github.com/vllm-project/vllm-ascend/blob/main/tests/e2e/singlecard/test_offline_inference.py)
|
||||
- Online test examples: [`tests/e2e/singlecard/test_prompt_embedding.py`](https://github.com/vllm-project/vllm-ascend/blob/main/tests/e2e/singlecard/test_prompt_embedding.py)
|
||||
- Correctness test example: [`tests/e2e/singlecard/test_aclgraph.py`](https://github.com/vllm-project/vllm-ascend/blob/main/tests/e2e/singlecard/test_aclgraph.py)
|
||||
- Reduced Layer model test example: [test_torchair_graph_mode.py - DeepSeek-V3-Pruning](https://github.com/vllm-project/vllm-ascend/blob/20767a043cccb3764214930d4695e53941de87ec/tests/e2e/multicard/test_torchair_graph_mode.py#L48)
|
||||
For running nightly multi-node model test cases locally, refer to the `Running Locally` section in [Multi Node Test](./multi_node_test.md).
|
||||
|
||||
The CI resource is limited, you might need to reduce layer number of the model, below is an example of how to generate a reduced layer model:
|
||||
1. Fork the original model repo in modelscope, we need all the files in the repo except for weights.
|
||||
2. Set `num_hidden_layers` to the expected number of layers, e.g., `{"num_hidden_layers": 2,}`
|
||||
3. Copy the following python script as `generate_random_weight.py`. Set the relevant parameters `MODEL_LOCAL_PATH`, `DIST_DTYPE` and `DIST_MODEL_PATH` as needed:
|
||||
#### E2E test examples
|
||||
|
||||
```python
|
||||
import torch
|
||||
from transformers import AutoTokenizer, AutoConfig
|
||||
from modeling_deepseek import DeepseekV3ForCausalLM
|
||||
from modelscope import snapshot_download
|
||||
- Offline test example: [`tests/e2e/pull_request/one_card/test_camem.py`](https://github.com/vllm-project/vllm-ascend/blob/main/tests/e2e/pull_request/one_card/test_camem.py)
|
||||
|
||||
MODEL_LOCAL_PATH = "~/.cache/modelscope/models/vllm-ascend/DeepSeek-V3-Pruning"
|
||||
DIST_DTYPE = torch.bfloat16
|
||||
DIST_MODEL_PATH = "./random_deepseek_v3_with_2_hidden_layer"
|
||||
The CI resource is limited, and you might need to reduce the number of layers of a model. Below is an example of how to generate a reduced layer model:
|
||||
|
||||
config = AutoConfig.from_pretrained(MODEL_LOCAL_PATH, trust_remote_code=True)
|
||||
model = DeepseekV3ForCausalLM(config)
|
||||
model = model.to(DIST_DTYPE)
|
||||
model.save_pretrained(DIST_MODEL_PATH)
|
||||
```
|
||||
1. Fork the original model repo in modelscope. All the files in the repo except for weights are required.
|
||||
2. Set `num_hidden_layers` to the expected number of layers, e.g., `{"num_hidden_layers": 2,}`
|
||||
3. Copy the following python script as `generate_random_weight.py`. Set the relevant parameters `MODEL_LOCAL_PATH`, `DIST_DTYPE` and `DIST_MODEL_PATH` as needed:
|
||||
|
||||
```python
|
||||
import torch
|
||||
from transformers import AutoTokenizer, AutoConfig
|
||||
from modeling_deepseek import DeepseekV3ForCausalLM
|
||||
from modelscope import snapshot_download
|
||||
|
||||
MODEL_LOCAL_PATH = "~/.cache/modelscope/models/vllm-ascend/DeepSeek-V3-Pruning"
|
||||
DIST_DTYPE = torch.bfloat16
|
||||
DIST_MODEL_PATH = "./random_deepseek_v3_with_2_hidden_layer"
|
||||
|
||||
config = AutoConfig.from_pretrained(MODEL_LOCAL_PATH, trust_remote_code=True)
|
||||
model = DeepseekV3ForCausalLM(config)
|
||||
model = model.to(DIST_DTYPE)
|
||||
model.save_pretrained(DIST_MODEL_PATH)
|
||||
```
|
||||
|
||||
### Run doctest
|
||||
|
||||
vllm-ascend provides a `vllm-ascend/tests/e2e/run_doctests.sh` command to run all doctests in the doc files.
|
||||
The doctest is a good way to make sure the docs are up to date and the examples are executable, you can run it locally as follows:
|
||||
The doctest is a good way to make sure docs stay current and examples remain executable, which can be run locally as follows:
|
||||
|
||||
```bash
|
||||
# Run doctest
|
||||
/vllm-workspace/vllm-ascend/tests/e2e/run_doctests.sh
|
||||
```
|
||||
|
||||
This will reproduce the same environment as the CI: [vllm_ascend_doctest.yaml](https://github.com/vllm-project/vllm-ascend/blob/main/.github/workflows/vllm_ascend_doctest.yaml).
|
||||
This will reproduce the same environment as the CI. See [labeled_doctest.yaml](https://github.com/vllm-project/vllm-ascend/blob/main/.github/workflows/labeled_doctest.yaml).
|
||||
|
||||
### Run docs link check
|
||||
|
||||
You can validate external links in the Sphinx docs locally with:
|
||||
|
||||
```bash
|
||||
make -C docs linkcheck SPHINXOPTS="-W --keep-going"
|
||||
```
|
||||
|
||||
To check links in a specific Markdown file, pass the file to `sphinx-build`.
|
||||
For example, to check only `docs/source/user_guide/release_notes.md`:
|
||||
|
||||
```bash
|
||||
cd docs
|
||||
sphinx-build -b linkcheck -W --keep-going \
|
||||
source _build/linkcheck source/user_guide/release_notes.md
|
||||
```
|
||||
|
||||
The detailed report will be written to:
|
||||
|
||||
- `docs/_build/linkcheck/output.txt`
|
||||
- `docs/_build/linkcheck/output.json`
|
||||
|
||||
@@ -5,6 +5,6 @@
|
||||
:maxdepth: 1
|
||||
using_evalscope
|
||||
using_lm_eval
|
||||
using_ais_bench
|
||||
using_opencompass
|
||||
accuracy_report/index
|
||||
:::
|
||||
|
||||
333
docs/source/developer_guide/evaluation/using_ais_bench.md
Normal file
333
docs/source/developer_guide/evaluation/using_ais_bench.md
Normal file
@@ -0,0 +1,333 @@
|
||||
# Using AISBench
|
||||
|
||||
This document guides you to conduct accuracy testing using [AISBench](https://github.com/AISBench/benchmark/tree/master). AISBench provides accuracy and performance evaluation for many datasets.
|
||||
|
||||
## Online Server
|
||||
|
||||
### 1. Start the vLLM server
|
||||
|
||||
You can run docker container to start the vLLM server on a single NPU:
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# Update DEVICE according to your device (/dev/davinci[0-7])
|
||||
export DEVICE=/dev/davinci7
|
||||
# Update the vllm-ascend image
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--device $DEVICE \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-p 8000:8000 \
|
||||
-e VLLM_USE_MODELSCOPE=True \
|
||||
-e PYTORCH_NPU_ALLOC_CONF=max_split_size_mb:256 \
|
||||
-it $IMAGE \
|
||||
/bin/bash
|
||||
```
|
||||
|
||||
Run the vLLM server in the docker.
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
vllm serve Qwen/Qwen2.5-0.5B-Instruct --max-model-len 35000 &
|
||||
```
|
||||
|
||||
:::{note}
|
||||
`--max-model-len` should be greater than `35000`, this will be suitable for most datasets. Otherwise the accuracy evaluation may be affected.
|
||||
:::
|
||||
|
||||
The vLLM server is started successfully, if you see logs as below:
|
||||
|
||||
```shell
|
||||
INFO: Started server process [9446]
|
||||
INFO: Waiting for application startup.
|
||||
INFO: Application startup complete.
|
||||
```
|
||||
|
||||
### 2. Run different datasets using AISBench
|
||||
|
||||
#### Install AISBench
|
||||
|
||||
Refer to [AISBench](https://github.com/AISBench/benchmark/tree/master) for details.
|
||||
Install AISBench from source.
|
||||
|
||||
```shell
|
||||
git clone https://github.com/AISBench/benchmark.git
|
||||
cd benchmark/
|
||||
pip3 install -e ./ --use-pep517
|
||||
```
|
||||
|
||||
Install extra AISBench dependencies.
|
||||
|
||||
```shell
|
||||
pip3 install -r requirements/api.txt
|
||||
pip3 install -r requirements/extra.txt
|
||||
```
|
||||
|
||||
Run `ais_bench -h` to check the installation.
|
||||
|
||||
#### Download Dataset
|
||||
|
||||
You can choose one or multiple datasets to execute accuracy evaluation.
|
||||
|
||||
1. `C-Eval` dataset.
|
||||
|
||||
Take `C-Eval` dataset as an example. You can refer to [Datasets](https://github.com/AISBench/benchmark/tree/master/ais_bench/benchmark/configs/datasets) for more datasets. Each dataset has a `README.md` with detailed download and installation instructions.
|
||||
|
||||
Download dataset and install it to specific path.
|
||||
|
||||
```shell
|
||||
cd ais_bench/datasets
|
||||
mkdir ceval/
|
||||
mkdir ceval/formal_ceval
|
||||
cd ceval/formal_ceval
|
||||
wget https://www.modelscope.cn/datasets/opencompass/ceval-exam/resolve/master/ceval-exam.zip
|
||||
unzip ceval-exam.zip
|
||||
rm ceval-exam.zip
|
||||
```
|
||||
|
||||
2. `MMLU` dataset.
|
||||
|
||||
```shell
|
||||
cd ais_bench/datasets
|
||||
wget http://opencompass.oss-cn-shanghai.aliyuncs.com/datasets/data/mmlu.zip
|
||||
unzip mmlu.zip
|
||||
rm mmlu.zip
|
||||
```
|
||||
|
||||
3. `GPQA` dataset.
|
||||
|
||||
```shell
|
||||
cd ais_bench/datasets
|
||||
wget http://opencompass.oss-cn-shanghai.aliyuncs.com/datasets/data/gpqa.zip
|
||||
unzip gpqa.zip
|
||||
rm gpqa.zip
|
||||
```
|
||||
|
||||
4. `MATH` dataset.
|
||||
|
||||
```shell
|
||||
cd ais_bench/datasets
|
||||
wget http://opencompass.oss-cn-shanghai.aliyuncs.com/datasets/data/math.zip
|
||||
unzip math.zip
|
||||
rm math.zip
|
||||
```
|
||||
|
||||
5. `LiveCodeBench` dataset.
|
||||
|
||||
```shell
|
||||
cd ais_bench/datasets
|
||||
git lfs install
|
||||
git clone https://huggingface.co/datasets/livecodebench/code_generation_lite
|
||||
```
|
||||
|
||||
6. `AIME 2024` dataset.
|
||||
|
||||
```shell
|
||||
cd ais_bench/datasets
|
||||
mkdir aime/
|
||||
cd aime/
|
||||
wget http://opencompass.oss-cn-shanghai.aliyuncs.com/datasets/data/aime.zip
|
||||
unzip aime.zip
|
||||
rm aime.zip
|
||||
```
|
||||
|
||||
7. `GSM8K` dataset.
|
||||
|
||||
```shell
|
||||
cd ais_bench/datasets
|
||||
wget http://opencompass.oss-cn-shanghai.aliyuncs.com/datasets/data/gsm8k.zip
|
||||
unzip gsm8k.zip
|
||||
rm gsm8k.zip
|
||||
```
|
||||
|
||||
#### Configuration
|
||||
|
||||
Update the file `benchmark/ais_bench/benchmark/configs/models/vllm_api/vllm_api_general_chat.py`.
|
||||
There are several arguments that you should update according to your environment.
|
||||
|
||||
- `attr`: Identifier for the inference backend type, fixed as `service` (serving-based inference) or `local` (local model).
|
||||
- `type`: Used to select different backend API types.
|
||||
- `abbr`: Unique identifier for a local task, used to distinguish between multiple tasks.
|
||||
- `path`: Update to your model weight path.
|
||||
- `model`: Update to your model name in vLLM.
|
||||
- `host_ip` and `host_port`: Update to your vLLM server ip and port.
|
||||
- `max_out_len`: Note `max_out_len` + LLM input length should be less than `max_model_len` (config in your vllm server), `32768` will be suitable for most datasets.
|
||||
- `batch_size`: Update according to your dataset.
|
||||
- `temperature`: Update inference argument.
|
||||
|
||||
```python
|
||||
from ais_bench.benchmark.models import VLLMCustomAPIChat
|
||||
from ais_bench.benchmark.utils.model_postprocessors import extract_non_reasoning_content
|
||||
|
||||
models = [
|
||||
dict(
|
||||
attr="service",
|
||||
type=VLLMCustomAPIChat,
|
||||
abbr='vllm-api-general-chat',
|
||||
path="xxxx",
|
||||
model="xxxx",
|
||||
request_rate = 0,
|
||||
retry = 2,
|
||||
host_ip = "localhost",
|
||||
host_port = 8000,
|
||||
max_out_len = xxx,
|
||||
batch_size = xxx,
|
||||
trust_remote_code=False,
|
||||
generation_kwargs = dict(
|
||||
temperature = 0.6,
|
||||
top_k = 10,
|
||||
top_p = 0.95,
|
||||
seed = None,
|
||||
repetition_penalty = 1.03,
|
||||
),
|
||||
pred_postprocessor=dict(type=extract_non_reasoning_content)
|
||||
)
|
||||
]
|
||||
|
||||
```
|
||||
|
||||
#### Execute Accuracy Evaluation
|
||||
|
||||
Run the following code to execute different accuracy evaluation.
|
||||
|
||||
```shell
|
||||
# run C-Eval dataset
|
||||
ais_bench --models vllm_api_general_chat --datasets ceval_gen_0_shot_cot_chat_prompt.py --mode all --dump-eval-details --merge-ds
|
||||
|
||||
# run MMLU dataset
|
||||
ais_bench --models vllm_api_general_chat --datasets mmlu_gen_0_shot_cot_chat_prompt.py --mode all --dump-eval-details --merge-ds
|
||||
|
||||
# run GPQA dataset
|
||||
ais_bench --models vllm_api_general_chat --datasets gpqa_gen_0_shot_str.py --mode all --dump-eval-details --merge-ds
|
||||
|
||||
# run MATH-500 dataset
|
||||
ais_bench --models vllm_api_general_chat --datasets math500_gen_0_shot_cot_chat_prompt.py --mode all --dump-eval-details --merge-ds
|
||||
|
||||
# run LiveCodeBench dataset
|
||||
ais_bench --models vllm_api_general_chat --datasets livecodebench_code_generate_lite_gen_0_shot_chat.py --mode all --dump-eval-details --merge-ds
|
||||
|
||||
# run AIME 2024 dataset
|
||||
ais_bench --models vllm_api_general_chat --datasets aime2024_gen_0_shot_chat_prompt.py --mode all --dump-eval-details --merge-ds
|
||||
|
||||
# run GSM8K dataset
|
||||
ais_bench --models vllm_api_general_chat --datasets gsm8k_gen_0_shot_cot_chat_prompt.py --mode all --dump-eval-details --merge-ds
|
||||
|
||||
```
|
||||
|
||||
After each dataset execution, you can get the result from saved files such as `outputs/default/20250628_151326`, there is an example as follows:
|
||||
|
||||
```shell
|
||||
20250628_151326/
|
||||
├── configs # Combined configuration file for model tasks, dataset tasks, and result presentation tasks
|
||||
│ └── 20250628_151326_29317.py
|
||||
├── logs # Execution logs; if --debug is added to the command, no intermediate logs are saved to disk (all are printed directly to the screen)
|
||||
│ ├── eval
|
||||
│ │ └── vllm-api-general-chat
|
||||
│ │ └── demo_gsm8k.out # Logs of the accuracy evaluation process based on inference results in the predictions/ folder
|
||||
│ └── infer
|
||||
│ └── vllm-api-general-chat
|
||||
│ └── demo_gsm8k.out # Logs of the inference process
|
||||
├── predictions
|
||||
│ └── vllm-api-general-chat
|
||||
│ └── demo_gsm8k.json # Inference results (all outputs returned by the inference service)
|
||||
├── results
|
||||
│ └── vllm-api-general-chat
|
||||
│ └── demo_gsm8k.json # Raw scores calculated from the accuracy evaluation
|
||||
└── summary
|
||||
├── summary_20250628_151326.csv # Final accuracy scores (in table format)
|
||||
├── summary_20250628_151326.md # Final accuracy scores (in Markdown format)
|
||||
└── summary_20250628_151326.txt # Final accuracy scores (in text format)
|
||||
```
|
||||
|
||||
#### Execute Performance Evaluation
|
||||
|
||||
Text-only benchmarks:
|
||||
|
||||
```shell
|
||||
# run C-Eval dataset
|
||||
ais_bench --models vllm_api_general_chat --datasets ceval_gen_0_shot_cot_chat_prompt.py --summarizer default_perf --mode perf
|
||||
|
||||
# run MMLU dataset
|
||||
ais_bench --models vllm_api_general_chat --datasets mmlu_gen_0_shot_cot_chat_prompt.py --summarizer default_perf --mode perf
|
||||
|
||||
# run GPQA dataset
|
||||
ais_bench --models vllm_api_general_chat --datasets gpqa_gen_0_shot_str.py --summarizer default_perf --mode perf
|
||||
|
||||
# run MATH-500 dataset
|
||||
ais_bench --models vllm_api_general_chat --datasets math500_gen_0_shot_cot_chat_prompt.py --summarizer default_perf --mode perf
|
||||
|
||||
# run LiveCodeBench dataset
|
||||
ais_bench --models vllm_api_general_chat --datasets livecodebench_code_generate_lite_gen_0_shot_chat.py --summarizer default_perf --mode perf
|
||||
|
||||
# run AIME 2024 dataset
|
||||
ais_bench --models vllm_api_general_chat --datasets aime2024_gen_0_shot_chat_prompt.py --summarizer default_perf --mode perf
|
||||
|
||||
# run GSM8K dataset
|
||||
ais_bench --models vllm_api_general_chat --datasets gsm8k_gen_0_shot_cot_str_perf.py --summarizer default_perf --mode perf
|
||||
```
|
||||
|
||||
Multi-modal benchmarks (text + images):
|
||||
|
||||
```shell
|
||||
# run textvqa dataset
|
||||
ais_bench --models vllm_api_stream_chat --datasets textvqa_gen_base64 --summarizer default_perf --mode perf
|
||||
```
|
||||
|
||||
After execution, you can get the result from saved files, there is an example as follows:
|
||||
|
||||
```shell
|
||||
20251031_070226/
|
||||
|-- configs # Combined configuration file for model tasks, dataset tasks, and result presentation tasks
|
||||
| `-- 20251031_070226_122485.py
|
||||
|-- logs
|
||||
| `-- performances
|
||||
| `-- vllm-api-general-chat
|
||||
| `-- cevaldataset.out # Logs of the performance evaluation process
|
||||
`-- performances
|
||||
`-- vllm-api-general-chat
|
||||
|-- cevaldataset.csv # Final performance results (in table format)
|
||||
|-- cevaldataset.json # Final performance results (in json format)
|
||||
|-- cevaldataset_details.h5 # Final performance results in details
|
||||
|-- cevaldataset_details.json # Final performance results in details
|
||||
|-- cevaldataset_plot.html # Final performance results (in html format)
|
||||
`-- cevaldataset_rps_distribution_plot_with_actual_rps.html # Final performance results (in html format)
|
||||
```
|
||||
|
||||
### 3. Troubleshooting
|
||||
|
||||
#### Invalid Image Path Error
|
||||
|
||||
If you download the TextVQA dataset following the AISBench documentation:
|
||||
|
||||
```bash
|
||||
cd ais_bench/datasets
|
||||
git lfs install
|
||||
git clone https://huggingface.co/datasets/maoxx241/textvqa_subset
|
||||
mv textvqa_subset/ textvqa/
|
||||
mkdir textvqa/textvqa_json/
|
||||
mv textvqa/*.json textvqa/textvqa_json/
|
||||
mv textvqa/*.jsonl textvqa/textvqa_json/
|
||||
```
|
||||
|
||||
you may encounter the following error:
|
||||
|
||||
```bash
|
||||
AISBench - ERROR - /vllm-workspace/benchmark/ais_bench/benchmark/clients/base_client.py - raise_error - 35 - [AisBenchClientException] Request failed: HTTP status 400. Server response: {"error":{"message":"1 validation error for ChatCompletionContentPartImageParam\nimage_url\n Input should be a valid dictionary [type=dict_type, input_value='data/textvqa/train_images/b2ae0f96dfbea5d8.jpg', input_type=str]\n For further information visit https://errors.pydantic.dev/2.12/v/dict_type None","type":"BadRequestError","param":null,"code":400}}
|
||||
```
|
||||
|
||||
You need to manually replace the dataset image paths with absolute paths, changing `/path/to/benchmark/ais_bench/datasets/textvqa/train_images/` to the actual absolute directory where the images are stored:
|
||||
|
||||
```bash
|
||||
cd ais_bench/datasets/textvqa/textvqa_json
|
||||
sed -i 's#data/textvqa/train_images/#/path/to/benchmark/ais_bench/datasets/textvqa/train_images/#g' textvqa_val.json
|
||||
```
|
||||
@@ -1,8 +1,8 @@
|
||||
# Using EvalScope
|
||||
|
||||
This document will guide you have model inference stress testing and accuracy testing using [EvalScope](https://github.com/modelscope/evalscope).
|
||||
This document will guide you through model inference stress testing and accuracy testing using [EvalScope](https://github.com/modelscope/evalscope).
|
||||
|
||||
## 1. Online serving
|
||||
## 1. Online server
|
||||
|
||||
You can run docker container to start the vLLM server on a single NPU:
|
||||
|
||||
@@ -13,6 +13,7 @@ export DEVICE=/dev/davinci7
|
||||
# Update the vllm-ascend image
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
docker run --rm \
|
||||
--shm-size=1g \
|
||||
--name vllm-ascend \
|
||||
--device $DEVICE \
|
||||
--device /dev/davinci_manager \
|
||||
@@ -31,30 +32,30 @@ docker run --rm \
|
||||
vllm serve Qwen/Qwen2.5-7B-Instruct --max_model_len 26240
|
||||
```
|
||||
|
||||
If your service start successfully, you can see the info shown below:
|
||||
If the vLLM server is started successfully, you can see information shown below:
|
||||
|
||||
```
|
||||
```shell
|
||||
INFO: Started server process [6873]
|
||||
INFO: Waiting for application startup.
|
||||
INFO: Application startup complete.
|
||||
```
|
||||
|
||||
Once your server is started, you can query the model with input prompts in new terminal:
|
||||
Once your server is started, you can query the model with input prompts in a new terminal:
|
||||
|
||||
```
|
||||
```shell
|
||||
curl http://localhost:8000/v1/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "Qwen/Qwen2.5-7B-Instruct",
|
||||
"prompt": "The future of AI is",
|
||||
"max_tokens": 7,
|
||||
"max_completion_tokens": 7,
|
||||
"temperature": 0
|
||||
}'
|
||||
```
|
||||
|
||||
## 2. Install EvalScope using pip
|
||||
|
||||
You can install EvalScope by using:
|
||||
You can install EvalScope as follows:
|
||||
|
||||
```bash
|
||||
python3 -m venv .venv-evalscope
|
||||
@@ -62,21 +63,21 @@ source .venv-evalscope/bin/activate
|
||||
pip install gradio plotly evalscope
|
||||
```
|
||||
|
||||
## 3. Run gsm8k accuracy test using EvalScope
|
||||
## 3. Run GSM8K using EvalScope for accuracy testing
|
||||
|
||||
You can `evalscope eval` run gsm8k accuracy test:
|
||||
You can use `evalscope eval` to run GSM8K (a grade-school math benchmark dataset) for accuracy testing:
|
||||
|
||||
```
|
||||
```shell
|
||||
evalscope eval \
|
||||
--model Qwen/Qwen2.5-7B-Instruct \
|
||||
--api-url http://localhost:8000/v1 \
|
||||
--api-key EMPTY \
|
||||
--eval-type service \
|
||||
--eval-type server \
|
||||
--datasets gsm8k \
|
||||
--limit 10
|
||||
```
|
||||
|
||||
After 1-2 mins, the output is as shown below:
|
||||
After 1 to 2 minutes, the output is shown below:
|
||||
|
||||
```shell
|
||||
+---------------------+-----------+-----------------+----------+-------+---------+---------+
|
||||
@@ -86,7 +87,7 @@ After 1-2 mins, the output is as shown below:
|
||||
+---------------------+-----------+-----------------+----------+-------+---------+---------+
|
||||
```
|
||||
|
||||
See more detail in: [EvalScope doc - Model API Service Evaluation](https://evalscope.readthedocs.io/en/latest/get_started/basic_usage.html#model-api-service-evaluation).
|
||||
See more details in [EvalScope doc - Model API Service Evaluation](https://evalscope.readthedocs.io/en/latest/get_started/basic_usage.html#model-api-service-evaluation).
|
||||
|
||||
## 4. Run model inference stress testing using EvalScope
|
||||
|
||||
@@ -98,9 +99,9 @@ pip install evalscope[perf] -U
|
||||
|
||||
### Basic usage
|
||||
|
||||
You can use `evalscope perf` run perf test:
|
||||
You can use `evalscope perf` to run perf testing:
|
||||
|
||||
```
|
||||
```shell
|
||||
evalscope perf \
|
||||
--url "http://localhost:8000/v1/chat/completions" \
|
||||
--parallel 5 \
|
||||
@@ -113,7 +114,7 @@ evalscope perf \
|
||||
|
||||
### Output results
|
||||
|
||||
After 1-2 mins, the output is as shown below:
|
||||
After 1 to 2 minutes, the output is shown below:
|
||||
|
||||
```shell
|
||||
Benchmarking summary:
|
||||
@@ -172,4 +173,4 @@ Percentile results:
|
||||
+------------+----------+---------+-------------+--------------+---------------+----------------------+
|
||||
```
|
||||
|
||||
See more detail in: [EvalScope doc - Model Inference Stress Testing](https://evalscope.readthedocs.io/en/latest/user_guides/stress_test/quick_start.html#basic-usage).
|
||||
See more detail in [EvalScope doc - Model Inference Stress Testing](https://evalscope.readthedocs.io/en/latest/user_guides/stress_test/quick_start.html#basic-usage).
|
||||
|
||||
@@ -1,9 +1,12 @@
|
||||
# Using lm-eval
|
||||
This document will guide you have a accuracy testing using [lm-eval][1].
|
||||
|
||||
This document guides you to conduct accuracy testing using [lm-eval][1].
|
||||
|
||||
## Online Server
|
||||
### 1. start the vLLM server
|
||||
You can run docker container to start the vLLM server on a single NPU:
|
||||
|
||||
### 1. Start the vLLM server
|
||||
|
||||
You can run a docker container to start the vLLM server on a single NPU:
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
@@ -13,6 +16,7 @@ export DEVICE=/dev/davinci7
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--device $DEVICE \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
@@ -31,46 +35,54 @@ docker run --rm \
|
||||
vllm serve Qwen/Qwen2.5-0.5B-Instruct --max_model_len 4096 &
|
||||
```
|
||||
|
||||
Started the vLLM server successfully,if you see log as below:
|
||||
The vLLM server is started successfully, if you see logs as below:
|
||||
|
||||
```
|
||||
```shell
|
||||
INFO: Started server process [9446]
|
||||
INFO: Waiting for application startup.
|
||||
INFO: Application startup complete.
|
||||
```
|
||||
|
||||
### 2. Run gsm8k accuracy test using lm-eval
|
||||
### 2. Run GSM8K using the vLLM server (curl) and then run lm-eval for accuracy testing
|
||||
|
||||
You can query result with input prompts:
|
||||
You can query the result with input prompts:
|
||||
|
||||
```shell
|
||||
PROMPT='<|im_start|>system
|
||||
You are a professional accountant. Answer questions using accounting knowledge, output only the option letter (A/B/C/D).<|im_end|>
|
||||
<|im_start|>user
|
||||
Question: A company'"'"'s balance sheet as of December 31, 2023 shows:
|
||||
Current assets: Cash and equivalents 5 million yuan, Accounts receivable 8 million yuan, Inventory 6 million yuan
|
||||
Non-current assets: Net fixed assets 12 million yuan
|
||||
Current liabilities: Short-term loans 4 million yuan, Accounts payable 3 million yuan
|
||||
Non-current liabilities: Long-term loans 9 million yuan
|
||||
Owner'"'"'s equity: Paid-in capital 10 million yuan, Retained earnings ?
|
||||
Requirement: Calculate the company'"'"'s Asset-Liability Ratio and Current Ratio (round to two decimal places).
|
||||
Options:
|
||||
A. Asset-Liability Ratio=58.33%, Current Ratio=1.90
|
||||
B. Asset-Liability Ratio=62.50%, Current Ratio=2.17
|
||||
C. Asset-Liability Ratio=65.22%, Current Ratio=1.75
|
||||
D. Asset-Liability Ratio=68.00%, Current Ratio=2.50<|im_end|>
|
||||
<|im_start|>assistant
|
||||
'
|
||||
|
||||
```
|
||||
curl http://localhost:8000/v1/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "Qwen/Qwen2.5-0.5B-Instruct",
|
||||
"prompt": "'"<|im_start|>system\nYou are a professional accountant. Answer questions using accounting knowledge, output only the option letter (A/B/C/D).<|im_end|>\n"\
|
||||
"<|im_start|>user\nQuestion: A company's balance sheet as of December 31, 2023 shows:\n"\
|
||||
" Current assets: Cash and equivalents 5 million yuan, Accounts receivable 8 million yuan, Inventory 6 million yuan\n"\
|
||||
" Non-current assets: Net fixed assets 12 million yuan\n"\
|
||||
" Current liabilities: Short-term loans 4 million yuan, Accounts payable 3 million yuan\n"\
|
||||
" Non-current liabilities: Long-term loans 9 million yuan\n"\
|
||||
" Owner's equity: Paid-in capital 10 million yuan, Retained earnings ?\n"\
|
||||
"Requirement: Calculate the company's Asset-Liability Ratio and Current Ratio (round to two decimal places).\n"\
|
||||
"Options:\n"\
|
||||
"A. Asset-Liability Ratio=58.33%, Current Ratio=1.90\n"\
|
||||
"B. Asset-Liability Ratio=62.50%, Current Ratio=2.17\n"\
|
||||
"C. Asset-Liability Ratio=65.22%, Current Ratio=1.75\n"\
|
||||
"D. Asset-Liability Ratio=68.00%, Current Ratio=2.50<|im_end|>\n"\
|
||||
"<|im_start|>assistant\n"'",
|
||||
"max_tokens": 1,
|
||||
"temperature": 0,
|
||||
"stop": ["<|im_end|>"]
|
||||
}' | python3 -m json.tool
|
||||
-d "$(jq -n \
|
||||
--arg model "Qwen/Qwen2.5-0.5B-Instruct" \
|
||||
--arg prompt "$PROMPT" \
|
||||
'{
|
||||
model: $model,
|
||||
prompt: $prompt,
|
||||
max_completion_tokens: 1,
|
||||
temperature: 0,
|
||||
stop: ["<|im_end|>"]
|
||||
}')" | python3 -m json.tool
|
||||
```
|
||||
|
||||
The output format matches the following:
|
||||
|
||||
```
|
||||
```json
|
||||
{
|
||||
"id": "cmpl-2f678e8bdf5a4b209a3f2c1fa5832e25",
|
||||
"object": "text_completion",
|
||||
@@ -98,16 +110,24 @@ The output format matches the following:
|
||||
}
|
||||
```
|
||||
|
||||
Install lm-eval in the container.
|
||||
Install lm-eval in the container:
|
||||
|
||||
```bash
|
||||
export HF_ENDPOINT="https://hf-mirror.com"
|
||||
export USE_MODELSCOPE_HUB=0
|
||||
pip install lm-eval[api]
|
||||
```
|
||||
|
||||
:::{note}
|
||||
The Docker container is launched with `VLLM_USE_MODELSCOPE=True`, which may
|
||||
cause lm-eval to download datasets from ModelScope instead of HuggingFace.
|
||||
Setting `USE_MODELSCOPE_HUB=0` disables this behavior so that lm-eval can
|
||||
fetch datasets from HuggingFace correctly.
|
||||
:::
|
||||
|
||||
Run the following command:
|
||||
|
||||
```
|
||||
```shell
|
||||
# Only test gsm8k dataset in this demo
|
||||
lm_eval \
|
||||
--model local-completions \
|
||||
@@ -116,19 +136,20 @@ lm_eval \
|
||||
--output_path ./
|
||||
```
|
||||
|
||||
After 30 mins, the output is as shown below:
|
||||
After 30 minutes, the output is as shown below:
|
||||
|
||||
```
|
||||
The markdown format results is as below:
|
||||
```shell
|
||||
The results in Markdown format are as follows:
|
||||
|
||||
Tasks|Version| Filter |n-shot| Metric | |Value | |Stderr|
|
||||
|Tasks|Version| Filter |n-shot| Metric | |Value | |Stderr|
|
||||
|-----|------:|----------------|-----:|-----------|---|-----:|---|-----:|
|
||||
|gsm8k| 3|flexible-extract| 5|exact_match|↑ |0.3215|± |0.0129|
|
||||
| | |strict-match | 5|exact_match|↑ |0.2077|± |0.0112|
|
||||
|gsm8k| 3|strict-match | 5|exact_match|↑ |0.2077|± |0.0112|
|
||||
|
||||
```
|
||||
|
||||
## Offline Server
|
||||
|
||||
### 1. Run docker container
|
||||
|
||||
You can run docker container on a single NPU:
|
||||
@@ -141,6 +162,7 @@ export DEVICE=/dev/davinci7
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--device $DEVICE \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
@@ -158,17 +180,26 @@ docker run --rm \
|
||||
/bin/bash
|
||||
```
|
||||
|
||||
### 2. Run gsm8k accuracy test using lm-eval
|
||||
Install lm-eval in the container.
|
||||
### 2. Run GSM8K using lm-eval for accuracy testing
|
||||
|
||||
Install lm-eval in the container:
|
||||
|
||||
```bash
|
||||
export HF_ENDPOINT="https://hf-mirror.com"
|
||||
export USE_MODELSCOPE_HUB=0
|
||||
pip install lm-eval
|
||||
```
|
||||
|
||||
:::{note}
|
||||
The Docker container is launched with `VLLM_USE_MODELSCOPE=True`, which may
|
||||
cause lm-eval to download datasets from ModelScope instead of HuggingFace.
|
||||
Setting `USE_MODELSCOPE_HUB=0` disables this behavior so that lm-eval can
|
||||
fetch datasets from HuggingFace correctly.
|
||||
:::
|
||||
|
||||
Run the following command:
|
||||
|
||||
```
|
||||
```shell
|
||||
# Only test gsm8k dataset in this demo
|
||||
lm_eval \
|
||||
--model vllm \
|
||||
@@ -177,21 +208,21 @@ lm_eval \
|
||||
--batch_size auto
|
||||
```
|
||||
|
||||
After 1-2 mins, the output is as shown below:
|
||||
After 1 to 2 minutes, the output is shown below:
|
||||
|
||||
```
|
||||
The markdown format results is as below:
|
||||
```shell
|
||||
The markdown format results are as below:
|
||||
|
||||
Tasks|Version| Filter |n-shot| Metric | |Value | |Stderr|
|
||||
|Tasks|Version| Filter |n-shot| Metric | |Value | |Stderr|
|
||||
|-----|------:|----------------|-----:|-----------|---|-----:|---|-----:|
|
||||
|gsm8k| 3|flexible-extract| 5|exact_match|↑ |0.3412|± |0.0131|
|
||||
| | |strict-match | 5|exact_match|↑ |0.3139|± |0.0128|
|
||||
|gsm8k| 3|strict-match | 5|exact_match|↑ |0.3139|± |0.0128|
|
||||
|
||||
```
|
||||
|
||||
## Use offline Datasets
|
||||
## Use Offline Datasets
|
||||
|
||||
Take gsm8k(single dataset) and mmlu(multi-subject dataset) as examples, and you can see more from [here][2].
|
||||
Take GSM8K (single dataset) and MMLU (multi-subject dataset) as examples, and you can see more from [using-local-datasets][2].
|
||||
|
||||
```bash
|
||||
# set HF_DATASETS_OFFLINE when using offline datasets
|
||||
@@ -205,7 +236,7 @@ cd lm_eval/tasks/gsm8k
|
||||
cd lm_eval/tasks/mmlu/default
|
||||
```
|
||||
|
||||
set [gsm8k.yaml][3] as follows:
|
||||
Set [gsm8k.yaml][3] as follows:
|
||||
|
||||
```yaml
|
||||
tag:
|
||||
@@ -230,7 +261,7 @@ training_split: train
|
||||
fewshot_split: train
|
||||
test_split: test
|
||||
doc_to_text: 'Q: {{question}}
|
||||
A(Please follow the summarize the result at the end with the format of "The answer is xxx", where xx is the result.):'
|
||||
A(Please follow the summarized result at the end with the format of "The answer is xxx", where xx is the result.):'
|
||||
doc_to_target: "{{answer}}" #" {{answer.split('### ')[-1].rstrip()}}"
|
||||
metric_list:
|
||||
- metric: exact_match
|
||||
@@ -268,7 +299,7 @@ metadata:
|
||||
version: 3.0
|
||||
```
|
||||
|
||||
set [_default_template_yaml][4] as follows:
|
||||
Set [_default_template_yaml][4] as follows:
|
||||
|
||||
```yaml
|
||||
# set dataset_path according to the downloaded dataset
|
||||
|
||||
@@ -1,9 +1,10 @@
|
||||
# Using OpenCompass
|
||||
This document will guide you have a accuracy testing using [OpenCompass](https://github.com/open-compass/opencompass).
|
||||
|
||||
## 1. Online Serving
|
||||
This document guides you to conduct accuracy testing using [OpenCompass](https://github.com/open-compass/opencompass).
|
||||
|
||||
You can run docker container to start the vLLM server on a single NPU:
|
||||
## 1. Online Server
|
||||
|
||||
You can run a docker container to start the vLLM server on a single NPU:
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
@@ -13,6 +14,7 @@ export DEVICE=/dev/davinci7
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--device $DEVICE \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
@@ -30,29 +32,30 @@ docker run --rm \
|
||||
vllm serve Qwen/Qwen2.5-7B-Instruct --max_model_len 26240
|
||||
```
|
||||
|
||||
If your service start successfully, you can see the info shown below:
|
||||
The vLLM server is started successfully, if you see information as below:
|
||||
|
||||
```
|
||||
```shell
|
||||
INFO: Started server process [6873]
|
||||
INFO: Waiting for application startup.
|
||||
INFO: Application startup complete.
|
||||
```
|
||||
|
||||
Once your server is started, you can query the model with input prompts in new terminal:
|
||||
Once your server is started, you can query the model with input prompts in a new terminal.
|
||||
|
||||
```
|
||||
```shell
|
||||
curl http://localhost:8000/v1/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "Qwen/Qwen2.5-7B-Instruct",
|
||||
"prompt": "The future of AI is",
|
||||
"max_tokens": 7,
|
||||
"max_completion_tokens": 7,
|
||||
"temperature": 0
|
||||
}'
|
||||
```
|
||||
|
||||
## 2. Run ceval accuracy test using OpenCompass
|
||||
Install OpenCompass and configure the environment variables in the container.
|
||||
## 2. Run C-Eval (a Chinese language model evaluation benchmark) using OpenCompass for accuracy testing
|
||||
|
||||
Install OpenCompass and configure the environment variables in the container:
|
||||
|
||||
```bash
|
||||
# Pin Python 3.10 due to:
|
||||
@@ -64,7 +67,7 @@ export DATASET_SOURCE=ModelScope
|
||||
git clone https://github.com/open-compass/opencompass.git
|
||||
```
|
||||
|
||||
Add `opencompass/configs/eval_vllm_ascend_demo.py` with the following content:
|
||||
Add the following content to `opencompass/configs/eval_vllm_ascend_demo.py`:
|
||||
|
||||
```python
|
||||
from mmengine.config import read_base
|
||||
@@ -106,14 +109,14 @@ models = [
|
||||
|
||||
Run the following command:
|
||||
|
||||
```
|
||||
```shell
|
||||
python3 run.py opencompass/configs/eval_vllm_ascend_demo.py --debug
|
||||
```
|
||||
|
||||
After 1-2 mins, the output is as shown below:
|
||||
After 1 to 2 minutes, the output is shown below:
|
||||
|
||||
```
|
||||
The markdown format results is as below:
|
||||
```shell
|
||||
The markdown format results are as below:
|
||||
|
||||
| dataset | version | metric | mode | Qwen2.5-7B-Instruct-vLLM-API |
|
||||
|----- | ----- | ----- | ----- | -----|
|
||||
|
||||
10
docs/source/developer_guide/performance_and_debug/index.md
Normal file
10
docs/source/developer_guide/performance_and_debug/index.md
Normal file
@@ -0,0 +1,10 @@
|
||||
# Performance and Debug
|
||||
|
||||
::::{toctree}
|
||||
:caption: Performance and Debug
|
||||
:maxdepth: 1
|
||||
performance_benchmark
|
||||
optimization_and_tuning
|
||||
service_profiling_guide
|
||||
msprobe_guide
|
||||
::::
|
||||
@@ -0,0 +1,450 @@
|
||||
# MSProbe Debugging Guide
|
||||
|
||||
During inference or training runs we often encounter accuracy anomalies such as outputs drifting away from the expectation, unstable numerical behavior (NaN/Inf), or predictions that no longer match the labels. To pinpoint the root cause we have to monitor and capture intermediate data produced while the model executes—feature maps, weights, activations, and layer outputs. By capturing key tensors at specific stages, logging I/O pairs for the core layers, and retaining contextual metadata (prompts, tensor dtypes, hardware configuration, etc.), we can systematically trace where the accuracy degradation or numerical error started. This guide describes the end-to-end workflow for diagnosing accuracy issues for AI models (with a focus on vllm-ascend services): preparation, data capture, and analysis & verification.
|
||||
|
||||
For more details, see [Ascend/msprobe](https://gitcode.com/Ascend/msprobe).
|
||||
|
||||
## 0. Background Concepts
|
||||
|
||||
`msprobe` supports three accuracy levels:
|
||||
|
||||
- **L0**: dumps tensors at the module level and generates `construct.json` so that visualization tools can rebuild the network structure. A model or submodule handle must be passed in.
|
||||
- **L1**: collects operator-level statistics only, which is suitable for lightweight troubleshooting.
|
||||
- **mix**: captures both structural information and operator statistics, which is useful when you need both graph reconstruction and numerical comparisons.
|
||||
|
||||
## 1. Prerequisites
|
||||
|
||||
### 1.1 Install `msprobe`
|
||||
|
||||
Install msprobe with pip:
|
||||
|
||||
```bash
|
||||
pip install mindstudio-probe
|
||||
```
|
||||
|
||||
### 1.2 Graph mode dump (optional)
|
||||
|
||||
If you need to dump cudagraph graphs, you need to install from source code:
|
||||
|
||||
1. Install `aclgraph_dump` from source code:
|
||||
|
||||
```bash
|
||||
git clone https://gitcode.com/Ascend/msprobe.git
|
||||
cd msprobe
|
||||
python3 setup.py bdist_wheel --include-mod=aclgraph_dump --no-check
|
||||
pip install dist/*.whl
|
||||
```
|
||||
|
||||
## 2. Collecting Data with `msprobe`
|
||||
|
||||
We generally follow a coarse-to-fine strategy when capturing data. First, identify the token where the issue shows up, and then decide which range needs to be sampled around that token. The typical workflow is described below.
|
||||
|
||||
### 2.1 Prepare the dump configuration content
|
||||
|
||||
Prepare configuration content that can be parsed by `PrecisionDebugger`. You can use either of the following ways:
|
||||
|
||||
- Pass the config object directly through `--additional-config.dump_config`.
|
||||
- Pass a config file path through `--additional-config.dump_config_path`.
|
||||
|
||||
Common fields are:
|
||||
|
||||
| Field | Description | Required | Eager Mode | Graph Mode |
|
||||
|:-----------:|:-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:--------:|:-------------:|:-------------:|
|
||||
| `task` | Type of dump task. Common PyTorch values include `"statistics"` and `"tensor"`. A statistics task collects tensor statistics (mean, variance, max, min, etc.) while a tensor task captures arbitrary tensors. | Yes | ✅ | ✅ |
|
||||
| `dump_path` | Directory where dump results are stored. When omitted, `msprobe` uses its default path. | No | ✅ | ✅ |
|
||||
| `rank` | Ranks to sample. An empty list collects every rank. For single-card tasks, you must set this field to `[]`. | No | ✅ | ✅ |
|
||||
| `step` | Token iteration(s) to sample. An empty list means every iteration. | No | ✅ | ❌ |
|
||||
| `level` | Dump level string (`"L0"`, `"L1"`, or `"mix"`). `L0` targets `nn.Module`, `L1` targets `torch.api`, and `mix` collects both. | Yes | ✅ | ✅ |
|
||||
| `async_dump`| Whether to enable asynchronous dump (supported for PyTorch `statistics`/`tensor` tasks). Defaults to `false`. | No | ✅ | ❌ |
|
||||
| `scope` | Module range to sample. An empty list collects every module. | No | ✅ | ❌ |
|
||||
| `dump_enable` | Dynamic switch for enabling/disabling dump in `PrecisionDebugger` during one running training/inference job. This allows turning dump on or off on demand in the same job. | No | ✅ | ❌ |
|
||||
| `list` | Operator range to sample. An empty list collects every operator. | No | ✅ | ✅ |
|
||||
|
||||
To restrict the operators that are captured, configure the `list` block:
|
||||
|
||||
- `scope` (list[str]): In PyTorch PyNative scenarios this field restricts the dump range. Provide two module or API names that follow the tool's naming convention to lock a range; only data between the two names will be dumped. Examples:
|
||||
|
||||
```json
|
||||
"scope": ["Module.conv1.Conv2d.forward.0", "Module.fc2.Linear.forward.0"]
|
||||
"scope": ["Cell.conv1.Conv2d.forward.0", "Cell.fc2.Dense.forward.0"]
|
||||
"scope": ["Tensor.add.0.forward", "Functional.square.2.forward"]
|
||||
```
|
||||
|
||||
The `level` setting determines what can be provided—modules when `level=L0`, APIs when `level=L1`, and either modules or APIs when `level=mix`.
|
||||
|
||||
- `list` (list[str]): Custom operator list. Options include:
|
||||
- Supply the full names of specific APIs in PyTorch pynative scenarios to only dump those APIs. Example: `"list": ["Tensor.permute.1.forward", "Tensor.transpose.2.forward", "Torch.relu.3.forward"]`.
|
||||
- When `level=mix`, you can provide module names so that the dump expands to everything produced while the module is running. Example: `"list": ["Module.module.language_model.encoder.layers.0.mlp.ParallelMlp.forward.0"]`.
|
||||
- Provide a substring such as `"list": ["relu"]` to dump every API whose name contains the substring. When `level=mix`, modules whose names contain the substring are also expanded.
|
||||
|
||||
Example configuration:
|
||||
eager mode:
|
||||
|
||||
```json
|
||||
{
|
||||
"task": "statistics",
|
||||
"dump_path": "/home/data_dump",
|
||||
"rank": [],
|
||||
"step": [],
|
||||
"level": "L1",
|
||||
"async_dump": false,
|
||||
|
||||
"statistics": {
|
||||
"scope": [],
|
||||
"list": [],
|
||||
"tensor_list": [],
|
||||
"data_mode": ["all"],
|
||||
"summary_mode": "statistics"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Graph mode:
|
||||
|
||||
```json
|
||||
{
|
||||
"task": "statistics",
|
||||
"level": "L1",
|
||||
"dump_path": "/home/data_dump",
|
||||
"statistics": {
|
||||
"list": []
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## 3. Enable `msprobe` in vllm-ascend
|
||||
|
||||
1. Start vLLM and pass the dump config content through `--additional-config`:
|
||||
|
||||
```bash
|
||||
vllm serve Qwen/Qwen2.5-0.5B-Instruct \
|
||||
--dtype bfloat16 \
|
||||
--host 0.0.0.0 \
|
||||
--port 8000 \
|
||||
--additional-config '{
|
||||
"dump_config": {
|
||||
"task": "statistics",
|
||||
"level": "L1",
|
||||
"dump_path": "/data/msprobe_dump",
|
||||
"statistics": {
|
||||
"list": []
|
||||
}
|
||||
}
|
||||
}' &
|
||||
```
|
||||
|
||||
Compatibility mode (legacy) is still supported:
|
||||
|
||||
```bash
|
||||
vllm serve Qwen/Qwen2.5-0.5B-Instruct \
|
||||
--dtype bfloat16 \
|
||||
--host 0.0.0.0 \
|
||||
--port 8000 \
|
||||
--additional-config '{"dump_config_path": "/data/msprobe_config.json"}' &
|
||||
```
|
||||
|
||||
## 4. Send requests and collect dumps
|
||||
|
||||
1. Send inference requests as usual, for example:
|
||||
|
||||
```bash
|
||||
curl http://localhost:8000/v1/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "Qwen/Qwen2.5-0.5B-Instruct",
|
||||
"prompt": "Explain gravity in one sentence.",
|
||||
"max_completion_tokens": 32,
|
||||
"temperature": 0
|
||||
}' | python -m json.tool
|
||||
```
|
||||
|
||||
2. Each request drives the sequence `msprobe: start -> forward -> stop -> step`. The runner invokes `step()` on every code path, so you always get a complete dataset even if inference returns early.
|
||||
|
||||
3. Dump files are written into `dump_path`. They usually contain:
|
||||
- Tensor files grouped by operator/module.
|
||||
- `dump.json`, which records metadata such as dtype, shape, min/max, and `requires_grad`.
|
||||
- `construct.json`, which is generated when `level` is `L0` or `mix` (required for visualization).
|
||||
|
||||
Example directory layout:
|
||||
eager mode:
|
||||
|
||||
```text
|
||||
├── dump_path
|
||||
│ ├── step0
|
||||
│ │ ├── rank0
|
||||
│ │ │ ├── dump_tensor_data
|
||||
│ │ │ │ ├── Tensor.permute.1.forward.pt # Format: {api_type}.{api_name}.{call_count}.forward.{input/output}.{arg_index}.
|
||||
│ │ │ │ │ # arg_index is the nth input or output of the API. If an input is a list, keep numbering with decimals (e.g., 1.1 is the first element of the first argument).
|
||||
│ │ │ │ ├── Module.conv1.Conv2d.forward.0.input.0.pt # Format: {Module}.{module_name}.{class_name}.forward.{call_count}.{input/output}.{arg_index}.
|
||||
│ │ │ │ └── Module.conv1.Conv2d.forward.0.parameters.bias.pt # Module parameter data: {Module}.{module_name}.{class_name}.forward.{call_count}.parameters.{parameter_name}.
|
||||
│ │ │ │ # When the `model` argument passed to dump is a List[torch.nn.Module] or Tuple[torch.nn.Module], module-level data names also include the index inside the list ({Module}.{index}.*), e.g., Module.0.conv1.Conv2d.forward.0.input.0.pt.
|
||||
│ │ │ ├── dump.json
|
||||
│ │ │ ├── stack.json
|
||||
│ │ │ ├── dump_error_info.log
|
||||
│ │ │ └── construct.json
|
||||
│ │ ├── rank1
|
||||
│ │ │ ├── dump_tensor_data
|
||||
│ │ │ │ └── ...
|
||||
│ │ │ ├── dump.json
|
||||
│ │ │ ├── stack.json
|
||||
│ │ │ ├── dump_error_info.log
|
||||
│ │ │ └── construct.json
|
||||
│ │ ├── ...
|
||||
│ │ │
|
||||
│ │ └── rank7
|
||||
│ ├── step1
|
||||
│ │ ├── ...
|
||||
│ ├── step2
|
||||
```
|
||||
|
||||
- `rank`: Device ID. Each card writes its data to the corresponding `rank{ID}` directory. In non-distributed scenarios the directory is simply named `rank`.
|
||||
- `dump_tensor_data`: Tensor payloads that were collected.
|
||||
- `dump.json`: Statistics for the forward data of each API or module, including names, dtype, shape, max, min, mean, L2 norm (square root of the L2 variance), and CRC-32 when `summary_mode="md5"`. See [dump.json file description](#dumpjson-file-description) for details.
|
||||
- `dump_error_info.log`: Present only when the dump tool encountered an error and records the failure log.
|
||||
- `stack.json`: Call stacks for APIs/modules.
|
||||
- `construct.json`: Hierarchical structure description. Empty when `level=L1`.
|
||||
|
||||
graph mode:
|
||||
|
||||
```text
|
||||
L0_dump
|
||||
├── step0
|
||||
│ └── rank0
|
||||
│ └── dump.json
|
||||
├── step1
|
||||
│ └── rank0
|
||||
│ └── dump.json
|
||||
├── step2
|
||||
│ └── rank0
|
||||
│ └── dump.json
|
||||
├── step3
|
||||
│ └── rank0
|
||||
│ └── dump.json
|
||||
├── step4
|
||||
│ └── rank0
|
||||
│ └── dump.json
|
||||
└── step5
|
||||
└── rank0
|
||||
└── dump.json
|
||||
```
|
||||
|
||||
- `dump.json`: See [dump.json file description](#dumpjson-file-description) for details.
|
||||
|
||||
## 5. Analyze the results
|
||||
|
||||
### 5.1 Prerequisites
|
||||
|
||||
You typically need two dump datasets: one from the "problem side" (the run that exposes the accuracy or numerical error) and another from the "benchmark side" (a good baseline). These datasets do not have to be identical—they can come from different branches, framework versions, or even alternative implementations (operator substitutions, different graph-optimization switches, etc.). As long as they use the same or similar inputs, hardware topology, and sampling points (step/token), `msprobe` can compare them and locate the divergent nodes. If you cannot find a perfectly clean benchmark, start by capturing the problem-side data, craft the smallest reproducible case by hand, and perform a self-comparison. Below we assume the problem dump is `problem_dump` and the benchmark dump is `bench_dump`.
|
||||
|
||||
### 5.2 Visualization
|
||||
|
||||
Use `msprobe graph_visualize` to build or compare graphs, then open the generated `*.vis.db` file(s) with TensorBoard (`tb_graph_ascend` plugin).
|
||||
|
||||
1. Ensure dump data is visualization-ready:
|
||||
- Dump level must be `L0` or `mix` so `construct.json` is non-empty.
|
||||
- Each rank directory should contain `dump.json`, `stack.json`, and `construct.json`.
|
||||
|
||||
2. Choose command mode:
|
||||
- Single-graph build:
|
||||
|
||||
```bash
|
||||
msprobe graph_visualize -tp <target_path> -o <output_path>
|
||||
```
|
||||
|
||||
- Graph comparison:
|
||||
|
||||
```bash
|
||||
msprobe graph_visualize -tp <target_path> -gp <golden_path> -o <output_path>
|
||||
```
|
||||
|
||||
- Common optional flags:
|
||||
- `-oc` / `--overflow_check`: enable overflow marking
|
||||
- `-fm` / `--fuzzy_match`: enable fuzzy matching for node mapping
|
||||
- `-lm` / `--layer_mapping [mapping.yaml]`: cross-framework/layer mapping compare
|
||||
- `-tensor_log`: print per-node compare log (tensor dump scenarios)
|
||||
- `-progress_log`: print detailed progress log
|
||||
|
||||
3. Path granularity is auto-detected by `graph_visualize`:
|
||||
- Single-rank: `.../step0/rank0`
|
||||
- Multi-rank (batch): `.../step0`
|
||||
- Multi-step (batch): dump root path containing `step*`
|
||||
|
||||
4. Output files:
|
||||
- Single-graph build: `build_{timestamp}.vis.db`
|
||||
- Graph comparison: `compare_{timestamp}.vis.db`
|
||||
|
||||
5. Launch TensorBoard with the output directory:
|
||||
|
||||
```bash
|
||||
tensorboard --logdir <output_path> --bind_all --port <optional_port>
|
||||
```
|
||||
|
||||
6. In the visualization UI, inspect structure and numeric differences:
|
||||
- Switch rank/step to locate unstable nodes quickly.
|
||||
- Use search/filter to focus on target ops/modules.
|
||||
- For compare mode, prioritize highlighted high-difference nodes and trace surrounding I/O/parameters.
|
||||
|
||||
## 6. Troubleshooting
|
||||
|
||||
- `RuntimeError: Please enforce eager mode`: Restart vLLM and add the `--enforce-eager` flag.
|
||||
- No dump files: Confirm that the JSON path is correct and every node has write permission. In distributed scenarios set `keep_all_ranks` so that every rank writes its own dump.
|
||||
- Dumps are too large: Start with a `statistics` task to locate abnormal tensors, then narrow the scope with `scope`/`list`/`tensor_list`, `filters`, `token_range`, etc.
|
||||
|
||||
---
|
||||
|
||||
## Appendix
|
||||
|
||||
### dump.json file description
|
||||
|
||||
#### L0 level
|
||||
|
||||
An L0 `dump.json` contains forward I/O for modules together with parameters. Using PyTorch's `Conv2d` as an example, the network code looks like:
|
||||
|
||||
`output = self.conv2(input) # self.conv2 = torch.nn.Conv2d(64, 128, 5, padding=2, bias=True)`
|
||||
|
||||
`dump.json` contains the following entries:
|
||||
|
||||
- `Module.conv2.Conv2d.forward.0`: Forward data of the module. `input_args` represents positional inputs, `input_kwargs` represents keyword inputs, `output` stores forward outputs, and `parameters` stores weights/biases.
|
||||
|
||||
**Note**: When the `model` parameter passed to the dump API is `List[torch.nn.Module]` or `Tuple[torch.nn.Module]`, module-level names include the index inside the list (`{Module}.{index}.*`). Example: `Module.0.conv1.Conv2d.forward.0`.
|
||||
|
||||
```json
|
||||
{
|
||||
"task": "tensor",
|
||||
"level": "L0",
|
||||
"framework": "pytorch",
|
||||
"dump_data_dir": "/dump/path",
|
||||
"data": {
|
||||
"Module.conv2.Conv2d.forward.0": {
|
||||
"input_args": [
|
||||
{
|
||||
"type": "torch.Tensor",
|
||||
"dtype": "torch.float32",
|
||||
"shape": [
|
||||
8,
|
||||
16,
|
||||
14,
|
||||
14
|
||||
],
|
||||
"Max": 1.638758659362793,
|
||||
"Min": 0.0,
|
||||
"Mean": 0.2544615864753723,
|
||||
"Norm": 70.50277709960938,
|
||||
"requires_grad": true,
|
||||
"data_name": "Module.conv2.Conv2d.forward.0.input.0.pt"
|
||||
}
|
||||
],
|
||||
"input_kwargs": {},
|
||||
"output": [
|
||||
{
|
||||
"type": "torch.Tensor",
|
||||
"dtype": "torch.float32",
|
||||
"shape": [
|
||||
8,
|
||||
32,
|
||||
10,
|
||||
10
|
||||
],
|
||||
"Max": 1.6815717220306396,
|
||||
"Min": -1.5120246410369873,
|
||||
"Mean": -0.025344856083393097,
|
||||
"Norm": 149.65576171875,
|
||||
"requires_grad": true,
|
||||
"data_name": "Module.conv2.Conv2d.forward.0.output.0.pt"
|
||||
}
|
||||
],
|
||||
"parameters": {
|
||||
"weight": {
|
||||
"type": "torch.Tensor",
|
||||
"dtype": "torch.float32",
|
||||
"shape": [
|
||||
32,
|
||||
16,
|
||||
5,
|
||||
5
|
||||
],
|
||||
"Max": 0.05992485210299492,
|
||||
"Min": -0.05999220535159111,
|
||||
"Mean": -0.0006165213999338448,
|
||||
"Norm": 3.421217441558838,
|
||||
"requires_grad": true,
|
||||
"data_name": "Module.conv2.Conv2d.forward.0.parameters.weight.pt"
|
||||
},
|
||||
"bias": {
|
||||
"type": "torch.Tensor",
|
||||
"dtype": "torch.float32",
|
||||
"shape": [
|
||||
32
|
||||
],
|
||||
"Max": 0.05744686722755432,
|
||||
"Min": -0.04894155263900757,
|
||||
"Mean": 0.006410328671336174,
|
||||
"Norm": 0.17263513803482056,
|
||||
"requires_grad": true,
|
||||
"data_name": "Module.conv2.Conv2d.forward.0.parameters.bias.pt"
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
#### L1 level
|
||||
|
||||
An L1 `dump.json` records forward I/O for APIs. Using PyTorch's `relu` function as an example (`output = torch.nn.functional.relu(input)`), the file contains:
|
||||
|
||||
- `Functional.relu.0.forward`: Forward data of the API. `input_args` are positional inputs, `input_kwargs` are keyword inputs, and `output` stores the forward outputs.
|
||||
|
||||
```json
|
||||
{
|
||||
"task": "tensor",
|
||||
"level": "L1",
|
||||
"framework": "pytorch",
|
||||
"dump_data_dir":"/dump/path",
|
||||
"data": {
|
||||
"Functional.relu.0.forward": {
|
||||
"input_args": [
|
||||
{
|
||||
"type": "torch.Tensor",
|
||||
"dtype": "torch.float32",
|
||||
"shape": [
|
||||
32,
|
||||
16,
|
||||
28,
|
||||
28
|
||||
],
|
||||
"Max": 1.3864083290100098,
|
||||
"Min": -1.3364859819412231,
|
||||
"Mean": 0.03711778670549393,
|
||||
"Norm": 236.20692443847656,
|
||||
"requires_grad": true,
|
||||
"data_name": "Functional.relu.0.forward.input.0.pt"
|
||||
}
|
||||
],
|
||||
"input_kwargs": {},
|
||||
"output": [
|
||||
{
|
||||
"type": "torch.Tensor",
|
||||
"dtype": "torch.float32",
|
||||
"shape": [
|
||||
32,
|
||||
16,
|
||||
28,
|
||||
28
|
||||
],
|
||||
"Max": 1.3864083290100098,
|
||||
"Min": 0.0,
|
||||
"Mean": 0.16849493980407715,
|
||||
"Norm": 175.23345947265625,
|
||||
"requires_grad": true,
|
||||
"data_name": "Functional.relu.0.forward.output.0.pt"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
#### mix level
|
||||
|
||||
A `mix` dump.json contains both L0 and L1 level data; the file format is the same as the examples above.
|
||||
@@ -0,0 +1,248 @@
|
||||
# Optimization and Tuning
|
||||
|
||||
This guide aims to help users improve vLLM Ascend performance at the system level. It includes OS configuration, library optimization, deployment guide, and so on. Any feedback is welcome.
|
||||
|
||||
## Preparation
|
||||
|
||||
### 1.Run the container
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# Update DEVICE according to your device (/dev/davinci[0-7])
|
||||
export DEVICE=/dev/davinci0
|
||||
# Update the cann base image
|
||||
export IMAGE=m.daocloud.io/quay.io/ascend/cann:|cann_image_tag|
|
||||
docker run --rm \
|
||||
--name performance-test \
|
||||
--shm-size=1g \
|
||||
--device $DEVICE \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
### 2.Configure your environment
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# Configure the mirror
|
||||
echo "deb https://mirrors.tuna.tsinghua.edu.cn/ubuntu-ports/ jammy main restricted universe multiverse" > /etc/apt/sources.list && \
|
||||
echo "deb-src https://mirrors.tuna.tsinghua.edu.cn/ubuntu-ports/ jammy main restricted universe multiverse" >> /etc/apt/sources.list && \
|
||||
echo "deb https://mirrors.tuna.tsinghua.edu.cn/ubuntu-ports/ jammy-updates main restricted universe multiverse" >> /etc/apt/sources.list && \
|
||||
echo "deb-src https://mirrors.tuna.tsinghua.edu.cn/ubuntu-ports/ jammy-updates main restricted universe multiverse" >> /etc/apt/sources.list && \
|
||||
echo "deb https://mirrors.tuna.tsinghua.edu.cn/ubuntu-ports/ jammy-backports main restricted universe multiverse" >> /etc/apt/sources.list && \
|
||||
echo "deb-src https://mirrors.tuna.tsinghua.edu.cn/ubuntu-ports/ jammy-backports main restricted universe multiverse" >> /etc/apt/sources.list && \
|
||||
echo "deb https://mirrors.tuna.tsinghua.edu.cn/ubuntu-ports/ jammy-security main restricted universe multiverse" >> /etc/apt/sources.list && \
|
||||
echo "deb-src https://mirrors.tuna.tsinghua.edu.cn/ubuntu-ports/ jammy-security main restricted universe multiverse" >> /etc/apt/sources.list
|
||||
|
||||
# Install os packages
|
||||
apt update && apt install wget gcc g++ libnuma-dev git vim -y
|
||||
```
|
||||
|
||||
### 3.Install vLLM and vLLM Ascend
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# Install necessary dependencies
|
||||
pip config set global.index-url https://pypi.tuna.tsinghua.edu.cn/simple
|
||||
pip install modelscope pandas datasets gevent sacrebleu rouge_score pybind11 pytest
|
||||
|
||||
# Configure this var to speed up model download
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
```
|
||||
|
||||
Please follow the [Installation Guide](https://docs.vllm.ai/projects/ascend/en/latest/installation.html) to make sure vLLM and vLLM Ascend are installed correctly.
|
||||
|
||||
:::{note}
|
||||
Make sure your vLLM and vLLM Ascend are installed after your Python configuration is completed, because these packages will build binary files using python in current environment. If you install vLLM and vLLM Ascend before completing [Configure your environment](#2configure-your-environment), the binary files will not use the optimized python.
|
||||
:::
|
||||
|
||||
## Optimizations
|
||||
|
||||
### 1. Memory Allocator Optimization
|
||||
|
||||
#### 1.1. jemalloc
|
||||
|
||||
**jemalloc** is a memory allocator that improves performance for multi-threaded scenarios and can reduce memory fragmentation. jemalloc uses a local thread memory manager to allocate variables, which can avoid lock contention between threads and can hugely optimize performance.
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# Install jemalloc
|
||||
sudo apt update
|
||||
sudo apt install libjemalloc2
|
||||
|
||||
# Configure jemalloc
|
||||
export LD_PRELOAD=/usr/lib/"$(uname -i)"-linux-gnu/libjemalloc.so.2:$LD_PRELOAD
|
||||
```
|
||||
|
||||
#### 1.2. Tcmalloc
|
||||
|
||||
**TCMalloc (Thread Caching Malloc)** is a universal memory allocator that improves overall performance while ensuring low latency by introducing a multi-level cache structure, reducing lock contention and optimizing large object processing flow. Find more [details](https://www.hiascend.com/document/detail/zh/Pytorch/700/ptmoddevg/trainingmigrguide/performance_tuning_0068.html).
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# Install tcmalloc
|
||||
sudo apt update
|
||||
sudo apt install libgoogle-perftools4 libgoogle-perftools-dev
|
||||
|
||||
# Get the location of libtcmalloc.so*
|
||||
find /usr -name libtcmalloc.so*
|
||||
|
||||
# Make the priority of tcmalloc higher
|
||||
# The <path> is the location of libtcmalloc.so we get from the upper command
|
||||
# Example: "$LD_PRELOAD:/usr/lib/aarch64-linux-gnu/libtcmalloc.so"
|
||||
export LD_PRELOAD="$LD_PRELOAD:<path>"
|
||||
|
||||
# Verify your configuration
|
||||
# The path of libtcmalloc.so will be contained in the result if your configuration is valid
|
||||
ldd `which python`
|
||||
```
|
||||
|
||||
### 2. `torch_npu` Optimization
|
||||
|
||||
Some performance tuning features in `torch_npu` are controlled by environment variables. Some features and their related environment variables are shown below.
|
||||
|
||||
Memory optimization:
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# Upper limit of memory block splitting allowed (MB): Setting this parameter can prevent large memory blocks from being split.
|
||||
export PYTORCH_NPU_ALLOC_CONF="max_split_size_mb:250"
|
||||
```
|
||||
|
||||
or
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# When operators on the communication stream have dependencies, they all need to be ended before being released for reuse. The logic of multi-stream reuse is to release the memory on the communication stream in advance so that the computing stream can be reused.
|
||||
export PYTORCH_NPU_ALLOC_CONF="expandable_segments:True"
|
||||
```
|
||||
|
||||
Scheduling optimization:
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# Optimize operator delivery queue. This will affect the memory peak value, and may degrade if the memory is tight.
|
||||
export TASK_QUEUE_ENABLE=2
|
||||
|
||||
# This will greatly improve the CPU bottleneck model and ensure the same performance for the NPU bottleneck model.
|
||||
export CPU_AFFINITY_CONF=1
|
||||
```
|
||||
|
||||
### 3. CANN Optimization
|
||||
|
||||
#### 3.1. HCCL Optimization
|
||||
|
||||
There are some performance tuning features in HCCL, which are controlled by environment variables.
|
||||
|
||||
You can configure HCCL to use "AIV" mode to optimize performance by setting the environment variable shown below. In "AIV" mode, the communication is scheduled by AI vector core directly with RoCE, instead of being scheduled by AI CPU.
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
```
|
||||
|
||||
Plus, there are more features for performance optimization in specific scenarios, which are shown below.
|
||||
|
||||
- `HCCL_INTRA_ROCE_ENABLE`: Use RDMA link instead of SDMA link between two 8Ps as the mesh interconnect link. Find more [details](https://www.hiascend.com/document/detail/zh/Pytorch/600/ptmoddevg/trainingmigrguide/performance_tuning_0044.html).
|
||||
- `HCCL_RDMA_TC`: Use this var to configure traffic class of RDMA NIC. Find more [details](https://www.hiascend.com/document/detail/zh/Pytorch/600/ptmoddevg/trainingmigrguide/performance_tuning_0045.html).
|
||||
- `HCCL_RDMA_SL`: Use this var to configure service level of RDMA NIC. Find more [details](https://www.hiascend.com/document/detail/zh/Pytorch/600/ptmoddevg/trainingmigrguide/performance_tuning_0046.html).
|
||||
- `HCCL_BUFFSIZE`: Use this var to control the cache size for sharing data between two NPUs. Find more [details](https://www.hiascend.com/document/detail/zh/Pytorch/600/ptmoddevg/trainingmigrguide/performance_tuning_0047.html).
|
||||
|
||||
### 4. Kernel Optimization
|
||||
|
||||
This section describes operating system–level optimizations applied on the host machine (bare metal or Kubernetes node) to improve performance stability, latency, and throughput for inference workloads.
|
||||
|
||||
:::{note}
|
||||
These settings must be applied on the host OS and with root privileges. Not inside containers.
|
||||
:::
|
||||
|
||||
#### 4.1 Set CPU Frequency Governor to `performance`
|
||||
|
||||
Set CPU Frequency Governor to `performance`
|
||||
|
||||
```shell
|
||||
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
|
||||
```
|
||||
|
||||
Purpose
|
||||
|
||||
- Forces all CPU cores to run under the `performance` governor
|
||||
- Disables dynamic frequency scaling (e.g., `ondemand`, `powersave`)
|
||||
|
||||
Benefits
|
||||
|
||||
- Keeps CPU cores at maximum frequency
|
||||
- Reduces latency jitter
|
||||
- Improves predictability for inference workloads
|
||||
|
||||
#### 4.2 Disable Swap Usage
|
||||
|
||||
```shell
|
||||
sysctl -w vm.swappiness=0
|
||||
```
|
||||
|
||||
Purpose
|
||||
|
||||
- Minimizes the kernel’s tendency to swap memory pages to disk
|
||||
|
||||
Benefits
|
||||
|
||||
- Prevents severe latency spikes caused by swapping
|
||||
- Improves stability for large in-memory models
|
||||
|
||||
Notes
|
||||
|
||||
- For inference workloads, swap can introduce second-level latency
|
||||
- Recommended values are `0` or `1`
|
||||
|
||||
#### 4.3 Disable Automatic NUMA Balancing
|
||||
|
||||
```shell
|
||||
sysctl -w kernel.numa_balancing=0
|
||||
```
|
||||
|
||||
Purpose
|
||||
|
||||
- Disables the kernel’s automatic NUMA page migration mechanism
|
||||
|
||||
Benefits
|
||||
|
||||
- Prevents background memory page migrations
|
||||
- Reduces unpredictable memory access latency
|
||||
- Improves performance stability on NUMA systems
|
||||
|
||||
Recommended For
|
||||
|
||||
- Multi-socket servers
|
||||
- Ascend / NPU deployments with explicit NUMA binding
|
||||
- Systems with manually managed CPU and memory affinity
|
||||
|
||||
#### 4.4 Increase Scheduler Migration Cost
|
||||
|
||||
```shell
|
||||
sysctl -w kernel.sched_migration_cost_ns=50000
|
||||
```
|
||||
|
||||
Purpose
|
||||
|
||||
- Increases the cost for the scheduler to migrate tasks between CPU cores
|
||||
|
||||
Benefits
|
||||
|
||||
- Reduces frequent thread migration
|
||||
- Improves CPU cache locality
|
||||
- Lowers latency jitter for inference workloads
|
||||
|
||||
Parameter Details
|
||||
|
||||
- Unit: nanoseconds (ns)
|
||||
- Typical recommended range: 50000–100000
|
||||
- Higher values encourage threads to stay on the same CPU core
|
||||
@@ -0,0 +1,249 @@
|
||||
# Performance Benchmark
|
||||
|
||||
This document details the benchmark methodology for vllm-ascend, aimed at evaluating the performance under a variety of workloads. To maintain alignment with vLLM, we use the [benchmark](https://github.com/vllm-project/vllm/tree/main/benchmarks) script provided by the vllm project.
|
||||
|
||||
**Benchmark Coverage**: We measure offline E2E latency and throughput, and fixed-QPS online serving benchmarks. For more details, see [vllm-ascend benchmark scripts](https://github.com/vllm-project/vllm-ascend/tree/main/benchmarks).
|
||||
|
||||
**Legend Description**:
|
||||
|
||||
- ✅ = Supported
|
||||
- 🟡 = Partial / Work in progress
|
||||
- 🚧 = Under development
|
||||
|
||||
## 1. Run docker container
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# Update DEVICE according to your device (/dev/davinci[0-7])
|
||||
export DEVICE=/dev/davinci7
|
||||
export IMAGE=m.daocloud.io/quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--device $DEVICE \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-p 8000:8000 \
|
||||
-e VLLM_USE_MODELSCOPE=True \
|
||||
-it $IMAGE \
|
||||
/bin/bash
|
||||
```
|
||||
|
||||
## 2. Install dependencies
|
||||
|
||||
```bash
|
||||
cd /workspace/vllm-ascend
|
||||
pip config set global.index-url https://mirrors.tuna.tsinghua.edu.cn/pypi/web/simple
|
||||
pip install -r benchmarks/requirements-bench.txt
|
||||
```
|
||||
|
||||
## 3. Run basic benchmarks
|
||||
|
||||
This section introduces how to perform performance testing using the benchmark suite built into vLLM.
|
||||
|
||||
### 3.1 Dataset
|
||||
|
||||
VLLM supports a variety of [datasets](https://github.com/vllm-project/vllm/blob/main/vllm/benchmarks/datasets/datasets.py).
|
||||
|
||||
<style>
|
||||
th {
|
||||
min-width: 0 !important;
|
||||
}
|
||||
</style>
|
||||
|
||||
| Dataset | Online | Offline | Data Path |
|
||||
|---------|--------|---------|-----------|
|
||||
| ShareGPT | ✅ | ✅ | `wget https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered/resolve/main/ShareGPT_V3_unfiltered_cleaned_split.json` |
|
||||
| ShareGPT4V (Image) | ✅ | ✅ | `wget https://huggingface.co/datasets/Lin-Chen/ShareGPT4V/resolve/main/sharegpt4v_instruct_gpt4-vision_cap100k.json`<br>Note that the images need to be downloaded separately. For example, to download COCO's 2017 Train images:<br>`wget http://images.cocodataset.org/zips/train2017.zip` |
|
||||
| ShareGPT4Video (Video) | ✅ | ✅ | `git clone https://huggingface.co/datasets/ShareGPT4Video/ShareGPT4Video` |
|
||||
| BurstGPT | ✅ | ✅ | `wget https://github.com/HPMLL/BurstGPT/releases/download/v1.1/BurstGPT_without_fails_2.csv` |
|
||||
| Sonnet (deprecated) | ✅ | ✅ | Local file: `benchmarks/sonnet.txt` |
|
||||
| Random | ✅ | ✅ | `synthetic` |
|
||||
| RandomMultiModal (Image/Video) | 🟡 | 🚧 | `synthetic` |
|
||||
| RandomForReranking | ✅ | ✅ | `synthetic` |
|
||||
| Prefix Repetition | ✅ | ✅ | `synthetic` |
|
||||
| HuggingFace-VisionArena | ✅ | ✅ | `lmarena-ai/VisionArena-Chat` |
|
||||
| HuggingFace-MMVU | ✅ | ✅ | `yale-nlp/MMVU` |
|
||||
| HuggingFace-InstructCoder | ✅ | ✅ | `likaixin/InstructCoder` |
|
||||
| HuggingFace-AIMO | ✅ | ✅ | `AI-MO/aimo-validation-aime`, `AI-MO/NuminaMath-1.5`, `AI-MO/NuminaMath-CoT` |
|
||||
| HuggingFace-Other | ✅ | ✅ | `lmms-lab/LLaVA-OneVision-Data`, `Aeala/ShareGPT_Vicuna_unfiltered` |
|
||||
| HuggingFace-MTBench | ✅ | ✅ | `philschmid/mt-bench` |
|
||||
| HuggingFace-Blazedit | ✅ | ✅ | `vdaita/edit_5k_char`, `vdaita/edit_10k_char` |
|
||||
| Spec Bench | ✅ | ✅ | `wget https://raw.githubusercontent.com/hemingkx/Spec-Bench/refs/heads/main/data/spec_bench/question.jsonl` |
|
||||
| Custom | ✅ | ✅ | Local file: `data.jsonl` |
|
||||
|
||||
:::{note}
|
||||
The datasets mentioned above are all links to datasets on huggingface.
|
||||
The dataset's `dataset-name` should be set to `hf`.
|
||||
For local `dataset-path`, please set `hf-name` to its Hugging Face ID like
|
||||
|
||||
```bash
|
||||
--dataset-path /datasets/VisionArena-Chat/ --hf-name lmarena-ai/VisionArena-Chat
|
||||
```
|
||||
|
||||
:::
|
||||
|
||||
### 3.2 Run basic benchmark
|
||||
|
||||
#### 3.2.1 Online serving
|
||||
|
||||
First start serving your model:
|
||||
|
||||
```bash
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
vllm serve Qwen/Qwen3-8B
|
||||
```
|
||||
|
||||
Then run the benchmarking script:
|
||||
|
||||
```bash
|
||||
# download dataset
|
||||
# wget https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered/resolve/main/ShareGPT_V3_unfiltered_cleaned_split.json
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
vllm bench serve \
|
||||
--backend vllm \
|
||||
--model Qwen/Qwen3-8B \
|
||||
--endpoint /v1/completions \
|
||||
--dataset-name sharegpt \
|
||||
--dataset-path <your data path>/ShareGPT_V3_unfiltered_cleaned_split.json \
|
||||
--num-prompts 10
|
||||
```
|
||||
|
||||
If successful, you will see the following output:
|
||||
|
||||
```shell
|
||||
============ Serving Benchmark Result ============
|
||||
Successful requests: 10
|
||||
Failed requests: 0
|
||||
Benchmark duration (s): 19.92
|
||||
Total input tokens: 1374
|
||||
Total generated tokens: 2663
|
||||
Request throughput (req/s): 0.50
|
||||
Output token throughput (tok/s): 133.67
|
||||
Peak output token throughput (tok/s): 312.00
|
||||
Peak concurrent requests: 10.00
|
||||
Total Token throughput (tok/s): 202.64
|
||||
---------------Time to First Token----------------
|
||||
Mean TTFT (ms): 127.10
|
||||
Median TTFT (ms): 136.29
|
||||
P99 TTFT (ms): 137.83
|
||||
-----Time per Output Token (excl. 1st token)------
|
||||
Mean TPOT (ms): 25.85
|
||||
Median TPOT (ms): 25.78
|
||||
P99 TPOT (ms): 26.64
|
||||
---------------Inter-token Latency----------------
|
||||
Mean ITL (ms): 25.78
|
||||
Median ITL (ms): 25.74
|
||||
P99 ITL (ms): 28.85
|
||||
==================================================
|
||||
```
|
||||
|
||||
#### 3.2.2 Offline Throughput Benchmark
|
||||
|
||||
```bash
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
vllm bench throughput \
|
||||
--model Qwen/Qwen3-8B \
|
||||
--dataset-name random \
|
||||
--input-len 128 \
|
||||
--output-len 128
|
||||
```
|
||||
|
||||
If successful, you will see the following output
|
||||
|
||||
```shell
|
||||
Processed prompts: 100%|█| 10/10 [00:03<00:00, 2.74it/s, est. speed input: 351.02 toks/s, output: 351.02 toks/s]
|
||||
Throughput: 2.73 requests/s, 699.93 total tokens/s, 349.97 output tokens/s
|
||||
Total num prompt tokens: 1280
|
||||
Total num output tokens: 1280
|
||||
```
|
||||
|
||||
#### 3.2.3 Multi-Modal Benchmark
|
||||
|
||||
```shell
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
vllm serve Qwen/Qwen2.5-VL-7B-Instruct \
|
||||
--dtype bfloat16 \
|
||||
--limit-mm-per-prompt '{"image": 1}' \
|
||||
--allowed-local-media-path /path/to/sharegpt4v/images
|
||||
```
|
||||
|
||||
```shell
|
||||
export HF_ENDPOINT="https://hf-mirror.com"
|
||||
vllm bench serve --model Qwen/Qwen2.5-VL-7B-Instruct \
|
||||
--backend "openai-chat" \
|
||||
--dataset-name hf \
|
||||
--hf-split train \
|
||||
--endpoint "/v1/chat/completions" \
|
||||
--dataset-path "lmarena-ai/vision-arena-bench-v0.1" \
|
||||
--num-prompts 10 \
|
||||
--no-stream
|
||||
```
|
||||
|
||||
```shell
|
||||
============ Serving Benchmark Result ============
|
||||
Successful requests: 10
|
||||
Failed requests: 0
|
||||
Benchmark duration (s): 4.89
|
||||
Total input tokens: 7191
|
||||
Total generated tokens: 951
|
||||
Request throughput (req/s): 2.05
|
||||
Output token throughput (tok/s): 194.63
|
||||
Peak output token throughput (tok/s): 290.00
|
||||
Peak concurrent requests: 10.00
|
||||
Total Token throughput (tok/s): 1666.35
|
||||
---------------Time to First Token----------------
|
||||
Mean TTFT (ms): 722.22
|
||||
Median TTFT (ms): 589.81
|
||||
P99 TTFT (ms): 1377.02
|
||||
-----Time per Output Token (excl. 1st token)------
|
||||
Mean TPOT (ms): 44.13
|
||||
Median TPOT (ms): 34.58
|
||||
P99 TPOT (ms): 124.72
|
||||
---------------Inter-token Latency----------------
|
||||
Mean ITL (ms): 33.14
|
||||
Median ITL (ms): 28.01
|
||||
P99 ITL (ms): 182.28
|
||||
==================================================
|
||||
```
|
||||
|
||||
#### 3.2.4 Embedding Benchmark
|
||||
|
||||
```shell
|
||||
vllm serve Qwen/Qwen3-Embedding-8B --trust-remote-code
|
||||
```
|
||||
|
||||
```shell
|
||||
# download dataset
|
||||
# wget https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered/resolve/main/ShareGPT_V3_unfiltered_cleaned_split.json
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
vllm bench serve \
|
||||
--model Qwen/Qwen3-Embedding-8B \
|
||||
--backend openai-embeddings \
|
||||
--endpoint /v1/embeddings \
|
||||
--dataset-name sharegpt \
|
||||
--num-prompts 10 \
|
||||
--dataset-path <your dataset path>/datasets/ShareGPT_V3_unfiltered_cleaned_split.json
|
||||
```
|
||||
|
||||
```shell
|
||||
============ Serving Benchmark Result ============
|
||||
Successful requests: 10
|
||||
Failed requests: 0
|
||||
Benchmark duration (s): 0.18
|
||||
Total input tokens: 1372
|
||||
Request throughput (req/s): 56.32
|
||||
Total Token throughput (tok/s): 7726.76
|
||||
----------------End-to-end Latency----------------
|
||||
Mean E2EL (ms): 154.06
|
||||
Median E2EL (ms): 165.57
|
||||
P99 E2EL (ms): 166.66
|
||||
==================================================
|
||||
```
|
||||
@@ -0,0 +1,371 @@
|
||||
# Service Profiling Guide
|
||||
|
||||
In an inference service process, it is sometimes necessary to monitor the internal execution flow of the inference service framework to identify performance issues. By collecting start and end timestamps of key processes, identifying key functions or iterations, recording critical events, and gathering various types of information, performance bottlenecks can be quickly located.
|
||||
|
||||
This guide will walk you through the process of collecting performance data from the vLLM-Ascend service framework and operators. It covers the complete workflow from preparation, collection, analysis, to visualization, helping you quickly get started with performance collection tools.
|
||||
|
||||
Two performance collection solutions are provided below: Ascend PyTorch Profiler and MS Service Profiler. You can choose the appropriate tool for performance analysis and troubleshooting based on your actual requirements.
|
||||
|
||||
## Solution Comparison
|
||||
|
||||
| Feature | Ascend PyTorch Profiler | MS Service Profiler |
|
||||
|:-----|:------------------------|:------------------|
|
||||
| Installation Method | Built-in, no additional installation required | Requires building msserviceprofiler from source |
|
||||
| Collection Granularity | PyTorch operator level | Service framework function level |
|
||||
| Control Method | API request control | Configuration file control |
|
||||
| Applicable Scenarios | Model operator performance analysis | Service framework workflow analysis |
|
||||
| Data Format | ascend_pt format | Chrome Tracing + CSV |
|
||||
| Main Advantage | Operator-level performance analysis | Service framework workflow visualization |
|
||||
| Supported Collection Capabilities | PyTorch operator level | PyTorch operator level and Service framework function level |
|
||||
|
||||
## Quick Selection Guide
|
||||
|
||||
- [**Model Operator Performance** → Use Ascend PyTorch Profiler](#ascend-pytorch-profiler)
|
||||
- [**Service Framework Workflow** → Use MS Service Profiler](#ms-service-profiler)
|
||||
|
||||
---
|
||||
|
||||
## Ascend PyTorch Profiler
|
||||
|
||||
### 0. Installation and Configuration
|
||||
|
||||
No additional packages need to be installed; it can be enabled through command-line configuration. Currently, vLLM enables **python stack** by default, which can significantly inflate the collected performance data. If you do not wish to collect **python stack**, you can disable it using `torch_profiler_with_stack=false`.
|
||||
|
||||
### 1. Preparation for Collection
|
||||
|
||||
Start the online service and set the `--profiler-config` parameter to control the path for saving performance files. After the parameter is set, the collection function is enabled.
|
||||
|
||||
```bash
|
||||
VLLM_PROMPT_SEQ_BUCKET_MAX=128
|
||||
VLLM_PROMPT_SEQ_BUCKET_MIN=128
|
||||
python3 -m vllm.entrypoints.openai.api_server \
|
||||
--port 8080 \
|
||||
--model "facebook/opt-125m" \
|
||||
--tensor-parallel-size 1 \
|
||||
--max-num-seqs 128 \
|
||||
--profiler-config '{"profiler": "torch", "torch_profiler_dir": "./vllm_profile", "torch_profiler_with_stack": false}' \
|
||||
--dtype bfloat16 \
|
||||
--max-model-len 256
|
||||
```
|
||||
|
||||
> Note:**January 19, 2026: The vLLM mainline has deprecated the VLLM_TORCH_PROFILER_DIR environment variable.**[Related PR](https://github.com/vllm-project/vllm-ascend/pull/5928) When using the vLLM Ascend mainline code to collect profiler data, remember to use the `--profiler-config` (online) parameter or the `profiler_config` (offline) parameter.
|
||||
|
||||
### 2. Start Collection
|
||||
|
||||
Performance collection is controlled by sending API requests. You can start collection after stabilizing the actual business data and collect profiling for a few seconds before stopping; or you can start collection first, then send business requests, and finally stop.
|
||||
|
||||
Send the following request to start the profiling service:
|
||||
|
||||
```bash
|
||||
curl -X POST http://localhost:8080/start_profile
|
||||
```
|
||||
|
||||
Send the following request to stop the profiling service:
|
||||
|
||||
```bash
|
||||
curl -X POST http://localhost:8080/stop_profile
|
||||
```
|
||||
|
||||
### 3. Send Requests
|
||||
|
||||
Send requests according to your actual business data. After sending the requests, stop the profiling service, and the data will be automatically saved to the previously configured path:
|
||||
|
||||
```bash
|
||||
curl http://localhost:8080/v1/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "facebook/opt-125m",
|
||||
"prompt": "San Francisco is a",
|
||||
"max_tokens": 7,
|
||||
"temperature": 0
|
||||
}'
|
||||
|
||||
curl -X POST http://localhost:8080/stop_profile
|
||||
```
|
||||
|
||||
### 4. Analyze Data
|
||||
|
||||
Navigate to the `./vllm_profile` directory and locate the generated `*ascend_pt` folder. This folder needs to be analyzed before profiling data can be examined.
|
||||
|
||||
```python
|
||||
from torch_npu.profiler.profiler import analyse
|
||||
analyse("./vllm_profile/localhost.localdomain_*_ascend_pt/")
|
||||
```
|
||||
|
||||
### 5. View Results
|
||||
|
||||
After analysis, the `*ascend_pt` directory will contain many files, with the main analysis focus being the `ASCEND_PROFILER_OUTPUT` folder. This directory will include the following files:
|
||||
|
||||
- `analysis.db`: Performance data in database format
|
||||
|
||||
- `api_statistic.csv`: API call statistics
|
||||
|
||||
- `ascend_pytorch_profiler_0.db`: Performance data in database format
|
||||
|
||||
- `kernel_details.csv`: Kernel-level related data
|
||||
|
||||
- `operator_details.csv`: Operator-level related data
|
||||
|
||||
- `op_statistic.csv`: Operator utilization data
|
||||
|
||||
- `step_trace_time.csv`: Scheduling data
|
||||
|
||||
- `trace_view.json`: Chrome tracing format data, can be opened with [MindStudio Insight](https://www.hiascend.com/document/detail/zh/mindstudio/81RC1/GUI_baseddevelopmenttool/msascendinsightug/Insight_userguide_0002.html)
|
||||
|
||||
[↑ Back to Top](#service-profiling-guide)
|
||||
|
||||
---
|
||||
|
||||
## MS Service Profiler
|
||||
|
||||
### 0. Build from Source and Upgrade
|
||||
|
||||
The `msserviceprofiler` tool is pre-installed with the CANN Toolkit package. Use the following commands to install or upgrade from source.
|
||||
|
||||
```bash
|
||||
git clone https://gitcode.com/Ascend/msserviceprofiler.git
|
||||
cd msserviceprofiler
|
||||
bash scripts/build_and_upgrade.sh
|
||||
```
|
||||
|
||||
### 1. Preparation
|
||||
|
||||
Before starting the service, set the environment variable `SERVICE_PROF_CONFIG_PATH` to point to the profiling configuration file, and set the environment variable `PROFILING_SYMBOLS_PATH` to specify the YAML configuration file for the symbols that need to be imported. After that, start the vLLM service according to your deployment method.
|
||||
|
||||
```bash
|
||||
cd ${path_to_store_profiling_files}
|
||||
# Set environment variable
|
||||
export SERVICE_PROF_CONFIG_PATH=ms_service_profiler_config.json
|
||||
export PROFILING_SYMBOLS_PATH=service_profiling_symbols.yaml
|
||||
|
||||
# Start vLLM service
|
||||
vllm serve Qwen/Qwen2.5-0.5B-Instruct &
|
||||
```
|
||||
|
||||
The file `ms_service_profiler_config.json` is the profiling configuration. If it does not exist at the specified path, a default configuration will be generated automatically. If needed, you can customize it in advance according to the instructions in the `Profiling Configuration File` section below.
|
||||
|
||||
`service_profiling_symbols.yaml` is the configuration file containing the profiling points to be imported. You can choose **not** to set the `PROFILING_SYMBOLS_PATH` environment variable, in which case the default configuration file will be used. If the file does not exist at the path you specified, likewise, the system will generate a configuration file at your specified path for future configuration. You can customize it according to the instructions in the `Symbols Configuration File` section below.
|
||||
|
||||
### 2. Enable Profiling
|
||||
|
||||
To enable the performance data collection switch, change the `enable` field from `0` to `1` in the configuration file `ms_service_profiler_config.json`. This can be accomplished by executing the following sed command:
|
||||
|
||||
```bash
|
||||
sed -i 's/"enable":\s*0/"enable": 1/' ./ms_service_profiler_config.json
|
||||
```
|
||||
|
||||
### 3. Send Requests
|
||||
|
||||
Choose a request-sending method that suits your actual profiling needs:
|
||||
|
||||
```bash
|
||||
curl http://localhost:8000/v1/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "Qwen/Qwen2.5-0.5B-Instruct",
|
||||
"prompt": "Beijing is a",
|
||||
"max_tokens": 5,
|
||||
"temperature": 0
|
||||
}' | python3 -m json.tool
|
||||
```
|
||||
|
||||
### 4. Parse Data
|
||||
|
||||
```bash
|
||||
# xxxx-xxxx is the directory automatically created based on vLLM startup time
|
||||
cd /root/.ms_server_profiler/xxxx-xxxx
|
||||
|
||||
# parse data
|
||||
msserviceprofiler parse --input-path=./ --output-path output
|
||||
```
|
||||
|
||||
### 5. View Results
|
||||
|
||||
After parsing, the `output` directory will contain:
|
||||
|
||||
- `chrome_tracing.json`: Chrome tracing format data, which can be opened in [MindStudio Insight](https://www.hiascend.com/document/detail/zh/mindstudio/830/GUI_baseddevelopmenttool/msascendinsightug/Insight_userguide_0002.html?framework=mindspore).
|
||||
- `profiler.db`: Performance data in database format.
|
||||
- `request.csv`: Request-related data.
|
||||
- `kvcache.csv`: KV Cache-related data.
|
||||
- `batch.csv`: Batch scheduling-related data.
|
||||
|
||||
---
|
||||
|
||||
### 6. Appendix related to MS Service Profiler
|
||||
|
||||
(profiling-configuration-file)=
|
||||
|
||||
#### 6.1 Profiling Configuration File
|
||||
|
||||
The profiling configuration file controls profiling parameters and behavior.
|
||||
|
||||
##### File Format
|
||||
|
||||
The configuration is in JSON format. Main parameters:
|
||||
|
||||
| Parameter | Description | Required |
|
||||
|:------:|:------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:-----:|
|
||||
| enable | Switch for profiling: <br />0: disable<br />1: enable<br />Default: 0 | Yes |
|
||||
| prof_dir | Directory to store collected performance data. <br />Default: `${HOME}/.ms_server_profiler` | No |
|
||||
| profiler_level | Data collection level. Default is "INFO" (normal level). | No |
|
||||
| acl_task_time | Switch to collect operator dispatch latency and execution latency. Values: <br />0: off. Default; 0 or any invalid value means off.<br />1: on. When enabled, calls `aclprofCreateConfig` with `ACL_PROF_TASK_TIME_L0`.<br />2: on. MSPTI-based dump. When enabled, set before starting the service: `export LD_PRELOAD={INSTALL_DIR}/lib64/libmspti.so`, where `{INSTALL_DIR}` is the CANN installation root (e.g. `/usr/local/Ascend/cann` for a typical root install).<br />3: on. Torch Profiler–based dump. | No |
|
||||
| acl_prof_task_time_level | Profiling level and duration. Values: <br />L0: collect operator dispatch and execution latency only; lower overhead (no operator basic info).<br />L1: collect AscendCL interface performance (host–device and inter-device sync/async memory copy latencies), plus operator dispatch, execution, and basic info for comprehensive analysis.<br />`{time}`: optional duration segment; integer 1–999, unit seconds.<br />If unset, defaults to L0 until program exit; invalid values fall back to defaults.<br />Level and duration can be combined, e.g., `"acl_prof_task_time_level": "L1;10"`.<br />**Note:** When Torch Profiler is used (`acl_task_time` set to `3`), `{time}` duration is not supported. | No |
|
||||
| timelimit | Profiling duration for the service. The process stops automatically after this time. Range: integer 0–7200, unit: seconds. Default 0 means unlimited. Recommend at least 120s; shorter runs may lack data for parsed outputs and trigger warnings. | No |
|
||||
| domain | Limit profiling to the specified domains to reduce data volume. String, separated by semicolons, case-sensitive, e.g., "Request; KVCache".<br />Empty means all available domains.<br />Available domains: Request, KVCache, ModelExecute, BatchSchedule, Communication.<br />Note: If the selected domains are incomplete, analysis output may show warnings due to missing data. See [Reference Table 1](https://www.hiascend.com/document/detail/zh/canncommercial/850/devaids/Profiling/mindieprofiling_0010.html). | No |
|
||||
| torch_prof_stack | Collect operator call stacks (framework and CPU operators). Values: `false` (default, off), `true` (on). Requires `acl_task_time` set to `3`. **Note:** Enabling this configuration introduces additional performance overhead. | No |
|
||||
| torch_prof_step_num | Torch Profiler step limit. Integer ≥ 0. Default `0` means collect all steps.<br />Requires `acl_task_time` set to `3`. | No |
|
||||
| profiler_step_num | Step limit for operator and service framework profiling. Integer ≥ 0.<br />`0` or invalid values stop the entire service profiling process.<br />The number of steps actually recorded depends on `modelRunnerExec` events. | No |
|
||||
|
||||
##### Example Configuration
|
||||
|
||||
```json
|
||||
{
|
||||
"enable": 1,
|
||||
"prof_dir": "./vllm_prof",
|
||||
"acl_task_time": 0,
|
||||
"acl_prof_task_time_level": ""
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
(symbols-configuration-file)=
|
||||
|
||||
#### 6.2 Symbols Configuration File
|
||||
|
||||
The symbols configuration file defines which functions/methods to profile and supports flexible configuration with custom attribute collection.
|
||||
|
||||
##### File Name and Loading
|
||||
|
||||
- Default load path: `~/.config/vllm_ascend/service_profiling_symbols.MAJOR.MINOR.PATCH.yaml` (According to the installed version of vLLM)
|
||||
|
||||
If you need to customize the profiling points, it is highly recommended to copy a symbol configuration file to your working directory and point to it with the `PROFILING_SYMBOLS_PATH` environment variable.
|
||||
|
||||
##### Configuration file updates
|
||||
|
||||
After you change profiling symbols, restart the vLLM service so the updated configuration file is loaded.
|
||||
|
||||
##### Field Descriptions
|
||||
|
||||
| Field | Description | Example |
|
||||
|:-----:|:-----|:-----|
|
||||
| symbol | Python import path + attribute chain | `"vllm.v1.core.kv_cache_manager:KVCacheManager.free"` |
|
||||
| handler | Handler type | `"timer"` (default) or `"pkg.mod:func"` (custom) |
|
||||
| domain | Domain tag | `"KVCache"`, `"ModelExecute"` |
|
||||
| name | Event name | `"EngineCoreExecute"` |
|
||||
| min_version | Minimum supported vLLM version | `"0.9.1"` |
|
||||
| max_version | Maximum supported vLLM version | `"0.11.0"` |
|
||||
| attributes | Custom attribute collection | Only supported for `"timer"` handler. See the section below |
|
||||
|
||||
##### Configuration Examples
|
||||
|
||||
- Example 1: Custom handler
|
||||
|
||||
```yaml
|
||||
- symbol: vllm.v1.core.kv_cache_manager:KVCacheManager.free
|
||||
handler: ms_service_profiler.patcher.config.custom_handler_example.kvcache_manager_free_example_handler
|
||||
domain: Example
|
||||
name: example_custom
|
||||
```
|
||||
|
||||
- Example 2: Default timer
|
||||
|
||||
```yaml
|
||||
- symbol: vllm.v1.engine.core:EngineCore.execute_model
|
||||
domain: ModelExecute
|
||||
name: EngineCoreExecute
|
||||
```
|
||||
|
||||
- Example 3: Version constraint
|
||||
|
||||
```yaml
|
||||
- symbol: vllm.v1.executor.abstract:Executor.execute_model
|
||||
min_version: "0.9.1"
|
||||
# No handler specified -> default timer
|
||||
```
|
||||
|
||||
##### Custom Attribute Collection
|
||||
|
||||
The `attributes` field supports flexible custom attribute collection and allows operations and transformations on function arguments and return values.
|
||||
|
||||
###### Basic Syntax
|
||||
|
||||
- Argument access: use the parameter name directly, e.g., `input_ids`
|
||||
- Return value access: use the `return` keyword
|
||||
- Pipeline operations: use `|` to chain multiple operations
|
||||
- Attribute access: use `attr` to access object attributes
|
||||
|
||||
###### Example
|
||||
|
||||
```yaml
|
||||
- symbol: vllm_ascend.worker.model_runner_v1:NPUModelRunner.execute_model
|
||||
name: ModelRunnerExecuteModel
|
||||
domain: ModelExecute
|
||||
attributes:
|
||||
- name: device
|
||||
expr: args[0] | attr device | str
|
||||
- name: dp
|
||||
expr: args[0] | attr dp_rank | str
|
||||
- name: batch_size
|
||||
expr: args[0] | attr input_batch | attr _req_ids | len
|
||||
```
|
||||
|
||||
###### Expression Notes
|
||||
|
||||
1. `len(input_ids)`: get the length of parameter `input_ids`.
|
||||
2. `len(return) | str`: get the length of the return value and convert to string (equivalent to `str(len(return))`).
|
||||
3. `return[0] | attr input_ids | len`: get the length of the `input_ids` attribute of the first element in the return value.
|
||||
|
||||
###### Supported Expression Types
|
||||
|
||||
- Basic operations: `len()`, `str()`, `int()`, `float()`
|
||||
- Index access: `return[0]`, `return['key']`
|
||||
- Attribute access: `return | attr attr_name`
|
||||
- Pipeline composition: chain operations with `|`
|
||||
|
||||
###### Advanced Examples
|
||||
|
||||
```yaml
|
||||
attributes:
|
||||
# Get tensor shape
|
||||
- name: tensor_shape
|
||||
expr: input_tensor | attr shape | str
|
||||
|
||||
# Get specific value from a dict
|
||||
- name: batch_size
|
||||
expr: kwargs['batch_size']
|
||||
|
||||
# Conditional expression (requires custom handler support)
|
||||
- name: is_training_mode
|
||||
expr: training | bool
|
||||
|
||||
# Complex data processing
|
||||
- name: processed_data_len
|
||||
expr: data | attr items | len | str
|
||||
```
|
||||
|
||||
##### Custom Handler
|
||||
|
||||
When `handler` specifies a custom function, it must match the following signature:
|
||||
|
||||
```python
|
||||
def custom_handler(original_func, this, *args, **kwargs):
|
||||
"""
|
||||
Custom handler
|
||||
|
||||
Args:
|
||||
original_func: the original function object
|
||||
this: the bound object (for methods)
|
||||
*args: positional arguments
|
||||
**kwargs: keyword arguments
|
||||
|
||||
Returns:
|
||||
processing result
|
||||
"""
|
||||
# Custom logic
|
||||
pass
|
||||
```
|
||||
|
||||
If the custom handler fails to import, the system will automatically fall back to the default timer mode.
|
||||
|
||||
[↑ Back to Top](#service-profiling-guide)
|
||||
Reference in New Issue
Block a user