10
docs/source/developer_guide/performance_and_debug/index.md
Normal file
10
docs/source/developer_guide/performance_and_debug/index.md
Normal file
@@ -0,0 +1,10 @@
|
||||
# Performance and Debug
|
||||
|
||||
::::{toctree}
|
||||
:caption: Performance and Debug
|
||||
:maxdepth: 1
|
||||
performance_benchmark
|
||||
optimization_and_tuning
|
||||
service_profiling_guide
|
||||
msprobe_guide
|
||||
::::
|
||||
@@ -0,0 +1,450 @@
|
||||
# MSProbe Debugging Guide
|
||||
|
||||
During inference or training runs we often encounter accuracy anomalies such as outputs drifting away from the expectation, unstable numerical behavior (NaN/Inf), or predictions that no longer match the labels. To pinpoint the root cause we have to monitor and capture intermediate data produced while the model executes—feature maps, weights, activations, and layer outputs. By capturing key tensors at specific stages, logging I/O pairs for the core layers, and retaining contextual metadata (prompts, tensor dtypes, hardware configuration, etc.), we can systematically trace where the accuracy degradation or numerical error started. This guide describes the end-to-end workflow for diagnosing accuracy issues for AI models (with a focus on vllm-ascend services): preparation, data capture, and analysis & verification.
|
||||
|
||||
For more details, see [Ascend/msprobe](https://gitcode.com/Ascend/msprobe).
|
||||
|
||||
## 0. Background Concepts
|
||||
|
||||
`msprobe` supports three accuracy levels:
|
||||
|
||||
- **L0**: dumps tensors at the module level and generates `construct.json` so that visualization tools can rebuild the network structure. A model or submodule handle must be passed in.
|
||||
- **L1**: collects operator-level statistics only, which is suitable for lightweight troubleshooting.
|
||||
- **mix**: captures both structural information and operator statistics, which is useful when you need both graph reconstruction and numerical comparisons.
|
||||
|
||||
## 1. Prerequisites
|
||||
|
||||
### 1.1 Install `msprobe`
|
||||
|
||||
Install msprobe with pip:
|
||||
|
||||
```bash
|
||||
pip install mindstudio-probe
|
||||
```
|
||||
|
||||
### 1.2 Graph mode dump (optional)
|
||||
|
||||
If you need to dump cudagraph graphs, you need to install from source code:
|
||||
|
||||
1. Install `aclgraph_dump` from source code:
|
||||
|
||||
```bash
|
||||
git clone https://gitcode.com/Ascend/msprobe.git
|
||||
cd msprobe
|
||||
python3 setup.py bdist_wheel --include-mod=aclgraph_dump --no-check
|
||||
pip install dist/*.whl
|
||||
```
|
||||
|
||||
## 2. Collecting Data with `msprobe`
|
||||
|
||||
We generally follow a coarse-to-fine strategy when capturing data. First, identify the token where the issue shows up, and then decide which range needs to be sampled around that token. The typical workflow is described below.
|
||||
|
||||
### 2.1 Prepare the dump configuration content
|
||||
|
||||
Prepare configuration content that can be parsed by `PrecisionDebugger`. You can use either of the following ways:
|
||||
|
||||
- Pass the config object directly through `--additional-config.dump_config`.
|
||||
- Pass a config file path through `--additional-config.dump_config_path`.
|
||||
|
||||
Common fields are:
|
||||
|
||||
| Field | Description | Required | Eager Mode | Graph Mode |
|
||||
|:-----------:|:-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:--------:|:-------------:|:-------------:|
|
||||
| `task` | Type of dump task. Common PyTorch values include `"statistics"` and `"tensor"`. A statistics task collects tensor statistics (mean, variance, max, min, etc.) while a tensor task captures arbitrary tensors. | Yes | ✅ | ✅ |
|
||||
| `dump_path` | Directory where dump results are stored. When omitted, `msprobe` uses its default path. | No | ✅ | ✅ |
|
||||
| `rank` | Ranks to sample. An empty list collects every rank. For single-card tasks, you must set this field to `[]`. | No | ✅ | ✅ |
|
||||
| `step` | Token iteration(s) to sample. An empty list means every iteration. | No | ✅ | ❌ |
|
||||
| `level` | Dump level string (`"L0"`, `"L1"`, or `"mix"`). `L0` targets `nn.Module`, `L1` targets `torch.api`, and `mix` collects both. | Yes | ✅ | ✅ |
|
||||
| `async_dump`| Whether to enable asynchronous dump (supported for PyTorch `statistics`/`tensor` tasks). Defaults to `false`. | No | ✅ | ❌ |
|
||||
| `scope` | Module range to sample. An empty list collects every module. | No | ✅ | ❌ |
|
||||
| `dump_enable` | Dynamic switch for enabling/disabling dump in `PrecisionDebugger` during one running training/inference job. This allows turning dump on or off on demand in the same job. | No | ✅ | ❌ |
|
||||
| `list` | Operator range to sample. An empty list collects every operator. | No | ✅ | ✅ |
|
||||
|
||||
To restrict the operators that are captured, configure the `list` block:
|
||||
|
||||
- `scope` (list[str]): In PyTorch PyNative scenarios this field restricts the dump range. Provide two module or API names that follow the tool's naming convention to lock a range; only data between the two names will be dumped. Examples:
|
||||
|
||||
```json
|
||||
"scope": ["Module.conv1.Conv2d.forward.0", "Module.fc2.Linear.forward.0"]
|
||||
"scope": ["Cell.conv1.Conv2d.forward.0", "Cell.fc2.Dense.forward.0"]
|
||||
"scope": ["Tensor.add.0.forward", "Functional.square.2.forward"]
|
||||
```
|
||||
|
||||
The `level` setting determines what can be provided—modules when `level=L0`, APIs when `level=L1`, and either modules or APIs when `level=mix`.
|
||||
|
||||
- `list` (list[str]): Custom operator list. Options include:
|
||||
- Supply the full names of specific APIs in PyTorch pynative scenarios to only dump those APIs. Example: `"list": ["Tensor.permute.1.forward", "Tensor.transpose.2.forward", "Torch.relu.3.forward"]`.
|
||||
- When `level=mix`, you can provide module names so that the dump expands to everything produced while the module is running. Example: `"list": ["Module.module.language_model.encoder.layers.0.mlp.ParallelMlp.forward.0"]`.
|
||||
- Provide a substring such as `"list": ["relu"]` to dump every API whose name contains the substring. When `level=mix`, modules whose names contain the substring are also expanded.
|
||||
|
||||
Example configuration:
|
||||
eager mode:
|
||||
|
||||
```json
|
||||
{
|
||||
"task": "statistics",
|
||||
"dump_path": "/home/data_dump",
|
||||
"rank": [],
|
||||
"step": [],
|
||||
"level": "L1",
|
||||
"async_dump": false,
|
||||
|
||||
"statistics": {
|
||||
"scope": [],
|
||||
"list": [],
|
||||
"tensor_list": [],
|
||||
"data_mode": ["all"],
|
||||
"summary_mode": "statistics"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Graph mode:
|
||||
|
||||
```json
|
||||
{
|
||||
"task": "statistics",
|
||||
"level": "L1",
|
||||
"dump_path": "/home/data_dump",
|
||||
"statistics": {
|
||||
"list": []
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## 3. Enable `msprobe` in vllm-ascend
|
||||
|
||||
1. Start vLLM and pass the dump config content through `--additional-config`:
|
||||
|
||||
```bash
|
||||
vllm serve Qwen/Qwen2.5-0.5B-Instruct \
|
||||
--dtype bfloat16 \
|
||||
--host 0.0.0.0 \
|
||||
--port 8000 \
|
||||
--additional-config '{
|
||||
"dump_config": {
|
||||
"task": "statistics",
|
||||
"level": "L1",
|
||||
"dump_path": "/data/msprobe_dump",
|
||||
"statistics": {
|
||||
"list": []
|
||||
}
|
||||
}
|
||||
}' &
|
||||
```
|
||||
|
||||
Compatibility mode (legacy) is still supported:
|
||||
|
||||
```bash
|
||||
vllm serve Qwen/Qwen2.5-0.5B-Instruct \
|
||||
--dtype bfloat16 \
|
||||
--host 0.0.0.0 \
|
||||
--port 8000 \
|
||||
--additional-config '{"dump_config_path": "/data/msprobe_config.json"}' &
|
||||
```
|
||||
|
||||
## 4. Send requests and collect dumps
|
||||
|
||||
1. Send inference requests as usual, for example:
|
||||
|
||||
```bash
|
||||
curl http://localhost:8000/v1/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "Qwen/Qwen2.5-0.5B-Instruct",
|
||||
"prompt": "Explain gravity in one sentence.",
|
||||
"max_completion_tokens": 32,
|
||||
"temperature": 0
|
||||
}' | python -m json.tool
|
||||
```
|
||||
|
||||
2. Each request drives the sequence `msprobe: start -> forward -> stop -> step`. The runner invokes `step()` on every code path, so you always get a complete dataset even if inference returns early.
|
||||
|
||||
3. Dump files are written into `dump_path`. They usually contain:
|
||||
- Tensor files grouped by operator/module.
|
||||
- `dump.json`, which records metadata such as dtype, shape, min/max, and `requires_grad`.
|
||||
- `construct.json`, which is generated when `level` is `L0` or `mix` (required for visualization).
|
||||
|
||||
Example directory layout:
|
||||
eager mode:
|
||||
|
||||
```text
|
||||
├── dump_path
|
||||
│ ├── step0
|
||||
│ │ ├── rank0
|
||||
│ │ │ ├── dump_tensor_data
|
||||
│ │ │ │ ├── Tensor.permute.1.forward.pt # Format: {api_type}.{api_name}.{call_count}.forward.{input/output}.{arg_index}.
|
||||
│ │ │ │ │ # arg_index is the nth input or output of the API. If an input is a list, keep numbering with decimals (e.g., 1.1 is the first element of the first argument).
|
||||
│ │ │ │ ├── Module.conv1.Conv2d.forward.0.input.0.pt # Format: {Module}.{module_name}.{class_name}.forward.{call_count}.{input/output}.{arg_index}.
|
||||
│ │ │ │ └── Module.conv1.Conv2d.forward.0.parameters.bias.pt # Module parameter data: {Module}.{module_name}.{class_name}.forward.{call_count}.parameters.{parameter_name}.
|
||||
│ │ │ │ # When the `model` argument passed to dump is a List[torch.nn.Module] or Tuple[torch.nn.Module], module-level data names also include the index inside the list ({Module}.{index}.*), e.g., Module.0.conv1.Conv2d.forward.0.input.0.pt.
|
||||
│ │ │ ├── dump.json
|
||||
│ │ │ ├── stack.json
|
||||
│ │ │ ├── dump_error_info.log
|
||||
│ │ │ └── construct.json
|
||||
│ │ ├── rank1
|
||||
│ │ │ ├── dump_tensor_data
|
||||
│ │ │ │ └── ...
|
||||
│ │ │ ├── dump.json
|
||||
│ │ │ ├── stack.json
|
||||
│ │ │ ├── dump_error_info.log
|
||||
│ │ │ └── construct.json
|
||||
│ │ ├── ...
|
||||
│ │ │
|
||||
│ │ └── rank7
|
||||
│ ├── step1
|
||||
│ │ ├── ...
|
||||
│ ├── step2
|
||||
```
|
||||
|
||||
- `rank`: Device ID. Each card writes its data to the corresponding `rank{ID}` directory. In non-distributed scenarios the directory is simply named `rank`.
|
||||
- `dump_tensor_data`: Tensor payloads that were collected.
|
||||
- `dump.json`: Statistics for the forward data of each API or module, including names, dtype, shape, max, min, mean, L2 norm (square root of the L2 variance), and CRC-32 when `summary_mode="md5"`. See [dump.json file description](#dumpjson-file-description) for details.
|
||||
- `dump_error_info.log`: Present only when the dump tool encountered an error and records the failure log.
|
||||
- `stack.json`: Call stacks for APIs/modules.
|
||||
- `construct.json`: Hierarchical structure description. Empty when `level=L1`.
|
||||
|
||||
graph mode:
|
||||
|
||||
```text
|
||||
L0_dump
|
||||
├── step0
|
||||
│ └── rank0
|
||||
│ └── dump.json
|
||||
├── step1
|
||||
│ └── rank0
|
||||
│ └── dump.json
|
||||
├── step2
|
||||
│ └── rank0
|
||||
│ └── dump.json
|
||||
├── step3
|
||||
│ └── rank0
|
||||
│ └── dump.json
|
||||
├── step4
|
||||
│ └── rank0
|
||||
│ └── dump.json
|
||||
└── step5
|
||||
└── rank0
|
||||
└── dump.json
|
||||
```
|
||||
|
||||
- `dump.json`: See [dump.json file description](#dumpjson-file-description) for details.
|
||||
|
||||
## 5. Analyze the results
|
||||
|
||||
### 5.1 Prerequisites
|
||||
|
||||
You typically need two dump datasets: one from the "problem side" (the run that exposes the accuracy or numerical error) and another from the "benchmark side" (a good baseline). These datasets do not have to be identical—they can come from different branches, framework versions, or even alternative implementations (operator substitutions, different graph-optimization switches, etc.). As long as they use the same or similar inputs, hardware topology, and sampling points (step/token), `msprobe` can compare them and locate the divergent nodes. If you cannot find a perfectly clean benchmark, start by capturing the problem-side data, craft the smallest reproducible case by hand, and perform a self-comparison. Below we assume the problem dump is `problem_dump` and the benchmark dump is `bench_dump`.
|
||||
|
||||
### 5.2 Visualization
|
||||
|
||||
Use `msprobe graph_visualize` to build or compare graphs, then open the generated `*.vis.db` file(s) with TensorBoard (`tb_graph_ascend` plugin).
|
||||
|
||||
1. Ensure dump data is visualization-ready:
|
||||
- Dump level must be `L0` or `mix` so `construct.json` is non-empty.
|
||||
- Each rank directory should contain `dump.json`, `stack.json`, and `construct.json`.
|
||||
|
||||
2. Choose command mode:
|
||||
- Single-graph build:
|
||||
|
||||
```bash
|
||||
msprobe graph_visualize -tp <target_path> -o <output_path>
|
||||
```
|
||||
|
||||
- Graph comparison:
|
||||
|
||||
```bash
|
||||
msprobe graph_visualize -tp <target_path> -gp <golden_path> -o <output_path>
|
||||
```
|
||||
|
||||
- Common optional flags:
|
||||
- `-oc` / `--overflow_check`: enable overflow marking
|
||||
- `-fm` / `--fuzzy_match`: enable fuzzy matching for node mapping
|
||||
- `-lm` / `--layer_mapping [mapping.yaml]`: cross-framework/layer mapping compare
|
||||
- `-tensor_log`: print per-node compare log (tensor dump scenarios)
|
||||
- `-progress_log`: print detailed progress log
|
||||
|
||||
3. Path granularity is auto-detected by `graph_visualize`:
|
||||
- Single-rank: `.../step0/rank0`
|
||||
- Multi-rank (batch): `.../step0`
|
||||
- Multi-step (batch): dump root path containing `step*`
|
||||
|
||||
4. Output files:
|
||||
- Single-graph build: `build_{timestamp}.vis.db`
|
||||
- Graph comparison: `compare_{timestamp}.vis.db`
|
||||
|
||||
5. Launch TensorBoard with the output directory:
|
||||
|
||||
```bash
|
||||
tensorboard --logdir <output_path> --bind_all --port <optional_port>
|
||||
```
|
||||
|
||||
6. In the visualization UI, inspect structure and numeric differences:
|
||||
- Switch rank/step to locate unstable nodes quickly.
|
||||
- Use search/filter to focus on target ops/modules.
|
||||
- For compare mode, prioritize highlighted high-difference nodes and trace surrounding I/O/parameters.
|
||||
|
||||
## 6. Troubleshooting
|
||||
|
||||
- `RuntimeError: Please enforce eager mode`: Restart vLLM and add the `--enforce-eager` flag.
|
||||
- No dump files: Confirm that the JSON path is correct and every node has write permission. In distributed scenarios set `keep_all_ranks` so that every rank writes its own dump.
|
||||
- Dumps are too large: Start with a `statistics` task to locate abnormal tensors, then narrow the scope with `scope`/`list`/`tensor_list`, `filters`, `token_range`, etc.
|
||||
|
||||
---
|
||||
|
||||
## Appendix
|
||||
|
||||
### dump.json file description
|
||||
|
||||
#### L0 level
|
||||
|
||||
An L0 `dump.json` contains forward I/O for modules together with parameters. Using PyTorch's `Conv2d` as an example, the network code looks like:
|
||||
|
||||
`output = self.conv2(input) # self.conv2 = torch.nn.Conv2d(64, 128, 5, padding=2, bias=True)`
|
||||
|
||||
`dump.json` contains the following entries:
|
||||
|
||||
- `Module.conv2.Conv2d.forward.0`: Forward data of the module. `input_args` represents positional inputs, `input_kwargs` represents keyword inputs, `output` stores forward outputs, and `parameters` stores weights/biases.
|
||||
|
||||
**Note**: When the `model` parameter passed to the dump API is `List[torch.nn.Module]` or `Tuple[torch.nn.Module]`, module-level names include the index inside the list (`{Module}.{index}.*`). Example: `Module.0.conv1.Conv2d.forward.0`.
|
||||
|
||||
```json
|
||||
{
|
||||
"task": "tensor",
|
||||
"level": "L0",
|
||||
"framework": "pytorch",
|
||||
"dump_data_dir": "/dump/path",
|
||||
"data": {
|
||||
"Module.conv2.Conv2d.forward.0": {
|
||||
"input_args": [
|
||||
{
|
||||
"type": "torch.Tensor",
|
||||
"dtype": "torch.float32",
|
||||
"shape": [
|
||||
8,
|
||||
16,
|
||||
14,
|
||||
14
|
||||
],
|
||||
"Max": 1.638758659362793,
|
||||
"Min": 0.0,
|
||||
"Mean": 0.2544615864753723,
|
||||
"Norm": 70.50277709960938,
|
||||
"requires_grad": true,
|
||||
"data_name": "Module.conv2.Conv2d.forward.0.input.0.pt"
|
||||
}
|
||||
],
|
||||
"input_kwargs": {},
|
||||
"output": [
|
||||
{
|
||||
"type": "torch.Tensor",
|
||||
"dtype": "torch.float32",
|
||||
"shape": [
|
||||
8,
|
||||
32,
|
||||
10,
|
||||
10
|
||||
],
|
||||
"Max": 1.6815717220306396,
|
||||
"Min": -1.5120246410369873,
|
||||
"Mean": -0.025344856083393097,
|
||||
"Norm": 149.65576171875,
|
||||
"requires_grad": true,
|
||||
"data_name": "Module.conv2.Conv2d.forward.0.output.0.pt"
|
||||
}
|
||||
],
|
||||
"parameters": {
|
||||
"weight": {
|
||||
"type": "torch.Tensor",
|
||||
"dtype": "torch.float32",
|
||||
"shape": [
|
||||
32,
|
||||
16,
|
||||
5,
|
||||
5
|
||||
],
|
||||
"Max": 0.05992485210299492,
|
||||
"Min": -0.05999220535159111,
|
||||
"Mean": -0.0006165213999338448,
|
||||
"Norm": 3.421217441558838,
|
||||
"requires_grad": true,
|
||||
"data_name": "Module.conv2.Conv2d.forward.0.parameters.weight.pt"
|
||||
},
|
||||
"bias": {
|
||||
"type": "torch.Tensor",
|
||||
"dtype": "torch.float32",
|
||||
"shape": [
|
||||
32
|
||||
],
|
||||
"Max": 0.05744686722755432,
|
||||
"Min": -0.04894155263900757,
|
||||
"Mean": 0.006410328671336174,
|
||||
"Norm": 0.17263513803482056,
|
||||
"requires_grad": true,
|
||||
"data_name": "Module.conv2.Conv2d.forward.0.parameters.bias.pt"
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
#### L1 level
|
||||
|
||||
An L1 `dump.json` records forward I/O for APIs. Using PyTorch's `relu` function as an example (`output = torch.nn.functional.relu(input)`), the file contains:
|
||||
|
||||
- `Functional.relu.0.forward`: Forward data of the API. `input_args` are positional inputs, `input_kwargs` are keyword inputs, and `output` stores the forward outputs.
|
||||
|
||||
```json
|
||||
{
|
||||
"task": "tensor",
|
||||
"level": "L1",
|
||||
"framework": "pytorch",
|
||||
"dump_data_dir":"/dump/path",
|
||||
"data": {
|
||||
"Functional.relu.0.forward": {
|
||||
"input_args": [
|
||||
{
|
||||
"type": "torch.Tensor",
|
||||
"dtype": "torch.float32",
|
||||
"shape": [
|
||||
32,
|
||||
16,
|
||||
28,
|
||||
28
|
||||
],
|
||||
"Max": 1.3864083290100098,
|
||||
"Min": -1.3364859819412231,
|
||||
"Mean": 0.03711778670549393,
|
||||
"Norm": 236.20692443847656,
|
||||
"requires_grad": true,
|
||||
"data_name": "Functional.relu.0.forward.input.0.pt"
|
||||
}
|
||||
],
|
||||
"input_kwargs": {},
|
||||
"output": [
|
||||
{
|
||||
"type": "torch.Tensor",
|
||||
"dtype": "torch.float32",
|
||||
"shape": [
|
||||
32,
|
||||
16,
|
||||
28,
|
||||
28
|
||||
],
|
||||
"Max": 1.3864083290100098,
|
||||
"Min": 0.0,
|
||||
"Mean": 0.16849493980407715,
|
||||
"Norm": 175.23345947265625,
|
||||
"requires_grad": true,
|
||||
"data_name": "Functional.relu.0.forward.output.0.pt"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
#### mix level
|
||||
|
||||
A `mix` dump.json contains both L0 and L1 level data; the file format is the same as the examples above.
|
||||
@@ -0,0 +1,248 @@
|
||||
# Optimization and Tuning
|
||||
|
||||
This guide aims to help users improve vLLM Ascend performance at the system level. It includes OS configuration, library optimization, deployment guide, and so on. Any feedback is welcome.
|
||||
|
||||
## Preparation
|
||||
|
||||
### 1.Run the container
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# Update DEVICE according to your device (/dev/davinci[0-7])
|
||||
export DEVICE=/dev/davinci0
|
||||
# Update the cann base image
|
||||
export IMAGE=m.daocloud.io/quay.io/ascend/cann:|cann_image_tag|
|
||||
docker run --rm \
|
||||
--name performance-test \
|
||||
--shm-size=1g \
|
||||
--device $DEVICE \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-it $IMAGE bash
|
||||
```
|
||||
|
||||
### 2.Configure your environment
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# Configure the mirror
|
||||
echo "deb https://mirrors.tuna.tsinghua.edu.cn/ubuntu-ports/ jammy main restricted universe multiverse" > /etc/apt/sources.list && \
|
||||
echo "deb-src https://mirrors.tuna.tsinghua.edu.cn/ubuntu-ports/ jammy main restricted universe multiverse" >> /etc/apt/sources.list && \
|
||||
echo "deb https://mirrors.tuna.tsinghua.edu.cn/ubuntu-ports/ jammy-updates main restricted universe multiverse" >> /etc/apt/sources.list && \
|
||||
echo "deb-src https://mirrors.tuna.tsinghua.edu.cn/ubuntu-ports/ jammy-updates main restricted universe multiverse" >> /etc/apt/sources.list && \
|
||||
echo "deb https://mirrors.tuna.tsinghua.edu.cn/ubuntu-ports/ jammy-backports main restricted universe multiverse" >> /etc/apt/sources.list && \
|
||||
echo "deb-src https://mirrors.tuna.tsinghua.edu.cn/ubuntu-ports/ jammy-backports main restricted universe multiverse" >> /etc/apt/sources.list && \
|
||||
echo "deb https://mirrors.tuna.tsinghua.edu.cn/ubuntu-ports/ jammy-security main restricted universe multiverse" >> /etc/apt/sources.list && \
|
||||
echo "deb-src https://mirrors.tuna.tsinghua.edu.cn/ubuntu-ports/ jammy-security main restricted universe multiverse" >> /etc/apt/sources.list
|
||||
|
||||
# Install os packages
|
||||
apt update && apt install wget gcc g++ libnuma-dev git vim -y
|
||||
```
|
||||
|
||||
### 3.Install vLLM and vLLM Ascend
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# Install necessary dependencies
|
||||
pip config set global.index-url https://pypi.tuna.tsinghua.edu.cn/simple
|
||||
pip install modelscope pandas datasets gevent sacrebleu rouge_score pybind11 pytest
|
||||
|
||||
# Configure this var to speed up model download
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
```
|
||||
|
||||
Please follow the [Installation Guide](https://docs.vllm.ai/projects/ascend/en/latest/installation.html) to make sure vLLM and vLLM Ascend are installed correctly.
|
||||
|
||||
:::{note}
|
||||
Make sure your vLLM and vLLM Ascend are installed after your Python configuration is completed, because these packages will build binary files using python in current environment. If you install vLLM and vLLM Ascend before completing [Configure your environment](#2configure-your-environment), the binary files will not use the optimized python.
|
||||
:::
|
||||
|
||||
## Optimizations
|
||||
|
||||
### 1. Memory Allocator Optimization
|
||||
|
||||
#### 1.1. jemalloc
|
||||
|
||||
**jemalloc** is a memory allocator that improves performance for multi-threaded scenarios and can reduce memory fragmentation. jemalloc uses a local thread memory manager to allocate variables, which can avoid lock contention between threads and can hugely optimize performance.
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# Install jemalloc
|
||||
sudo apt update
|
||||
sudo apt install libjemalloc2
|
||||
|
||||
# Configure jemalloc
|
||||
export LD_PRELOAD=/usr/lib/"$(uname -i)"-linux-gnu/libjemalloc.so.2:$LD_PRELOAD
|
||||
```
|
||||
|
||||
#### 1.2. Tcmalloc
|
||||
|
||||
**TCMalloc (Thread Caching Malloc)** is a universal memory allocator that improves overall performance while ensuring low latency by introducing a multi-level cache structure, reducing lock contention and optimizing large object processing flow. Find more [details](https://www.hiascend.com/document/detail/zh/Pytorch/700/ptmoddevg/trainingmigrguide/performance_tuning_0068.html).
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# Install tcmalloc
|
||||
sudo apt update
|
||||
sudo apt install libgoogle-perftools4 libgoogle-perftools-dev
|
||||
|
||||
# Get the location of libtcmalloc.so*
|
||||
find /usr -name libtcmalloc.so*
|
||||
|
||||
# Make the priority of tcmalloc higher
|
||||
# The <path> is the location of libtcmalloc.so we get from the upper command
|
||||
# Example: "$LD_PRELOAD:/usr/lib/aarch64-linux-gnu/libtcmalloc.so"
|
||||
export LD_PRELOAD="$LD_PRELOAD:<path>"
|
||||
|
||||
# Verify your configuration
|
||||
# The path of libtcmalloc.so will be contained in the result if your configuration is valid
|
||||
ldd `which python`
|
||||
```
|
||||
|
||||
### 2. `torch_npu` Optimization
|
||||
|
||||
Some performance tuning features in `torch_npu` are controlled by environment variables. Some features and their related environment variables are shown below.
|
||||
|
||||
Memory optimization:
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# Upper limit of memory block splitting allowed (MB): Setting this parameter can prevent large memory blocks from being split.
|
||||
export PYTORCH_NPU_ALLOC_CONF="max_split_size_mb:250"
|
||||
```
|
||||
|
||||
or
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# When operators on the communication stream have dependencies, they all need to be ended before being released for reuse. The logic of multi-stream reuse is to release the memory on the communication stream in advance so that the computing stream can be reused.
|
||||
export PYTORCH_NPU_ALLOC_CONF="expandable_segments:True"
|
||||
```
|
||||
|
||||
Scheduling optimization:
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# Optimize operator delivery queue. This will affect the memory peak value, and may degrade if the memory is tight.
|
||||
export TASK_QUEUE_ENABLE=2
|
||||
|
||||
# This will greatly improve the CPU bottleneck model and ensure the same performance for the NPU bottleneck model.
|
||||
export CPU_AFFINITY_CONF=1
|
||||
```
|
||||
|
||||
### 3. CANN Optimization
|
||||
|
||||
#### 3.1. HCCL Optimization
|
||||
|
||||
There are some performance tuning features in HCCL, which are controlled by environment variables.
|
||||
|
||||
You can configure HCCL to use "AIV" mode to optimize performance by setting the environment variable shown below. In "AIV" mode, the communication is scheduled by AI vector core directly with RoCE, instead of being scheduled by AI CPU.
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
```
|
||||
|
||||
Plus, there are more features for performance optimization in specific scenarios, which are shown below.
|
||||
|
||||
- `HCCL_INTRA_ROCE_ENABLE`: Use RDMA link instead of SDMA link between two 8Ps as the mesh interconnect link. Find more [details](https://www.hiascend.com/document/detail/zh/Pytorch/600/ptmoddevg/trainingmigrguide/performance_tuning_0044.html).
|
||||
- `HCCL_RDMA_TC`: Use this var to configure traffic class of RDMA NIC. Find more [details](https://www.hiascend.com/document/detail/zh/Pytorch/600/ptmoddevg/trainingmigrguide/performance_tuning_0045.html).
|
||||
- `HCCL_RDMA_SL`: Use this var to configure service level of RDMA NIC. Find more [details](https://www.hiascend.com/document/detail/zh/Pytorch/600/ptmoddevg/trainingmigrguide/performance_tuning_0046.html).
|
||||
- `HCCL_BUFFSIZE`: Use this var to control the cache size for sharing data between two NPUs. Find more [details](https://www.hiascend.com/document/detail/zh/Pytorch/600/ptmoddevg/trainingmigrguide/performance_tuning_0047.html).
|
||||
|
||||
### 4. Kernel Optimization
|
||||
|
||||
This section describes operating system–level optimizations applied on the host machine (bare metal or Kubernetes node) to improve performance stability, latency, and throughput for inference workloads.
|
||||
|
||||
:::{note}
|
||||
These settings must be applied on the host OS and with root privileges. Not inside containers.
|
||||
:::
|
||||
|
||||
#### 4.1 Set CPU Frequency Governor to `performance`
|
||||
|
||||
Set CPU Frequency Governor to `performance`
|
||||
|
||||
```shell
|
||||
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
|
||||
```
|
||||
|
||||
Purpose
|
||||
|
||||
- Forces all CPU cores to run under the `performance` governor
|
||||
- Disables dynamic frequency scaling (e.g., `ondemand`, `powersave`)
|
||||
|
||||
Benefits
|
||||
|
||||
- Keeps CPU cores at maximum frequency
|
||||
- Reduces latency jitter
|
||||
- Improves predictability for inference workloads
|
||||
|
||||
#### 4.2 Disable Swap Usage
|
||||
|
||||
```shell
|
||||
sysctl -w vm.swappiness=0
|
||||
```
|
||||
|
||||
Purpose
|
||||
|
||||
- Minimizes the kernel’s tendency to swap memory pages to disk
|
||||
|
||||
Benefits
|
||||
|
||||
- Prevents severe latency spikes caused by swapping
|
||||
- Improves stability for large in-memory models
|
||||
|
||||
Notes
|
||||
|
||||
- For inference workloads, swap can introduce second-level latency
|
||||
- Recommended values are `0` or `1`
|
||||
|
||||
#### 4.3 Disable Automatic NUMA Balancing
|
||||
|
||||
```shell
|
||||
sysctl -w kernel.numa_balancing=0
|
||||
```
|
||||
|
||||
Purpose
|
||||
|
||||
- Disables the kernel’s automatic NUMA page migration mechanism
|
||||
|
||||
Benefits
|
||||
|
||||
- Prevents background memory page migrations
|
||||
- Reduces unpredictable memory access latency
|
||||
- Improves performance stability on NUMA systems
|
||||
|
||||
Recommended For
|
||||
|
||||
- Multi-socket servers
|
||||
- Ascend / NPU deployments with explicit NUMA binding
|
||||
- Systems with manually managed CPU and memory affinity
|
||||
|
||||
#### 4.4 Increase Scheduler Migration Cost
|
||||
|
||||
```shell
|
||||
sysctl -w kernel.sched_migration_cost_ns=50000
|
||||
```
|
||||
|
||||
Purpose
|
||||
|
||||
- Increases the cost for the scheduler to migrate tasks between CPU cores
|
||||
|
||||
Benefits
|
||||
|
||||
- Reduces frequent thread migration
|
||||
- Improves CPU cache locality
|
||||
- Lowers latency jitter for inference workloads
|
||||
|
||||
Parameter Details
|
||||
|
||||
- Unit: nanoseconds (ns)
|
||||
- Typical recommended range: 50000–100000
|
||||
- Higher values encourage threads to stay on the same CPU core
|
||||
@@ -0,0 +1,249 @@
|
||||
# Performance Benchmark
|
||||
|
||||
This document details the benchmark methodology for vllm-ascend, aimed at evaluating the performance under a variety of workloads. To maintain alignment with vLLM, we use the [benchmark](https://github.com/vllm-project/vllm/tree/main/benchmarks) script provided by the vllm project.
|
||||
|
||||
**Benchmark Coverage**: We measure offline E2E latency and throughput, and fixed-QPS online serving benchmarks. For more details, see [vllm-ascend benchmark scripts](https://github.com/vllm-project/vllm-ascend/tree/main/benchmarks).
|
||||
|
||||
**Legend Description**:
|
||||
|
||||
- ✅ = Supported
|
||||
- 🟡 = Partial / Work in progress
|
||||
- 🚧 = Under development
|
||||
|
||||
## 1. Run docker container
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
# Update DEVICE according to your device (/dev/davinci[0-7])
|
||||
export DEVICE=/dev/davinci7
|
||||
export IMAGE=m.daocloud.io/quay.io/ascend/vllm-ascend:|vllm_ascend_version|
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--device $DEVICE \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
--device /dev/hisi_hdc \
|
||||
-v /usr/local/dcmi:/usr/local/dcmi \
|
||||
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
|
||||
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
|
||||
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
|
||||
-v /etc/ascend_install.info:/etc/ascend_install.info \
|
||||
-v /root/.cache:/root/.cache \
|
||||
-p 8000:8000 \
|
||||
-e VLLM_USE_MODELSCOPE=True \
|
||||
-it $IMAGE \
|
||||
/bin/bash
|
||||
```
|
||||
|
||||
## 2. Install dependencies
|
||||
|
||||
```bash
|
||||
cd /workspace/vllm-ascend
|
||||
pip config set global.index-url https://mirrors.tuna.tsinghua.edu.cn/pypi/web/simple
|
||||
pip install -r benchmarks/requirements-bench.txt
|
||||
```
|
||||
|
||||
## 3. Run basic benchmarks
|
||||
|
||||
This section introduces how to perform performance testing using the benchmark suite built into vLLM.
|
||||
|
||||
### 3.1 Dataset
|
||||
|
||||
VLLM supports a variety of [datasets](https://github.com/vllm-project/vllm/blob/main/vllm/benchmarks/datasets/datasets.py).
|
||||
|
||||
<style>
|
||||
th {
|
||||
min-width: 0 !important;
|
||||
}
|
||||
</style>
|
||||
|
||||
| Dataset | Online | Offline | Data Path |
|
||||
|---------|--------|---------|-----------|
|
||||
| ShareGPT | ✅ | ✅ | `wget https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered/resolve/main/ShareGPT_V3_unfiltered_cleaned_split.json` |
|
||||
| ShareGPT4V (Image) | ✅ | ✅ | `wget https://huggingface.co/datasets/Lin-Chen/ShareGPT4V/resolve/main/sharegpt4v_instruct_gpt4-vision_cap100k.json`<br>Note that the images need to be downloaded separately. For example, to download COCO's 2017 Train images:<br>`wget http://images.cocodataset.org/zips/train2017.zip` |
|
||||
| ShareGPT4Video (Video) | ✅ | ✅ | `git clone https://huggingface.co/datasets/ShareGPT4Video/ShareGPT4Video` |
|
||||
| BurstGPT | ✅ | ✅ | `wget https://github.com/HPMLL/BurstGPT/releases/download/v1.1/BurstGPT_without_fails_2.csv` |
|
||||
| Sonnet (deprecated) | ✅ | ✅ | Local file: `benchmarks/sonnet.txt` |
|
||||
| Random | ✅ | ✅ | `synthetic` |
|
||||
| RandomMultiModal (Image/Video) | 🟡 | 🚧 | `synthetic` |
|
||||
| RandomForReranking | ✅ | ✅ | `synthetic` |
|
||||
| Prefix Repetition | ✅ | ✅ | `synthetic` |
|
||||
| HuggingFace-VisionArena | ✅ | ✅ | `lmarena-ai/VisionArena-Chat` |
|
||||
| HuggingFace-MMVU | ✅ | ✅ | `yale-nlp/MMVU` |
|
||||
| HuggingFace-InstructCoder | ✅ | ✅ | `likaixin/InstructCoder` |
|
||||
| HuggingFace-AIMO | ✅ | ✅ | `AI-MO/aimo-validation-aime`, `AI-MO/NuminaMath-1.5`, `AI-MO/NuminaMath-CoT` |
|
||||
| HuggingFace-Other | ✅ | ✅ | `lmms-lab/LLaVA-OneVision-Data`, `Aeala/ShareGPT_Vicuna_unfiltered` |
|
||||
| HuggingFace-MTBench | ✅ | ✅ | `philschmid/mt-bench` |
|
||||
| HuggingFace-Blazedit | ✅ | ✅ | `vdaita/edit_5k_char`, `vdaita/edit_10k_char` |
|
||||
| Spec Bench | ✅ | ✅ | `wget https://raw.githubusercontent.com/hemingkx/Spec-Bench/refs/heads/main/data/spec_bench/question.jsonl` |
|
||||
| Custom | ✅ | ✅ | Local file: `data.jsonl` |
|
||||
|
||||
:::{note}
|
||||
The datasets mentioned above are all links to datasets on huggingface.
|
||||
The dataset's `dataset-name` should be set to `hf`.
|
||||
For local `dataset-path`, please set `hf-name` to its Hugging Face ID like
|
||||
|
||||
```bash
|
||||
--dataset-path /datasets/VisionArena-Chat/ --hf-name lmarena-ai/VisionArena-Chat
|
||||
```
|
||||
|
||||
:::
|
||||
|
||||
### 3.2 Run basic benchmark
|
||||
|
||||
#### 3.2.1 Online serving
|
||||
|
||||
First start serving your model:
|
||||
|
||||
```bash
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
vllm serve Qwen/Qwen3-8B
|
||||
```
|
||||
|
||||
Then run the benchmarking script:
|
||||
|
||||
```bash
|
||||
# download dataset
|
||||
# wget https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered/resolve/main/ShareGPT_V3_unfiltered_cleaned_split.json
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
vllm bench serve \
|
||||
--backend vllm \
|
||||
--model Qwen/Qwen3-8B \
|
||||
--endpoint /v1/completions \
|
||||
--dataset-name sharegpt \
|
||||
--dataset-path <your data path>/ShareGPT_V3_unfiltered_cleaned_split.json \
|
||||
--num-prompts 10
|
||||
```
|
||||
|
||||
If successful, you will see the following output:
|
||||
|
||||
```shell
|
||||
============ Serving Benchmark Result ============
|
||||
Successful requests: 10
|
||||
Failed requests: 0
|
||||
Benchmark duration (s): 19.92
|
||||
Total input tokens: 1374
|
||||
Total generated tokens: 2663
|
||||
Request throughput (req/s): 0.50
|
||||
Output token throughput (tok/s): 133.67
|
||||
Peak output token throughput (tok/s): 312.00
|
||||
Peak concurrent requests: 10.00
|
||||
Total Token throughput (tok/s): 202.64
|
||||
---------------Time to First Token----------------
|
||||
Mean TTFT (ms): 127.10
|
||||
Median TTFT (ms): 136.29
|
||||
P99 TTFT (ms): 137.83
|
||||
-----Time per Output Token (excl. 1st token)------
|
||||
Mean TPOT (ms): 25.85
|
||||
Median TPOT (ms): 25.78
|
||||
P99 TPOT (ms): 26.64
|
||||
---------------Inter-token Latency----------------
|
||||
Mean ITL (ms): 25.78
|
||||
Median ITL (ms): 25.74
|
||||
P99 ITL (ms): 28.85
|
||||
==================================================
|
||||
```
|
||||
|
||||
#### 3.2.2 Offline Throughput Benchmark
|
||||
|
||||
```bash
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
vllm bench throughput \
|
||||
--model Qwen/Qwen3-8B \
|
||||
--dataset-name random \
|
||||
--input-len 128 \
|
||||
--output-len 128
|
||||
```
|
||||
|
||||
If successful, you will see the following output
|
||||
|
||||
```shell
|
||||
Processed prompts: 100%|█| 10/10 [00:03<00:00, 2.74it/s, est. speed input: 351.02 toks/s, output: 351.02 toks/s]
|
||||
Throughput: 2.73 requests/s, 699.93 total tokens/s, 349.97 output tokens/s
|
||||
Total num prompt tokens: 1280
|
||||
Total num output tokens: 1280
|
||||
```
|
||||
|
||||
#### 3.2.3 Multi-Modal Benchmark
|
||||
|
||||
```shell
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
vllm serve Qwen/Qwen2.5-VL-7B-Instruct \
|
||||
--dtype bfloat16 \
|
||||
--limit-mm-per-prompt '{"image": 1}' \
|
||||
--allowed-local-media-path /path/to/sharegpt4v/images
|
||||
```
|
||||
|
||||
```shell
|
||||
export HF_ENDPOINT="https://hf-mirror.com"
|
||||
vllm bench serve --model Qwen/Qwen2.5-VL-7B-Instruct \
|
||||
--backend "openai-chat" \
|
||||
--dataset-name hf \
|
||||
--hf-split train \
|
||||
--endpoint "/v1/chat/completions" \
|
||||
--dataset-path "lmarena-ai/vision-arena-bench-v0.1" \
|
||||
--num-prompts 10 \
|
||||
--no-stream
|
||||
```
|
||||
|
||||
```shell
|
||||
============ Serving Benchmark Result ============
|
||||
Successful requests: 10
|
||||
Failed requests: 0
|
||||
Benchmark duration (s): 4.89
|
||||
Total input tokens: 7191
|
||||
Total generated tokens: 951
|
||||
Request throughput (req/s): 2.05
|
||||
Output token throughput (tok/s): 194.63
|
||||
Peak output token throughput (tok/s): 290.00
|
||||
Peak concurrent requests: 10.00
|
||||
Total Token throughput (tok/s): 1666.35
|
||||
---------------Time to First Token----------------
|
||||
Mean TTFT (ms): 722.22
|
||||
Median TTFT (ms): 589.81
|
||||
P99 TTFT (ms): 1377.02
|
||||
-----Time per Output Token (excl. 1st token)------
|
||||
Mean TPOT (ms): 44.13
|
||||
Median TPOT (ms): 34.58
|
||||
P99 TPOT (ms): 124.72
|
||||
---------------Inter-token Latency----------------
|
||||
Mean ITL (ms): 33.14
|
||||
Median ITL (ms): 28.01
|
||||
P99 ITL (ms): 182.28
|
||||
==================================================
|
||||
```
|
||||
|
||||
#### 3.2.4 Embedding Benchmark
|
||||
|
||||
```shell
|
||||
vllm serve Qwen/Qwen3-Embedding-8B --trust-remote-code
|
||||
```
|
||||
|
||||
```shell
|
||||
# download dataset
|
||||
# wget https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered/resolve/main/ShareGPT_V3_unfiltered_cleaned_split.json
|
||||
export VLLM_USE_MODELSCOPE=True
|
||||
vllm bench serve \
|
||||
--model Qwen/Qwen3-Embedding-8B \
|
||||
--backend openai-embeddings \
|
||||
--endpoint /v1/embeddings \
|
||||
--dataset-name sharegpt \
|
||||
--num-prompts 10 \
|
||||
--dataset-path <your dataset path>/datasets/ShareGPT_V3_unfiltered_cleaned_split.json
|
||||
```
|
||||
|
||||
```shell
|
||||
============ Serving Benchmark Result ============
|
||||
Successful requests: 10
|
||||
Failed requests: 0
|
||||
Benchmark duration (s): 0.18
|
||||
Total input tokens: 1372
|
||||
Request throughput (req/s): 56.32
|
||||
Total Token throughput (tok/s): 7726.76
|
||||
----------------End-to-end Latency----------------
|
||||
Mean E2EL (ms): 154.06
|
||||
Median E2EL (ms): 165.57
|
||||
P99 E2EL (ms): 166.66
|
||||
==================================================
|
||||
```
|
||||
@@ -0,0 +1,371 @@
|
||||
# Service Profiling Guide
|
||||
|
||||
In an inference service process, it is sometimes necessary to monitor the internal execution flow of the inference service framework to identify performance issues. By collecting start and end timestamps of key processes, identifying key functions or iterations, recording critical events, and gathering various types of information, performance bottlenecks can be quickly located.
|
||||
|
||||
This guide will walk you through the process of collecting performance data from the vLLM-Ascend service framework and operators. It covers the complete workflow from preparation, collection, analysis, to visualization, helping you quickly get started with performance collection tools.
|
||||
|
||||
Two performance collection solutions are provided below: Ascend PyTorch Profiler and MS Service Profiler. You can choose the appropriate tool for performance analysis and troubleshooting based on your actual requirements.
|
||||
|
||||
## Solution Comparison
|
||||
|
||||
| Feature | Ascend PyTorch Profiler | MS Service Profiler |
|
||||
|:-----|:------------------------|:------------------|
|
||||
| Installation Method | Built-in, no additional installation required | Requires building msserviceprofiler from source |
|
||||
| Collection Granularity | PyTorch operator level | Service framework function level |
|
||||
| Control Method | API request control | Configuration file control |
|
||||
| Applicable Scenarios | Model operator performance analysis | Service framework workflow analysis |
|
||||
| Data Format | ascend_pt format | Chrome Tracing + CSV |
|
||||
| Main Advantage | Operator-level performance analysis | Service framework workflow visualization |
|
||||
| Supported Collection Capabilities | PyTorch operator level | PyTorch operator level and Service framework function level |
|
||||
|
||||
## Quick Selection Guide
|
||||
|
||||
- [**Model Operator Performance** → Use Ascend PyTorch Profiler](#ascend-pytorch-profiler)
|
||||
- [**Service Framework Workflow** → Use MS Service Profiler](#ms-service-profiler)
|
||||
|
||||
---
|
||||
|
||||
## Ascend PyTorch Profiler
|
||||
|
||||
### 0. Installation and Configuration
|
||||
|
||||
No additional packages need to be installed; it can be enabled through command-line configuration. Currently, vLLM enables **python stack** by default, which can significantly inflate the collected performance data. If you do not wish to collect **python stack**, you can disable it using `torch_profiler_with_stack=false`.
|
||||
|
||||
### 1. Preparation for Collection
|
||||
|
||||
Start the online service and set the `--profiler-config` parameter to control the path for saving performance files. After the parameter is set, the collection function is enabled.
|
||||
|
||||
```bash
|
||||
VLLM_PROMPT_SEQ_BUCKET_MAX=128
|
||||
VLLM_PROMPT_SEQ_BUCKET_MIN=128
|
||||
python3 -m vllm.entrypoints.openai.api_server \
|
||||
--port 8080 \
|
||||
--model "facebook/opt-125m" \
|
||||
--tensor-parallel-size 1 \
|
||||
--max-num-seqs 128 \
|
||||
--profiler-config '{"profiler": "torch", "torch_profiler_dir": "./vllm_profile", "torch_profiler_with_stack": false}' \
|
||||
--dtype bfloat16 \
|
||||
--max-model-len 256
|
||||
```
|
||||
|
||||
> Note:**January 19, 2026: The vLLM mainline has deprecated the VLLM_TORCH_PROFILER_DIR environment variable.**[Related PR](https://github.com/vllm-project/vllm-ascend/pull/5928) When using the vLLM Ascend mainline code to collect profiler data, remember to use the `--profiler-config` (online) parameter or the `profiler_config` (offline) parameter.
|
||||
|
||||
### 2. Start Collection
|
||||
|
||||
Performance collection is controlled by sending API requests. You can start collection after stabilizing the actual business data and collect profiling for a few seconds before stopping; or you can start collection first, then send business requests, and finally stop.
|
||||
|
||||
Send the following request to start the profiling service:
|
||||
|
||||
```bash
|
||||
curl -X POST http://localhost:8080/start_profile
|
||||
```
|
||||
|
||||
Send the following request to stop the profiling service:
|
||||
|
||||
```bash
|
||||
curl -X POST http://localhost:8080/stop_profile
|
||||
```
|
||||
|
||||
### 3. Send Requests
|
||||
|
||||
Send requests according to your actual business data. After sending the requests, stop the profiling service, and the data will be automatically saved to the previously configured path:
|
||||
|
||||
```bash
|
||||
curl http://localhost:8080/v1/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "facebook/opt-125m",
|
||||
"prompt": "San Francisco is a",
|
||||
"max_tokens": 7,
|
||||
"temperature": 0
|
||||
}'
|
||||
|
||||
curl -X POST http://localhost:8080/stop_profile
|
||||
```
|
||||
|
||||
### 4. Analyze Data
|
||||
|
||||
Navigate to the `./vllm_profile` directory and locate the generated `*ascend_pt` folder. This folder needs to be analyzed before profiling data can be examined.
|
||||
|
||||
```python
|
||||
from torch_npu.profiler.profiler import analyse
|
||||
analyse("./vllm_profile/localhost.localdomain_*_ascend_pt/")
|
||||
```
|
||||
|
||||
### 5. View Results
|
||||
|
||||
After analysis, the `*ascend_pt` directory will contain many files, with the main analysis focus being the `ASCEND_PROFILER_OUTPUT` folder. This directory will include the following files:
|
||||
|
||||
- `analysis.db`: Performance data in database format
|
||||
|
||||
- `api_statistic.csv`: API call statistics
|
||||
|
||||
- `ascend_pytorch_profiler_0.db`: Performance data in database format
|
||||
|
||||
- `kernel_details.csv`: Kernel-level related data
|
||||
|
||||
- `operator_details.csv`: Operator-level related data
|
||||
|
||||
- `op_statistic.csv`: Operator utilization data
|
||||
|
||||
- `step_trace_time.csv`: Scheduling data
|
||||
|
||||
- `trace_view.json`: Chrome tracing format data, can be opened with [MindStudio Insight](https://www.hiascend.com/document/detail/zh/mindstudio/81RC1/GUI_baseddevelopmenttool/msascendinsightug/Insight_userguide_0002.html)
|
||||
|
||||
[↑ Back to Top](#service-profiling-guide)
|
||||
|
||||
---
|
||||
|
||||
## MS Service Profiler
|
||||
|
||||
### 0. Build from Source and Upgrade
|
||||
|
||||
The `msserviceprofiler` tool is pre-installed with the CANN Toolkit package. Use the following commands to install or upgrade from source.
|
||||
|
||||
```bash
|
||||
git clone https://gitcode.com/Ascend/msserviceprofiler.git
|
||||
cd msserviceprofiler
|
||||
bash scripts/build_and_upgrade.sh
|
||||
```
|
||||
|
||||
### 1. Preparation
|
||||
|
||||
Before starting the service, set the environment variable `SERVICE_PROF_CONFIG_PATH` to point to the profiling configuration file, and set the environment variable `PROFILING_SYMBOLS_PATH` to specify the YAML configuration file for the symbols that need to be imported. After that, start the vLLM service according to your deployment method.
|
||||
|
||||
```bash
|
||||
cd ${path_to_store_profiling_files}
|
||||
# Set environment variable
|
||||
export SERVICE_PROF_CONFIG_PATH=ms_service_profiler_config.json
|
||||
export PROFILING_SYMBOLS_PATH=service_profiling_symbols.yaml
|
||||
|
||||
# Start vLLM service
|
||||
vllm serve Qwen/Qwen2.5-0.5B-Instruct &
|
||||
```
|
||||
|
||||
The file `ms_service_profiler_config.json` is the profiling configuration. If it does not exist at the specified path, a default configuration will be generated automatically. If needed, you can customize it in advance according to the instructions in the `Profiling Configuration File` section below.
|
||||
|
||||
`service_profiling_symbols.yaml` is the configuration file containing the profiling points to be imported. You can choose **not** to set the `PROFILING_SYMBOLS_PATH` environment variable, in which case the default configuration file will be used. If the file does not exist at the path you specified, likewise, the system will generate a configuration file at your specified path for future configuration. You can customize it according to the instructions in the `Symbols Configuration File` section below.
|
||||
|
||||
### 2. Enable Profiling
|
||||
|
||||
To enable the performance data collection switch, change the `enable` field from `0` to `1` in the configuration file `ms_service_profiler_config.json`. This can be accomplished by executing the following sed command:
|
||||
|
||||
```bash
|
||||
sed -i 's/"enable":\s*0/"enable": 1/' ./ms_service_profiler_config.json
|
||||
```
|
||||
|
||||
### 3. Send Requests
|
||||
|
||||
Choose a request-sending method that suits your actual profiling needs:
|
||||
|
||||
```bash
|
||||
curl http://localhost:8000/v1/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "Qwen/Qwen2.5-0.5B-Instruct",
|
||||
"prompt": "Beijing is a",
|
||||
"max_tokens": 5,
|
||||
"temperature": 0
|
||||
}' | python3 -m json.tool
|
||||
```
|
||||
|
||||
### 4. Parse Data
|
||||
|
||||
```bash
|
||||
# xxxx-xxxx is the directory automatically created based on vLLM startup time
|
||||
cd /root/.ms_server_profiler/xxxx-xxxx
|
||||
|
||||
# parse data
|
||||
msserviceprofiler parse --input-path=./ --output-path output
|
||||
```
|
||||
|
||||
### 5. View Results
|
||||
|
||||
After parsing, the `output` directory will contain:
|
||||
|
||||
- `chrome_tracing.json`: Chrome tracing format data, which can be opened in [MindStudio Insight](https://www.hiascend.com/document/detail/zh/mindstudio/830/GUI_baseddevelopmenttool/msascendinsightug/Insight_userguide_0002.html?framework=mindspore).
|
||||
- `profiler.db`: Performance data in database format.
|
||||
- `request.csv`: Request-related data.
|
||||
- `kvcache.csv`: KV Cache-related data.
|
||||
- `batch.csv`: Batch scheduling-related data.
|
||||
|
||||
---
|
||||
|
||||
### 6. Appendix related to MS Service Profiler
|
||||
|
||||
(profiling-configuration-file)=
|
||||
|
||||
#### 6.1 Profiling Configuration File
|
||||
|
||||
The profiling configuration file controls profiling parameters and behavior.
|
||||
|
||||
##### File Format
|
||||
|
||||
The configuration is in JSON format. Main parameters:
|
||||
|
||||
| Parameter | Description | Required |
|
||||
|:------:|:------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:-----:|
|
||||
| enable | Switch for profiling: <br />0: disable<br />1: enable<br />Default: 0 | Yes |
|
||||
| prof_dir | Directory to store collected performance data. <br />Default: `${HOME}/.ms_server_profiler` | No |
|
||||
| profiler_level | Data collection level. Default is "INFO" (normal level). | No |
|
||||
| acl_task_time | Switch to collect operator dispatch latency and execution latency. Values: <br />0: off. Default; 0 or any invalid value means off.<br />1: on. When enabled, calls `aclprofCreateConfig` with `ACL_PROF_TASK_TIME_L0`.<br />2: on. MSPTI-based dump. When enabled, set before starting the service: `export LD_PRELOAD={INSTALL_DIR}/lib64/libmspti.so`, where `{INSTALL_DIR}` is the CANN installation root (e.g. `/usr/local/Ascend/cann` for a typical root install).<br />3: on. Torch Profiler–based dump. | No |
|
||||
| acl_prof_task_time_level | Profiling level and duration. Values: <br />L0: collect operator dispatch and execution latency only; lower overhead (no operator basic info).<br />L1: collect AscendCL interface performance (host–device and inter-device sync/async memory copy latencies), plus operator dispatch, execution, and basic info for comprehensive analysis.<br />`{time}`: optional duration segment; integer 1–999, unit seconds.<br />If unset, defaults to L0 until program exit; invalid values fall back to defaults.<br />Level and duration can be combined, e.g., `"acl_prof_task_time_level": "L1;10"`.<br />**Note:** When Torch Profiler is used (`acl_task_time` set to `3`), `{time}` duration is not supported. | No |
|
||||
| timelimit | Profiling duration for the service. The process stops automatically after this time. Range: integer 0–7200, unit: seconds. Default 0 means unlimited. Recommend at least 120s; shorter runs may lack data for parsed outputs and trigger warnings. | No |
|
||||
| domain | Limit profiling to the specified domains to reduce data volume. String, separated by semicolons, case-sensitive, e.g., "Request; KVCache".<br />Empty means all available domains.<br />Available domains: Request, KVCache, ModelExecute, BatchSchedule, Communication.<br />Note: If the selected domains are incomplete, analysis output may show warnings due to missing data. See [Reference Table 1](https://www.hiascend.com/document/detail/zh/canncommercial/850/devaids/Profiling/mindieprofiling_0010.html). | No |
|
||||
| torch_prof_stack | Collect operator call stacks (framework and CPU operators). Values: `false` (default, off), `true` (on). Requires `acl_task_time` set to `3`. **Note:** Enabling this configuration introduces additional performance overhead. | No |
|
||||
| torch_prof_step_num | Torch Profiler step limit. Integer ≥ 0. Default `0` means collect all steps.<br />Requires `acl_task_time` set to `3`. | No |
|
||||
| profiler_step_num | Step limit for operator and service framework profiling. Integer ≥ 0.<br />`0` or invalid values stop the entire service profiling process.<br />The number of steps actually recorded depends on `modelRunnerExec` events. | No |
|
||||
|
||||
##### Example Configuration
|
||||
|
||||
```json
|
||||
{
|
||||
"enable": 1,
|
||||
"prof_dir": "./vllm_prof",
|
||||
"acl_task_time": 0,
|
||||
"acl_prof_task_time_level": ""
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
(symbols-configuration-file)=
|
||||
|
||||
#### 6.2 Symbols Configuration File
|
||||
|
||||
The symbols configuration file defines which functions/methods to profile and supports flexible configuration with custom attribute collection.
|
||||
|
||||
##### File Name and Loading
|
||||
|
||||
- Default load path: `~/.config/vllm_ascend/service_profiling_symbols.MAJOR.MINOR.PATCH.yaml` (According to the installed version of vLLM)
|
||||
|
||||
If you need to customize the profiling points, it is highly recommended to copy a symbol configuration file to your working directory and point to it with the `PROFILING_SYMBOLS_PATH` environment variable.
|
||||
|
||||
##### Configuration file updates
|
||||
|
||||
After you change profiling symbols, restart the vLLM service so the updated configuration file is loaded.
|
||||
|
||||
##### Field Descriptions
|
||||
|
||||
| Field | Description | Example |
|
||||
|:-----:|:-----|:-----|
|
||||
| symbol | Python import path + attribute chain | `"vllm.v1.core.kv_cache_manager:KVCacheManager.free"` |
|
||||
| handler | Handler type | `"timer"` (default) or `"pkg.mod:func"` (custom) |
|
||||
| domain | Domain tag | `"KVCache"`, `"ModelExecute"` |
|
||||
| name | Event name | `"EngineCoreExecute"` |
|
||||
| min_version | Minimum supported vLLM version | `"0.9.1"` |
|
||||
| max_version | Maximum supported vLLM version | `"0.11.0"` |
|
||||
| attributes | Custom attribute collection | Only supported for `"timer"` handler. See the section below |
|
||||
|
||||
##### Configuration Examples
|
||||
|
||||
- Example 1: Custom handler
|
||||
|
||||
```yaml
|
||||
- symbol: vllm.v1.core.kv_cache_manager:KVCacheManager.free
|
||||
handler: ms_service_profiler.patcher.config.custom_handler_example.kvcache_manager_free_example_handler
|
||||
domain: Example
|
||||
name: example_custom
|
||||
```
|
||||
|
||||
- Example 2: Default timer
|
||||
|
||||
```yaml
|
||||
- symbol: vllm.v1.engine.core:EngineCore.execute_model
|
||||
domain: ModelExecute
|
||||
name: EngineCoreExecute
|
||||
```
|
||||
|
||||
- Example 3: Version constraint
|
||||
|
||||
```yaml
|
||||
- symbol: vllm.v1.executor.abstract:Executor.execute_model
|
||||
min_version: "0.9.1"
|
||||
# No handler specified -> default timer
|
||||
```
|
||||
|
||||
##### Custom Attribute Collection
|
||||
|
||||
The `attributes` field supports flexible custom attribute collection and allows operations and transformations on function arguments and return values.
|
||||
|
||||
###### Basic Syntax
|
||||
|
||||
- Argument access: use the parameter name directly, e.g., `input_ids`
|
||||
- Return value access: use the `return` keyword
|
||||
- Pipeline operations: use `|` to chain multiple operations
|
||||
- Attribute access: use `attr` to access object attributes
|
||||
|
||||
###### Example
|
||||
|
||||
```yaml
|
||||
- symbol: vllm_ascend.worker.model_runner_v1:NPUModelRunner.execute_model
|
||||
name: ModelRunnerExecuteModel
|
||||
domain: ModelExecute
|
||||
attributes:
|
||||
- name: device
|
||||
expr: args[0] | attr device | str
|
||||
- name: dp
|
||||
expr: args[0] | attr dp_rank | str
|
||||
- name: batch_size
|
||||
expr: args[0] | attr input_batch | attr _req_ids | len
|
||||
```
|
||||
|
||||
###### Expression Notes
|
||||
|
||||
1. `len(input_ids)`: get the length of parameter `input_ids`.
|
||||
2. `len(return) | str`: get the length of the return value and convert to string (equivalent to `str(len(return))`).
|
||||
3. `return[0] | attr input_ids | len`: get the length of the `input_ids` attribute of the first element in the return value.
|
||||
|
||||
###### Supported Expression Types
|
||||
|
||||
- Basic operations: `len()`, `str()`, `int()`, `float()`
|
||||
- Index access: `return[0]`, `return['key']`
|
||||
- Attribute access: `return | attr attr_name`
|
||||
- Pipeline composition: chain operations with `|`
|
||||
|
||||
###### Advanced Examples
|
||||
|
||||
```yaml
|
||||
attributes:
|
||||
# Get tensor shape
|
||||
- name: tensor_shape
|
||||
expr: input_tensor | attr shape | str
|
||||
|
||||
# Get specific value from a dict
|
||||
- name: batch_size
|
||||
expr: kwargs['batch_size']
|
||||
|
||||
# Conditional expression (requires custom handler support)
|
||||
- name: is_training_mode
|
||||
expr: training | bool
|
||||
|
||||
# Complex data processing
|
||||
- name: processed_data_len
|
||||
expr: data | attr items | len | str
|
||||
```
|
||||
|
||||
##### Custom Handler
|
||||
|
||||
When `handler` specifies a custom function, it must match the following signature:
|
||||
|
||||
```python
|
||||
def custom_handler(original_func, this, *args, **kwargs):
|
||||
"""
|
||||
Custom handler
|
||||
|
||||
Args:
|
||||
original_func: the original function object
|
||||
this: the bound object (for methods)
|
||||
*args: positional arguments
|
||||
**kwargs: keyword arguments
|
||||
|
||||
Returns:
|
||||
processing result
|
||||
"""
|
||||
# Custom logic
|
||||
pass
|
||||
```
|
||||
|
||||
If the custom handler fails to import, the system will automatically fall back to the default timer mode.
|
||||
|
||||
[↑ Back to Top](#service-profiling-guide)
|
||||
Reference in New Issue
Block a user