init v0.23.0

Signed-off-by: Sun Ruoxi <sunruoxi@4paradigm.com>
This commit is contained in:
2026-08-27 15:11:51 +08:00
parent b582a8e7d1
commit 7f8a1b1f7a
2849 changed files with 712887 additions and 22001 deletions

View File

@@ -0,0 +1,469 @@
# Doc writing guide
## Guide to Writing Model Tutorial Doc
`docs/source/_templates/Model-Deployment-Tutorial-Template.md` is a template for writing model deployment tutorials. You can copy and modify it to create new docs.
## Testable doc code block generation (``model-code``)
- For **documentation authors**: how to insert testable command blocks into docs
- For **developers**: how to add a new converter
Built-in supported `converter_tag` values:
| converter_tag | Renders | YAML source |
| --- | --- | --- |
| `single_node` | A single node's env exports + `vllm serve` script | `test_cases[case_index]` |
| `multi_node` | One host's env exports + `vllm serve` script | `deployment[host_index]` |
| `external_dp_template` | One external-DP node's env exports + `vllm serve` command | `templates[host_index]` |
| `external_dp_launch` | One `launch_online_dp.py` line per node | `config[]` |
| `external_dp_proxy` | The load-balance proxy launch command | `config[]` + `routing` |
### For authors: add a block
:::{important}
By default, the generator scans only `.md` files under `docs/source/tutorials/models/` and produces artifacts.
If you put ``model-code`` blocks in other directories, Sphinx builds will not automatically generate the corresponding scripts.
:::
All ``model-code`` blocks need:
| Option | Required | Description |
| --- | --- | --- |
| `block_name` | Yes | Block name; must be unique within the current document |
| `converter_tag` | Yes | Selects one of the built-in converters |
| `test_case_path` | Yes | Repository-relative YAML path that stays within the repo; the file must exist |
Use the body of the block to add shell wrapper lines such as `set -eux`. Always
place the `{{ generated }}` placeholder where the converter output should be
inserted.
#### converter_tag: `single_node`
`single_node` reads one item from `test_cases`. The optional `case_index`
metadata selects the item; when omitted, it defaults to `0`.
Only the fields read by this converter are expanded below. Other test metadata
can be left in the YAML and is ignored by this converter.
```yaml
test_cases:
- name: qwen3-8b-single
model: Qwen/Qwen3-8B
envs:
HCCL_BUFFSIZE: "1024"
SERVER_PORT: DEFAULT_PORT
server_cmd:
- --tensor-parallel-size
- "1"
- --port
- $SERVER_PORT
- --trust-remote-code
server_cmd_extra:
- --enable-expert-parallel
benchmarks: ...
```
`envs` is rendered as `export` lines. `SERVER_PORT: DEFAULT_PORT` is resolved
to the default single-node port `8000`. `model` becomes `vllm serve <model>`,
and `server_cmd` plus optional `server_cmd_extra` become command arguments.
Both command fields can be either a shell string or a flat token list.
Write the doc block like this:
````md
```{model-code}
:block_name: qwen3_8b_single_node
:converter_tag: single_node
:test_case_path: tests/e2e/nightly/single_node/models/configs/your_model.yaml
:case_index: 0
set -eux
{{ generated }}
```
````
Generated shell script:
```bash
set -eux
export HCCL_BUFFSIZE=1024
export SERVER_PORT=8000
vllm serve Qwen/Qwen3-8B \
--tensor-parallel-size 1 \
--port $SERVER_PORT \
--trust-remote-code \
--enable-expert-parallel
```
#### converter_tag: `multi_node`
`multi_node` reads one item from `deployment`. The required `host_index`
metadata selects which host to render.
```yaml
deployment:
- envs:
SERVER_PORT: "8000"
server_cmd: >
vllm serve Qwen/Qwen3-235B-A22B
--host 0.0.0.0
--port $SERVER_PORT
--data-parallel-size 2
--tensor-parallel-size 8
--data-parallel-address $LOCAL_IP
- envs:
SERVER_PORT: "8000"
server_cmd: >
vllm serve Qwen/Qwen3-235B-A22B
--headless
--port $SERVER_PORT
--data-parallel-size 2
--tensor-parallel-size 8
--data-parallel-start-rank 1
--data-parallel-address $MASTER_IP
benchmarks: ...
```
`server_cmd` must be a complete command starting with `vllm serve <model>`.
It can be written as a shell string or a flat token list.
Write the doc block like this:
````md
```{model-code}
:block_name: qwen3_235b_worker_1
:converter_tag: multi_node
:test_case_path: tests/e2e/nightly/multi_node/internal_dp/config/your_model.yaml
:host_index: 1
set -eux
{{ generated }}
```
````
Generated shell script for `host_index: 1`:
```bash
set -eux
export MASTER_IP=192.168.1.10
export SERVER_PORT=8000
vllm serve Qwen/Qwen3-235B-A22B \
--headless \
--port $SERVER_PORT \
--data-parallel-size 2 \
--tensor-parallel-size 8 \
--data-parallel-start-rank 1 \
--data-parallel-address $MASTER_IP
```
#### converter_tag: `external_dp_template`
`external_dp_template` reads one item from `templates`. The required
`host_index` metadata selects which template to render. The top-level `model`
field is also required because the converter builds `vllm serve <model>`.
```yaml
model: Eco-Tech/GLM-Test
templates:
- node_index: 0
envs:
HCCL_BUFFSIZE: "1024"
ASCEND_RT_VISIBLE_DEVICES: "${VISIBLE_DEVICES}"
server_cmd_template:
- --host
- 0.0.0.0
- --port
- ${PORT}
- --data-parallel-size
- ${DP_SIZE}
- --data-parallel-rank
- ${DP_RANK}
- --data-parallel-address
- ${DP_ADDRESS}
- --data-parallel-rpc-port
- ${DP_RPC_PORT}
- --tensor-parallel-size
- ${TP_SIZE}
- --trust-remote-code
config: ...
routing: ...
```
Known braced template variables are rewritten to the positional shell arguments
that `run_dp_template.sh` receives from `launch_online_dp.py`:
| Template variable | Rendered positional |
| --- | --- |
| `${VISIBLE_DEVICES}` | `$1` |
| `${PORT}` | `$2` |
| `${DP_SIZE}` | `$3` |
| `${DP_RANK}` | `$4` |
| `${DP_ADDRESS}` | `$5` |
| `${DP_RPC_PORT}` | `$6` |
| `${TP_SIZE}` | `$7` |
Unknown braced variables and unbraced shell references such as `$SERVER_PORT`
are left unchanged.
Write the doc block like this:
````md
```{model-code}
:block_name: glm_external_dp_template_node0
:converter_tag: external_dp_template
:test_case_path: tests/e2e/nightly/multi_node/external_dp/config/your_model.yaml
:host_index: 0
set -eux
export HCCL_IF_IP=$local_ip
export GLOO_SOCKET_IFNAME=$nic_name
export TP_SOCKET_IFNAME=$nic_name
export HCCL_SOCKET_IFNAME=$nic_name
{{ generated }}
```
````
Generated shell script for `host_index: 0`:
```bash
set -eux
export HCCL_IF_IP=$local_ip
export GLOO_SOCKET_IFNAME=$nic_name
export TP_SOCKET_IFNAME=$nic_name
export HCCL_SOCKET_IFNAME=$nic_name
export HCCL_BUFFSIZE=1024
export ASCEND_RT_VISIBLE_DEVICES=$1
vllm serve Eco-Tech/GLM-Test \
--host 0.0.0.0 \
--port $2 \
--data-parallel-size $3 \
--data-parallel-rank $4 \
--data-parallel-address $5 \
--data-parallel-rpc-port $6 \
--tensor-parallel-size $7 \
--trust-remote-code
```
#### converter_tag: `external_dp_launch`
`external_dp_launch` reads the full `config` list and renders one
`launch_online_dp.py` command per node. It does not take an index option.
```yaml
config:
- node_index: 0
port_start: 7100
dp_rpc_port: 12321
dp_size: 2
dp_size_local: 2
dp_rank_start: 0
tp_size: 8
dp_address: "${NODE_0_IP}"
- node_index: 1
port_start: 7200
dp_rpc_port: 12321
dp_size: 4
dp_size_local: 4
dp_rank_start: 0
tp_size: 4
dp_address: "${NODE_1_IP}"
templates: ...
routing: ...
```
Write the doc block like this:
````md
```{model-code}
:block_name: glm_external_dp_launch
:converter_tag: external_dp_launch
:test_case_path: tests/e2e/nightly/multi_node/external_dp/config/your_model.yaml
set -eux
{{ generated }}
```
````
Generated shell script:
```bash
set -eux
python launch_online_dp.py --dp-size 2 --tp-size 8 --dp-size-local 2 --dp-rank-start 0 --dp-address ${NODE_0_IP} --dp-rpc-port 12321 --vllm-start-port 7100
python launch_online_dp.py --dp-size 4 --tp-size 4 --dp-size-local 4 --dp-rank-start 0 --dp-address ${NODE_1_IP} --dp-rpc-port 12321 --vllm-start-port 7200
```
#### converter_tag: `external_dp_proxy`
`external_dp_proxy` reads `config` and `routing`. It renders the
`load_balance_proxy_server_example.py` command for `routing.type:
disaggregated_prefill`. It does not take an index option.
```yaml
routing:
type: disaggregated_prefill
groups:
prefiller: [0]
decoder: [1]
config:
- node_index: 0
port_start: 7100
dp_size_local: 2
dp_rpc_port: 12321
dp_size: 2
dp_rank_start: 0
tp_size: 8
dp_address: "${NODE_0_IP}"
- node_index: 1
port_start: 7200
dp_size_local: 4
dp_rpc_port: 12321
dp_size: 4
dp_rank_start: 0
tp_size: 4
dp_address: "${NODE_1_IP}"
templates: ...
```
`routing.groups.prefiller` and `routing.groups.decoder` contain indices into
`config`. Each referenced node expands to `dp_size_local` host and port entries.
The proxy itself is rendered on `${NODE_0_IP}:1999`.
Write the doc block like this:
````md
```{model-code}
:block_name: glm_external_dp_proxy
:converter_tag: external_dp_proxy
:test_case_path: tests/e2e/nightly/multi_node/external_dp/config/your_model.yaml
set -eux
{{ generated }}
```
````
Generated shell script:
```bash
set -eux
python load_balance_proxy_server_example.py \
--host ${NODE_0_IP} \
--port 1999 \
--prefiller-hosts \
${NODE_0_IP} \
${NODE_0_IP} \
--prefiller-ports \
7100 \
7101 \
--decoder-hosts \
${NODE_1_IP} \
${NODE_1_IP} \
${NODE_1_IP} \
${NODE_1_IP} \
--decoder-ports \
7200 \
7201 \
7202 \
7203
```
### Local debugging and generation
#### Generate only (without building the full site)
```bash
# Generate all model-code artifacts under docs/source/tutorials/models/
python3 tools/docs_codegen/cli.py
# Generate artifacts for a single document
python3 tools/docs_codegen/cli.py --doc docs/source/tutorials/models/Kimi-K2-Thinking.md
# Generate a single block and print it (no files written)
python3 tools/docs_codegen/cli.py \
--block docs/source/tutorials/models/Kimi-K2-Thinking.md::kimi_k2_thinking_single_node \
--dry-run --stdout
```
By default, artifacts are written to: `docs/_build/doc_codegen/<doc_stem>/<block_name>.sh`.
:::{note}
After the script is generated, please make sure to check whether the generated content is runnable, especially key parts such as environment variables and command-line parameters.
:::
#### Build the site & preview locally
```bash
# Install documentation build dependencies
python3 -m pip install -r docs/requirements-docs.txt
# (Optional) Clean previous builds
make -C docs clean
# Build the English site
make -C docs html
# (Optional) Build the Chinese site
make -C docs intl
# Preview locally
python3 -m http.server -d docs/_build/html 8000
# Then open in a browser:
# http://localhost:8000
```
### For developers: add a new converter
A converter turns one loaded YAML file plus one parsed `ModelCodeBlock` into a
`GeneratedScript`. The current pipeline is:
1. `BlockScanner` parses ``model-code`` fences and accepts only options listed
in `MODEL_CODE_OPTION_NAMES`.
2. `YamlLoader` loads `test_case_path`.
3. `get_converter()` looks up `block.converter_tag` from
`build_default_converters()`.
4. The selected converter returns `GeneratedScript(content=..., language="shell")`.
5. `GeneratorService` replaces `{{ generated }}` in the block body, validates
that the final script is non-empty, and writes
`docs/_build/doc_codegen/<doc_stem>/<block_name>.sh`.
To add a converter:
1. In `tools/docs_codegen/converters.py`, add a `BaseConverter` subclass with a
unique `name`. That name is the value authors put in `:converter_tag:`.
2. Implement `convert(self, loaded_yaml, *, block) -> GeneratedScript`. Use
`make_docs_codegen_error(..., block=block)` for user-facing validation
errors so the CLI and Sphinx output include document context.
3. Reuse helpers from `tools/docs_codegen/utils.py`, such as
`require_mapping`, `require_mapping_list`, `require_scalar_mapping`,
`require_indexed_mapping`, `require_node_field`, `parse_command_tokens`,
`substitute_template_positionals`, and `render_cli_command`.
4. Register the converter in `build_default_converters()`. If it is not
registered, `get_converter()` will reject the new `converter_tag`.
5. If the converter needs new directive metadata, add the option name to
`MODEL_CODE_OPTION_NAMES` in `tools/docs_codegen/scanner.py` and to
`ModelCodeDirective.option_spec` in
`tools/docs_codegen/sphinx_extension.py`. Read the option with
`block.get_option("<option_name>")`.
6. Add or update tests in `tests/ut/tools/test_docs_codegen.py`. Cover the
successful render path, required option validation, YAML shape validation,
and any CLI/Sphinx scanner behavior affected by new metadata.
7. Add a real ``model-code`` example in a model tutorial, preferably under
`docs/source/tutorials/models/`, and point it to an existing YAML file under
`tests/`.
8. Validate with the CLI:
```bash
python3 tools/docs_codegen/cli.py --doc <your_doc> --dry-run
python3 tools/docs_codegen/cli.py --block <your_doc>::<block_name> --dry-run --stdout
```
If a converter should render something other than shell, set
`GeneratedScript.language` accordingly so Sphinx can highlight the generated
literal block correctly.

View File

@@ -0,0 +1,154 @@
# E2E CI Test
This document explains how to trigger specific E2E tests against your PR code via a
comment command, without running the full E2E test suite.
## Background
The `E2E-Full` workflow ([`pr_test.yaml`](https://github.com/vllm-project/vllm-ascend/blob/main/.github/workflows/pr_test.yaml)) normally runs the complete E2E test suite
when a PR has `ready` label. This is expensive in CI resources
and time.
Authorized users can trigger only the specific test files they care about by posting a
`/e2e` comment on the PR, then adding the `ready` label.
## How to Trigger
### 1. Post a comment
First, post a comment on the PR specifying which test paths to run:
```text
/e2e [test-path-1] [test-path-2] ...
```
- Each path must be a valid pytest path relative to the repository root.
- Multiple paths can be listed in a single comment, separated by spaces.
- A specific test case can be targeted using `::` notation.
| Comment format | Effect |
|---|---|
| `/e2e tests/e2e/pull_request/one_card/test_foo.py` | Run one test file on one_card |
| `/e2e tests/e2e/pull_request/two_card/test_bar.py` | Run one test file on two_card |
| `/e2e path1 path2 path3` | Run multiple files, routed by path pattern |
| `/e2e tests/e2e/pull_request/one_card/test_foo.py::test_case` | Run a specific test case |
### 2. Add the label
After posting the comment, add the **`ready`** label to your PR.
Adding the label is what actually **triggers** the workflow — at that point the workflow
reads the existing comments to find the `/e2e` command.
:::{note}
Only repository **Contributors** (Triage role) and **Maintainers** (Write role) can add
labels. If you do not have this permission, ask a maintainer to add the label for you.
You can find the list of maintainers and contributors by checking the
[CODEOWNERS](https://github.com/vllm-project/vllm-ascend/blob/main/.github/CODEOWNERS)
file.
:::
:::{important}
The comment must be posted **before** the label is added. If you add the label first,
the workflow will find no `/e2e` comment and will not trigger any per-test runs.
:::
:::{note}
Additionally, only the **PR author** or collaborators with **write or admin** repository
access can trigger tests via comment. The workflow validates the commenter's permission
before proceeding.
:::
### 3. Wait for results
GitHub Actions will trigger the `E2E-Full` workflow. Only the hardware jobs matching
the provided test paths will run, which saves CI resources.
## Path Routing Rules
The workflow automatically routes each test path to the correct hardware runner based
on path patterns:
| Path pattern | Hardware | Runner |
|---|---|---|
| `two_card` in path | two_card A3 NPU | `linux-aarch64-a3-2` |
| `four_card` in path | four_card A3 NPU | `linux-aarch64-a3-4` |
| `_310p` in filename under one/two_card | Ascend 310P x1 | `linux-aarch64-310p-*` |
| `_310p` in filename under four_card | Ascend 310P x4 | `linux-aarch64-310p-*` |
| All other paths | one_card A2 NPU | `linux-aarch64-a2b3-1` |
When paths from multiple categories are listed in a single comment, each category's
tests run on its respective hardware in parallel.
## Test Path Reference
The `tests/e2e/pull_request/` directory is organized by hardware category:
```text
tests/e2e/pull_request/
├── one_card/ # Single card tests → A2 NPU x1 runner
├── two_card/ # Two card tests → A3 NPU x2 runner
├── four_card/ # Four card tests → A3 NPU x4 runner
```
310P tests use `_310p` subdirectories or `_310p.py` filename suffix under the
corresponding card directory:
```text
tests/e2e/pull_request/one_card/_310p/ # 310P single card
tests/e2e/pull_request/four_card/_310p/ # 310P four card
```
## Comparison with Full E2E Suite
| Aspect | Full E2E suite | Per-test comment trigger |
|---|---|---|
| Trigger | `ready` labels | `/e2e` comment + `ready` label |
| Scope | All E2E tests | Only specified test paths |
| Who can trigger | Anyone who can add labels | PR author or write/admin collaborator |
| Use case | Pre-merge validation | Iterative debugging of specific tests |
## Examples
Run a single one_card test:
```text
/e2e tests/e2e/pull_request/one_card/test_offline_inference.py
```
Run a two_card test:
```text
/e2e tests/e2e/pull_request/two_card/test_data_parallel.py
```
Run tests across multiple hardware categories in one comment:
```text
/e2e tests/e2e/pull_request/one_card/test_offline_inference.py tests/e2e/pull_request/two_card/test_data_parallel.py
```
Re-trigger after fixing an issue: just push a new commit. The `synchronize` event
re-runs the workflow and picks up the existing `/e2e` comment automatically — no need
to post a new comment.
## Troubleshooting
**The workflow did not start after I added the label.**
- Make sure the `/e2e` comment was posted **before** the label was added.
If the label was added first, remove it and re-add it after posting the comment.
- Check that the comment starts exactly with `/e2e` followed by at least one path,
with no leading spaces or extra characters before the slash.
- To re-trigger after fixing an issue, simply push a new commit — the workflow will
reuse the existing `/e2e` comment automatically.
**Tests ran on the wrong hardware.**
- Check that the path includes the expected directory segment (`one_card`, `two_card`,
`four_card`, or `_310p`). Paths that do not match any of these patterns are routed to
the one_card runner by default.
**The `parse-comment` job skipped with a permission error.**
- Only the PR author or write/admin collaborators can use the comment trigger.
Ask a maintainer to post the `/e2e` comment instead.

View File

@@ -1,16 +1,17 @@
# Contributing
## Building and testing
It's recommended to set up a local development environment to build and test
## Building and Testing
It's recommended to set up a local development environment to build vllm-ascend and run tests
before you submit a PR.
### Setup development environment
### Set up a development environment
Theoretically, the vllm-ascend build is only supported on Linux because
`vllm-ascend` dependency `torch_npu` only supports Linux.
But you can still set up dev env on Linux/Windows/macOS for linting and basic
test as following commands:
But you can still set up a development environment on Linux/Windows/macOS for linting and running basic
tests.
#### Run lint locally
@@ -27,20 +28,19 @@ cd vllm-ascend
# Install lint requirement and enable pre-commit hook
pip install -r requirements-lint.txt
# Run lint (You need install pre-commits deps via proxy network at first time)
# Run lint (You need to install pre-commits deps via proxy network at first time)
bash format.sh
```
#### Run CI locally
After complete "Run lint" setup, you can run CI locally:
After completing "Run lint" setup, you can run CI (Continuous integration) locally:
```{code-block} bash
:substitutions:
cd ~/vllm-project/
# Run CI need vLLM installed
# Run CI needs vLLM installed
git clone --branch |vllm_version| https://github.com/vllm-project/vllm.git
cd vllm
pip install -r requirements/build.txt
@@ -51,7 +51,7 @@ cd ..
cd vllm-ascend
# For Linux:
pip install -r requirements-dev.txt
# For non Linux:
# For non-Linux:
cat requirements-dev.txt | grep -Ev '^#|^--|^$|^-r' | while read PACKAGE; do pip install "$PACKAGE"; done
cat requirements.txt | grep -Ev '^#|^--|^$|^-r' | while read PACKAGE; do pip install "$PACKAGE"; done
@@ -68,13 +68,13 @@ git commit -sm "your commit info"
🎉 Congratulations! You have completed the development environment setup.
### Test locally
### Testing locally
You can refer to [Testing](./testing.md) doc to help you setup testing environment and running tests locally.
You can refer to [Testing](./testing.md) to set up a testing environment and running tests locally.
## DCO and Signed-off-by
When contributing changes to this project, you must agree to the DCO. Commits must include a `Signed-off-by:` header which certifies agreement with the terms of the DCO.
When contributing changes to this project, you must agree to the DCO. Commits must include a `Signed-off-by:` header which certifies agreement with the terms of the DCO (Developer Certificate of Origin).
Using `-s` with `git commit` will automatically add this header.
@@ -88,8 +88,8 @@ Only specific types of PRs will be reviewed. The PR title is prefixed appropriat
- `[Platform]` for new features or optimization in platform.
- `[Worker]` for new features or optimization in worker.
- `[Core]` for new features or optimization in the core vllm-ascend logic (such as platform, attention, communicators, model runner)
- `[Kernel]` changes affecting compute kernels and ops.
- `[Bugfix]` for bug fixes.
- `[Kernel]` for changes affecting compute kernels and ops.
- `[BugFix]` for bug fixes.
- `[Doc]` for documentation fixes and improvements.
- `[Test]` for tests (such as unit tests).
- `[CI]` for build or continuous integration improvements.
@@ -101,11 +101,15 @@ If the PR spans more than one category, please include all relevant prefixes.
## Others
You may find more information about contributing to vLLM Ascend backend plugin on [<u>docs.vllm.ai</u>](https://docs.vllm.ai/en/latest/contributing/overview.html).
If you find any problem when contributing, you can feel free to submit a PR to improve the doc to help other developers.
You may find more information about contributing to vLLM Ascend backend plugin on [<u>docs.vllm.ai</u>](https://docs.vllm.ai/en/latest/contributing).
If you encounter any problems while contributing, feel free to submit a PR to improve the documentation to help other developers.
:::{toctree}
:caption: Index
:maxdepth: 1
testing
doc_writing
multi_node_test
nightly_ci_test
e2e_ci_test
:::

View File

@@ -0,0 +1,553 @@
# Multi Node Test
Multi-Node CI is designed to test distributed scenarios of very large models, for example, disaggregated_prefill multi DP across multi nodes and so on.
## How it works
The following picture shows the basic deployment view of the multi-node CI mechanism. It shows how the GitHub action interacts with [lws](https://lws.sigs.k8s.io/docs/overview/) (a kind of kubernetes crd resource).
![alt text](../../assets/deployment.png)
From the workflow perspective, we can see how the final test script is executed. The key point is that the shared files `tests/e2e/nightly/multi_node/scripts/lws.yaml.jinja2` and `tests/e2e/nightly/multi_node/scripts/run.sh` define the cluster template and pod entry script. Each node executes different logic according to the [LWS_WORKER_INDEX](https://lws.sigs.k8s.io/docs/reference/labels-annotations-and-environment-variables/) environment variable, so that multiple nodes can form a distributed cluster to perform tasks. `run.sh` selects the pytest entrypoint from the config path: internal DP configs use `internal_dp/scripts/test_multi_node.py`, while external DP configs use `external_dp/scripts/test_external_dp.py`.
![alt text](../../assets/workflow.png)
## How to contribute
1. Upload custom weights
If you need customized weights, for example, you quantized a w8a8 weight for DeepSeek-V3 and you want your weight to run on CI, uploading weights to ModelScope's [vllm-ascend](https://www.modelscope.cn/organization/vllm-ascend) organization is welcome. If you do not have permission to upload, please contact @Potabk
2. Add config yaml
For the normal internal DP multi-node flow, add the config yaml to `tests/e2e/nightly/multi_node/internal_dp/config/`, like `DeepSeek-V3.yaml`. External DP cases use the separate `tests/e2e/nightly/multi_node/external_dp/config/` directory and should pass that directory through `config_base_path` in workflow or `CONFIG_BASE_PATH` locally.
Suppose you have **2 nodes** running a 1P1D setup (1 Prefillers + 1 Decoder):
you may add a config file looks like:
```yaml
test_name: "test DeepSeek-V3 disaggregated_prefill"
# the model being tested
model: "vllm-ascend/DeepSeek-V3-W8A8"
# how large the cluster is
num_nodes: 2
npu_per_node: 16
# All env vars you need should add it here
env_common: &env_common
VLLM_USE_MODELSCOPE: true
OMP_PROC_BIND: false
OMP_NUM_THREADS: 100
HCCL_BUFFSIZE: 1024
SERVER_PORT: 8080
disaggregated_prefill:
enabled: true
# node index(a list) which meet all the conditions:
# - prefiller
# - no headless(have api server)
prefiller_host_index: [0]
# node index(a list) which meet all the conditions:
# - decoder
decoder_host_index: [1]
# Add each node's vllm serve cli command just like you run locally
# Add each node's individual envs like follow
deployment:
- name: prefiller node # optional: just for description, not used in code
envs:
<<: *env_common
VLLM_ASCEND_ENABLE_FLASHCOMM1: 1
# Continue to add other envs if needed
server_cmd: >
vllm serve ...
- name: decoder node # optional: just for description, not used in code
envs:
<<: *env_common
VLLM_ASCEND_ENABLE_FLASHCOMM1: 1
# Continue to add other envs if needed
server_cmd: >
vllm serve ...
benchmarks:
perf:
# fill with performance test kwargs
acc:
# fill with accuracy test kwargs
```
3. Add the case to nightly workflow
Currently, the multi-node test workflow is defined in `.github/workflows/schedule_nightly_test_a3.yaml`.
```yaml
multi-node-tests:
name: multi-node
if: always() && (github.event_name == 'schedule' || github.event_name == 'workflow_dispatch')
strategy:
fail-fast: false
max-parallel: 1
matrix:
test_config:
- name: multi-node-deepseek-pd
config_file_path: DeepSeek-V3.yaml
size: 2
- name: multi-node-qwen3-dp
config_file_path: Qwen3-235B-A22B.yaml
size: 2
- name: GLM5_1-W8A8-EP-external
config_file_path: GLM5_1-W8A8-EP-external.yaml
config_base_path: tests/e2e/nightly/multi_node/external_dp/config/
size: 4
uses: ./.github/workflows/_e2e_nightly_multi_node.yaml
with:
soc_version: a3
runner: linux-aarch64-a3-0
image: 'swr.cn-southwest-2.myhuaweicloud.com/base_image/ascend-ci/vllm-ascend:nightly-a3'
replicas: 1
size: ${{ matrix.test_config.size }}
config_file_path: ${{ matrix.test_config.config_file_path }}
config_base_path: ${{ matrix.test_config.config_base_path || '' }}
name: ${{ matrix.test_config.name }}
secrets:
KUBECONFIG_B64: ${{ secrets.KUBECONFIG_B64 }}
```
The matrix above defines all the parameters required to add a multi-machine use
case. The parameters worth noting are `size`, `config_file_path`, and
`config_base_path`. `size` defines the number of nodes required for your use
case. `config_file_path` is the yaml file name, and `config_base_path` tells the
loader which config directory to use. For internal DP cases, use an empty
`config_base_path` so the loader uses its default internal DP config directory.
For external DP cases, set it to
`tests/e2e/nightly/multi_node/external_dp/config/`.
## Run Multi-Node tests locally
### 1. Use kubernetes
This section assumes that you already have a [Kubernetes](https://kubernetes.io/docs/setup/) NPU cluster environment locally. Then you can easily start our test with one click.
- Step 1. Install LWS CRD resources
See <https://lws.sigs.k8s.io/docs/installation/> Which can be used as a reference
- Step 2. Deploy the following yaml file `lws.yaml` as needed
```yaml
apiVersion: leaderworkerset.x-k8s.io/v1
kind: LeaderWorkerSet
metadata:
name: test-server
namespace: vllm-project
spec:
replicas: 1
leaderWorkerTemplate:
size: 2
restartPolicy: None
leaderTemplate:
metadata:
labels:
role: leader
spec:
containers:
- name: vllm-leader
imagePullPolicy: Always
image: swr.cn-southwest-2.myhuaweicloud.com/base_image/ascend-ci/vllm-ascend:nightly-a3
env:
- name: CONFIG_YAML_PATH
value: DeepSeek-V3.yaml
- name: CONFIG_BASE_PATH
value: tests/e2e/nightly/multi_node/internal_dp/config/
- name: WORKSPACE
value: "/vllm-workspace"
- name: FAIL_TAG
value: FAIL_TAG
command:
- sh
- -c
- |
bash /vllm-workspace/vllm-ascend/tests/e2e/nightly/multi_node/scripts/run.sh
resources:
limits:
huawei.com/ascend-1980: 16
memory: 512Gi
ephemeral-storage: 100Gi
requests:
huawei.com/ascend-1980: 16
memory: 512Gi
ephemeral-storage: 100Gi
cpu: 125
ports:
- containerPort: 8080
# readinessProbe:
# tcpSocket:
# port: 8080
# initialDelaySeconds: 15
# periodSeconds: 10
volumeMounts:
- mountPath: /root/.cache
name: shared-volume
- mountPath: /usr/local/Ascend/driver/tools
name: driver-tools
- mountPath: /dev/shm
name: dshm
volumes:
- name: dshm
emptyDir:
medium: Memory
sizeLimit: 15Gi
- name: shared-volume
persistentVolumeClaim:
claimName: nv-action-vllm-benchmarks-v2
- name: driver-tools
hostPath:
path: /usr/local/Ascend/driver/tools
workerTemplate:
spec:
containers:
- name: vllm-worker
imagePullPolicy: Always
image: swr.cn-southwest-2.myhuaweicloud.com/base_image/ascend-ci/vllm-ascend:nightly-a3
env:
- name: CONFIG_YAML_PATH
value: DeepSeek-V3.yaml
- name: CONFIG_BASE_PATH
value: tests/e2e/nightly/multi_node/internal_dp/config/
- name: WORKSPACE
value: "/vllm-workspace"
- name: FAIL_TAG
value: FAIL_TAG
command:
- sh
- -c
- |
bash /vllm-workspace/vllm-ascend/tests/e2e/nightly/multi_node/scripts/run.sh
resources:
limits:
huawei.com/ascend-1980: 16
memory: 512Gi
ephemeral-storage: 100Gi
requests:
huawei.com/ascend-1980: 16
ephemeral-storage: 100Gi
cpu: 125
volumeMounts:
- mountPath: /root/.cache
name: shared-volume
- mountPath: /usr/local/Ascend/driver/tools
name: driver-tools
- mountPath: /dev/shm
name: dshm
volumes:
- name: dshm
emptyDir:
medium: Memory
sizeLimit: 15Gi
- name: shared-volume
persistentVolumeClaim:
claimName: nv-action-vllm-benchmarks-v2
- name: driver-tools
hostPath:
path: /usr/local/Ascend/driver/tools
---
apiVersion: v1
kind: Service
metadata:
name: vllm-leader
namespace: vllm-project
spec:
ports:
- name: http
port: 8080
protocol: TCP
targetPort: 8080
selector:
leaderworkerset.sigs.k8s.io/name: vllm
role: leader
type: ClusterIP
```
```bash
kubectl apply -f lws.yaml
```
Verify the status of the pods:
```bash
kubectl get pods -n vllm-project
```
Should get an output similar to this:
```bash
NAME READY STATUS RESTARTS AGE
vllm-0 1/1 Running 0 2s
vllm-0-1 1/1 Running 0 2s
```
Verify that the distributed inference works:
```bash
kubectl logs -f vllm-0 -n vllm-project
```
Should get something similar to this:
```shell
INFO 12-30 11:00:57 [__init__.py:43] Available plugins for group vllm.platform_plugins:
INFO 12-30 11:00:57 [__init__.py:45] - ascend -> vllm_ascend:register
INFO 12-30 11:00:57 [__init__.py:48] All plugins in this group will be loaded. Set `VLLM_PLUGINS` to control which plugins to load.
INFO 12-30 11:00:57 [__init__.py:217] Platform plugin ascend is activated
INFO 12-30 11:00:57 [importing.py:68] Triton not installed or not compatible; certain GPU-related functions will not be available.
================================================================================================== test session starts ===================================================================================================
platform linux -- Python 3.12.13, pytest-8.4.2, pluggy-1.6.0 -- /usr/local/python3.12.13/bin/python3
cachedir: .pytest_cache
rootdir: /vllm-workspace/vllm-ascend
configfile: pyproject.toml
plugins: cov-7.0.0, asyncio-1.3.0, mock-3.15.1, anyio-4.12.0
asyncio: mode=Mode.STRICT, debug=False, asyncio_default_fixture_loop_scope=None, asyncio_default_test_loop_scope=function
collected 1 item
tests/e2e/nightly/multi_node/internal_dp/scripts/test_multi_node.py::test_multi_node [2025-12-30 11:01:01] INFO multi_node_config.py:294: Loading config yaml: tests/e2e/nightly/multi_node/internal_dp/config/DeepSeek-V3.yaml
[2025-12-30 11:01:01] INFO multi_node_config.py:348: Resolving cluster IPs via DNS...
[2025-12-30 11:01:01] INFO multi_node_config.py:212: Node 0 envs: {'VLLM_USE_MODELSCOPE': 'True', 'OMP_PROC_BIND': 'False', 'OMP_NUM_THREADS': '100', 'HCCL_BUFFSIZE': '1024', 'SERVER_PORT': '8080', 'NUMEXPR_MAX_THREADS': '128', 'DISAGGREGATED_PREFILL_PROXY_SCRIPT': 'examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py', 'HCCL_IF_IP': '10.0.0.102', 'HCCL_SOCKET_IFNAME': 'eth0', 'GLOO_SOCKET_IFNAME': 'eth0', 'TP_SOCKET_IFNAME': 'eth0', 'LOCAL_IP': '10.0.0.102', 'NIC_NAME': 'eth0', 'MASTER_IP': '10.0.0.102'}
[2025-12-30 11:01:01] INFO multi_node_config.py:159: Launching proxy: python examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py --host 10.0.0.102 --port 6000 --prefiller-hosts 10.0.0.102 --prefiller-ports 8080 --decoder-hosts 10.0.0.138 --decoder-ports 8080
[2025-12-30 11:01:01] INFO conftest.py:107: Starting server with command: vllm serve vllm-ascend/DeepSeek-V3-W8A8 --host 0.0.0.0 --port 8080 --data-parallel-size 2 --data-parallel-size-local 2 --tensor-parallel-size 8 --seed 1024 --enforce-eager --enable-expert-parallel --max-num-seqs 16 --max-model-len 8192 --max-num-batched-tokens 8192 --quantization ascend --trust-remote-code --no-enable-prefix-caching --gpu-memory-utilization 0.9 --kv-transfer-config {"kv_connector": "MooncakeConnectorV1", "kv_role": "kv_producer", "kv_port": "30000",
"kv_connector_extra_config": {
"prefill": {
"dp_size": 2,
"tp_size": 8
},
"decode": {
"dp_size": 2,
"tp_size": 8
}
}
}
```
### 2. Test without Kubernetes
The same `tests/e2e/nightly/multi_node/scripts/run.sh` entrypoint can be used
on prepared bare-metal or container hosts. Without LWS, set the values that
Kubernetes normally injects yourself:
- `cluster_hosts` in the config yaml, using IPs reachable from every node.
- `LWS_WORKER_INDEX` on each node, starting from `0`.
- `CONFIG_YAML_PATH` as the config file name and `CONFIG_BASE_PATH` as the
config directory.
Use the host NIC IPs that can reach each other, for example addresses shown by
`ip addr` or `ifconfig` on the active network interface. Do not use per-host
Docker bridge addresses such as `172.17.0.1`, because each host has its own
local bridge.
Local `cluster_hosts` edits should be removed before submitting a PR unless the
hosts are part of a committed test environment.
#### 2.1 Internal DP local run
##### 2.1.1 Add cluster hosts
Edit the internal DP config you want to run, for example:
```text
tests/e2e/nightly/multi_node/internal_dp/config/DeepSeek-V3.yaml
```
Add `cluster_hosts` as a top-level field, for example near `num_nodes` and
`npu_per_node`:
```yaml
cluster_hosts:
- "172.22.0.xxx"
- "172.22.0.xxx"
```
##### 2.1.2 Prepare the environment
Install vllm-ascend development dependencies on every cluster host:
```bash
cd /vllm-workspace/vllm-ascend
python3 -m pip install -r requirements-dev.txt
```
Install AISBench on the first host, which is the node with
`LWS_WORKER_INDEX=0`:
```bash
export AIS_BENCH_TAG="v3.1-20260330-master"
export AIS_BENCH_URL="https://github.com/AISBench/benchmark.git"
export BENCHMARK_HOME=/vllm-workspace/vllm-ascend/benchmark
git clone -b ${AIS_BENCH_TAG} --depth 1 ${AIS_BENCH_URL} $BENCHMARK_HOME
cd $BENCHMARK_HOME
pip install -e . -r requirements/api.txt -r requirements/extra.txt
```
If your local image already contains the model, benchmark data, Ascend runtime,
and AISBench, you only need the run-time exports in the next step.
##### 2.1.3 Start each node
Run the script on each node separately. Start worker nodes first, then start
node 0.
On node 1:
```bash
export WORKSPACE=/vllm-workspace
export IS_PR_TEST=false
export CONFIG_YAML_PATH=DeepSeek-V3.yaml
export CONFIG_BASE_PATH=tests/e2e/nightly/multi_node/internal_dp/config/
export LWS_WORKER_INDEX=1
cd $WORKSPACE/vllm-ascend
bash tests/e2e/nightly/multi_node/scripts/run.sh
```
On node 0:
```bash
export WORKSPACE=/vllm-workspace
export IS_PR_TEST=false
export CONFIG_YAML_PATH=DeepSeek-V3.yaml
export CONFIG_BASE_PATH=tests/e2e/nightly/multi_node/internal_dp/config/
export LWS_WORKER_INDEX=0
cd $WORKSPACE/vllm-ascend
bash tests/e2e/nightly/multi_node/scripts/run.sh
```
Internal DP logs are mainly printed to the terminal running `run.sh`. When
`LOG_PREFIX` is set, the shared script also backs up Ascend logs to:
```text
$LOG_PREFIX/node_<LWS_WORKER_INDEX>_plogs/
```
#### 2.2 External DP local run
##### 2.2.1 Add cluster hosts
Edit the external DP config you want to run. For example:
```text
tests/e2e/nightly/multi_node/external_dp/config/GLM5_1-W8A8-EP-external.yaml
```
Add `cluster_hosts` as a top-level field, for example near `num_nodes` and
`npu_per_node`:
```yaml
cluster_hosts:
- "172.22.0.xxx"
- "172.22.0.xxx"
- "172.22.0.xxx"
- "172.22.0.xxx"
```
##### 2.2.2 Prepare the environment
Install vllm-ascend development dependencies on every cluster host:
```bash
cd /vllm-workspace/vllm-ascend
python3 -m pip install -r requirements-dev.txt
```
Install AISBench on node 0:
```bash
export AIS_BENCH_TAG="v3.1-20260330-master"
export AIS_BENCH_URL="https://github.com/AISBench/benchmark.git"
export BENCHMARK_HOME=/vllm-workspace/vllm-ascend/benchmark
git clone -b ${AIS_BENCH_TAG} --depth 1 ${AIS_BENCH_URL} $BENCHMARK_HOME
cd $BENCHMARK_HOME
pip install -e . -r requirements/api.txt -r requirements/extra.txt
```
If your local image already contains the model, benchmark data, Ascend runtime,
and AISBench, you only need the run-time exports in the next step.
##### 2.2.3 Start each node
External DP uses the same shared `run.sh`. Set `CONFIG_BASE_PATH` to the
external DP config directory so the script chooses
`external_dp/scripts/test_external_dp.py`.
Then start non-master nodes first, and start node 0 last. The following example
uses `GLM5_1-W8A8-EP-external.yaml`, which is a 4-node disaggregated prefill
case.
On node 1, node 2, and node 3, set the matching `LWS_WORKER_INDEX`:
```bash
export WORKSPACE=/vllm-workspace
export IS_PR_TEST=false
export CONFIG_BASE_PATH=tests/e2e/nightly/multi_node/external_dp/config/
export CONFIG_YAML_PATH=GLM5_1-W8A8-EP-external.yaml
export LWS_WORKER_INDEX=1 # Use 2 on node 2, and 3 on node 3.
cd $WORKSPACE/vllm-ascend
bash tests/e2e/nightly/multi_node/scripts/run.sh
```
On node 0:
```bash
export WORKSPACE=/vllm-workspace
export IS_PR_TEST=false
export CONFIG_BASE_PATH=tests/e2e/nightly/multi_node/external_dp/config/
export CONFIG_YAML_PATH=GLM5_1-W8A8-EP-external.yaml
export LWS_WORKER_INDEX=0
cd $WORKSPACE/vllm-ascend
bash tests/e2e/nightly/multi_node/scripts/run.sh
```
For `GLM5_1-W8A8-EP-external.yaml`, node 0 and node 1 start prefiller ranks,
node 2 and node 3 start decoder ranks, and node 0 also starts the proxy and
benchmark.
##### 2.2.4 Read logs while the test is running
The terminal running `run.sh` prints pytest orchestration logs. For external DP,
AISBench output is also printed on node 0, while rank and proxy stdout/stderr
are written to `EXTERNAL_DP_LOG_DIR`. The default layout is:
```text
/tmp/external_dp_logs/
node-0/
rank-0.log
rank-1.log
proxy.log
node-1/
rank-0.log
rank-1.log
```
The first line of each rank log records the exact command and environment used
to start that rank. `proxy.log` exists only on the configured proxy node,
usually node 0.
Use a separate log directory when running multiple local experiments:
```bash
export EXTERNAL_DP_LOG_DIR=/tmp/external_dp_logs_pd_local
```
To watch logs in real time, run these commands in another terminal on the
corresponding node:
```bash
# node 0: ranks and proxy
tail -F /tmp/external_dp_logs/node-0/rank-0.log \
/tmp/external_dp_logs/node-0/rank-1.log \
/tmp/external_dp_logs/node-0/proxy.log
# node 1: ranks
tail -F /tmp/external_dp_logs/node-1/rank-0.log \
/tmp/external_dp_logs/node-1/rank-1.log
```

View File

@@ -0,0 +1,317 @@
# Nightly CI Test
This document explains how to trigger nightly hardware CI tests against your own PR code
on Ascend NPU hardware (A2/A3), without waiting for the scheduled nightly run.
## Background
By default, nightly CI tests run on a fixed schedule using pre-built nightly images.
Contributors can self-service trigger these tests directly against their PR changes
by combining a GitHub label with a comment command.
## How to Trigger
### 1. Post a comment
Post one of the following comments in the PR to specify which tests to run.
The comment itself triggers the workflow — no label is required.
| Comment | Effect |
|---------|--------|
| `/nightly` | Run **all** nightly tests |
| `/nightly all` | Run **all** nightly tests (same as above) |
| `/nightly test1 test2 ...` | Run only the **named** tests |
:::{note}
Only repository **Contributors** (Triage role) and **Maintainers** (Write role) can
trigger the `/nightly` command. If you do not have this permission, ask a maintainer
to post the comment for you. You can find the list of maintainers and contributors in
the project's [Governance](../../community/governance.md) page or by checking the
[CODEOWNERS](https://github.com/vllm-project/vllm-ascend/blob/main/.github/CODEOWNERS)
file.
:::
### 2. Wait for results
GitHub Actions will trigger the `Nightly-A2` or `Nightly-A3` workflow. Only tests
matching the filter will be dispatched, which saves hardware resources.
## Differences Between PR and Scheduled Runs
| | Scheduled / Manual Dispatch | PR-triggered |
|---|----------------------------|---|
| Trigger | Cron (daily) or `workflow_dispatch` | `/nightly` comment |
| Code tested | Pre-built nightly image | Your PR's HEAD commit (source installed fresh) |
| Test scope | All tests | Configurable via `/nightly <names>` |
| vLLM + vllm-ascend | From image | Checked out and installed from source |
| Test matrix | From main branch's matrix YAML | From PR branch's matrix YAML |
When a PR run is detected (`is_pr_test: true`), the workflow additionally:
1. Uninstalls any existing vllm packages in the container.
2. Checks out the specific vllm version and your PR's vllm-ascend commit from source.
3. Installs all dependencies from source.
4. Installs the `aisbench` benchmark suite.
## Test Matrix Data Source
The set of nightly test cases (their names, runners, test paths, model configs) is
declared in a single data file:
```text
.github/workflows/configs/nightly_config.yaml
```
The file is organized as `a2:` and `a3:` top-level keys (one per SoC). Under each
SoC, tests are grouped by execution shape (single-node, multi-node, double-node,
multi-card, accuracy) and each group holds a `test_config` (or `nightly` / `pr_only`
for accuracy) list whose entries carry a `name` plus the fields consumed by the
downstream reusable workflows (`os`, `tests`, `config_file_path`, `size`, etc.).
Both the `Nightly-A2` and `Nightly-A3` workflows dynamically read this file at run
time — there is no hardcoded test matrix in the workflow YAMLs. The
`/nightly <name>` slash command resolves names by walking the same file from the
PR branch, so newly added entries can be exercised on a PR before they land on
main.
## Adding a New Nightly Test Case
To add a new test case (no need to touch the workflow YAMLs):
1. Append an entry under the appropriate section in
`.github/workflows/configs/nightly_config.yaml`. Each entry needs at least:
- `name`: unique identifier used in `/nightly <name>` filters
- `os` (for single-node / multi-card pytest+yaml tests) or `runner` is inferred
- one of `tests:` (pytest directory) or `config_file_path:` (YAML-driven model config)
- `size` (multi-node / double-node only)
2. Add the actual test files (pytest modules under `tests/e2e/nightly/...` or
YAML model configs in `tests/e2e/nightly/.../configs/`).
3. Open a PR. Once CI is green, you can validate the new entry against real NPU
hardware **without** merging the PR — see *Examples* below.
## Available Test Names
The test names you can pass to `/nightly` correspond to the `name` fields under
the matching section in `.github/workflows/configs/nightly_config.yaml`. The
tables below mirror the current contents of that file.
### A2 workflow (`.github/workflows/schedule_nightly_test_a2.yaml`)
**Single-node tests** (`a2.single_node.test_config`):
| Test name | Description |
|-----------|-------------|
| `test_custom_op_multi_card` | Custom operator tests (multi card) |
| `qwen3-vl-32b-instruct-w8a8` | Qwen3-VL-32B-Instruct W8A8 |
| `qwen3-32b-int8` | Qwen3-32B INT8 quantization |
| `Qwen3.5-27B-w8a8-A2` | Qwen3.5-27B W8A8 |
| `Qwen3.5-397B-A17B-w4a8-mtp` | Qwen3.5-397B-A17B W4A8 + MTP |
**Multi-node tests** (`a2.multi_node.test_config`):
| Test name | Description |
|-----------|-------------|
| `multi-node-qwen3-235b-dp` | Qwen3-235B-A22B, 2-node DP |
| `multi-node-GLM-5.1-w8a8-A2` | GLM-5.1 W8A8, 2 nodes |
| `multi-node-Kimi-K2.5-W4A8-A2` | Kimi-K2.5 W4A8, 2 nodes |
**Accuracy tests** (`a2.accuracy.nightly` and `a2.accuracy.pr_only`):
| Test name | Description | Scope |
|-----------|-------------|-------|
| `accuracy-group-1` | Qwen3-VL-8B, Qwen3-8B, Qwen2-Audio-7B, etc. | nightly |
| `accuracy-group-2` | ERNIE-4.5, Molmo-7B, Llama-3.2-3B, etc. | nightly |
| `accuracy-group-3` | Qwen3-30B-A3B, Qwen3-VL-30B-A3B, etc. | nightly |
| `accuracy-group-4` | Qwen3-Next-80B-A3B, Qwen3-Omni-30B-A3B, etc. | nightly |
| `pr-accuracy-group-1` | gemma-3-4b-it, internlm3-8b-instruct, etc. | pr_only |
| `pr-accuracy-group-2` | Qwen2.5-Math-RM-72B, Hunyuan-A13B-Instruct | pr_only |
The `pr-accuracy-group-*` entries only run on `/nightly` (PR-triggered) runs;
`/nightly all` on the schedule skips them.
### A3 workflow (`.github/workflows/schedule_nightly_test_a3.yaml`)
**Multi-node tests** (`a3.multi_node.test_config`, 4-node):
| Test name | Description |
|-----------|-------------|
| `multi-node-deepseek-v3.2-W8A8-EP` | DeepSeek-V3.2-W8A8 with EP, 4-node |
**Double-node tests** (`a3.double_node.test_config`, 2-node, run after multi-node):
| Test name | Description |
|-----------|-------------|
| `multi-node-deepseek-r1-w8a8-longseq` | DeepSeek-R1-W8A8 long sequence, 2-node |
| `multi-node-qwen3-dp` | Qwen3-235B-A22B, 2-node DP |
| `multi-node-qwenw8a8-2node-eplb` | Qwen3-235B-W8A8 with EPLB, 2-node |
| `multi-node-dpsk3.2-2node` | DeepSeek-V3.2-W8A8, 2-node |
| `multi-node-qwenw8a8-2node-longseq` | Qwen3-235B-W8A8 long sequence, 2-node |
| `multi-node-qwen-disagg-pd` | Qwen3-235B disaggregated PD, 2-node |
| `multi-node-qwen-vl-disagg-pd` | Qwen3-VL-235B disaggregated PD, 2-node |
| `multi-node-deepseek-v3.1` | DeepSeek-V3.1-BF16, 2-node |
| `multi-node-deepseek-v3.2-W8A8-EP` | DeepSeek-V3.2-W8A8 with EP, 4-node |
| `multi-node-glm-5.2` | GLM-5.1-W8A8, 2-node |
**Single-node tests** (`a3.single_node.test_config`):
| Test name | Description |
|-----------|-------------|
| `mtpx-deepseek-r1-0528-w8a8` | MTP-X + DeepSeek-R1-0528-W8A8 |
| `deepseek-r1-0528-w8a8` | DeepSeek-R1-0528-W8A8 |
| `kimi-k2-thinking` | Kimi-K2-Thinking |
| `qwen3-vl-235b-a22b-instruct-w8a8` | Qwen3-VL-235B-A22B-Instruct-W8A8 |
| `deepseek-r1-0528-w8a8-prefix-cache` | DeepSeek-R1-0528-W8A8 prefix cache |
| `deepseek-v3-2-w8a8` | DeepSeek-V3.2-W8A8 |
| `glm-4.7-w8a8` | GLM-4.7 W8A8 |
| `kimi-k2.5` | Kimi-K2.5 |
| `qwen3-235b-a22b-w8a8` | Qwen3-235B-A22B-W8A8 |
| `Qwen3.5-397B-A17B-w8a8-mtp` | Qwen3.5-397B-A17B W8A8 + MTP |
| `MiniMax-M2.5-w8a8-QuaRot-A3` | MiniMax-M2.5 W8A8 + QuaRot |
| `Qwen3.5-27B-w8a8-A3` | Qwen3.5-27B W8A8 |
| `Qwen3.5-122B-A10B-W8A8-A3` | Qwen3.5-122B-A10B W8A8 |
| `DeepSeek-V4-Flash-W8A8-A3` | DeepSeek-V4-Flash W8A8 |
**Multi-card tests** (`a3.multi_card.test_config`):
| Test name | Description |
|-----------|-------------|
| `qwen3-30b-acc` | Qwen3-30B accuracy test |
| `qwen3-30b-a3b-w8a8` | Qwen3-30B-A3B-W8A8 |
| `qwen3-32b-int8` | Qwen3-32B-Int8 |
| `qwen3-32b-int8-prefix-cache` | Qwen3-32B-Int8 prefix cache |
| `Qwen3-30B-A3B-W4A8-llm-compressor` | Qwen3-30B-A3B W4A8 via llm-compressor |
| `Qwen3-30B-QuaRot` | Qwen3-30B QuaRot + eagle3 |
| `Qwen3-32B-QuaRot` | Qwen3-32B QuaRot + eagle3 |
:::{warning}
The A3 resource pool has a maximum concurrency of **5×16 NPUs**. Multi-node tests
run with `max-parallel: 2` to avoid resource exhaustion. Running `/nightly all` on
A3 will queue a large number of jobs — prefer targeting specific test names when
possible.
:::
## Examples
Run all available nightly tests against your PR:
```text
/nightly
```
Run only the custom operator multi-card test:
```text
/nightly test_custom_op_multi_card
```
Run two specific tests at once (one per SoC):
```text
/nightly test_custom_op_multi_card mtpx-deepseek-r1-0528-w8a8
```
Run a single accuracy group (with all of its models):
```text
/nightly accuracy-group-1
```
Run a single accuracy model (only that model from a group):
```text
/nightly accuracy-group-1/Qwen3-8B
```
Re-trigger after fixing an issue: just push a new commit. The `synchronize` event
re-runs the workflow and picks up the existing `/nightly` comment automatically — no
need to post a new comment.
## Adding a New Test Case — Worked Example
To add `my-new-test` to the A2 single-node section:
1. Edit `.github/workflows/configs/nightly_config.yaml`, append under
`a2.single_node.test_config`:
```yaml
- name: my-new-test
os: linux-aarch64-a2b3-4
tests: tests/e2e/nightly/single_node/ops/multicard_ops_a2/test_my_new.py
```
2. Commit the new pytest file (`test_my_new.py`) in the same PR.
3. Trigger from the PR:
```text
/nightly my-new-test
```
The workflow will:
- `pr_nightly_command.yml` reads your PR's `nightly_config.yaml` and resolves
`my-new-test` → dispatch A2 only.
- `Nightly-A2` is dispatched at `main`, but `generate-a2-matrix` checks out your
PR commit and reads the new entry from the matrix.
- `single-node-tests` runs one matrix job for `my-new-test`, with
`should_run=true`. The reusable workflow checks out your PR code (via
`vllm_ascend_ref`) and runs your pytest.
## Troubleshooting
**The workflow didn't start after I posted the comment.**
- Check that the comment starts exactly with `/nightly` with no leading spaces or
extra characters before the slash.
- Confirm you have at least Triage permission on the repository; unauthorized
users' comments are ignored.
- To re-trigger after fixing an issue, simply push a new commit — the workflow will
reuse the existing `/nightly` comment automatically.
**Only some tests ran, not the ones I expected.**
- Test names are case-sensitive and must match the `name` field in
`.github/workflows/configs/nightly_config.yaml` exactly (see the tables above).
- For a PR-triggered run, the matrix is loaded from your PR's
`nightly_config.yaml`, not main. If a name isn't in your PR's file, it won't
be recognized and the dispatch will be skipped.
- Check the `parse-trigger` job output in GitHub Actions for the resolved
`test_filter` value.
**The workflow ran with the scheduled image, not my PR code.**
- Confirm the workflow was triggered by `repository_dispatch` (slash command),
not bare `workflow_dispatch`. The `pr_nightly_command.yml` workflow is what
actually dispatches `schedule_nightly_test_a2.yaml` / `_a3.yaml` with
`vllm_ascend_ref` pointing at your PR SHA.
**A new test I added isn't being recognized.**
- Confirm the entry is well-formed YAML under
`.github/workflows/configs/nightly_config.yaml`. The `name` field is required
and must be unique within the SoC's section.
- The matrix is loaded from your PR branch, so make sure the file is committed
to the same branch the `/nightly` comment was posted on.
**How to obtain more detailed logs to pinpoint problems for multi-node tests**
- For most issues, the stdout pop-up logs from GitHub actions are sufficient (this log always represents the logs from the first node).
- If the logs from a first node are no longer sufficient to provide effective logging information, see the summary of your jobs to download log archive for the corresponding test, which includes the framework-side logs and plog information for each node, structured as follows:
```shell
.
├── node0
│ ├── root
│ │ └── ascend
│ │ └── log
│ └── var
│ └── log
│ └── vllm-deepseek-v3-0f233d-0_logs.txt
└── node1
├── root
│ └── ascend
│ └── log
└── var
└── log
└── vllm-deepseek-v3-0f233d-0-1_logs.txt
```

View File

@@ -1,10 +1,10 @@
# Testing
This secition explains how to write e2e tests and unit tests to verify the implementation of your feature.
This document explains how to write unit tests, E2E tests, and nightly tests to verify your feature implementation.
## Setup test environment
## Set up a test environment
The fastest way to setup test environment is to use the main branch container image:
The fastest way to set up a test environment is to use the main branch's container image:
:::::{tab-set}
:sync-group: e2e
@@ -13,7 +13,7 @@ The fastest way to setup test environment is to use the main branch container im
:selected:
:sync: cpu
You can run the unit tests on CPU with the following steps:
You can run the unit tests on CPUs with the following steps:
```{code-block} bash
:substitutions:
@@ -22,39 +22,50 @@ cd ~/vllm-project/
# ls
# vllm vllm-ascend
# Use mirror to speedup download
# docker pull quay.nju.edu.cn/ascend/cann:|cann_image_tag|
# Use mirror to speed up download
# docker pull m.daocloud.io/quay.io/ascend/cann:|cann_image_tag|
export IMAGE=quay.io/ascend/cann:|cann_image_tag|
docker run --rm --name vllm-ascend-ut \
-v $(pwd):/vllm-project \
-v ~/.cache:/root/.cache \
-ti $IMAGE bash
# (Optional) Configure mirror to speedup download
# (Optional) Configure mirror to speed up download
sed -i 's|ports.ubuntu.com|mirrors.huaweicloud.com|g' /etc/apt/sources.list
pip config set global.index-url https://mirrors.huaweicloud.com/repository/pypi/simple/
# For torch-npu dev version or x86 machine
# For TorchNPU dev version or x86 machine
export PIP_EXTRA_INDEX_URL="https://download.pytorch.org/whl/cpu/ https://mirrors.huaweicloud.com/ascend/repos/pypi"
# src path
export SRC_WORKSPACE=/vllm-workspace
mkdir -p $SRC_WORKSPACE
cd $SRC_WORKSPACE
apt-get update -y
apt-get install -y python3-pip git vim wget net-tools gcc g++ cmake libnuma-dev curl gnupg2
# Install vllm
cd /vllm-project/vllm
VLLM_TARGET_DEVICE=empty python3 -m pip -v install .
git clone -b |vllm_ascend_version| --depth 1 https://github.com/vllm-project/vllm-ascend.git
git clone --depth 1 https://github.com/vllm-project/vllm.git
# Install vllm-ascend
cd /vllm-project/vllm-ascend
# [IMPORTANT] Import LD_LIBRARY_PATH to enumerate the CANN environment under CPU
# vllm
cd $SRC_WORKSPACE/vllm
VLLM_TARGET_DEVICE=empty python3 -m pip install .
python3 -m pip uninstall -y triton
# vllm-ascend
cd $SRC_WORKSPACE/vllm-ascend
export LD_LIBRARY_PATH=$LD_LIBRARY_PATH:/usr/local/Ascend/ascend-toolkit/latest/$(uname -m)-linux/devlib
# For cpu environment, set SOC_VERSION for different chips.
# See https://github.com/vllm-project/vllm-ascend/blob/3cb0af0bcf3299089ca7e72159fa36e825a470f8/setup.py#L132 for detail.
export SOC_VERSION="ascend910b1"
python3 -m pip install .
python3 -m pip install -r requirements-dev.txt
python3 -m pip install -v .
```
::::
::::{tab-item} Single card
::::{tab-item} Single-card
:sync: single
```{code-block} bash
@@ -66,6 +77,7 @@ export DEVICE=/dev/davinci0
export IMAGE=quay.io/ascend/vllm-ascend:main
docker run --rm \
--name vllm-ascend \
--shm-size=1g \
--device $DEVICE \
--device /dev/davinci_manager \
--device /dev/devmm_svm \
@@ -86,13 +98,16 @@ After starting the container, you should install the required packages:
# Prepare
pip config set global.index-url https://mirrors.tuna.tsinghua.edu.cn/pypi/web/simple
# Switch to the /vllm-workspace/vllm-ascend directory
cd /vllm-workspace/vllm-ascend/
# Install required packages
pip install -r requirements-dev.txt
```
::::
::::{tab-item} Multi cards
::::{tab-item} Multi-cards
:sync: multi
```{code-block} bash
@@ -101,6 +116,7 @@ pip install -r requirements-dev.txt
export IMAGE=quay.io/ascend/vllm-ascend:main
docker run --rm \
--name vllm-ascend \
--shm-size=1g \
--device /dev/davinci0 \
--device /dev/davinci1 \
--device /dev/davinci2 \
@@ -136,13 +152,13 @@ pip install -r requirements-dev.txt
## Running tests
### Unit test
### Unit tests
There are several principles to follow when writing unit tests:
- The test file path should be consistent with source file and start with `test_` prefix, such as: `vllm_ascend/worker/worker_v1.py` --> `tests/ut/worker/test_worker_v1.py`
- The vLLM Ascend test are using unittest framework, see [here](https://docs.python.org/3/library/unittest.html#module-unittest) to understand how to write unit tests.
- All unit tests can be run on CPU, so you must mock the device-related function to host.
- The test file path should be consistent with the source file and start with the `test_` prefix, such as: `vllm_ascend/worker/worker.py` --> `tests/ut/worker/test_worker.py`
- The vLLM Ascend test uses unittest framework. See [the Python unittest documentation](https://docs.python.org/3/library/unittest.html#module-unittest) to understand how to write unit tests.
- All unit tests can be run on CPUs, so you must mock the device-related functions on the host.
- Example: [tests/ut/test_ascend_config.py](https://github.com/vllm-project/vllm-ascend/blob/main/tests/ut/test_ascend_config.py).
- You can run the unit tests using `pytest`:
@@ -161,12 +177,12 @@ TORCH_DEVICE_BACKEND_AUTOLOAD=0 pytest -sv tests/ut
::::
::::{tab-item} Single card
::::{tab-item} Single-card
:sync: single
```bash
cd /vllm-workspace/vllm-ascend/
# Run all single card the tests
# Run all single-card tests
pytest -sv tests/ut
# Run single test
@@ -175,12 +191,12 @@ pytest -sv tests/ut/test_ascend_config.py
::::
::::{tab-item} Multi cards test
::::{tab-item} Multi-card
:sync: multi
```bash
cd /vllm-workspace/vllm-ascend/
# Run all single card the tests
# Run all multi-card tests
pytest -sv tests/ut
# Run single test
@@ -193,8 +209,63 @@ pytest -sv tests/ut/test_ascend_config.py
### E2E test
Although vllm-ascend CI provide [e2e test](https://github.com/vllm-project/vllm-ascend/blob/main/.github/workflows/vllm_ascend_test.yaml) on Ascend CI, you can run it
locally.
Although vllm-ascend CI provides E2E tests on Ascend CI (for example,
[schedule_nightly_test_a2.yaml](https://github.com/vllm-project/vllm-ascend/blob/main/.github/workflows/schedule_nightly_test_a2.yaml), [schedule_nightly_test_a3.yaml](https://github.com/vllm-project/vllm-ascend/blob/main/.github/workflows/schedule_nightly_test_a3.yaml), [pr_test.yaml](https://github.com/vllm-project/vllm-ascend/blob/main/.github/workflows/pr_test.yaml)), you can run them locally.
#### PR-triggered E2E test
You can run tests with `pytest` as well. Typical examples:
:::::{tab-set}
:sync-group: e2e
::::{tab-item} Local (CPU)
:sync: cpu
You can't run the E2E test on CPUs.
::::
::::{tab-item} Single-card
:selected:
:sync: single
```bash
cd /vllm-workspace/vllm-ascend/
# Run all single-card tests
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/pull_request/one_card/
# Run a certain test script
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/pull_request/one_card/test_camem.py
# Run a certain case in test script
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/pull_request/one_card/test_camem.py::test_end_to_end
```
::::
::::{tab-item} Multi-card
:sync: multi
```bash
cd /vllm-workspace/vllm-ascend/
# Run all multi-card tests
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/pull_request/two_card/
# Run a certain test script
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/pull_request/two_card/test_qwen3_moe_eplb.py
# Run a certain case in test script
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/pull_request/two_card/test_qwen3_moe_eplb.py::test_qwen3_moe_w8a8_distributed_tp2_ep_dynamic_eplb
```
::::
:::::
This will reproduce the E2E test behavior.
#### Nightly-triggered E2E test
You can run tests with `pytest` as well. Typical examples:
:::::{tab-set}
:sync-group: e2e
@@ -202,84 +273,102 @@ locally.
::::{tab-item} Local (CPU)
:sync: cpu
You can't run e2e test on CPU.
You can't run the E2E test on CPUs.
::::
::::{tab-item} Single card
::::{tab-item} Single-card
:selected:
:sync: single
```bash
cd /vllm-workspace/vllm-ascend/
# Run all single card the tests
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/singlecard/
# Run a certain test script
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/singlecard/test_offline_inference.py
# Run a certain case in test script
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/singlecard/test_offline_inference.py::test_models
# run all single-card op tests
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/nightly/single_node/ops/singlecard_ops/
```
::::
::::{tab-item} Multi cards test
::::{tab-item} Multi-card
:sync: multi
```bash
cd /vllm-workspace/vllm-ascend/
# Run all single card the tests
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/multicard/
# run all multi-card op tests on A2
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/nightly/single_node/ops/multicard_ops_a2/
# Run a certain test script
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/multicard/test_dynamic_npugraph_batchsize.py
# Run a certain case in test script
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/multicard/test_offline_inference.py::test_models
# run all multi-card op tests on A3
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/nightly/single_node/ops/multicard_ops_a3/
```
::::
:::::
This will reproduce e2e test: [vllm_ascend_test.yaml](https://github.com/vllm-project/vllm-ascend/blob/main/.github/workflows/vllm_ascend_test.yaml).
For running nightly single-node model test cases locally, refer to the following example.
#### E2E test example:
```bash
export CONFIG_YAML_PATH=Qwen3-32B.yaml
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/nightly/single_node/models/scripts/test_single_node.py
```
- Offline test example: [`tests/e2e/singlecard/test_offline_inference.py`](https://github.com/vllm-project/vllm-ascend/blob/main/tests/e2e/singlecard/test_offline_inference.py)
- Online test examples: [`tests/e2e/singlecard/test_prompt_embedding.py`](https://github.com/vllm-project/vllm-ascend/blob/main/tests/e2e/singlecard/test_prompt_embedding.py)
- Correctness test example: [`tests/e2e/singlecard/test_aclgraph.py`](https://github.com/vllm-project/vllm-ascend/blob/main/tests/e2e/singlecard/test_aclgraph.py)
- Reduced Layer model test example: [test_torchair_graph_mode.py - DeepSeek-V3-Pruning](https://github.com/vllm-project/vllm-ascend/blob/20767a043cccb3764214930d4695e53941de87ec/tests/e2e/multicard/test_torchair_graph_mode.py#L48)
For running nightly multi-node model test cases locally, refer to the `Running Locally` section in [Multi Node Test](./multi_node_test.md).
The CI resource is limited, you might need to reduce layer number of the model, below is an example of how to generate a reduced layer model:
1. Fork the original model repo in modelscope, we need all the files in the repo except for weights.
2. Set `num_hidden_layers` to the expected number of layers, e.g., `{"num_hidden_layers": 2,}`
3. Copy the following python script as `generate_random_weight.py`. Set the relevant parameters `MODEL_LOCAL_PATH`, `DIST_DTYPE` and `DIST_MODEL_PATH` as needed:
#### E2E test examples
```python
import torch
from transformers import AutoTokenizer, AutoConfig
from modeling_deepseek import DeepseekV3ForCausalLM
from modelscope import snapshot_download
- Offline test example: [`tests/e2e/pull_request/one_card/test_camem.py`](https://github.com/vllm-project/vllm-ascend/blob/main/tests/e2e/pull_request/one_card/test_camem.py)
MODEL_LOCAL_PATH = "~/.cache/modelscope/models/vllm-ascend/DeepSeek-V3-Pruning"
DIST_DTYPE = torch.bfloat16
DIST_MODEL_PATH = "./random_deepseek_v3_with_2_hidden_layer"
The CI resource is limited, and you might need to reduce the number of layers of a model. Below is an example of how to generate a reduced layer model:
config = AutoConfig.from_pretrained(MODEL_LOCAL_PATH, trust_remote_code=True)
model = DeepseekV3ForCausalLM(config)
model = model.to(DIST_DTYPE)
model.save_pretrained(DIST_MODEL_PATH)
```
1. Fork the original model repo in modelscope. All the files in the repo except for weights are required.
2. Set `num_hidden_layers` to the expected number of layers, e.g., `{"num_hidden_layers": 2,}`
3. Copy the following python script as `generate_random_weight.py`. Set the relevant parameters `MODEL_LOCAL_PATH`, `DIST_DTYPE` and `DIST_MODEL_PATH` as needed:
```python
import torch
from transformers import AutoTokenizer, AutoConfig
from modeling_deepseek import DeepseekV3ForCausalLM
from modelscope import snapshot_download
MODEL_LOCAL_PATH = "~/.cache/modelscope/models/vllm-ascend/DeepSeek-V3-Pruning"
DIST_DTYPE = torch.bfloat16
DIST_MODEL_PATH = "./random_deepseek_v3_with_2_hidden_layer"
config = AutoConfig.from_pretrained(MODEL_LOCAL_PATH, trust_remote_code=True)
model = DeepseekV3ForCausalLM(config)
model = model.to(DIST_DTYPE)
model.save_pretrained(DIST_MODEL_PATH)
```
### Run doctest
vllm-ascend provides a `vllm-ascend/tests/e2e/run_doctests.sh` command to run all doctests in the doc files.
The doctest is a good way to make sure the docs are up to date and the examples are executable, you can run it locally as follows:
The doctest is a good way to make sure docs stay current and examples remain executable, which can be run locally as follows:
```bash
# Run doctest
/vllm-workspace/vllm-ascend/tests/e2e/run_doctests.sh
```
This will reproduce the same environment as the CI: [vllm_ascend_doctest.yaml](https://github.com/vllm-project/vllm-ascend/blob/main/.github/workflows/vllm_ascend_doctest.yaml).
This will reproduce the same environment as the CI. See [labeled_doctest.yaml](https://github.com/vllm-project/vllm-ascend/blob/main/.github/workflows/labeled_doctest.yaml).
### Run docs link check
You can validate external links in the Sphinx docs locally with:
```bash
make -C docs linkcheck SPHINXOPTS="-W --keep-going"
```
To check links in a specific Markdown file, pass the file to `sphinx-build`.
For example, to check only `docs/source/user_guide/release_notes.md`:
```bash
cd docs
sphinx-build -b linkcheck -W --keep-going \
source _build/linkcheck source/user_guide/release_notes.md
```
The detailed report will be written to:
- `docs/_build/linkcheck/output.txt`
- `docs/_build/linkcheck/output.json`