469
docs/source/developer_guide/contribution/doc_writing.md
Normal file
469
docs/source/developer_guide/contribution/doc_writing.md
Normal file
@@ -0,0 +1,469 @@
|
||||
# Doc writing guide
|
||||
|
||||
## Guide to Writing Model Tutorial Doc
|
||||
|
||||
`docs/source/_templates/Model-Deployment-Tutorial-Template.md` is a template for writing model deployment tutorials. You can copy and modify it to create new docs.
|
||||
|
||||
## Testable doc code block generation (``model-code``)
|
||||
|
||||
- For **documentation authors**: how to insert testable command blocks into docs
|
||||
- For **developers**: how to add a new converter
|
||||
|
||||
Built-in supported `converter_tag` values:
|
||||
|
||||
| converter_tag | Renders | YAML source |
|
||||
| --- | --- | --- |
|
||||
| `single_node` | A single node's env exports + `vllm serve` script | `test_cases[case_index]` |
|
||||
| `multi_node` | One host's env exports + `vllm serve` script | `deployment[host_index]` |
|
||||
| `external_dp_template` | One external-DP node's env exports + `vllm serve` command | `templates[host_index]` |
|
||||
| `external_dp_launch` | One `launch_online_dp.py` line per node | `config[]` |
|
||||
| `external_dp_proxy` | The load-balance proxy launch command | `config[]` + `routing` |
|
||||
|
||||
### For authors: add a block
|
||||
|
||||
:::{important}
|
||||
By default, the generator scans only `.md` files under `docs/source/tutorials/models/` and produces artifacts.
|
||||
If you put ``model-code`` blocks in other directories, Sphinx builds will not automatically generate the corresponding scripts.
|
||||
:::
|
||||
|
||||
All ``model-code`` blocks need:
|
||||
|
||||
| Option | Required | Description |
|
||||
| --- | --- | --- |
|
||||
| `block_name` | Yes | Block name; must be unique within the current document |
|
||||
| `converter_tag` | Yes | Selects one of the built-in converters |
|
||||
| `test_case_path` | Yes | Repository-relative YAML path that stays within the repo; the file must exist |
|
||||
|
||||
Use the body of the block to add shell wrapper lines such as `set -eux`. Always
|
||||
place the `{{ generated }}` placeholder where the converter output should be
|
||||
inserted.
|
||||
|
||||
#### converter_tag: `single_node`
|
||||
|
||||
`single_node` reads one item from `test_cases`. The optional `case_index`
|
||||
metadata selects the item; when omitted, it defaults to `0`.
|
||||
|
||||
Only the fields read by this converter are expanded below. Other test metadata
|
||||
can be left in the YAML and is ignored by this converter.
|
||||
|
||||
```yaml
|
||||
test_cases:
|
||||
- name: qwen3-8b-single
|
||||
model: Qwen/Qwen3-8B
|
||||
envs:
|
||||
HCCL_BUFFSIZE: "1024"
|
||||
SERVER_PORT: DEFAULT_PORT
|
||||
server_cmd:
|
||||
- --tensor-parallel-size
|
||||
- "1"
|
||||
- --port
|
||||
- $SERVER_PORT
|
||||
- --trust-remote-code
|
||||
server_cmd_extra:
|
||||
- --enable-expert-parallel
|
||||
benchmarks: ...
|
||||
```
|
||||
|
||||
`envs` is rendered as `export` lines. `SERVER_PORT: DEFAULT_PORT` is resolved
|
||||
to the default single-node port `8000`. `model` becomes `vllm serve <model>`,
|
||||
and `server_cmd` plus optional `server_cmd_extra` become command arguments.
|
||||
Both command fields can be either a shell string or a flat token list.
|
||||
|
||||
Write the doc block like this:
|
||||
|
||||
````md
|
||||
```{model-code}
|
||||
:block_name: qwen3_8b_single_node
|
||||
:converter_tag: single_node
|
||||
:test_case_path: tests/e2e/nightly/single_node/models/configs/your_model.yaml
|
||||
:case_index: 0
|
||||
|
||||
set -eux
|
||||
{{ generated }}
|
||||
```
|
||||
````
|
||||
|
||||
Generated shell script:
|
||||
|
||||
```bash
|
||||
set -eux
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export SERVER_PORT=8000
|
||||
|
||||
vllm serve Qwen/Qwen3-8B \
|
||||
--tensor-parallel-size 1 \
|
||||
--port $SERVER_PORT \
|
||||
--trust-remote-code \
|
||||
--enable-expert-parallel
|
||||
```
|
||||
|
||||
#### converter_tag: `multi_node`
|
||||
|
||||
`multi_node` reads one item from `deployment`. The required `host_index`
|
||||
metadata selects which host to render.
|
||||
|
||||
```yaml
|
||||
deployment:
|
||||
- envs:
|
||||
SERVER_PORT: "8000"
|
||||
server_cmd: >
|
||||
vllm serve Qwen/Qwen3-235B-A22B
|
||||
--host 0.0.0.0
|
||||
--port $SERVER_PORT
|
||||
--data-parallel-size 2
|
||||
--tensor-parallel-size 8
|
||||
--data-parallel-address $LOCAL_IP
|
||||
- envs:
|
||||
SERVER_PORT: "8000"
|
||||
server_cmd: >
|
||||
vllm serve Qwen/Qwen3-235B-A22B
|
||||
--headless
|
||||
--port $SERVER_PORT
|
||||
--data-parallel-size 2
|
||||
--tensor-parallel-size 8
|
||||
--data-parallel-start-rank 1
|
||||
--data-parallel-address $MASTER_IP
|
||||
benchmarks: ...
|
||||
```
|
||||
|
||||
`server_cmd` must be a complete command starting with `vllm serve <model>`.
|
||||
It can be written as a shell string or a flat token list.
|
||||
|
||||
Write the doc block like this:
|
||||
|
||||
````md
|
||||
```{model-code}
|
||||
:block_name: qwen3_235b_worker_1
|
||||
:converter_tag: multi_node
|
||||
:test_case_path: tests/e2e/nightly/multi_node/internal_dp/config/your_model.yaml
|
||||
:host_index: 1
|
||||
|
||||
set -eux
|
||||
{{ generated }}
|
||||
```
|
||||
````
|
||||
|
||||
Generated shell script for `host_index: 1`:
|
||||
|
||||
```bash
|
||||
set -eux
|
||||
export MASTER_IP=192.168.1.10
|
||||
export SERVER_PORT=8000
|
||||
|
||||
vllm serve Qwen/Qwen3-235B-A22B \
|
||||
--headless \
|
||||
--port $SERVER_PORT \
|
||||
--data-parallel-size 2 \
|
||||
--tensor-parallel-size 8 \
|
||||
--data-parallel-start-rank 1 \
|
||||
--data-parallel-address $MASTER_IP
|
||||
```
|
||||
|
||||
#### converter_tag: `external_dp_template`
|
||||
|
||||
`external_dp_template` reads one item from `templates`. The required
|
||||
`host_index` metadata selects which template to render. The top-level `model`
|
||||
field is also required because the converter builds `vllm serve <model>`.
|
||||
|
||||
```yaml
|
||||
model: Eco-Tech/GLM-Test
|
||||
templates:
|
||||
- node_index: 0
|
||||
envs:
|
||||
HCCL_BUFFSIZE: "1024"
|
||||
ASCEND_RT_VISIBLE_DEVICES: "${VISIBLE_DEVICES}"
|
||||
server_cmd_template:
|
||||
- --host
|
||||
- 0.0.0.0
|
||||
- --port
|
||||
- ${PORT}
|
||||
- --data-parallel-size
|
||||
- ${DP_SIZE}
|
||||
- --data-parallel-rank
|
||||
- ${DP_RANK}
|
||||
- --data-parallel-address
|
||||
- ${DP_ADDRESS}
|
||||
- --data-parallel-rpc-port
|
||||
- ${DP_RPC_PORT}
|
||||
- --tensor-parallel-size
|
||||
- ${TP_SIZE}
|
||||
- --trust-remote-code
|
||||
config: ...
|
||||
routing: ...
|
||||
```
|
||||
|
||||
Known braced template variables are rewritten to the positional shell arguments
|
||||
that `run_dp_template.sh` receives from `launch_online_dp.py`:
|
||||
|
||||
| Template variable | Rendered positional |
|
||||
| --- | --- |
|
||||
| `${VISIBLE_DEVICES}` | `$1` |
|
||||
| `${PORT}` | `$2` |
|
||||
| `${DP_SIZE}` | `$3` |
|
||||
| `${DP_RANK}` | `$4` |
|
||||
| `${DP_ADDRESS}` | `$5` |
|
||||
| `${DP_RPC_PORT}` | `$6` |
|
||||
| `${TP_SIZE}` | `$7` |
|
||||
|
||||
Unknown braced variables and unbraced shell references such as `$SERVER_PORT`
|
||||
are left unchanged.
|
||||
|
||||
Write the doc block like this:
|
||||
|
||||
````md
|
||||
```{model-code}
|
||||
:block_name: glm_external_dp_template_node0
|
||||
:converter_tag: external_dp_template
|
||||
:test_case_path: tests/e2e/nightly/multi_node/external_dp/config/your_model.yaml
|
||||
:host_index: 0
|
||||
|
||||
set -eux
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
|
||||
{{ generated }}
|
||||
```
|
||||
````
|
||||
|
||||
Generated shell script for `host_index: 0`:
|
||||
|
||||
```bash
|
||||
set -eux
|
||||
export HCCL_IF_IP=$local_ip
|
||||
export GLOO_SOCKET_IFNAME=$nic_name
|
||||
export TP_SOCKET_IFNAME=$nic_name
|
||||
export HCCL_SOCKET_IFNAME=$nic_name
|
||||
|
||||
export HCCL_BUFFSIZE=1024
|
||||
export ASCEND_RT_VISIBLE_DEVICES=$1
|
||||
|
||||
vllm serve Eco-Tech/GLM-Test \
|
||||
--host 0.0.0.0 \
|
||||
--port $2 \
|
||||
--data-parallel-size $3 \
|
||||
--data-parallel-rank $4 \
|
||||
--data-parallel-address $5 \
|
||||
--data-parallel-rpc-port $6 \
|
||||
--tensor-parallel-size $7 \
|
||||
--trust-remote-code
|
||||
```
|
||||
|
||||
#### converter_tag: `external_dp_launch`
|
||||
|
||||
`external_dp_launch` reads the full `config` list and renders one
|
||||
`launch_online_dp.py` command per node. It does not take an index option.
|
||||
|
||||
```yaml
|
||||
config:
|
||||
- node_index: 0
|
||||
port_start: 7100
|
||||
dp_rpc_port: 12321
|
||||
dp_size: 2
|
||||
dp_size_local: 2
|
||||
dp_rank_start: 0
|
||||
tp_size: 8
|
||||
dp_address: "${NODE_0_IP}"
|
||||
- node_index: 1
|
||||
port_start: 7200
|
||||
dp_rpc_port: 12321
|
||||
dp_size: 4
|
||||
dp_size_local: 4
|
||||
dp_rank_start: 0
|
||||
tp_size: 4
|
||||
dp_address: "${NODE_1_IP}"
|
||||
templates: ...
|
||||
routing: ...
|
||||
```
|
||||
|
||||
Write the doc block like this:
|
||||
|
||||
````md
|
||||
```{model-code}
|
||||
:block_name: glm_external_dp_launch
|
||||
:converter_tag: external_dp_launch
|
||||
:test_case_path: tests/e2e/nightly/multi_node/external_dp/config/your_model.yaml
|
||||
|
||||
set -eux
|
||||
{{ generated }}
|
||||
```
|
||||
````
|
||||
|
||||
Generated shell script:
|
||||
|
||||
```bash
|
||||
set -eux
|
||||
python launch_online_dp.py --dp-size 2 --tp-size 8 --dp-size-local 2 --dp-rank-start 0 --dp-address ${NODE_0_IP} --dp-rpc-port 12321 --vllm-start-port 7100
|
||||
|
||||
python launch_online_dp.py --dp-size 4 --tp-size 4 --dp-size-local 4 --dp-rank-start 0 --dp-address ${NODE_1_IP} --dp-rpc-port 12321 --vllm-start-port 7200
|
||||
```
|
||||
|
||||
#### converter_tag: `external_dp_proxy`
|
||||
|
||||
`external_dp_proxy` reads `config` and `routing`. It renders the
|
||||
`load_balance_proxy_server_example.py` command for `routing.type:
|
||||
disaggregated_prefill`. It does not take an index option.
|
||||
|
||||
```yaml
|
||||
routing:
|
||||
type: disaggregated_prefill
|
||||
groups:
|
||||
prefiller: [0]
|
||||
decoder: [1]
|
||||
config:
|
||||
- node_index: 0
|
||||
port_start: 7100
|
||||
dp_size_local: 2
|
||||
dp_rpc_port: 12321
|
||||
dp_size: 2
|
||||
dp_rank_start: 0
|
||||
tp_size: 8
|
||||
dp_address: "${NODE_0_IP}"
|
||||
- node_index: 1
|
||||
port_start: 7200
|
||||
dp_size_local: 4
|
||||
dp_rpc_port: 12321
|
||||
dp_size: 4
|
||||
dp_rank_start: 0
|
||||
tp_size: 4
|
||||
dp_address: "${NODE_1_IP}"
|
||||
templates: ...
|
||||
```
|
||||
|
||||
`routing.groups.prefiller` and `routing.groups.decoder` contain indices into
|
||||
`config`. Each referenced node expands to `dp_size_local` host and port entries.
|
||||
The proxy itself is rendered on `${NODE_0_IP}:1999`.
|
||||
|
||||
Write the doc block like this:
|
||||
|
||||
````md
|
||||
```{model-code}
|
||||
:block_name: glm_external_dp_proxy
|
||||
:converter_tag: external_dp_proxy
|
||||
:test_case_path: tests/e2e/nightly/multi_node/external_dp/config/your_model.yaml
|
||||
|
||||
set -eux
|
||||
{{ generated }}
|
||||
```
|
||||
````
|
||||
|
||||
Generated shell script:
|
||||
|
||||
```bash
|
||||
set -eux
|
||||
python load_balance_proxy_server_example.py \
|
||||
--host ${NODE_0_IP} \
|
||||
--port 1999 \
|
||||
--prefiller-hosts \
|
||||
${NODE_0_IP} \
|
||||
${NODE_0_IP} \
|
||||
--prefiller-ports \
|
||||
7100 \
|
||||
7101 \
|
||||
--decoder-hosts \
|
||||
${NODE_1_IP} \
|
||||
${NODE_1_IP} \
|
||||
${NODE_1_IP} \
|
||||
${NODE_1_IP} \
|
||||
--decoder-ports \
|
||||
7200 \
|
||||
7201 \
|
||||
7202 \
|
||||
7203
|
||||
```
|
||||
|
||||
### Local debugging and generation
|
||||
|
||||
#### Generate only (without building the full site)
|
||||
|
||||
```bash
|
||||
# Generate all model-code artifacts under docs/source/tutorials/models/
|
||||
python3 tools/docs_codegen/cli.py
|
||||
|
||||
# Generate artifacts for a single document
|
||||
python3 tools/docs_codegen/cli.py --doc docs/source/tutorials/models/Kimi-K2-Thinking.md
|
||||
|
||||
# Generate a single block and print it (no files written)
|
||||
python3 tools/docs_codegen/cli.py \
|
||||
--block docs/source/tutorials/models/Kimi-K2-Thinking.md::kimi_k2_thinking_single_node \
|
||||
--dry-run --stdout
|
||||
```
|
||||
|
||||
By default, artifacts are written to: `docs/_build/doc_codegen/<doc_stem>/<block_name>.sh`.
|
||||
|
||||
:::{note}
|
||||
After the script is generated, please make sure to check whether the generated content is runnable, especially key parts such as environment variables and command-line parameters.
|
||||
:::
|
||||
|
||||
#### Build the site & preview locally
|
||||
|
||||
```bash
|
||||
# Install documentation build dependencies
|
||||
python3 -m pip install -r docs/requirements-docs.txt
|
||||
|
||||
# (Optional) Clean previous builds
|
||||
make -C docs clean
|
||||
|
||||
# Build the English site
|
||||
make -C docs html
|
||||
|
||||
# (Optional) Build the Chinese site
|
||||
make -C docs intl
|
||||
|
||||
# Preview locally
|
||||
python3 -m http.server -d docs/_build/html 8000
|
||||
|
||||
# Then open in a browser:
|
||||
# http://localhost:8000
|
||||
```
|
||||
|
||||
### For developers: add a new converter
|
||||
|
||||
A converter turns one loaded YAML file plus one parsed `ModelCodeBlock` into a
|
||||
`GeneratedScript`. The current pipeline is:
|
||||
|
||||
1. `BlockScanner` parses ``model-code`` fences and accepts only options listed
|
||||
in `MODEL_CODE_OPTION_NAMES`.
|
||||
2. `YamlLoader` loads `test_case_path`.
|
||||
3. `get_converter()` looks up `block.converter_tag` from
|
||||
`build_default_converters()`.
|
||||
4. The selected converter returns `GeneratedScript(content=..., language="shell")`.
|
||||
5. `GeneratorService` replaces `{{ generated }}` in the block body, validates
|
||||
that the final script is non-empty, and writes
|
||||
`docs/_build/doc_codegen/<doc_stem>/<block_name>.sh`.
|
||||
|
||||
To add a converter:
|
||||
|
||||
1. In `tools/docs_codegen/converters.py`, add a `BaseConverter` subclass with a
|
||||
unique `name`. That name is the value authors put in `:converter_tag:`.
|
||||
2. Implement `convert(self, loaded_yaml, *, block) -> GeneratedScript`. Use
|
||||
`make_docs_codegen_error(..., block=block)` for user-facing validation
|
||||
errors so the CLI and Sphinx output include document context.
|
||||
3. Reuse helpers from `tools/docs_codegen/utils.py`, such as
|
||||
`require_mapping`, `require_mapping_list`, `require_scalar_mapping`,
|
||||
`require_indexed_mapping`, `require_node_field`, `parse_command_tokens`,
|
||||
`substitute_template_positionals`, and `render_cli_command`.
|
||||
4. Register the converter in `build_default_converters()`. If it is not
|
||||
registered, `get_converter()` will reject the new `converter_tag`.
|
||||
5. If the converter needs new directive metadata, add the option name to
|
||||
`MODEL_CODE_OPTION_NAMES` in `tools/docs_codegen/scanner.py` and to
|
||||
`ModelCodeDirective.option_spec` in
|
||||
`tools/docs_codegen/sphinx_extension.py`. Read the option with
|
||||
`block.get_option("<option_name>")`.
|
||||
6. Add or update tests in `tests/ut/tools/test_docs_codegen.py`. Cover the
|
||||
successful render path, required option validation, YAML shape validation,
|
||||
and any CLI/Sphinx scanner behavior affected by new metadata.
|
||||
7. Add a real ``model-code`` example in a model tutorial, preferably under
|
||||
`docs/source/tutorials/models/`, and point it to an existing YAML file under
|
||||
`tests/`.
|
||||
8. Validate with the CLI:
|
||||
|
||||
```bash
|
||||
python3 tools/docs_codegen/cli.py --doc <your_doc> --dry-run
|
||||
python3 tools/docs_codegen/cli.py --block <your_doc>::<block_name> --dry-run --stdout
|
||||
```
|
||||
|
||||
If a converter should render something other than shell, set
|
||||
`GeneratedScript.language` accordingly so Sphinx can highlight the generated
|
||||
literal block correctly.
|
||||
154
docs/source/developer_guide/contribution/e2e_ci_test.md
Normal file
154
docs/source/developer_guide/contribution/e2e_ci_test.md
Normal file
@@ -0,0 +1,154 @@
|
||||
# E2E CI Test
|
||||
|
||||
This document explains how to trigger specific E2E tests against your PR code via a
|
||||
comment command, without running the full E2E test suite.
|
||||
|
||||
## Background
|
||||
|
||||
The `E2E-Full` workflow ([`pr_test.yaml`](https://github.com/vllm-project/vllm-ascend/blob/main/.github/workflows/pr_test.yaml)) normally runs the complete E2E test suite
|
||||
when a PR has `ready` label. This is expensive in CI resources
|
||||
and time.
|
||||
|
||||
Authorized users can trigger only the specific test files they care about by posting a
|
||||
`/e2e` comment on the PR, then adding the `ready` label.
|
||||
|
||||
## How to Trigger
|
||||
|
||||
### 1. Post a comment
|
||||
|
||||
First, post a comment on the PR specifying which test paths to run:
|
||||
|
||||
```text
|
||||
/e2e [test-path-1] [test-path-2] ...
|
||||
```
|
||||
|
||||
- Each path must be a valid pytest path relative to the repository root.
|
||||
- Multiple paths can be listed in a single comment, separated by spaces.
|
||||
- A specific test case can be targeted using `::` notation.
|
||||
|
||||
| Comment format | Effect |
|
||||
|---|---|
|
||||
| `/e2e tests/e2e/pull_request/one_card/test_foo.py` | Run one test file on one_card |
|
||||
| `/e2e tests/e2e/pull_request/two_card/test_bar.py` | Run one test file on two_card |
|
||||
| `/e2e path1 path2 path3` | Run multiple files, routed by path pattern |
|
||||
| `/e2e tests/e2e/pull_request/one_card/test_foo.py::test_case` | Run a specific test case |
|
||||
|
||||
### 2. Add the label
|
||||
|
||||
After posting the comment, add the **`ready`** label to your PR.
|
||||
Adding the label is what actually **triggers** the workflow — at that point the workflow
|
||||
reads the existing comments to find the `/e2e` command.
|
||||
|
||||
:::{note}
|
||||
Only repository **Contributors** (Triage role) and **Maintainers** (Write role) can add
|
||||
labels. If you do not have this permission, ask a maintainer to add the label for you.
|
||||
You can find the list of maintainers and contributors by checking the
|
||||
[CODEOWNERS](https://github.com/vllm-project/vllm-ascend/blob/main/.github/CODEOWNERS)
|
||||
file.
|
||||
:::
|
||||
|
||||
:::{important}
|
||||
The comment must be posted **before** the label is added. If you add the label first,
|
||||
the workflow will find no `/e2e` comment and will not trigger any per-test runs.
|
||||
:::
|
||||
|
||||
:::{note}
|
||||
Additionally, only the **PR author** or collaborators with **write or admin** repository
|
||||
access can trigger tests via comment. The workflow validates the commenter's permission
|
||||
before proceeding.
|
||||
:::
|
||||
|
||||
### 3. Wait for results
|
||||
|
||||
GitHub Actions will trigger the `E2E-Full` workflow. Only the hardware jobs matching
|
||||
the provided test paths will run, which saves CI resources.
|
||||
|
||||
## Path Routing Rules
|
||||
|
||||
The workflow automatically routes each test path to the correct hardware runner based
|
||||
on path patterns:
|
||||
|
||||
| Path pattern | Hardware | Runner |
|
||||
|---|---|---|
|
||||
| `two_card` in path | two_card A3 NPU | `linux-aarch64-a3-2` |
|
||||
| `four_card` in path | four_card A3 NPU | `linux-aarch64-a3-4` |
|
||||
| `_310p` in filename under one/two_card | Ascend 310P x1 | `linux-aarch64-310p-*` |
|
||||
| `_310p` in filename under four_card | Ascend 310P x4 | `linux-aarch64-310p-*` |
|
||||
| All other paths | one_card A2 NPU | `linux-aarch64-a2b3-1` |
|
||||
|
||||
When paths from multiple categories are listed in a single comment, each category's
|
||||
tests run on its respective hardware in parallel.
|
||||
|
||||
## Test Path Reference
|
||||
|
||||
The `tests/e2e/pull_request/` directory is organized by hardware category:
|
||||
|
||||
```text
|
||||
tests/e2e/pull_request/
|
||||
├── one_card/ # Single card tests → A2 NPU x1 runner
|
||||
├── two_card/ # Two card tests → A3 NPU x2 runner
|
||||
├── four_card/ # Four card tests → A3 NPU x4 runner
|
||||
```
|
||||
|
||||
310P tests use `_310p` subdirectories or `_310p.py` filename suffix under the
|
||||
corresponding card directory:
|
||||
|
||||
```text
|
||||
tests/e2e/pull_request/one_card/_310p/ # 310P single card
|
||||
tests/e2e/pull_request/four_card/_310p/ # 310P four card
|
||||
```
|
||||
|
||||
## Comparison with Full E2E Suite
|
||||
|
||||
| Aspect | Full E2E suite | Per-test comment trigger |
|
||||
|---|---|---|
|
||||
| Trigger | `ready` labels | `/e2e` comment + `ready` label |
|
||||
| Scope | All E2E tests | Only specified test paths |
|
||||
| Who can trigger | Anyone who can add labels | PR author or write/admin collaborator |
|
||||
| Use case | Pre-merge validation | Iterative debugging of specific tests |
|
||||
|
||||
## Examples
|
||||
|
||||
Run a single one_card test:
|
||||
|
||||
```text
|
||||
/e2e tests/e2e/pull_request/one_card/test_offline_inference.py
|
||||
```
|
||||
|
||||
Run a two_card test:
|
||||
|
||||
```text
|
||||
/e2e tests/e2e/pull_request/two_card/test_data_parallel.py
|
||||
```
|
||||
|
||||
Run tests across multiple hardware categories in one comment:
|
||||
|
||||
```text
|
||||
/e2e tests/e2e/pull_request/one_card/test_offline_inference.py tests/e2e/pull_request/two_card/test_data_parallel.py
|
||||
```
|
||||
|
||||
Re-trigger after fixing an issue: just push a new commit. The `synchronize` event
|
||||
re-runs the workflow and picks up the existing `/e2e` comment automatically — no need
|
||||
to post a new comment.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
**The workflow did not start after I added the label.**
|
||||
|
||||
- Make sure the `/e2e` comment was posted **before** the label was added.
|
||||
If the label was added first, remove it and re-add it after posting the comment.
|
||||
- Check that the comment starts exactly with `/e2e` followed by at least one path,
|
||||
with no leading spaces or extra characters before the slash.
|
||||
- To re-trigger after fixing an issue, simply push a new commit — the workflow will
|
||||
reuse the existing `/e2e` comment automatically.
|
||||
|
||||
**Tests ran on the wrong hardware.**
|
||||
|
||||
- Check that the path includes the expected directory segment (`one_card`, `two_card`,
|
||||
`four_card`, or `_310p`). Paths that do not match any of these patterns are routed to
|
||||
the one_card runner by default.
|
||||
|
||||
**The `parse-comment` job skipped with a permission error.**
|
||||
|
||||
- Only the PR author or write/admin collaborators can use the comment trigger.
|
||||
Ask a maintainer to post the `/e2e` comment instead.
|
||||
@@ -1,16 +1,17 @@
|
||||
# Contributing
|
||||
|
||||
## Building and testing
|
||||
It's recommended to set up a local development environment to build and test
|
||||
## Building and Testing
|
||||
|
||||
It's recommended to set up a local development environment to build vllm-ascend and run tests
|
||||
before you submit a PR.
|
||||
|
||||
### Setup development environment
|
||||
### Set up a development environment
|
||||
|
||||
Theoretically, the vllm-ascend build is only supported on Linux because
|
||||
`vllm-ascend` dependency `torch_npu` only supports Linux.
|
||||
|
||||
But you can still set up dev env on Linux/Windows/macOS for linting and basic
|
||||
test as following commands:
|
||||
But you can still set up a development environment on Linux/Windows/macOS for linting and running basic
|
||||
tests.
|
||||
|
||||
#### Run lint locally
|
||||
|
||||
@@ -27,20 +28,19 @@ cd vllm-ascend
|
||||
# Install lint requirement and enable pre-commit hook
|
||||
pip install -r requirements-lint.txt
|
||||
|
||||
# Run lint (You need install pre-commits deps via proxy network at first time)
|
||||
# Run lint (You need to install pre-commits deps via proxy network at first time)
|
||||
bash format.sh
|
||||
```
|
||||
|
||||
#### Run CI locally
|
||||
|
||||
After complete "Run lint" setup, you can run CI locally:
|
||||
After completing "Run lint" setup, you can run CI (Continuous integration) locally:
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
|
||||
cd ~/vllm-project/
|
||||
|
||||
# Run CI need vLLM installed
|
||||
# Run CI needs vLLM installed
|
||||
git clone --branch |vllm_version| https://github.com/vllm-project/vllm.git
|
||||
cd vllm
|
||||
pip install -r requirements/build.txt
|
||||
@@ -51,7 +51,7 @@ cd ..
|
||||
cd vllm-ascend
|
||||
# For Linux:
|
||||
pip install -r requirements-dev.txt
|
||||
# For non Linux:
|
||||
# For non-Linux:
|
||||
cat requirements-dev.txt | grep -Ev '^#|^--|^$|^-r' | while read PACKAGE; do pip install "$PACKAGE"; done
|
||||
cat requirements.txt | grep -Ev '^#|^--|^$|^-r' | while read PACKAGE; do pip install "$PACKAGE"; done
|
||||
|
||||
@@ -68,13 +68,13 @@ git commit -sm "your commit info"
|
||||
|
||||
🎉 Congratulations! You have completed the development environment setup.
|
||||
|
||||
### Test locally
|
||||
### Testing locally
|
||||
|
||||
You can refer to [Testing](./testing.md) doc to help you setup testing environment and running tests locally.
|
||||
You can refer to [Testing](./testing.md) to set up a testing environment and running tests locally.
|
||||
|
||||
## DCO and Signed-off-by
|
||||
|
||||
When contributing changes to this project, you must agree to the DCO. Commits must include a `Signed-off-by:` header which certifies agreement with the terms of the DCO.
|
||||
When contributing changes to this project, you must agree to the DCO. Commits must include a `Signed-off-by:` header which certifies agreement with the terms of the DCO (Developer Certificate of Origin).
|
||||
|
||||
Using `-s` with `git commit` will automatically add this header.
|
||||
|
||||
@@ -88,8 +88,8 @@ Only specific types of PRs will be reviewed. The PR title is prefixed appropriat
|
||||
- `[Platform]` for new features or optimization in platform.
|
||||
- `[Worker]` for new features or optimization in worker.
|
||||
- `[Core]` for new features or optimization in the core vllm-ascend logic (such as platform, attention, communicators, model runner)
|
||||
- `[Kernel]` changes affecting compute kernels and ops.
|
||||
- `[Bugfix]` for bug fixes.
|
||||
- `[Kernel]` for changes affecting compute kernels and ops.
|
||||
- `[BugFix]` for bug fixes.
|
||||
- `[Doc]` for documentation fixes and improvements.
|
||||
- `[Test]` for tests (such as unit tests).
|
||||
- `[CI]` for build or continuous integration improvements.
|
||||
@@ -101,11 +101,15 @@ If the PR spans more than one category, please include all relevant prefixes.
|
||||
|
||||
## Others
|
||||
|
||||
You may find more information about contributing to vLLM Ascend backend plugin on [<u>docs.vllm.ai</u>](https://docs.vllm.ai/en/latest/contributing/overview.html).
|
||||
If you find any problem when contributing, you can feel free to submit a PR to improve the doc to help other developers.
|
||||
You may find more information about contributing to vLLM Ascend backend plugin on [<u>docs.vllm.ai</u>](https://docs.vllm.ai/en/latest/contributing).
|
||||
If you encounter any problems while contributing, feel free to submit a PR to improve the documentation to help other developers.
|
||||
|
||||
:::{toctree}
|
||||
:caption: Index
|
||||
:maxdepth: 1
|
||||
testing
|
||||
doc_writing
|
||||
multi_node_test
|
||||
nightly_ci_test
|
||||
e2e_ci_test
|
||||
:::
|
||||
|
||||
553
docs/source/developer_guide/contribution/multi_node_test.md
Normal file
553
docs/source/developer_guide/contribution/multi_node_test.md
Normal file
@@ -0,0 +1,553 @@
|
||||
# Multi Node Test
|
||||
|
||||
Multi-Node CI is designed to test distributed scenarios of very large models, for example, disaggregated_prefill multi DP across multi nodes and so on.
|
||||
|
||||
## How it works
|
||||
|
||||
The following picture shows the basic deployment view of the multi-node CI mechanism. It shows how the GitHub action interacts with [lws](https://lws.sigs.k8s.io/docs/overview/) (a kind of kubernetes crd resource).
|
||||
|
||||

|
||||
|
||||
From the workflow perspective, we can see how the final test script is executed. The key point is that the shared files `tests/e2e/nightly/multi_node/scripts/lws.yaml.jinja2` and `tests/e2e/nightly/multi_node/scripts/run.sh` define the cluster template and pod entry script. Each node executes different logic according to the [LWS_WORKER_INDEX](https://lws.sigs.k8s.io/docs/reference/labels-annotations-and-environment-variables/) environment variable, so that multiple nodes can form a distributed cluster to perform tasks. `run.sh` selects the pytest entrypoint from the config path: internal DP configs use `internal_dp/scripts/test_multi_node.py`, while external DP configs use `external_dp/scripts/test_external_dp.py`.
|
||||
|
||||

|
||||
|
||||
## How to contribute
|
||||
|
||||
1. Upload custom weights
|
||||
|
||||
If you need customized weights, for example, you quantized a w8a8 weight for DeepSeek-V3 and you want your weight to run on CI, uploading weights to ModelScope's [vllm-ascend](https://www.modelscope.cn/organization/vllm-ascend) organization is welcome. If you do not have permission to upload, please contact @Potabk
|
||||
|
||||
2. Add config yaml
|
||||
|
||||
For the normal internal DP multi-node flow, add the config yaml to `tests/e2e/nightly/multi_node/internal_dp/config/`, like `DeepSeek-V3.yaml`. External DP cases use the separate `tests/e2e/nightly/multi_node/external_dp/config/` directory and should pass that directory through `config_base_path` in workflow or `CONFIG_BASE_PATH` locally.
|
||||
|
||||
Suppose you have **2 nodes** running a 1P1D setup (1 Prefillers + 1 Decoder):
|
||||
|
||||
you may add a config file looks like:
|
||||
|
||||
```yaml
|
||||
test_name: "test DeepSeek-V3 disaggregated_prefill"
|
||||
# the model being tested
|
||||
model: "vllm-ascend/DeepSeek-V3-W8A8"
|
||||
# how large the cluster is
|
||||
num_nodes: 2
|
||||
npu_per_node: 16
|
||||
# All env vars you need should add it here
|
||||
env_common: &env_common
|
||||
VLLM_USE_MODELSCOPE: true
|
||||
OMP_PROC_BIND: false
|
||||
OMP_NUM_THREADS: 100
|
||||
HCCL_BUFFSIZE: 1024
|
||||
SERVER_PORT: 8080
|
||||
disaggregated_prefill:
|
||||
enabled: true
|
||||
# node index(a list) which meet all the conditions:
|
||||
# - prefiller
|
||||
# - no headless(have api server)
|
||||
prefiller_host_index: [0]
|
||||
# node index(a list) which meet all the conditions:
|
||||
# - decoder
|
||||
decoder_host_index: [1]
|
||||
|
||||
# Add each node's vllm serve cli command just like you run locally
|
||||
# Add each node's individual envs like follow
|
||||
deployment:
|
||||
- name: prefiller node # optional: just for description, not used in code
|
||||
envs:
|
||||
<<: *env_common
|
||||
VLLM_ASCEND_ENABLE_FLASHCOMM1: 1
|
||||
# Continue to add other envs if needed
|
||||
server_cmd: >
|
||||
vllm serve ...
|
||||
- name: decoder node # optional: just for description, not used in code
|
||||
envs:
|
||||
<<: *env_common
|
||||
VLLM_ASCEND_ENABLE_FLASHCOMM1: 1
|
||||
# Continue to add other envs if needed
|
||||
server_cmd: >
|
||||
vllm serve ...
|
||||
benchmarks:
|
||||
perf:
|
||||
# fill with performance test kwargs
|
||||
acc:
|
||||
# fill with accuracy test kwargs
|
||||
```
|
||||
|
||||
3. Add the case to nightly workflow
|
||||
|
||||
Currently, the multi-node test workflow is defined in `.github/workflows/schedule_nightly_test_a3.yaml`.
|
||||
|
||||
```yaml
|
||||
multi-node-tests:
|
||||
name: multi-node
|
||||
if: always() && (github.event_name == 'schedule' || github.event_name == 'workflow_dispatch')
|
||||
strategy:
|
||||
fail-fast: false
|
||||
max-parallel: 1
|
||||
matrix:
|
||||
test_config:
|
||||
- name: multi-node-deepseek-pd
|
||||
config_file_path: DeepSeek-V3.yaml
|
||||
size: 2
|
||||
- name: multi-node-qwen3-dp
|
||||
config_file_path: Qwen3-235B-A22B.yaml
|
||||
size: 2
|
||||
- name: GLM5_1-W8A8-EP-external
|
||||
config_file_path: GLM5_1-W8A8-EP-external.yaml
|
||||
config_base_path: tests/e2e/nightly/multi_node/external_dp/config/
|
||||
size: 4
|
||||
uses: ./.github/workflows/_e2e_nightly_multi_node.yaml
|
||||
with:
|
||||
soc_version: a3
|
||||
runner: linux-aarch64-a3-0
|
||||
image: 'swr.cn-southwest-2.myhuaweicloud.com/base_image/ascend-ci/vllm-ascend:nightly-a3'
|
||||
replicas: 1
|
||||
size: ${{ matrix.test_config.size }}
|
||||
config_file_path: ${{ matrix.test_config.config_file_path }}
|
||||
config_base_path: ${{ matrix.test_config.config_base_path || '' }}
|
||||
name: ${{ matrix.test_config.name }}
|
||||
secrets:
|
||||
KUBECONFIG_B64: ${{ secrets.KUBECONFIG_B64 }}
|
||||
```
|
||||
|
||||
The matrix above defines all the parameters required to add a multi-machine use
|
||||
case. The parameters worth noting are `size`, `config_file_path`, and
|
||||
`config_base_path`. `size` defines the number of nodes required for your use
|
||||
case. `config_file_path` is the yaml file name, and `config_base_path` tells the
|
||||
loader which config directory to use. For internal DP cases, use an empty
|
||||
`config_base_path` so the loader uses its default internal DP config directory.
|
||||
For external DP cases, set it to
|
||||
`tests/e2e/nightly/multi_node/external_dp/config/`.
|
||||
|
||||
## Run Multi-Node tests locally
|
||||
|
||||
### 1. Use kubernetes
|
||||
|
||||
This section assumes that you already have a [Kubernetes](https://kubernetes.io/docs/setup/) NPU cluster environment locally. Then you can easily start our test with one click.
|
||||
|
||||
- Step 1. Install LWS CRD resources
|
||||
|
||||
See <https://lws.sigs.k8s.io/docs/installation/> Which can be used as a reference
|
||||
|
||||
- Step 2. Deploy the following yaml file `lws.yaml` as needed
|
||||
|
||||
```yaml
|
||||
apiVersion: leaderworkerset.x-k8s.io/v1
|
||||
kind: LeaderWorkerSet
|
||||
metadata:
|
||||
name: test-server
|
||||
namespace: vllm-project
|
||||
spec:
|
||||
replicas: 1
|
||||
leaderWorkerTemplate:
|
||||
size: 2
|
||||
restartPolicy: None
|
||||
leaderTemplate:
|
||||
metadata:
|
||||
labels:
|
||||
role: leader
|
||||
spec:
|
||||
containers:
|
||||
- name: vllm-leader
|
||||
imagePullPolicy: Always
|
||||
image: swr.cn-southwest-2.myhuaweicloud.com/base_image/ascend-ci/vllm-ascend:nightly-a3
|
||||
env:
|
||||
- name: CONFIG_YAML_PATH
|
||||
value: DeepSeek-V3.yaml
|
||||
- name: CONFIG_BASE_PATH
|
||||
value: tests/e2e/nightly/multi_node/internal_dp/config/
|
||||
- name: WORKSPACE
|
||||
value: "/vllm-workspace"
|
||||
- name: FAIL_TAG
|
||||
value: FAIL_TAG
|
||||
command:
|
||||
- sh
|
||||
- -c
|
||||
- |
|
||||
bash /vllm-workspace/vllm-ascend/tests/e2e/nightly/multi_node/scripts/run.sh
|
||||
resources:
|
||||
limits:
|
||||
huawei.com/ascend-1980: 16
|
||||
memory: 512Gi
|
||||
ephemeral-storage: 100Gi
|
||||
requests:
|
||||
huawei.com/ascend-1980: 16
|
||||
memory: 512Gi
|
||||
ephemeral-storage: 100Gi
|
||||
cpu: 125
|
||||
ports:
|
||||
- containerPort: 8080
|
||||
# readinessProbe:
|
||||
# tcpSocket:
|
||||
# port: 8080
|
||||
# initialDelaySeconds: 15
|
||||
# periodSeconds: 10
|
||||
volumeMounts:
|
||||
- mountPath: /root/.cache
|
||||
name: shared-volume
|
||||
- mountPath: /usr/local/Ascend/driver/tools
|
||||
name: driver-tools
|
||||
- mountPath: /dev/shm
|
||||
name: dshm
|
||||
volumes:
|
||||
- name: dshm
|
||||
emptyDir:
|
||||
medium: Memory
|
||||
sizeLimit: 15Gi
|
||||
- name: shared-volume
|
||||
persistentVolumeClaim:
|
||||
claimName: nv-action-vllm-benchmarks-v2
|
||||
- name: driver-tools
|
||||
hostPath:
|
||||
path: /usr/local/Ascend/driver/tools
|
||||
workerTemplate:
|
||||
spec:
|
||||
containers:
|
||||
- name: vllm-worker
|
||||
imagePullPolicy: Always
|
||||
image: swr.cn-southwest-2.myhuaweicloud.com/base_image/ascend-ci/vllm-ascend:nightly-a3
|
||||
env:
|
||||
- name: CONFIG_YAML_PATH
|
||||
value: DeepSeek-V3.yaml
|
||||
- name: CONFIG_BASE_PATH
|
||||
value: tests/e2e/nightly/multi_node/internal_dp/config/
|
||||
- name: WORKSPACE
|
||||
value: "/vllm-workspace"
|
||||
- name: FAIL_TAG
|
||||
value: FAIL_TAG
|
||||
command:
|
||||
- sh
|
||||
- -c
|
||||
- |
|
||||
bash /vllm-workspace/vllm-ascend/tests/e2e/nightly/multi_node/scripts/run.sh
|
||||
resources:
|
||||
limits:
|
||||
huawei.com/ascend-1980: 16
|
||||
memory: 512Gi
|
||||
ephemeral-storage: 100Gi
|
||||
requests:
|
||||
huawei.com/ascend-1980: 16
|
||||
ephemeral-storage: 100Gi
|
||||
cpu: 125
|
||||
volumeMounts:
|
||||
- mountPath: /root/.cache
|
||||
name: shared-volume
|
||||
- mountPath: /usr/local/Ascend/driver/tools
|
||||
name: driver-tools
|
||||
- mountPath: /dev/shm
|
||||
name: dshm
|
||||
volumes:
|
||||
- name: dshm
|
||||
emptyDir:
|
||||
medium: Memory
|
||||
sizeLimit: 15Gi
|
||||
- name: shared-volume
|
||||
persistentVolumeClaim:
|
||||
claimName: nv-action-vllm-benchmarks-v2
|
||||
- name: driver-tools
|
||||
hostPath:
|
||||
path: /usr/local/Ascend/driver/tools
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: Service
|
||||
metadata:
|
||||
name: vllm-leader
|
||||
namespace: vllm-project
|
||||
spec:
|
||||
ports:
|
||||
- name: http
|
||||
port: 8080
|
||||
protocol: TCP
|
||||
targetPort: 8080
|
||||
selector:
|
||||
leaderworkerset.sigs.k8s.io/name: vllm
|
||||
role: leader
|
||||
type: ClusterIP
|
||||
```
|
||||
|
||||
```bash
|
||||
kubectl apply -f lws.yaml
|
||||
```
|
||||
|
||||
Verify the status of the pods:
|
||||
|
||||
```bash
|
||||
kubectl get pods -n vllm-project
|
||||
```
|
||||
|
||||
Should get an output similar to this:
|
||||
|
||||
```bash
|
||||
NAME READY STATUS RESTARTS AGE
|
||||
vllm-0 1/1 Running 0 2s
|
||||
vllm-0-1 1/1 Running 0 2s
|
||||
```
|
||||
|
||||
Verify that the distributed inference works:
|
||||
|
||||
```bash
|
||||
kubectl logs -f vllm-0 -n vllm-project
|
||||
```
|
||||
|
||||
Should get something similar to this:
|
||||
|
||||
```shell
|
||||
INFO 12-30 11:00:57 [__init__.py:43] Available plugins for group vllm.platform_plugins:
|
||||
INFO 12-30 11:00:57 [__init__.py:45] - ascend -> vllm_ascend:register
|
||||
INFO 12-30 11:00:57 [__init__.py:48] All plugins in this group will be loaded. Set `VLLM_PLUGINS` to control which plugins to load.
|
||||
INFO 12-30 11:00:57 [__init__.py:217] Platform plugin ascend is activated
|
||||
INFO 12-30 11:00:57 [importing.py:68] Triton not installed or not compatible; certain GPU-related functions will not be available.
|
||||
================================================================================================== test session starts ===================================================================================================
|
||||
platform linux -- Python 3.12.13, pytest-8.4.2, pluggy-1.6.0 -- /usr/local/python3.12.13/bin/python3
|
||||
cachedir: .pytest_cache
|
||||
rootdir: /vllm-workspace/vllm-ascend
|
||||
configfile: pyproject.toml
|
||||
plugins: cov-7.0.0, asyncio-1.3.0, mock-3.15.1, anyio-4.12.0
|
||||
asyncio: mode=Mode.STRICT, debug=False, asyncio_default_fixture_loop_scope=None, asyncio_default_test_loop_scope=function
|
||||
collected 1 item
|
||||
|
||||
tests/e2e/nightly/multi_node/internal_dp/scripts/test_multi_node.py::test_multi_node [2025-12-30 11:01:01] INFO multi_node_config.py:294: Loading config yaml: tests/e2e/nightly/multi_node/internal_dp/config/DeepSeek-V3.yaml
|
||||
[2025-12-30 11:01:01] INFO multi_node_config.py:348: Resolving cluster IPs via DNS...
|
||||
[2025-12-30 11:01:01] INFO multi_node_config.py:212: Node 0 envs: {'VLLM_USE_MODELSCOPE': 'True', 'OMP_PROC_BIND': 'False', 'OMP_NUM_THREADS': '100', 'HCCL_BUFFSIZE': '1024', 'SERVER_PORT': '8080', 'NUMEXPR_MAX_THREADS': '128', 'DISAGGREGATED_PREFILL_PROXY_SCRIPT': 'examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py', 'HCCL_IF_IP': '10.0.0.102', 'HCCL_SOCKET_IFNAME': 'eth0', 'GLOO_SOCKET_IFNAME': 'eth0', 'TP_SOCKET_IFNAME': 'eth0', 'LOCAL_IP': '10.0.0.102', 'NIC_NAME': 'eth0', 'MASTER_IP': '10.0.0.102'}
|
||||
[2025-12-30 11:01:01] INFO multi_node_config.py:159: Launching proxy: python examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py --host 10.0.0.102 --port 6000 --prefiller-hosts 10.0.0.102 --prefiller-ports 8080 --decoder-hosts 10.0.0.138 --decoder-ports 8080
|
||||
[2025-12-30 11:01:01] INFO conftest.py:107: Starting server with command: vllm serve vllm-ascend/DeepSeek-V3-W8A8 --host 0.0.0.0 --port 8080 --data-parallel-size 2 --data-parallel-size-local 2 --tensor-parallel-size 8 --seed 1024 --enforce-eager --enable-expert-parallel --max-num-seqs 16 --max-model-len 8192 --max-num-batched-tokens 8192 --quantization ascend --trust-remote-code --no-enable-prefix-caching --gpu-memory-utilization 0.9 --kv-transfer-config {"kv_connector": "MooncakeConnectorV1", "kv_role": "kv_producer", "kv_port": "30000",
|
||||
"kv_connector_extra_config": {
|
||||
"prefill": {
|
||||
"dp_size": 2,
|
||||
"tp_size": 8
|
||||
},
|
||||
"decode": {
|
||||
"dp_size": 2,
|
||||
"tp_size": 8
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### 2. Test without Kubernetes
|
||||
|
||||
The same `tests/e2e/nightly/multi_node/scripts/run.sh` entrypoint can be used
|
||||
on prepared bare-metal or container hosts. Without LWS, set the values that
|
||||
Kubernetes normally injects yourself:
|
||||
|
||||
- `cluster_hosts` in the config yaml, using IPs reachable from every node.
|
||||
- `LWS_WORKER_INDEX` on each node, starting from `0`.
|
||||
- `CONFIG_YAML_PATH` as the config file name and `CONFIG_BASE_PATH` as the
|
||||
config directory.
|
||||
|
||||
Use the host NIC IPs that can reach each other, for example addresses shown by
|
||||
`ip addr` or `ifconfig` on the active network interface. Do not use per-host
|
||||
Docker bridge addresses such as `172.17.0.1`, because each host has its own
|
||||
local bridge.
|
||||
|
||||
Local `cluster_hosts` edits should be removed before submitting a PR unless the
|
||||
hosts are part of a committed test environment.
|
||||
|
||||
#### 2.1 Internal DP local run
|
||||
|
||||
##### 2.1.1 Add cluster hosts
|
||||
|
||||
Edit the internal DP config you want to run, for example:
|
||||
|
||||
```text
|
||||
tests/e2e/nightly/multi_node/internal_dp/config/DeepSeek-V3.yaml
|
||||
```
|
||||
|
||||
Add `cluster_hosts` as a top-level field, for example near `num_nodes` and
|
||||
`npu_per_node`:
|
||||
|
||||
```yaml
|
||||
cluster_hosts:
|
||||
- "172.22.0.xxx"
|
||||
- "172.22.0.xxx"
|
||||
```
|
||||
|
||||
##### 2.1.2 Prepare the environment
|
||||
|
||||
Install vllm-ascend development dependencies on every cluster host:
|
||||
|
||||
```bash
|
||||
cd /vllm-workspace/vllm-ascend
|
||||
python3 -m pip install -r requirements-dev.txt
|
||||
```
|
||||
|
||||
Install AISBench on the first host, which is the node with
|
||||
`LWS_WORKER_INDEX=0`:
|
||||
|
||||
```bash
|
||||
export AIS_BENCH_TAG="v3.1-20260330-master"
|
||||
export AIS_BENCH_URL="https://github.com/AISBench/benchmark.git"
|
||||
export BENCHMARK_HOME=/vllm-workspace/vllm-ascend/benchmark
|
||||
|
||||
git clone -b ${AIS_BENCH_TAG} --depth 1 ${AIS_BENCH_URL} $BENCHMARK_HOME
|
||||
cd $BENCHMARK_HOME
|
||||
pip install -e . -r requirements/api.txt -r requirements/extra.txt
|
||||
```
|
||||
|
||||
If your local image already contains the model, benchmark data, Ascend runtime,
|
||||
and AISBench, you only need the run-time exports in the next step.
|
||||
|
||||
##### 2.1.3 Start each node
|
||||
|
||||
Run the script on each node separately. Start worker nodes first, then start
|
||||
node 0.
|
||||
|
||||
On node 1:
|
||||
|
||||
```bash
|
||||
export WORKSPACE=/vllm-workspace
|
||||
export IS_PR_TEST=false
|
||||
export CONFIG_YAML_PATH=DeepSeek-V3.yaml
|
||||
export CONFIG_BASE_PATH=tests/e2e/nightly/multi_node/internal_dp/config/
|
||||
export LWS_WORKER_INDEX=1
|
||||
|
||||
cd $WORKSPACE/vllm-ascend
|
||||
bash tests/e2e/nightly/multi_node/scripts/run.sh
|
||||
```
|
||||
|
||||
On node 0:
|
||||
|
||||
```bash
|
||||
export WORKSPACE=/vllm-workspace
|
||||
export IS_PR_TEST=false
|
||||
export CONFIG_YAML_PATH=DeepSeek-V3.yaml
|
||||
export CONFIG_BASE_PATH=tests/e2e/nightly/multi_node/internal_dp/config/
|
||||
export LWS_WORKER_INDEX=0
|
||||
|
||||
cd $WORKSPACE/vllm-ascend
|
||||
bash tests/e2e/nightly/multi_node/scripts/run.sh
|
||||
```
|
||||
|
||||
Internal DP logs are mainly printed to the terminal running `run.sh`. When
|
||||
`LOG_PREFIX` is set, the shared script also backs up Ascend logs to:
|
||||
|
||||
```text
|
||||
$LOG_PREFIX/node_<LWS_WORKER_INDEX>_plogs/
|
||||
```
|
||||
|
||||
#### 2.2 External DP local run
|
||||
|
||||
##### 2.2.1 Add cluster hosts
|
||||
|
||||
Edit the external DP config you want to run. For example:
|
||||
|
||||
```text
|
||||
tests/e2e/nightly/multi_node/external_dp/config/GLM5_1-W8A8-EP-external.yaml
|
||||
```
|
||||
|
||||
Add `cluster_hosts` as a top-level field, for example near `num_nodes` and
|
||||
`npu_per_node`:
|
||||
|
||||
```yaml
|
||||
cluster_hosts:
|
||||
- "172.22.0.xxx"
|
||||
- "172.22.0.xxx"
|
||||
- "172.22.0.xxx"
|
||||
- "172.22.0.xxx"
|
||||
```
|
||||
|
||||
##### 2.2.2 Prepare the environment
|
||||
|
||||
Install vllm-ascend development dependencies on every cluster host:
|
||||
|
||||
```bash
|
||||
cd /vllm-workspace/vllm-ascend
|
||||
python3 -m pip install -r requirements-dev.txt
|
||||
```
|
||||
|
||||
Install AISBench on node 0:
|
||||
|
||||
```bash
|
||||
export AIS_BENCH_TAG="v3.1-20260330-master"
|
||||
export AIS_BENCH_URL="https://github.com/AISBench/benchmark.git"
|
||||
export BENCHMARK_HOME=/vllm-workspace/vllm-ascend/benchmark
|
||||
|
||||
git clone -b ${AIS_BENCH_TAG} --depth 1 ${AIS_BENCH_URL} $BENCHMARK_HOME
|
||||
cd $BENCHMARK_HOME
|
||||
pip install -e . -r requirements/api.txt -r requirements/extra.txt
|
||||
```
|
||||
|
||||
If your local image already contains the model, benchmark data, Ascend runtime,
|
||||
and AISBench, you only need the run-time exports in the next step.
|
||||
|
||||
##### 2.2.3 Start each node
|
||||
|
||||
External DP uses the same shared `run.sh`. Set `CONFIG_BASE_PATH` to the
|
||||
external DP config directory so the script chooses
|
||||
`external_dp/scripts/test_external_dp.py`.
|
||||
|
||||
Then start non-master nodes first, and start node 0 last. The following example
|
||||
uses `GLM5_1-W8A8-EP-external.yaml`, which is a 4-node disaggregated prefill
|
||||
case.
|
||||
|
||||
On node 1, node 2, and node 3, set the matching `LWS_WORKER_INDEX`:
|
||||
|
||||
```bash
|
||||
export WORKSPACE=/vllm-workspace
|
||||
export IS_PR_TEST=false
|
||||
export CONFIG_BASE_PATH=tests/e2e/nightly/multi_node/external_dp/config/
|
||||
export CONFIG_YAML_PATH=GLM5_1-W8A8-EP-external.yaml
|
||||
export LWS_WORKER_INDEX=1 # Use 2 on node 2, and 3 on node 3.
|
||||
|
||||
cd $WORKSPACE/vllm-ascend
|
||||
bash tests/e2e/nightly/multi_node/scripts/run.sh
|
||||
```
|
||||
|
||||
On node 0:
|
||||
|
||||
```bash
|
||||
export WORKSPACE=/vllm-workspace
|
||||
export IS_PR_TEST=false
|
||||
export CONFIG_BASE_PATH=tests/e2e/nightly/multi_node/external_dp/config/
|
||||
export CONFIG_YAML_PATH=GLM5_1-W8A8-EP-external.yaml
|
||||
export LWS_WORKER_INDEX=0
|
||||
|
||||
cd $WORKSPACE/vllm-ascend
|
||||
bash tests/e2e/nightly/multi_node/scripts/run.sh
|
||||
```
|
||||
|
||||
For `GLM5_1-W8A8-EP-external.yaml`, node 0 and node 1 start prefiller ranks,
|
||||
node 2 and node 3 start decoder ranks, and node 0 also starts the proxy and
|
||||
benchmark.
|
||||
|
||||
##### 2.2.4 Read logs while the test is running
|
||||
|
||||
The terminal running `run.sh` prints pytest orchestration logs. For external DP,
|
||||
AISBench output is also printed on node 0, while rank and proxy stdout/stderr
|
||||
are written to `EXTERNAL_DP_LOG_DIR`. The default layout is:
|
||||
|
||||
```text
|
||||
/tmp/external_dp_logs/
|
||||
node-0/
|
||||
rank-0.log
|
||||
rank-1.log
|
||||
proxy.log
|
||||
node-1/
|
||||
rank-0.log
|
||||
rank-1.log
|
||||
```
|
||||
|
||||
The first line of each rank log records the exact command and environment used
|
||||
to start that rank. `proxy.log` exists only on the configured proxy node,
|
||||
usually node 0.
|
||||
|
||||
Use a separate log directory when running multiple local experiments:
|
||||
|
||||
```bash
|
||||
export EXTERNAL_DP_LOG_DIR=/tmp/external_dp_logs_pd_local
|
||||
```
|
||||
|
||||
To watch logs in real time, run these commands in another terminal on the
|
||||
corresponding node:
|
||||
|
||||
```bash
|
||||
# node 0: ranks and proxy
|
||||
tail -F /tmp/external_dp_logs/node-0/rank-0.log \
|
||||
/tmp/external_dp_logs/node-0/rank-1.log \
|
||||
/tmp/external_dp_logs/node-0/proxy.log
|
||||
|
||||
# node 1: ranks
|
||||
tail -F /tmp/external_dp_logs/node-1/rank-0.log \
|
||||
/tmp/external_dp_logs/node-1/rank-1.log
|
||||
```
|
||||
317
docs/source/developer_guide/contribution/nightly_ci_test.md
Normal file
317
docs/source/developer_guide/contribution/nightly_ci_test.md
Normal file
@@ -0,0 +1,317 @@
|
||||
# Nightly CI Test
|
||||
|
||||
This document explains how to trigger nightly hardware CI tests against your own PR code
|
||||
on Ascend NPU hardware (A2/A3), without waiting for the scheduled nightly run.
|
||||
|
||||
## Background
|
||||
|
||||
By default, nightly CI tests run on a fixed schedule using pre-built nightly images.
|
||||
Contributors can self-service trigger these tests directly against their PR changes
|
||||
by combining a GitHub label with a comment command.
|
||||
|
||||
## How to Trigger
|
||||
|
||||
### 1. Post a comment
|
||||
|
||||
Post one of the following comments in the PR to specify which tests to run.
|
||||
The comment itself triggers the workflow — no label is required.
|
||||
|
||||
| Comment | Effect |
|
||||
|---------|--------|
|
||||
| `/nightly` | Run **all** nightly tests |
|
||||
| `/nightly all` | Run **all** nightly tests (same as above) |
|
||||
| `/nightly test1 test2 ...` | Run only the **named** tests |
|
||||
|
||||
:::{note}
|
||||
Only repository **Contributors** (Triage role) and **Maintainers** (Write role) can
|
||||
trigger the `/nightly` command. If you do not have this permission, ask a maintainer
|
||||
to post the comment for you. You can find the list of maintainers and contributors in
|
||||
the project's [Governance](../../community/governance.md) page or by checking the
|
||||
[CODEOWNERS](https://github.com/vllm-project/vllm-ascend/blob/main/.github/CODEOWNERS)
|
||||
file.
|
||||
:::
|
||||
|
||||
### 2. Wait for results
|
||||
|
||||
GitHub Actions will trigger the `Nightly-A2` or `Nightly-A3` workflow. Only tests
|
||||
matching the filter will be dispatched, which saves hardware resources.
|
||||
|
||||
## Differences Between PR and Scheduled Runs
|
||||
|
||||
| | Scheduled / Manual Dispatch | PR-triggered |
|
||||
|---|----------------------------|---|
|
||||
| Trigger | Cron (daily) or `workflow_dispatch` | `/nightly` comment |
|
||||
| Code tested | Pre-built nightly image | Your PR's HEAD commit (source installed fresh) |
|
||||
| Test scope | All tests | Configurable via `/nightly <names>` |
|
||||
| vLLM + vllm-ascend | From image | Checked out and installed from source |
|
||||
| Test matrix | From main branch's matrix YAML | From PR branch's matrix YAML |
|
||||
|
||||
When a PR run is detected (`is_pr_test: true`), the workflow additionally:
|
||||
|
||||
1. Uninstalls any existing vllm packages in the container.
|
||||
2. Checks out the specific vllm version and your PR's vllm-ascend commit from source.
|
||||
3. Installs all dependencies from source.
|
||||
4. Installs the `aisbench` benchmark suite.
|
||||
|
||||
## Test Matrix Data Source
|
||||
|
||||
The set of nightly test cases (their names, runners, test paths, model configs) is
|
||||
declared in a single data file:
|
||||
|
||||
```text
|
||||
.github/workflows/configs/nightly_config.yaml
|
||||
```
|
||||
|
||||
The file is organized as `a2:` and `a3:` top-level keys (one per SoC). Under each
|
||||
SoC, tests are grouped by execution shape (single-node, multi-node, double-node,
|
||||
multi-card, accuracy) and each group holds a `test_config` (or `nightly` / `pr_only`
|
||||
for accuracy) list whose entries carry a `name` plus the fields consumed by the
|
||||
downstream reusable workflows (`os`, `tests`, `config_file_path`, `size`, etc.).
|
||||
|
||||
Both the `Nightly-A2` and `Nightly-A3` workflows dynamically read this file at run
|
||||
time — there is no hardcoded test matrix in the workflow YAMLs. The
|
||||
`/nightly <name>` slash command resolves names by walking the same file from the
|
||||
PR branch, so newly added entries can be exercised on a PR before they land on
|
||||
main.
|
||||
|
||||
## Adding a New Nightly Test Case
|
||||
|
||||
To add a new test case (no need to touch the workflow YAMLs):
|
||||
|
||||
1. Append an entry under the appropriate section in
|
||||
`.github/workflows/configs/nightly_config.yaml`. Each entry needs at least:
|
||||
- `name`: unique identifier used in `/nightly <name>` filters
|
||||
- `os` (for single-node / multi-card pytest+yaml tests) or `runner` is inferred
|
||||
- one of `tests:` (pytest directory) or `config_file_path:` (YAML-driven model config)
|
||||
- `size` (multi-node / double-node only)
|
||||
2. Add the actual test files (pytest modules under `tests/e2e/nightly/...` or
|
||||
YAML model configs in `tests/e2e/nightly/.../configs/`).
|
||||
3. Open a PR. Once CI is green, you can validate the new entry against real NPU
|
||||
hardware **without** merging the PR — see *Examples* below.
|
||||
|
||||
## Available Test Names
|
||||
|
||||
The test names you can pass to `/nightly` correspond to the `name` fields under
|
||||
the matching section in `.github/workflows/configs/nightly_config.yaml`. The
|
||||
tables below mirror the current contents of that file.
|
||||
|
||||
### A2 workflow (`.github/workflows/schedule_nightly_test_a2.yaml`)
|
||||
|
||||
**Single-node tests** (`a2.single_node.test_config`):
|
||||
|
||||
| Test name | Description |
|
||||
|-----------|-------------|
|
||||
| `test_custom_op_multi_card` | Custom operator tests (multi card) |
|
||||
| `qwen3-vl-32b-instruct-w8a8` | Qwen3-VL-32B-Instruct W8A8 |
|
||||
| `qwen3-32b-int8` | Qwen3-32B INT8 quantization |
|
||||
| `Qwen3.5-27B-w8a8-A2` | Qwen3.5-27B W8A8 |
|
||||
| `Qwen3.5-397B-A17B-w4a8-mtp` | Qwen3.5-397B-A17B W4A8 + MTP |
|
||||
|
||||
**Multi-node tests** (`a2.multi_node.test_config`):
|
||||
|
||||
| Test name | Description |
|
||||
|-----------|-------------|
|
||||
| `multi-node-qwen3-235b-dp` | Qwen3-235B-A22B, 2-node DP |
|
||||
| `multi-node-GLM-5.1-w8a8-A2` | GLM-5.1 W8A8, 2 nodes |
|
||||
| `multi-node-Kimi-K2.5-W4A8-A2` | Kimi-K2.5 W4A8, 2 nodes |
|
||||
|
||||
**Accuracy tests** (`a2.accuracy.nightly` and `a2.accuracy.pr_only`):
|
||||
|
||||
| Test name | Description | Scope |
|
||||
|-----------|-------------|-------|
|
||||
| `accuracy-group-1` | Qwen3-VL-8B, Qwen3-8B, Qwen2-Audio-7B, etc. | nightly |
|
||||
| `accuracy-group-2` | ERNIE-4.5, Molmo-7B, Llama-3.2-3B, etc. | nightly |
|
||||
| `accuracy-group-3` | Qwen3-30B-A3B, Qwen3-VL-30B-A3B, etc. | nightly |
|
||||
| `accuracy-group-4` | Qwen3-Next-80B-A3B, Qwen3-Omni-30B-A3B, etc. | nightly |
|
||||
| `pr-accuracy-group-1` | gemma-3-4b-it, internlm3-8b-instruct, etc. | pr_only |
|
||||
| `pr-accuracy-group-2` | Qwen2.5-Math-RM-72B, Hunyuan-A13B-Instruct | pr_only |
|
||||
|
||||
The `pr-accuracy-group-*` entries only run on `/nightly` (PR-triggered) runs;
|
||||
`/nightly all` on the schedule skips them.
|
||||
|
||||
### A3 workflow (`.github/workflows/schedule_nightly_test_a3.yaml`)
|
||||
|
||||
**Multi-node tests** (`a3.multi_node.test_config`, 4-node):
|
||||
|
||||
| Test name | Description |
|
||||
|-----------|-------------|
|
||||
| `multi-node-deepseek-v3.2-W8A8-EP` | DeepSeek-V3.2-W8A8 with EP, 4-node |
|
||||
|
||||
**Double-node tests** (`a3.double_node.test_config`, 2-node, run after multi-node):
|
||||
|
||||
| Test name | Description |
|
||||
|-----------|-------------|
|
||||
| `multi-node-deepseek-r1-w8a8-longseq` | DeepSeek-R1-W8A8 long sequence, 2-node |
|
||||
| `multi-node-qwen3-dp` | Qwen3-235B-A22B, 2-node DP |
|
||||
| `multi-node-qwenw8a8-2node-eplb` | Qwen3-235B-W8A8 with EPLB, 2-node |
|
||||
| `multi-node-dpsk3.2-2node` | DeepSeek-V3.2-W8A8, 2-node |
|
||||
| `multi-node-qwenw8a8-2node-longseq` | Qwen3-235B-W8A8 long sequence, 2-node |
|
||||
| `multi-node-qwen-disagg-pd` | Qwen3-235B disaggregated PD, 2-node |
|
||||
| `multi-node-qwen-vl-disagg-pd` | Qwen3-VL-235B disaggregated PD, 2-node |
|
||||
| `multi-node-deepseek-v3.1` | DeepSeek-V3.1-BF16, 2-node |
|
||||
| `multi-node-deepseek-v3.2-W8A8-EP` | DeepSeek-V3.2-W8A8 with EP, 4-node |
|
||||
| `multi-node-glm-5.2` | GLM-5.1-W8A8, 2-node |
|
||||
|
||||
**Single-node tests** (`a3.single_node.test_config`):
|
||||
|
||||
| Test name | Description |
|
||||
|-----------|-------------|
|
||||
| `mtpx-deepseek-r1-0528-w8a8` | MTP-X + DeepSeek-R1-0528-W8A8 |
|
||||
| `deepseek-r1-0528-w8a8` | DeepSeek-R1-0528-W8A8 |
|
||||
| `kimi-k2-thinking` | Kimi-K2-Thinking |
|
||||
| `qwen3-vl-235b-a22b-instruct-w8a8` | Qwen3-VL-235B-A22B-Instruct-W8A8 |
|
||||
| `deepseek-r1-0528-w8a8-prefix-cache` | DeepSeek-R1-0528-W8A8 prefix cache |
|
||||
| `deepseek-v3-2-w8a8` | DeepSeek-V3.2-W8A8 |
|
||||
| `glm-4.7-w8a8` | GLM-4.7 W8A8 |
|
||||
| `kimi-k2.5` | Kimi-K2.5 |
|
||||
| `qwen3-235b-a22b-w8a8` | Qwen3-235B-A22B-W8A8 |
|
||||
| `Qwen3.5-397B-A17B-w8a8-mtp` | Qwen3.5-397B-A17B W8A8 + MTP |
|
||||
| `MiniMax-M2.5-w8a8-QuaRot-A3` | MiniMax-M2.5 W8A8 + QuaRot |
|
||||
| `Qwen3.5-27B-w8a8-A3` | Qwen3.5-27B W8A8 |
|
||||
| `Qwen3.5-122B-A10B-W8A8-A3` | Qwen3.5-122B-A10B W8A8 |
|
||||
| `DeepSeek-V4-Flash-W8A8-A3` | DeepSeek-V4-Flash W8A8 |
|
||||
|
||||
**Multi-card tests** (`a3.multi_card.test_config`):
|
||||
|
||||
| Test name | Description |
|
||||
|-----------|-------------|
|
||||
| `qwen3-30b-acc` | Qwen3-30B accuracy test |
|
||||
| `qwen3-30b-a3b-w8a8` | Qwen3-30B-A3B-W8A8 |
|
||||
| `qwen3-32b-int8` | Qwen3-32B-Int8 |
|
||||
| `qwen3-32b-int8-prefix-cache` | Qwen3-32B-Int8 prefix cache |
|
||||
| `Qwen3-30B-A3B-W4A8-llm-compressor` | Qwen3-30B-A3B W4A8 via llm-compressor |
|
||||
| `Qwen3-30B-QuaRot` | Qwen3-30B QuaRot + eagle3 |
|
||||
| `Qwen3-32B-QuaRot` | Qwen3-32B QuaRot + eagle3 |
|
||||
|
||||
:::{warning}
|
||||
The A3 resource pool has a maximum concurrency of **5×16 NPUs**. Multi-node tests
|
||||
run with `max-parallel: 2` to avoid resource exhaustion. Running `/nightly all` on
|
||||
A3 will queue a large number of jobs — prefer targeting specific test names when
|
||||
possible.
|
||||
:::
|
||||
|
||||
## Examples
|
||||
|
||||
Run all available nightly tests against your PR:
|
||||
|
||||
```text
|
||||
/nightly
|
||||
```
|
||||
|
||||
Run only the custom operator multi-card test:
|
||||
|
||||
```text
|
||||
/nightly test_custom_op_multi_card
|
||||
```
|
||||
|
||||
Run two specific tests at once (one per SoC):
|
||||
|
||||
```text
|
||||
/nightly test_custom_op_multi_card mtpx-deepseek-r1-0528-w8a8
|
||||
```
|
||||
|
||||
Run a single accuracy group (with all of its models):
|
||||
|
||||
```text
|
||||
/nightly accuracy-group-1
|
||||
```
|
||||
|
||||
Run a single accuracy model (only that model from a group):
|
||||
|
||||
```text
|
||||
/nightly accuracy-group-1/Qwen3-8B
|
||||
```
|
||||
|
||||
Re-trigger after fixing an issue: just push a new commit. The `synchronize` event
|
||||
re-runs the workflow and picks up the existing `/nightly` comment automatically — no
|
||||
need to post a new comment.
|
||||
|
||||
## Adding a New Test Case — Worked Example
|
||||
|
||||
To add `my-new-test` to the A2 single-node section:
|
||||
|
||||
1. Edit `.github/workflows/configs/nightly_config.yaml`, append under
|
||||
`a2.single_node.test_config`:
|
||||
|
||||
```yaml
|
||||
- name: my-new-test
|
||||
os: linux-aarch64-a2b3-4
|
||||
tests: tests/e2e/nightly/single_node/ops/multicard_ops_a2/test_my_new.py
|
||||
```
|
||||
|
||||
2. Commit the new pytest file (`test_my_new.py`) in the same PR.
|
||||
|
||||
3. Trigger from the PR:
|
||||
|
||||
```text
|
||||
/nightly my-new-test
|
||||
```
|
||||
|
||||
The workflow will:
|
||||
|
||||
- `pr_nightly_command.yml` reads your PR's `nightly_config.yaml` and resolves
|
||||
`my-new-test` → dispatch A2 only.
|
||||
- `Nightly-A2` is dispatched at `main`, but `generate-a2-matrix` checks out your
|
||||
PR commit and reads the new entry from the matrix.
|
||||
- `single-node-tests` runs one matrix job for `my-new-test`, with
|
||||
`should_run=true`. The reusable workflow checks out your PR code (via
|
||||
`vllm_ascend_ref`) and runs your pytest.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
**The workflow didn't start after I posted the comment.**
|
||||
|
||||
- Check that the comment starts exactly with `/nightly` with no leading spaces or
|
||||
extra characters before the slash.
|
||||
- Confirm you have at least Triage permission on the repository; unauthorized
|
||||
users' comments are ignored.
|
||||
- To re-trigger after fixing an issue, simply push a new commit — the workflow will
|
||||
reuse the existing `/nightly` comment automatically.
|
||||
|
||||
**Only some tests ran, not the ones I expected.**
|
||||
|
||||
- Test names are case-sensitive and must match the `name` field in
|
||||
`.github/workflows/configs/nightly_config.yaml` exactly (see the tables above).
|
||||
- For a PR-triggered run, the matrix is loaded from your PR's
|
||||
`nightly_config.yaml`, not main. If a name isn't in your PR's file, it won't
|
||||
be recognized and the dispatch will be skipped.
|
||||
- Check the `parse-trigger` job output in GitHub Actions for the resolved
|
||||
`test_filter` value.
|
||||
|
||||
**The workflow ran with the scheduled image, not my PR code.**
|
||||
|
||||
- Confirm the workflow was triggered by `repository_dispatch` (slash command),
|
||||
not bare `workflow_dispatch`. The `pr_nightly_command.yml` workflow is what
|
||||
actually dispatches `schedule_nightly_test_a2.yaml` / `_a3.yaml` with
|
||||
`vllm_ascend_ref` pointing at your PR SHA.
|
||||
|
||||
**A new test I added isn't being recognized.**
|
||||
|
||||
- Confirm the entry is well-formed YAML under
|
||||
`.github/workflows/configs/nightly_config.yaml`. The `name` field is required
|
||||
and must be unique within the SoC's section.
|
||||
- The matrix is loaded from your PR branch, so make sure the file is committed
|
||||
to the same branch the `/nightly` comment was posted on.
|
||||
|
||||
**How to obtain more detailed logs to pinpoint problems for multi-node tests**
|
||||
|
||||
- For most issues, the stdout pop-up logs from GitHub actions are sufficient (this log always represents the logs from the first node).
|
||||
- If the logs from a first node are no longer sufficient to provide effective logging information, see the summary of your jobs to download log archive for the corresponding test, which includes the framework-side logs and plog information for each node, structured as follows:
|
||||
|
||||
```shell
|
||||
.
|
||||
├── node0
|
||||
│ ├── root
|
||||
│ │ └── ascend
|
||||
│ │ └── log
|
||||
│ └── var
|
||||
│ └── log
|
||||
│ └── vllm-deepseek-v3-0f233d-0_logs.txt
|
||||
└── node1
|
||||
├── root
|
||||
│ └── ascend
|
||||
│ └── log
|
||||
└── var
|
||||
└── log
|
||||
└── vllm-deepseek-v3-0f233d-0-1_logs.txt
|
||||
```
|
||||
@@ -1,10 +1,10 @@
|
||||
# Testing
|
||||
|
||||
This secition explains how to write e2e tests and unit tests to verify the implementation of your feature.
|
||||
This document explains how to write unit tests, E2E tests, and nightly tests to verify your feature implementation.
|
||||
|
||||
## Setup test environment
|
||||
## Set up a test environment
|
||||
|
||||
The fastest way to setup test environment is to use the main branch container image:
|
||||
The fastest way to set up a test environment is to use the main branch's container image:
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: e2e
|
||||
@@ -13,7 +13,7 @@ The fastest way to setup test environment is to use the main branch container im
|
||||
:selected:
|
||||
:sync: cpu
|
||||
|
||||
You can run the unit tests on CPU with the following steps:
|
||||
You can run the unit tests on CPUs with the following steps:
|
||||
|
||||
```{code-block} bash
|
||||
:substitutions:
|
||||
@@ -22,39 +22,50 @@ cd ~/vllm-project/
|
||||
# ls
|
||||
# vllm vllm-ascend
|
||||
|
||||
# Use mirror to speedup download
|
||||
# docker pull quay.nju.edu.cn/ascend/cann:|cann_image_tag|
|
||||
# Use mirror to speed up download
|
||||
# docker pull m.daocloud.io/quay.io/ascend/cann:|cann_image_tag|
|
||||
export IMAGE=quay.io/ascend/cann:|cann_image_tag|
|
||||
docker run --rm --name vllm-ascend-ut \
|
||||
-v $(pwd):/vllm-project \
|
||||
-v ~/.cache:/root/.cache \
|
||||
-ti $IMAGE bash
|
||||
|
||||
# (Optional) Configure mirror to speedup download
|
||||
# (Optional) Configure mirror to speed up download
|
||||
sed -i 's|ports.ubuntu.com|mirrors.huaweicloud.com|g' /etc/apt/sources.list
|
||||
pip config set global.index-url https://mirrors.huaweicloud.com/repository/pypi/simple/
|
||||
|
||||
# For torch-npu dev version or x86 machine
|
||||
# For TorchNPU dev version or x86 machine
|
||||
export PIP_EXTRA_INDEX_URL="https://download.pytorch.org/whl/cpu/ https://mirrors.huaweicloud.com/ascend/repos/pypi"
|
||||
|
||||
# src path
|
||||
export SRC_WORKSPACE=/vllm-workspace
|
||||
mkdir -p $SRC_WORKSPACE
|
||||
cd $SRC_WORKSPACE
|
||||
|
||||
apt-get update -y
|
||||
apt-get install -y python3-pip git vim wget net-tools gcc g++ cmake libnuma-dev curl gnupg2
|
||||
|
||||
# Install vllm
|
||||
cd /vllm-project/vllm
|
||||
VLLM_TARGET_DEVICE=empty python3 -m pip -v install .
|
||||
git clone -b |vllm_ascend_version| --depth 1 https://github.com/vllm-project/vllm-ascend.git
|
||||
git clone --depth 1 https://github.com/vllm-project/vllm.git
|
||||
|
||||
# Install vllm-ascend
|
||||
cd /vllm-project/vllm-ascend
|
||||
# [IMPORTANT] Import LD_LIBRARY_PATH to enumerate the CANN environment under CPU
|
||||
# vllm
|
||||
cd $SRC_WORKSPACE/vllm
|
||||
VLLM_TARGET_DEVICE=empty python3 -m pip install .
|
||||
python3 -m pip uninstall -y triton
|
||||
|
||||
# vllm-ascend
|
||||
cd $SRC_WORKSPACE/vllm-ascend
|
||||
export LD_LIBRARY_PATH=$LD_LIBRARY_PATH:/usr/local/Ascend/ascend-toolkit/latest/$(uname -m)-linux/devlib
|
||||
# For cpu environment, set SOC_VERSION for different chips.
|
||||
# See https://github.com/vllm-project/vllm-ascend/blob/3cb0af0bcf3299089ca7e72159fa36e825a470f8/setup.py#L132 for detail.
|
||||
export SOC_VERSION="ascend910b1"
|
||||
python3 -m pip install .
|
||||
python3 -m pip install -r requirements-dev.txt
|
||||
python3 -m pip install -v .
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Single card
|
||||
::::{tab-item} Single-card
|
||||
:sync: single
|
||||
|
||||
```{code-block} bash
|
||||
@@ -66,6 +77,7 @@ export DEVICE=/dev/davinci0
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:main
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--device $DEVICE \
|
||||
--device /dev/davinci_manager \
|
||||
--device /dev/devmm_svm \
|
||||
@@ -86,13 +98,16 @@ After starting the container, you should install the required packages:
|
||||
# Prepare
|
||||
pip config set global.index-url https://mirrors.tuna.tsinghua.edu.cn/pypi/web/simple
|
||||
|
||||
# Switch to the /vllm-workspace/vllm-ascend directory
|
||||
cd /vllm-workspace/vllm-ascend/
|
||||
|
||||
# Install required packages
|
||||
pip install -r requirements-dev.txt
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Multi cards
|
||||
::::{tab-item} Multi-cards
|
||||
:sync: multi
|
||||
|
||||
```{code-block} bash
|
||||
@@ -101,6 +116,7 @@ pip install -r requirements-dev.txt
|
||||
export IMAGE=quay.io/ascend/vllm-ascend:main
|
||||
docker run --rm \
|
||||
--name vllm-ascend \
|
||||
--shm-size=1g \
|
||||
--device /dev/davinci0 \
|
||||
--device /dev/davinci1 \
|
||||
--device /dev/davinci2 \
|
||||
@@ -136,13 +152,13 @@ pip install -r requirements-dev.txt
|
||||
|
||||
## Running tests
|
||||
|
||||
### Unit test
|
||||
### Unit tests
|
||||
|
||||
There are several principles to follow when writing unit tests:
|
||||
|
||||
- The test file path should be consistent with source file and start with `test_` prefix, such as: `vllm_ascend/worker/worker_v1.py` --> `tests/ut/worker/test_worker_v1.py`
|
||||
- The vLLM Ascend test are using unittest framework, see [here](https://docs.python.org/3/library/unittest.html#module-unittest) to understand how to write unit tests.
|
||||
- All unit tests can be run on CPU, so you must mock the device-related function to host.
|
||||
- The test file path should be consistent with the source file and start with the `test_` prefix, such as: `vllm_ascend/worker/worker.py` --> `tests/ut/worker/test_worker.py`
|
||||
- The vLLM Ascend test uses unittest framework. See [the Python unittest documentation](https://docs.python.org/3/library/unittest.html#module-unittest) to understand how to write unit tests.
|
||||
- All unit tests can be run on CPUs, so you must mock the device-related functions on the host.
|
||||
- Example: [tests/ut/test_ascend_config.py](https://github.com/vllm-project/vllm-ascend/blob/main/tests/ut/test_ascend_config.py).
|
||||
- You can run the unit tests using `pytest`:
|
||||
|
||||
@@ -161,12 +177,12 @@ TORCH_DEVICE_BACKEND_AUTOLOAD=0 pytest -sv tests/ut
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Single card
|
||||
::::{tab-item} Single-card
|
||||
:sync: single
|
||||
|
||||
```bash
|
||||
cd /vllm-workspace/vllm-ascend/
|
||||
# Run all single card the tests
|
||||
# Run all single-card tests
|
||||
pytest -sv tests/ut
|
||||
|
||||
# Run single test
|
||||
@@ -175,12 +191,12 @@ pytest -sv tests/ut/test_ascend_config.py
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Multi cards test
|
||||
::::{tab-item} Multi-card
|
||||
:sync: multi
|
||||
|
||||
```bash
|
||||
cd /vllm-workspace/vllm-ascend/
|
||||
# Run all single card the tests
|
||||
# Run all multi-card tests
|
||||
pytest -sv tests/ut
|
||||
|
||||
# Run single test
|
||||
@@ -193,8 +209,63 @@ pytest -sv tests/ut/test_ascend_config.py
|
||||
|
||||
### E2E test
|
||||
|
||||
Although vllm-ascend CI provide [e2e test](https://github.com/vllm-project/vllm-ascend/blob/main/.github/workflows/vllm_ascend_test.yaml) on Ascend CI, you can run it
|
||||
locally.
|
||||
Although vllm-ascend CI provides E2E tests on Ascend CI (for example,
|
||||
[schedule_nightly_test_a2.yaml](https://github.com/vllm-project/vllm-ascend/blob/main/.github/workflows/schedule_nightly_test_a2.yaml), [schedule_nightly_test_a3.yaml](https://github.com/vllm-project/vllm-ascend/blob/main/.github/workflows/schedule_nightly_test_a3.yaml), [pr_test.yaml](https://github.com/vllm-project/vllm-ascend/blob/main/.github/workflows/pr_test.yaml)), you can run them locally.
|
||||
|
||||
#### PR-triggered E2E test
|
||||
|
||||
You can run tests with `pytest` as well. Typical examples:
|
||||
:::::{tab-set}
|
||||
:sync-group: e2e
|
||||
|
||||
::::{tab-item} Local (CPU)
|
||||
:sync: cpu
|
||||
|
||||
You can't run the E2E test on CPUs.
|
||||
::::
|
||||
|
||||
::::{tab-item} Single-card
|
||||
:selected:
|
||||
:sync: single
|
||||
|
||||
```bash
|
||||
cd /vllm-workspace/vllm-ascend/
|
||||
# Run all single-card tests
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/pull_request/one_card/
|
||||
|
||||
# Run a certain test script
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/pull_request/one_card/test_camem.py
|
||||
|
||||
# Run a certain case in test script
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/pull_request/one_card/test_camem.py::test_end_to_end
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Multi-card
|
||||
:sync: multi
|
||||
|
||||
```bash
|
||||
cd /vllm-workspace/vllm-ascend/
|
||||
# Run all multi-card tests
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/pull_request/two_card/
|
||||
|
||||
# Run a certain test script
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/pull_request/two_card/test_qwen3_moe_eplb.py
|
||||
|
||||
# Run a certain case in test script
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/pull_request/two_card/test_qwen3_moe_eplb.py::test_qwen3_moe_w8a8_distributed_tp2_ep_dynamic_eplb
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
:::::
|
||||
|
||||
This will reproduce the E2E test behavior.
|
||||
|
||||
#### Nightly-triggered E2E test
|
||||
|
||||
You can run tests with `pytest` as well. Typical examples:
|
||||
|
||||
:::::{tab-set}
|
||||
:sync-group: e2e
|
||||
@@ -202,84 +273,102 @@ locally.
|
||||
::::{tab-item} Local (CPU)
|
||||
:sync: cpu
|
||||
|
||||
You can't run e2e test on CPU.
|
||||
You can't run the E2E test on CPUs.
|
||||
::::
|
||||
|
||||
::::{tab-item} Single card
|
||||
::::{tab-item} Single-card
|
||||
:selected:
|
||||
:sync: single
|
||||
|
||||
```bash
|
||||
cd /vllm-workspace/vllm-ascend/
|
||||
# Run all single card the tests
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/singlecard/
|
||||
|
||||
# Run a certain test script
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/singlecard/test_offline_inference.py
|
||||
|
||||
# Run a certain case in test script
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/singlecard/test_offline_inference.py::test_models
|
||||
# run all single-card op tests
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/nightly/single_node/ops/singlecard_ops/
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
::::{tab-item} Multi cards test
|
||||
::::{tab-item} Multi-card
|
||||
:sync: multi
|
||||
|
||||
```bash
|
||||
cd /vllm-workspace/vllm-ascend/
|
||||
# Run all single card the tests
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/multicard/
|
||||
# run all multi-card op tests on A2
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/nightly/single_node/ops/multicard_ops_a2/
|
||||
|
||||
# Run a certain test script
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/multicard/test_dynamic_npugraph_batchsize.py
|
||||
|
||||
# Run a certain case in test script
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/multicard/test_offline_inference.py::test_models
|
||||
# run all multi-card op tests on A3
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/nightly/single_node/ops/multicard_ops_a3/
|
||||
```
|
||||
|
||||
::::
|
||||
|
||||
:::::
|
||||
|
||||
This will reproduce e2e test: [vllm_ascend_test.yaml](https://github.com/vllm-project/vllm-ascend/blob/main/.github/workflows/vllm_ascend_test.yaml).
|
||||
For running nightly single-node model test cases locally, refer to the following example.
|
||||
|
||||
#### E2E test example:
|
||||
```bash
|
||||
export CONFIG_YAML_PATH=Qwen3-32B.yaml
|
||||
VLLM_USE_MODELSCOPE=true pytest -sv tests/e2e/nightly/single_node/models/scripts/test_single_node.py
|
||||
```
|
||||
|
||||
- Offline test example: [`tests/e2e/singlecard/test_offline_inference.py`](https://github.com/vllm-project/vllm-ascend/blob/main/tests/e2e/singlecard/test_offline_inference.py)
|
||||
- Online test examples: [`tests/e2e/singlecard/test_prompt_embedding.py`](https://github.com/vllm-project/vllm-ascend/blob/main/tests/e2e/singlecard/test_prompt_embedding.py)
|
||||
- Correctness test example: [`tests/e2e/singlecard/test_aclgraph.py`](https://github.com/vllm-project/vllm-ascend/blob/main/tests/e2e/singlecard/test_aclgraph.py)
|
||||
- Reduced Layer model test example: [test_torchair_graph_mode.py - DeepSeek-V3-Pruning](https://github.com/vllm-project/vllm-ascend/blob/20767a043cccb3764214930d4695e53941de87ec/tests/e2e/multicard/test_torchair_graph_mode.py#L48)
|
||||
For running nightly multi-node model test cases locally, refer to the `Running Locally` section in [Multi Node Test](./multi_node_test.md).
|
||||
|
||||
The CI resource is limited, you might need to reduce layer number of the model, below is an example of how to generate a reduced layer model:
|
||||
1. Fork the original model repo in modelscope, we need all the files in the repo except for weights.
|
||||
2. Set `num_hidden_layers` to the expected number of layers, e.g., `{"num_hidden_layers": 2,}`
|
||||
3. Copy the following python script as `generate_random_weight.py`. Set the relevant parameters `MODEL_LOCAL_PATH`, `DIST_DTYPE` and `DIST_MODEL_PATH` as needed:
|
||||
#### E2E test examples
|
||||
|
||||
```python
|
||||
import torch
|
||||
from transformers import AutoTokenizer, AutoConfig
|
||||
from modeling_deepseek import DeepseekV3ForCausalLM
|
||||
from modelscope import snapshot_download
|
||||
- Offline test example: [`tests/e2e/pull_request/one_card/test_camem.py`](https://github.com/vllm-project/vllm-ascend/blob/main/tests/e2e/pull_request/one_card/test_camem.py)
|
||||
|
||||
MODEL_LOCAL_PATH = "~/.cache/modelscope/models/vllm-ascend/DeepSeek-V3-Pruning"
|
||||
DIST_DTYPE = torch.bfloat16
|
||||
DIST_MODEL_PATH = "./random_deepseek_v3_with_2_hidden_layer"
|
||||
The CI resource is limited, and you might need to reduce the number of layers of a model. Below is an example of how to generate a reduced layer model:
|
||||
|
||||
config = AutoConfig.from_pretrained(MODEL_LOCAL_PATH, trust_remote_code=True)
|
||||
model = DeepseekV3ForCausalLM(config)
|
||||
model = model.to(DIST_DTYPE)
|
||||
model.save_pretrained(DIST_MODEL_PATH)
|
||||
```
|
||||
1. Fork the original model repo in modelscope. All the files in the repo except for weights are required.
|
||||
2. Set `num_hidden_layers` to the expected number of layers, e.g., `{"num_hidden_layers": 2,}`
|
||||
3. Copy the following python script as `generate_random_weight.py`. Set the relevant parameters `MODEL_LOCAL_PATH`, `DIST_DTYPE` and `DIST_MODEL_PATH` as needed:
|
||||
|
||||
```python
|
||||
import torch
|
||||
from transformers import AutoTokenizer, AutoConfig
|
||||
from modeling_deepseek import DeepseekV3ForCausalLM
|
||||
from modelscope import snapshot_download
|
||||
|
||||
MODEL_LOCAL_PATH = "~/.cache/modelscope/models/vllm-ascend/DeepSeek-V3-Pruning"
|
||||
DIST_DTYPE = torch.bfloat16
|
||||
DIST_MODEL_PATH = "./random_deepseek_v3_with_2_hidden_layer"
|
||||
|
||||
config = AutoConfig.from_pretrained(MODEL_LOCAL_PATH, trust_remote_code=True)
|
||||
model = DeepseekV3ForCausalLM(config)
|
||||
model = model.to(DIST_DTYPE)
|
||||
model.save_pretrained(DIST_MODEL_PATH)
|
||||
```
|
||||
|
||||
### Run doctest
|
||||
|
||||
vllm-ascend provides a `vllm-ascend/tests/e2e/run_doctests.sh` command to run all doctests in the doc files.
|
||||
The doctest is a good way to make sure the docs are up to date and the examples are executable, you can run it locally as follows:
|
||||
The doctest is a good way to make sure docs stay current and examples remain executable, which can be run locally as follows:
|
||||
|
||||
```bash
|
||||
# Run doctest
|
||||
/vllm-workspace/vllm-ascend/tests/e2e/run_doctests.sh
|
||||
```
|
||||
|
||||
This will reproduce the same environment as the CI: [vllm_ascend_doctest.yaml](https://github.com/vllm-project/vllm-ascend/blob/main/.github/workflows/vllm_ascend_doctest.yaml).
|
||||
This will reproduce the same environment as the CI. See [labeled_doctest.yaml](https://github.com/vllm-project/vllm-ascend/blob/main/.github/workflows/labeled_doctest.yaml).
|
||||
|
||||
### Run docs link check
|
||||
|
||||
You can validate external links in the Sphinx docs locally with:
|
||||
|
||||
```bash
|
||||
make -C docs linkcheck SPHINXOPTS="-W --keep-going"
|
||||
```
|
||||
|
||||
To check links in a specific Markdown file, pass the file to `sphinx-build`.
|
||||
For example, to check only `docs/source/user_guide/release_notes.md`:
|
||||
|
||||
```bash
|
||||
cd docs
|
||||
sphinx-build -b linkcheck -W --keep-going \
|
||||
source _build/linkcheck source/user_guide/release_notes.md
|
||||
```
|
||||
|
||||
The detailed report will be written to:
|
||||
|
||||
- `docs/_build/linkcheck/output.txt`
|
||||
- `docs/_build/linkcheck/output.json`
|
||||
|
||||
Reference in New Issue
Block a user