xc-llm-ascend

Author	SHA1	Message	Date
SILONG ZENG	a1f321a556	[Doc]Refresh model tutorial examples and serving commands (#7426 ) ### What this PR does / why we need it? Main updates include: - update model IDs and default model paths in serving / offline inference examples - adjust some command snippets and notes for better copy-paste usability - replace `SamplingParams` argument usage from `max_completion_tokens` to `max_tokens`（Offline inference currently does not support the "max_completion_tokens"） ``` bash Traceback (most recent call last): File "/vllm-workspace/vllm-ascend/qwen-next.py", line 18, in <module> sampling_params = SamplingParams(temperature=0.6, top_p=0.95, top_k=40, max_completion_tokens=32) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ TypeError: Unexpected keyword argument 'max_completion_tokens' [ERROR] 2026-03-17-09:57:40 (PID:276, Device:-1, RankID:-1) ERR99999 UNKNOWN applicaiton exception ``` - refresh Qwen3-Omni-30B-A3B-Thinking recommended environment variable ``` bash export HCCL_BUFFSIZE=512 export HCCL_OP_EXPANSION_MODE=AIV ``` ``` bash EZ9999[PID: 25038] 2026-03-17-08:21:12.001.372 (EZ9999): HCCL_BUFFSIZE is too SMALL, maxBs = 256, h = 2048, epWorldSize = 2, localMoeExpertNum = 64, sharedExpertNum = 0, tokenNeedSizeDispatch = 4608, tokenNeedSizeCombine = 4096, k = 8, NEEDED_HCCL_BUFFSIZE(((maxBs * tokenNeedSizeDispatch * ep_worldsize * localMoeExpertNum) + (maxBs * tokenNeedSizeCombine * (k + sharedExpertNum))) * 2) = 305MB, HCCL_BUFFSIZE=200MB. [FUNC:CheckWinSize][FILE:moe_distribute_dispatch_v2_tiling.cpp][LINE:984] ``` - fix Qwen3-reranker example usage to match the current pooling runner interface and score output access ``` python model = LLM( model=model_name, task="score", # need fix hf_overrides={ "architectures": ["Qwen3ForSequenceClassification"], "classifier_from_token": ["no", "yes"], ``` ---> ``` python model = LLM( model=model_name, runner="pooling", hf_overrides={ "architectures": ["Qwen3ForSequenceClassification"], "classifier_from_token": ["no", "yes"], ``` - modify PaddleOCR-VL parameter `TASK_QUEUE_ENABLE` from `2` to `1` ``` bash (EngineCore_DP0 pid=26273) RuntimeError: NPUModelRunner init failed, error is NPUModelRunner failed, error is Do not support TASK_QUEUE_ENABLE = 2 during NPU graph capture, please export TASK_QUEUE_ENABLE=1/0. ``` These changes are needed because several documentation examples had drifted from the current runtime behavior and recommended invocation patterns, which could confuse users when following the tutorials directly. ### Does this PR introduce _any_ user-facing change? No. ### How was this patch tested? - vLLM version: v0.17.0 - vLLM main: `4497431df6` Signed-off-by: MrZ20 <2609716663@qq.com>	2026-03-20 11:34:18 +08:00
aipaes	5e65062973	[doc] Fix issues in the GLM4.7 documentation (#7457 ) ### What this PR does / why we need it? Fix issues in the GLM4.7 documentation and add some missing explanations. ### Does this PR introduce _any_ user-facing change? no ### How was this patch tested? document test - vLLM version: v0.17.0 - vLLM main: `8a680463fa` --------- Signed-off-by: zjks98 <zhangjiakang4@huawei.com> Co-authored-by: zjks98 <zhangjiakang4@huawei.com>	2026-03-19 16:42:59 +08:00
pz1116	3effc4bc70	[Doc][KV Pool]Revision KV Pool User Guide (#7434 ) ### What this PR does / why we need it? Revise the KV Pool user guide: 1. Revise Mooncake environment variables and kvconnector extra configs. 2. Delete `use_ascend_direct` in kv connector extra config as it is deprecated 3. Delete `kv_buffer_device` and `kv_rank` in P2P mooncake config 4. Unifies default `max-model-len` and `max-num-batch-tokens` in examples given. ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM version: v0.17.0 - vLLM main: `4497431df6` --------- Signed-off-by: Pz1116 <zpbzpb123123@gmail.com> Co-authored-by: Chao Lei <leichao139636@163.com>	2026-03-19 10:13:13 +08:00
SparrowMu	fb8e22ec00	[DOC] MiniMax-M2.5 model intro (#7296 ) ### What this PR does / why we need it? 1. Add nightly test on MiniMax-M2.5 with deployment method on A3 2. Add MiniMax-M2.5 deployment introduction to vllm-ascend docs - vLLM version: v0.17.0 - vLLM main: `4034c3d32e` --------- Signed-off-by: limuyuan <limuyuan3@huawei.com> Signed-off-by: SparrowMu <52023119+SparrowMu@users.noreply.github.com> Co-authored-by: limuyuan <limuyuan3@huawei.com>	2026-03-18 20:14:36 +08:00
LoganJane	565868a2a6	[doc] add doc for Kimi-K2.5.md (#7371 ) ### What this PR does / why we need it? Upload doc for Kimi-K2.5 on Ascend Base on vllm-ascend:v0.17.0rc1 - vLLM version: v0.17.0 - vLLM main: `4034c3d32e` --------- Signed-off-by: g00887675/loganJane <g00887675/loganJane73@hotmail.com> Signed-off-by: LoganJane <loganJane73@hotmail.com> Co-authored-by: g00887675/loganJane <g00887675/loganJane73@hotmail.com>	2026-03-18 17:16:35 +08:00
liuhy1213-cell	58725b8b24	[doc] add Prefill-Decode Disaggregation doc for GLM5.md (#7300 ) ### What this PR does / why we need it? add Prefill-Decode Disaggregation doc for GLM5.md w8a8 65k-1.5k Concurrency: 80 prefixcache: 90% tps: 2054 - vLLM version: v0.17.0 - vLLM main: `4034c3d32e` --------- Signed-off-by: liuhaiyang27 <liuhaiyang27@huawei.com> Co-authored-by: liuhaiyang27 <liuhaiyang27@huawei.com>	2026-03-18 17:00:31 +08:00
Nagisa125	6bc68c55d0	[doc] Refresh the documentation for DeepSeek-V3.2 (#7403 ) ### What this PR does / why we need it? Updated the DSV32 document. 1. Changed the PD separation boot mode to layerwise. 2. Changed max-num-batched-tokens to a multiple of the TP to avoid triggering a verification error. 3. Added a link to help users adjust the configuration. - vLLM version: v0.17.0 - vLLM main: `4497431df6` Signed-off-by: wyh145 <1987244901@qq.com>	2026-03-18 14:59:48 +08:00
aipaes	3b3dd2a889	[doc] Refresh the documentation for GLM-4.7 (#7292 ) ### What this PR does / why we need it? Refresh the documentation for GLM4.7. --------- Signed-off-by: zjks98 <zhangjiakang4@huawei.com> Co-authored-by: zjks98 <zhangjiakang4@huawei.com>	2026-03-17 23:09:12 +08:00
pppeng	a457d0f0e8	[doc] Upload doc for qwen3.5-27B and qwen3.5-397B-A17B on Ascend (#7313 ) ### What this PR does / why we need it? Upload doc for qwen3.5-27B and qwen3.5-397B-A17B on Ascend Base on vllm-ascend:v0.17.0rc1 - vLLM version: v0.17.0 - vLLM main: `4034c3d32e` --------- Signed-off-by: pppeng <zepengliu912@qq.com> Signed-off-by: pppeng <60355449+ppppeng@users.noreply.github.com>	2026-03-17 22:54:57 +08:00
bazingazhou233-hub	9e6c547d98	[Doc] Replace deprecated full_cuda_graph with cudagraph_mode in Qwen2.5-Omni (#7286 ) ## Summary - Replace `full_cuda_graph: 1` with `cudagraph_mode: FULL_DECODE_ONLY` in both single-NPU and multi-NPU examples - `full_cuda_graph` is deprecated and falls back to `NONE` on NPU Fixes #4696 - vLLM version: v0.17.0 - vLLM main: `4034c3d32e` Signed-off-by: bazingazhou233-hub <bazingazhou233-hub@users.noreply.github.com> Co-authored-by: bazingazhou233-hub <bazingazhou233-hub@users.noreply.github.com>	2026-03-14 22:38:36 +08:00
MengLong Chen	bbffe58b63	[Doc] fix DSV3.1 PD configs (#7187 ) ### What this PR does / why we need it? Modify the `kv_port` and `engine_id` config of DeepSeek-V3.1/R1 in the 2P1D scenario - vLLM version: v0.16.0 - vLLM main: `4034c3d32e` Signed-off-by: chenmenglong <chenmenglong1@huawei.com>	2026-03-12 14:24:49 +08:00
ZKSU	bdad11e9a8	[doc] Update GLM4.x.md, add GLM4.x multi-node deploy tutorial (#6872 ) ### What this PR does / why we need it? This PR updates the GLM4.x documentation by adding multi-node like 2 × Atlas 800 A2 (64G × 8) deployment tutorial. - What changed: Added instructions for deploying GLM-4.X models across multiple nodes, including environment variables and example commands. - Why needed: Although the previous tutorial stated that multi-node deployment on Atlas 800 A2 (64GB × 8) is not recommended, but we still face some situation that must deploy GLM-4.7 on 2 × Atlas 800 A2 (64G × 8). And we successfully run GLM-4.7 on 2 nodes and it works fine, so we think it might be the time to update this part. ### Does this PR introduce _any_ user-facing change? No. ### How was this patch tested? - Verified that the new documentation renders correctly in Markdown format. - Tested the multi-node deployment steps on 2 × Atlas 800 A2 (64G × 8) to ensure the commands work as described. - Confirmed that existing GLM4.x documentation links and structure remain intact. - vLLM version: v0.16.0 - vLLM main: `15d76f74e2` --------- Signed-off-by: ZKSU <zksu@outlook.com>	2026-03-10 10:01:53 +08:00
zyz111222	81fb7d5779	[Doc] add 310P3 guidance of PaddleOCR-VL (#6837 ) ### What this PR does / why we need it? add 310P3 guidance of PaddleOCR-VL model, refresh PaddleOCR-VL.md in the docs/source/tutorials/ ### Does this PR introduce _any_ user-facing change? no ### How was this patch tested? by CI - vLLM version: v0.15.0 - vLLM main: `83b47f67b1` --------- Signed-off-by: zouyizhou <zouyizhou@huawei.com>	2026-02-28 16:03:07 +08:00
starmountain1997	80316c5824	[DOC] enable both flashcomm1 and cudagraph (#6807 ) ## What this PR does / why we need it? This PR updates the DeepSeek-V3.2 documentation to include the latest performance optimizations and configuration improvements. ### Changes - Enable FlashComm1: Added `VLLM_ASCEND_ENABLE_FLASHCOMM1=1` environment variable across all deployment scenarios to enable FlashComm1 for improved communication performance - Layer Sharding: Added `--additional-config '{"layer_sharding": ["q_b_proj", "o_proj"]}'` configuration to enable layer sharding for better memory distribution - CUDA Graph Optimization: Updated cudagraph capture sizes from `[3,6,9,12,15,18,21,24,27,30,33,36,39,42,45,48]` to `[8, 16, 24, 32, 40, 48]` - Speculative Decoding: Increased `num_speculative_tokens` from 2 to 3 - Documentation Links: Fixed request forwarding documentation to use proper GitHub repository links ## Does this PR introduce _any_ user-facing change? Yes, users can now follow the updated documentation to enable FlashComm1 and layer sharding for improved DeepSeek-V3.2 performance. ## How was this patch tested? Existing documentation examples have been validated to ensure configuration consistency across all deployment scenarios. --- - vLLM version: v0.15.0 - vLLM main: `83b47f67b1` Signed-off-by: guozr <guozr1997@hotmail.com> Co-authored-by: guozr <guozr1997@hotmail.com>	2026-02-27 14:52:55 +08:00
wangxiyuan	a95c0b8b82	[Doc] fix the nit in docs (#6826 ) Refresh the doc, fix the nit in the docs - vLLM version: v0.15.0 - vLLM main: `83b47f67b1` Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>	2026-02-27 11:50:27 +08:00
starmountain1997	2260af405f	[DOC] add request forwarding (#6780 ) ### What this PR does / why we need it? - New section: "Request Forwarding" documentation in docs/source/tutorials/models/DeepSeek-V3.2.md - Environment fix: Changed VLLM_ASCEND_ENABLE_FLASHCOMM1 from 0 to 1 in the DeepSeek-V3 configuration examples ### Does this PR introduce _any_ user-facing change? Documentation update only - provides new configuration guidance for request forwarding setups ### How was this patch tested? - vLLM version: v0.15.0 - vLLM main: `9562912cea` --------- Signed-off-by: guozr <guozr1997@hotmail.com> Co-authored-by: guozr <guozr1997@hotmail.com>	2026-02-25 14:43:51 +08:00
yydyzr	70e26551cf	[Doc] modify glm doc (#6770 ) ### What this PR does / why we need it? 1. add description of another version of glm5-w4a8 weight 2. update the introduction of installation 3. introduce a script to enable bf16 MTP ### Does this PR introduce _any_ user-facing change? N/A ### How was this patch tested? N/A - vLLM version: v0.15.0 - vLLM main: `9562912cea` --------- Signed-off-by: yydyzr <liuyuncong1@huawei.com>	2026-02-14 16:47:23 +08:00
Cao Yi	6de207de88	[main][Docs] Fix typos across documentation (#6728 ) ## Summary Fix typos and improve grammar consistency across 50 documentation files. ### Changes include: - Spelling corrections (e.g., "Facotory" → "Factory", "certainty" → "determinism") - Grammar improvements (e.g., "multi-thread" → "multi-threaded", "re-routed" → "re-run") - Punctuation fixes (semicolon consistency in filter parameters) - Code style fixes (correct flag name `--num-prompts` instead of `--num-prompt`) - Capitalization consistency (e.g., "python" → "Python", "ascend" → "Ascend") - vLLM version: v0.15.0 - vLLM main: `9562912cea` --------- Signed-off-by: SlightwindSec <slightwindsec@gmail.com>	2026-02-13 15:50:05 +08:00
taoyao1221	41d056f947	[doc] add A2 series doc for GLM5.md (#6717 ) ### What this PR does / why we need it? Added support for A2 in the GLM-5 doc. ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? vLLM version: v0.15.0 vLLM main: `9562912cea` - vLLM version: v0.15.0 - vLLM main: `9562912cea`	2026-02-12 16:08:17 +08:00
Canlin Guo	052cc4e61b	[Docs] Fix GLM-5 deploy command (#6711 ) This pull request refines the GLM-5 deployment documentation by updating the Docker run command to include a more comprehensive set of device mappings and by removing an extraneous quantization flag from the `vllm serve` commands. These changes aim to correct and clarify the deployment instructions, ensuring users can successfully set up and run the GLM-5 model as intended. - vLLM version: v0.15.0 - vLLM main: `9562912cea` Signed-off-by: Canlin Guo <961750412@qq.com>	2026-02-12 08:55:48 +08:00
rika	b86ea66b0a	[doc]add GLM5.md (#6709 ) ### What this PR does / why we need it? Add GLM5 doc ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? - vLLM version: v0.15.0 - vLLM main: `9562912cea` Signed-off-by: nakairika <982275964@qq.com>	2026-02-12 04:00:40 +08:00
wangxiyuan	7d4833bce9	[Doc][Misc] Restructure tutorial documentation (#6501 ) ### What this PR does / why we need it? This PR refactors the tutorial documentation by restructuring it into three categories: Models, Features, and Hardware. This improves the organization and navigation of the tutorials, making it easier for users to find relevant information. - The single `tutorials/index.md` is split into three separate index files: - `docs/source/tutorials/models/index.md` - `docs/source/tutorials/features/index.md` - `docs/source/tutorials/hardwares/index.md` - Existing tutorial markdown files have been moved into their respective new subdirectories (`models/`, `features/`, `hardwares/`). - The main `index.md` has been updated to link to these new tutorial sections. This change makes the documentation structure more logical and scalable for future additions. ### Does this PR introduce _any_ user-facing change? Yes, this PR changes the structure and URLs of the tutorial documentation pages. Users following old links to tutorials will encounter broken links. It is recommended to set up redirects if the documentation framework supports them. ### How was this patch tested? These are documentation-only changes. The documentation should be built and reviewed locally to ensure all links are correct and the pages render as expected. - vLLM version: v0.15.0 - vLLM main: https://github.com/vllm-project/vllm/commit/v0.15.0 Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>	2026-02-10 15:03:35 +08:00

22 Commits