xc-llm-ascend

Author	SHA1	Message	Date
Li Wang	1165b2c863	[1/N][CI] Refactor accuracy test (#5400 ) ### What this PR does / why we need it? 1. Accuracy testing no longer compares eager and graph modes; instead, it directly extracts the golden result under the graph mode configuration (the implicit purpose of this case is to verify whether modifications affect existing results) 2. Next step: finer-grained supervision of logits/sampler results ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM version: release/v0.13.0 - vLLM main: `254f6b9867` Signed-off-by: wangli <wangli858794774@gmail.com>	2026-01-07 20:58:15 +08:00
Icey	b94fc13d3f	[BugFix][Fusion] Fix graph fusion failure problem (#5676 ) Currently, the vllm pull request (https://github.com/vllm-project/vllm/pull/24252) is causing operator fusion to fail. This issue was previously fixed by patching the backend. The root cause has been identified, and the problem can be resolved with this pull request. - vLLM version: v0.13.0 - vLLM main: `2f4e6548ef` --------- Signed-off-by: wxsIcey <1790571317@qq.com>	2026-01-07 18:42:55 +08:00
Icey	137f28341d	[Tests] Add qwen3-8b nightly test (#5597 ) ### What this PR does / why we need it? Add qwen3-8b nightly test - vLLM version: v0.13.0 - vLLM main: `7157596103` --------- Signed-off-by: wxsIcey <1790571317@qq.com>	2026-01-07 18:42:05 +08:00
Mengqing Cao	3f4f2b4ae6	[Refactor] Import global var form vllm instead of overwirte it (#5469 ) ### What this PR does / why we need it? Import global var form vllm instead of overwirte it, so that we could use the correct global variant value - vLLM version: v0.13.0 - vLLM main: `5326c89803` --------- Signed-off-by: MengqingCao <cmq0113@163.com>	2026-01-07 18:41:45 +08:00
LICO67373	380f089fbf	[Refactor] Fix AttentionMaskBuilder singleton and remove redundant pcp_prefill_mask (#4870 ) ## What this PR does / why we need it? This PR fixes the `AttentionMaskBuilder` singleton initialization issue introduced in PR #4779 and removes the unused `pcp_prefill_mask` field. ### Background After PR #4779 made `AttentionMaskBuilder` a singleton with `@singleton` decorator, the class constructor now requires a `device` parameter. However, two initialization sites were still using the old parameterless constructor, causing failures. ### Changes 1. Fix singleton initialization - Fixed `AttentionMaskBuilder()` → `AttentionMaskBuilder(self.device)` in `AscendMLAMetadataBuilder.__init__()` - Fixed `AttentionMaskBuilder()` → `AttentionMaskBuilder(self.device)` in `AscendAttentionMetadataBuilder.__init__()` 2. Remove unused field - Removed `pcp_prefill_mask` field from `AscendPrefillContextParallelMetadata` (never used in codebase) - Updated related test assertions ### Related - Issue #5463 - PR #4779 (Unify all mask generation methods) - PR #5389 (Make AttentionMaskBuilder singleton) ## Does this PR introduce _any_ user-facing change? No. This is an internal refactoring. ## How was this patch tested? - ✅ Local testing: No linter errors - ✅ Unit tests for attention modules verified - ⏳ CI pipeline Signed-off-by: lico67373 <918688502@qq.com> Co-authored-by: weijinqian0 <1184188277@qq.com>	2026-01-07 17:09:52 +08:00
wangxiyuan	91790fd85a	[CI] move image and wheel job to schedule way (#5685 ) move image and wheel job to schedule way to save CI resource - vLLM version: v0.13.0 - vLLM main: `2f4e6548ef` Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>	2026-01-07 16:40:19 +08:00
无脸男	1140789e83	[Bugfix] Fix the graph capture failure issue in the eagle3+full scenario. (#5553 ) ### What this PR does / why we need it? When launching the service in the scenario where the cudagraph_mode is set to FULL and Eagle3 acceleration is enabled for inference, an error in fia will cause graph capture to fail. This PR fixes the issue. ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM version: v0.13.0 - vLLM main: `7157596103` Signed-off-by: WithHades <244036962@qq.com>	2026-01-07 15:57:16 +08:00
weiguihua2	2b8a9ce8bd	[Bugfix] fix resource are insufficient when pcp and piecewise (#5377 ) ### What this PR does / why we need it? Resolving the issue of insufficient resources during service operation when PCP is enabled in a piecewise scenario. When enabling PCP and executing in piecewise mode, the curl request fails due to insufficient resources, resulting in the error message "The resources are insufficient." Through profiling analysis, it was found that the PCP communication domain also occupies streams and consumes resources. Therefore, when updating aclgraph sizes, the PCP communication domain needs to be taken into account. ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM version: release/v0.13.0 - vLLM main: `bc0a5a0c08` --------- Signed-off-by: weiguihua2 <weiguihua2@huawei.com>	2026-01-07 15:39:52 +08:00
Paco Xu	4f9808002b	[CI] Add workflow to cancel running workflows on PR close (#5646 ) Example: https://github.com/vllm-project/vllm-ascend/actions/runs/20735955959/job/59533181655 is still running after https://github.com/vllm-project/vllm-ascend/pull/5612 is closed. And the action will be running for more than 2 hours, which needs to be cleanup. It seems that the Github Aciton will not cancel it automatically, so I add this to cannel those PR related actions once it is closed. Tested in https://github.com/pacoxu/pacoxu/actions/runs/20743173119. - vLLM version: v0.13.0 - vLLM main: `2f4e6548ef` Signed-off-by: Paco Xu <roollingstone@gmail.com>	2026-01-07 15:38:10 +08:00
Li Wang	d314ea8d3d	[CI] Bump lm-eval version to v0.4.9.2 (#5655 ) ### What this PR does / why we need it? fix https://github.com/vllm-project/vllm-ascend/issues/2865, lm-eval [got an official update last month](https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.9.2), so let's bump the version. ### How was this patch tested? - vLLM version: v0.13.0 - vLLM main: `2f4e6548ef` Signed-off-by: wangli <wangli858794774@gmail.com>	2026-01-07 14:15:53 +08:00
wangxiyuan	6f7a81cd9f	[CI] cleanup single/multi-card test (#5623 ) 1. speed up e2e light test. 2. create `2-cards` and `4-cards` folder in multicard 3. move ops to nightly 4. run test in Alphabetical Order - vLLM version: v0.13.0 - vLLM main: `8be6432bda` Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>	2026-01-07 14:13:34 +08:00
SILONG ZENG	1afbc01ed4	[misc]Add Kimi-K2 series to CI model list (#5656 ) ### What this PR does / why we need it? Add the model to CI for subsequent addition of nightly test cases: - moonshotai/Kimi-K2-Thinking - vllm-ascend/Kimi-K2-Instruct-W8A8 ### How was this patch tested? - vLLM version: v0.13.0 - vLLM main: `2f4e6548ef` --------- Signed-off-by: MrZ20 <2609716663@qq.com> Signed-off-by: wangli <wangli858794774@gmail.com> Co-authored-by: wangli <wangli858794774@gmail.com>	2026-01-07 11:32:48 +08:00
UnifiedCacheManager	d6bb17f10e	[Bugfix]Add register_kv_cache in ucm_connector (#5657 ) ### What this PR does / why we need it? To adapt different shapes of the KV cache, UCM optimized the initialization of store by moving it into `register_kv_caches`. Therefore, this update adds `register_kv_caches` interface to UCMConnectorV1. ### How was this patch tested? - vLLM version: v0.13.0 - vLLM main: `2f4e6548ef` Signed-off-by: UnifiedCacheManager <unifiedcachem@163.com>	2026-01-07 11:30:33 +08:00
LI SHENGYONG	cd59323e40	[Bugfix] Revert pr4214 multi-stream collect expert hotpot (#5529 ) ### What this PR does / why we need it? PR4214 was intended to collect expert heat by processing multiple streams, which could lead to memory overwriting and accuracy issues. After communicating with the PR submitter, this PR has been reverted. ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? qwen3-moe dynamic eplb Befor revert \| dataset \| version \| metric \| mode \| vllm-api-general-chat \| \|----- \| ----- \| ----- \| ----- \| -----\| \| aime2024 \| 604a78 \| accuracy \| gen \| 43.33 \| After revert \| dataset \| version \| metric \| mode \| vllm-api-general-chat \| \|----- \| ----- \| ----- \| ----- \| -----\| \| aime2024 \| 604a78 \| accuracy \| gen \| 86.67 \| baseline (without eplb) \| dataset \| version \| metric \| mode \| vllm-api-general-chat \| \|----- \| ----- \| ----- \| ----- \| -----\| \| aime2024 \| 604a78 \| accuracy \| gen \| 86.67 \| - vLLM version: v0.13.0 - vLLM main: `45c1ca1ca1` Signed-off-by: shenchuxiaofugui <1311027364@qq.com>	2026-01-07 11:26:47 +08:00
wangyibo1005	25baf6df09	[Feature]EPLB:Adapt DispatchGmmCombineDecode operator to eplb tensor list and expert token numbers (#5552 ) #### What this PR does / why we need it? This PR adapt DispatchGmmCombineDecode operator to eplb tensor list and expert token numbers. This operator support gmm1, gmm2, gmm1Scale and gmm2Scale in format of list. This operator support couting how many token each local expert recieves by expertTokensNum . - vLLM version: v0.13.0 - vLLM main: `7157596103` More info about this operator, please refer to RFC: issue https://github.com/vllm-project/vllm-ascend/issues/5476	2026-01-07 11:23:42 +08:00
starmountain1997	086c093347	[CI] Add DeepSeek-V3.2-W8A8 nightly ci test (#5371 ) # What this PR does / why we need it? Add DeepSeek-V3.2-W8A8 dual-node nightly CI test and update A3 nightly test configuration: 1. Add DeepSeek-V3.2-W8A8 dual-node test: tests/e2e/nightly/multi_node/config/DeepSeek-V3_2-W8A8-A3-dual-nodes.yaml - 2 nodes, 16 NPUs per node (32 NPUs total) - Configuration: 2P+1D (data-parallel-size=4, tensor-parallel-size=8, data-parallel-size-local=2) - Includes performance and accuracy benchmarks with GSM8K dataset 2. Update A3 nightly workflow: .github/workflows/nightly_test_a3.yaml - Added DeepSeek-V3.2-W8A8 dual-node test to the A3 nightly test matrix - Test name: multi-node-dpsk3.2-2node 3. Improve test scripts: Updated .github/workflows/_e2e_nightly_multi_node.yaml and related scripts for better multi-node testing support test on A3 instances - Performance baseline: 1 (threshold: 0.97) - Accuracy baseline: 95% (threshold: 5%) - Test dataset: GSM8K with 512 prompts for performance, gsm8k-lite for accuracy --------- Signed-off-by: guozr <guozr1997@hotmail.com> Co-authored-by: guozr <guozr1997@hotmail.com>	2026-01-07 10:02:02 +08:00
Feng Liu	cbc987db0b	[bugfix (pcp)] fix chunked prefill accurancy issue (#5647 ) ### What this PR does / why we need it? Purpose: initialize padded slot mapping buffer to prevent garbage values. In PCP mode, the `pcp_padded_slot_mapping` buffer is reused across invocations. Without explicit initialization, this buffer retain stale values from previous runs, which can lead to incorrect results. This change ensures the buffer is filled with -1. ### Does this PR introduce _any_ user-facing change? No. ### How was this patch tested? - vLLM version: v0.13.0 - vLLM main: `2f4e6548ef` --------- Signed-off-by: F.Liu <liufeng248@huawei.com> Co-authored-by: F.Liu <liufeng248@huawei.com>	2026-01-07 10:01:27 +08:00
wangxiyuan	1112208052	[Refactor] Cleanup platform (#5566 ) ### What this PR does / why we need it? 1. add `COMPILATION_PASS_KEY` constant 2. clean up useless platform interface `empty_cache`, `synchronize`, `mem_get_info`, `clear_npu_memory` 3. rename `CUSTOM_OP_REGISTERED` to `_CUSTOM_OP_REGISTERED` 4. remove uesless env `VLLM_ENABLE_CUDAGRAPH_GC` NPUPlatform is the interface called by vLLM. Do not call it inner vllm-ascend. ### Does this PR introduce _any_ user-facing change? This PR is just a cleanup. All CI should pass. ### How was this patch tested? - vLLM version: v0.13.0 - vLLM main: `7157596103` Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>	2026-01-07 09:25:55 +08:00
Ronald	6ea2afe5fa	[Feature] implement basic framework for batch invariant (#5517 ) ### What this PR does / why we need it? This PR implement the basic framework for batch invariant, please see https://github.com/vllm-project/vllm-ascend/issues/5487. ### Does this PR introduce _any_ user-facing change? we reuse the function `vllm_is_batch_invariant` in vllm to judge if batch invariant is enabled. - vLLM version: v0.13.0 - vLLM main: `45c1ca1ca1` --------- Signed-off-by: Ronald1995 <ronaldautomobile@163.com> Signed-off-by: Lord_of_Ironhill <suiweiyi@huawei.com> Signed-off-by: zjchenn <zjchenn@gmail.com> Signed-off-by: wangx700 <wangxin700@huawei.com> Co-authored-by: Lord_of_Ironhill <suiweiyi@huawei.com> Co-authored-by: zjchenn <zjchenn@gmail.com> Co-authored-by: wangx700 <wangxin700@huawei.com>	2026-01-07 09:11:26 +08:00
CodeCat	bdedf3c9f8	[Graph][Fusion] Add AddRMSNormSPPattern and AddRMSNormSPPatternWithBias (#5569 ) ### What this PR does / why we need it? This PR builds upon PR https://github.com/vllm-project/vllm-ascend/pull/5011 and aims to further enhance the npu_graph_ex_passes module. Based on prior work, we have added graph optimization support for the add_rms_quant fused operator in scenarios where a bias term is present—ensuring the fusion pattern is correctly registered and matched into the computation graph. For validation, we switched to the Qwen3-235B-A22B-W8A8 model for SPPatternWithBias and Qwen3-32B model for SPPattern. Benchmark results show that, compared to the unfused baseline, enabling this fusion pass significantly improves inference throughput for W8A8 quantized models. For more details can refer to the RFC:https://github.com/vllm-project/vllm-ascend/issues/4715 ### Does this PR introduce _any_ user-facing change? no ### How was this patch tested? ``` llm = LLM( model=model, tensor_parallel_size=GPUs_per_dp_rank, enforce_eager=False, enable_expert_parallel=enable_expert_parallel, trust_remote_code=trust_remote_code, gpu_memory_utilization=0.98, max_num_batched_tokens=512, # load_format="dummy", max_model_len=2048, max_num_seqs=16, quantization="ascend", additional_config={ "refresh": True, "enable_npugraph_ex": True }, compilation_config={ "cudagraph_capture_sizes": [8, 16], "cudagraph_mode": "FULL_DECODE_ONLY", }, ) if profile_dir: llm.start_profile() outputs = llm.generate(prompts, sampling_params) if profile_dir: llm.stop_profile() for i, output in enumerate(outputs): if i >= 5: break prompt = output.prompt generated_text = output.outputs[0].text print( f"DP rank {global_dp_rank}, Prompt: {prompt!r}, " f"Generated text: {generated_text!r}" ) ``` - vLLM version: v0.13.0 - vLLM main: `7157596103` Signed-off-by: cjian <2318164299@qq.com>	2026-01-07 09:03:45 +08:00
zhenwenqi2024	ad9b711f89	[Bugfix] fix dcp_only bug and add e2e accuracy test for dcp only and pcp only (#5565 ) ### What this PR does / why we need it? [Bugfix] fix dcp_only bug and add e2e accuracy test for dcp only and pcp only this pr fix the bug of accuracy test when decode_parallel_size>1 and prefill_context_parallel_size=1. ### Does this PR introduce _any_ user-facing change? NO ### How was this patch tested? - vLLM version: v0.13.0 - vLLM main: `7157596103` --------- Signed-off-by: zhenwenqi2024 <zhenwenqi_2022@qq.com>	2026-01-06 22:48:21 +08:00
Fager10086	77a029979e	Revert "[BugFix][Fusion] Fix graph fusion failure problem (#5253 )" (#5667 ) ### What this PR does / why we need it? Revert PR 5253 to fix the smoking problem ### Does this PR introduce _any_ user-facing change? Does not. ### How was this patch tested? It was tested in the failure case. Signed-off-by: Rifa <865071616@qq.com>	2026-01-06 21:55:47 +08:00
liziyu	330e25ab1d	[P/D] Performance enhancement of Layerwise connector in TP asymmetric scenarios (#5540 ) ### What this PR does / why we need it? [P/D] Performance enhancement of Layerwise connector in TP asymmetric scenarios 1. Session fusion: For transmission tasks at each layer, aggregate transmission tasks with the same destination and merge them into a single task for assignment. 2. Alltoall aggregation: For TP asymmetric scenarios, perform all alltoall operations at once according to the block granularity for all requests. [RFC]: CDCP Scheduling for Disaggregated Prefilling with KV Cache Layerwise Push Support https://github.com/vllm-project/vllm-ascend/issues/4842 ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM version: v0.13.0 - vLLM main: `45c1ca1ca1` --------- Signed-off-by: liziyu <liziyu16@huawei.com> Signed-off-by: nwpu-zxr <zhouxuerong2@huawei.com> Signed-off-by: wangxiaoteng <wangxiaoteng@huawei.com> Co-authored-by: nwpu-zxr <zhouxuerong2@huawei.com> Co-authored-by: wangxiaoteng <wangxiaoteng@huawei.com>	2026-01-06 20:25:36 +08:00
wangxiyuan	cd1162e25a	[Misc] Remove useless weight loader patch (#5619 ) The patch for weight loader is useless now. Let's remove it - vLLM version: v0.13.0 - vLLM main: `8be6432bda` Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>	2026-01-06 20:17:32 +08:00
InSec	089ca2ddcc	[Nightly][Test] Add Qwen3-Next-80B-A3B-Instruct-W8A8 nightly test (#5616 ) ### What this PR does / why we need it? There was an accuracy issue with Qwen3-Next-80B-A3B-Instruct-W8A8 model in the old version of Triton-Ascend, so, we are now adding one nightly test to maintain it. ### Does this PR introduce _any_ user-facing change? N/A ### How was this patch tested? - vLLM version: v0.13.0 - vLLM main: `7157596103` Signed-off-by: IncSec <1790766300@qq.com>	2026-01-06 17:36:00 +08:00
yeyifan	cc0110abb4	[Bugfix] Remove swa parameter of fia (#5602 ) ### What this PR does / why we need it? When using the swa parameter in fia, headDim does not currently support 256, and when gemma3's headDim is equal to 256, an error will occur. Therefore, code rollback is required, and it will be incorporated after cann supports it. ### Does this PR introduce _any_ user-facing change? Remove swa parameter of fia. ### How was this patch tested? - vLLM version: v0.13.0 - vLLM main: `7157596103` --------- Signed-off-by: nsdie <yeyifan@huawei.com> Co-authored-by: Mengqing Cao <cmq0113@163.com>	2026-01-06 17:24:43 +08:00
Mercykid-bash	29e2f9a43e	Bugfix: Align expert map shapes with redundant experts in EPLB adjustment (#5285 ) #### Overview This PR fixes a shape mismatch bug between `expert_placement_map` and `log2phy_expert_map` when redundant experts are enabled in the vLLM-Ascend platform. The issue occurred during the initialization of expert maps and their updates via EPLB (Expert Load Balancer) adjustment, leading to potential tensor shape errors and incorrect expert routing in distributed MoE deployments. #### Key Changes 1. Unify expert map shape calculation logic - Ensure the shape of `expert_placement_map` and `log2phy_expert_map` strictly aligns with the total number of experts (including redundant experts) during initialization. - Update the shape adjustment logic in EPLB dynamic update process to match the initial expert map dimensions. 2. Add shape consistency checks - Add assertion statements to verify the shape consistency of the two maps after initialization and EPLB adjustment, preventing silent shape mismatches in subsequent operations. #### Impact - Resolves tensor shape errors when using redundant experts with EPLB on Ascend platform. - Ensures correct expert routing and load balancing for MoE models with redundant expert configurations. - No breaking changes to existing functionality; compatible with non-redundant expert deployments. - vLLM version: release/v0.13.0 - vLLM main: `ad32e3e19c` --------- Signed-off-by: Che Ruan <cr623@ic.ac.uk> Signed-off-by: shenchuxiaofugui <1311027364@qq.com> Co-authored-by: Che Ruan <cr623@ic.ac.uk> Co-authored-by: shenchuxiaofugui <1311027364@qq.com>	2026-01-06 17:22:36 +08:00
Zetong Li	fe3f2c7702	[Refactor][EAGLE] 3/N delete redundant methods in mtp_proposer (#5420 ) ### What this PR does / why we need it? This PR aims to delete redundant methods in mtp_proposer. All the deleted methods now can be found in eagle_proposer. We also remove some methods in eagle_proposer since they are identical to those in vllm-eagle. ### Does this PR introduce _any_ user-facing change? N/A ### How was this patch tested? by ci - vLLM version: release/v0.13.0 - vLLM main: `81786c8774` --------- Signed-off-by: Zetong Li <slippersss@126.com>	2026-01-06 16:47:39 +08:00
Shanshan Shen	b94d589769	[MM][Bugfix] Update `hf_config` to `hf_text_config` (#5319 ) ### What this PR does / why we need it? Following https://github.com/vllm-project/vllm-ascend/pull/5205, update `hf_config` to `hf_text_config`. Find more details at https://github.com/vllm-project/vllm-ascend/pull/5205#issuecomment-3675417534 and https://github.com/vllm-project/vllm-ascend/pull/5205#issuecomment-3677920872. ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM version: release/v0.13.0 - vLLM main: `5fbfa8d9ef` Signed-off-by: shen-shanshan <467638484@qq.com>	2026-01-06 16:41:39 +08:00
Magnus	293b2275df	[CI] Specify the version of xlite (#5612 ) ### What this PR does / why we need it? Pin the xlite version to avoid CI failures during its upgrade. - vLLM version: v0.13.0 - vLLM main: `7157596103` Signed-off-by: changdawei1 <changdawei3@huawei.com>	2026-01-06 16:02:16 +08:00
wjunLu	b8f245792e	[Main2Main] Upgrade vllm commit to 0106 (#5617 ) ### What this PR does / why we need it? Upgrade vllm commit to 0106 - vLLM version: v0.13.0 - vLLM main: `8be6432bda` Signed-off-by: wjunLu <wjunlu217@gmail.com>	2026-01-06 15:50:40 +08:00
meihanc	c1dcddce3f	[CI]update bisheng version (#5621 ) ### What this PR does / why we need it? update bisheng version in 20260105 - vLLM version: v0.13.0 - vLLM main: `8be6432bda` Signed-off-by: Meihan-chen <jcccx.cmh@gmail.com>	2026-01-06 15:22:22 +08:00
Qiu	e07938047e	[UT][PCP&DCP] UT for block_table.py (#5032 ) ## Purpose This PR add unit test for `compute_slot_mapping` function in `block_table.py` with various `pcp_size` & `dcp_size` & `cp_kv_cache_interleave_size`. ## Test Plan ``` pytest tests/ut/worker/test_block_table.py ``` ## Test Result ``` ==== 3 passed, 2 warnings in 0.20s ==== ``` - vLLM version: v0.12.0 - vLLM main: `ad32e3e19c` --------- Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>	2026-01-06 11:19:25 +08:00
wjunLu	3cf059a72b	[Main2Main] Upgrade vllm commit to 0105 (#5595 ) ### What this PR does / why we need it? Upgrade vllm commit to 0105 (8be6432bdaf6275664d857b1e5e9bf8ed1ce299e) 1. Remove `maybe_padded_num_tokens` arg in `model_runner_v1.py` since https://github.com/vllm-project/vllm/pull/31517 deleted unused arg 2. Remove dense `Qwen/Qwen3-0.6B` in `tests/e2e/multicard/test_aclgraph_capture_replay.py` and `tests/e2e/multicard/test_data_parallel.py` due to https://github.com/vllm-project/vllm/pull/30739 where offline data parallel mode will not be supported/useful for dense models 3. Adapt `vllm_ascend/worker/worker.py` due to https://github.com/vllm-project/vllm/pull/31584 4. Adapt `self.block_size` calling due to https://github.com/vllm-project/vllm/pull/31540 5. Modify `test_mla_v1.py` due to https://github.com/vllm-project/vllm/pull/28454 , which refactorred `get_head_size()` ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM version: v0.13.0 - vLLM main: `7157596103` Signed-off-by: wjunLu <wjunlu217@gmail.com>	2026-01-06 08:44:29 +08:00
Li Wang	c5e2f48510	[CI] mv ops to correct path (#5615 ) ### What this PR does / why we need it? mv ops to correct path :`tests/e2e/nightly/single_node/ops/singlecard_ops/triton` Signed-off-by: wangli <wangli858794774@gmail.com>	2026-01-05 23:17:07 +08:00
dsxsteven	129ba9fe1b	[BugFix] Fix Smoke Testing Bug for DSR1 longseq (#5613 ) ### What this PR does / why we need it? Fix Smoke Testing Bug for DSR1 longseq We need to make this change because the daily smoke test case is throwing an error: "max_tokens or max_completion_tokens is too large: 32768.This model's maximum context length is 32768 tokens and your request has 128 input tokens". We encounter this error due to max-out-len equals to max-model-len. We can fix this error by increasing max-model-len argument in the script. ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM version: v0.13.0 - vLLM main: `7157596103` Signed-off-by: daishixun <dsxsteven@sina.com>	2026-01-05 22:40:28 +08:00
ZixuanWang	8eae949d11	Revert "[Feat] enable hierarchical mc2 ops on A2 by default (#5545 )" (#5611 ) This reverts commit `fb9fdcdbe4`. ### What this PR does / why we need it? this pr breaks the smoke test because of that leads the error of aclnnNeScalar:Kernel Run failed. opType: 25, NotEqual launch failed for NotEqual, errno:361001 <img width="1149" height="166" alt="A6C9453D-4F0B-4256-DD80-A9C181DAB2D9" src="https://github.com/user-attachments/assets/cab9c4b8-3fd1-4c6b-b424-474b46042726" /> ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM version: v0.13.0 - vLLM main: `7157596103` Signed-off-by: zxwang <1476209578@qq.com>	2026-01-05 22:39:05 +08:00
Angazenn	11e75494b1	[TRITON][TEST]Add nightly test for triton split_qkv_rmsnorm_rope (#5267 ) ### What this PR does / why we need it? Add nightly test for triton split_rmsnorm_rope ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM version: release/v0.13.0 - vLLM main: `ad32e3e19c` --------- Signed-off-by: Angazenn <supperccell@163.com>	2026-01-05 21:35:37 +08:00
Chen Chen	a2daacbd71	[perf] Fix MLAPO weight disposal for KV-consumer MLA in PD-mix deploy... (#5192 ) ### What this PR does / why we need it? - Problem: In MLA+MLAPO, KV-consumer deployments keep fused_qkv_a_proj/q_proj weights and quant params even though MLAPO uses the prepacked buffers, increasing memory footprint on decode nodes. - Fix: Conditionally drop those tensors only when `kv_transfer_config.is_kv_consumer` to reclaim memory (consistent with the SFA behavior #4774 ). ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM version: v0.12.0 - vLLM main: `ad32e3e19c` Signed-off-by: Chen Chen <0109chenchen@gmail.com>	2026-01-05 21:29:45 +08:00
Qiu	b10ef9b9f3	[docs] Correct image about prefill phase of PCP (#5598 ) ### What this PR does / why we need it? Remove the incorrectly depicted DCP all_gather operation in the prefill stage PCP for GQA diagram. Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>	2026-01-05 20:21:59 +08:00
meihanc	a034941d06	[CI] update triton-ascend version (#5584 ) ### What this PR does / why we need it? update triton-ascend version to 20260105 - vLLM version: v0.13.0 - vLLM main: `7157596103` --------- Signed-off-by: Meihan-chen <jcccx.cmh@gmail.com>	2026-01-05 20:20:11 +08:00
Chao Lei	473431e7e2	[P/D]Remove mooncake kvpool unused parameter `local_hostname` (#5574 ) ### What this PR does / why we need it? In mooncake kvpool, `local_hostname` is not used. Instead, the local IP is obtained directly via `get_ip()`. Therefore, remove this parameter to avoid confusion. ### How was this patch tested? - vLLM version: v0.13.0 - vLLM main: `7157596103` Signed-off-by: LCAIZJ <leichao139636@163.com>	2026-01-05 20:18:59 +08:00
Debonet	d86021f7b4	[Bugfix] record cos and sin cache in AscendRotaryEmbedding (#5516 ) ### What this PR does / why we need it? In scenarios where models like [Moonlight](https://modelscope.cn/models/moonshotai/Moonlight-16B-A3B-Instruct) (using MLA but without `rope_scaling` in config.json) invoke `AscendRotaryEmbedding`. `_cos_cache` and `_sin_cache` are not recorded correctly. ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? - vLLM version: v0.13.0 - vLLM main: `45c1ca1ca1` Signed-off-by: Debonex <719893090@qq.com>	2026-01-05 20:12:41 +08:00
meihanc	16b1bee804	[bugfix] fix test_camem failed with triton-ascend (#5492 ) ### What this PR does / why we need it? This fixes a bug that occurred when running `test_camem.py` in the triton-ascend environment `NPU function error: aclrtGetMemInfo(ACL_HBM_MEM, &device_free, &device_total)` - vLLM version: v0.13.0 - vLLM main: `5326c89803` --------- Signed-off-by: Meihan-chen <jcccx.cmh@gmail.com>	2026-01-05 20:10:54 +08:00
ZT-AIA	58e8d19c35	[UT]add triton ops ut : test_fused_qkvzba_split_reshape_cat (#5474 ) ### What this PR does / why we need it? [UT]add triton ops ut : test_fused_qkvzba_split_reshape_cat ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? pytest -sv tests/ut/ops/test_fused_qkvzba_split_reshape_cat.py - vLLM version: v0.13.0 - vLLM main: `5326c89803` --------- Signed-off-by: ZT-AIA <1028681969@qq.com>	2026-01-05 20:05:07 +08:00
Li Wang	1e6228d8cd	[CI] Download models from ms (#5405 ) ### What this PR does / why we need it? Add a new workflow to allow developers submit a pull_request downloading new models for CI cache - vLLM version: release/v0.13.0 - vLLM main: `254f6b9867` Signed-off-by: wangli <wangli858794774@gmail.com>	2026-01-05 19:59:13 +08:00
huqi	2d22700d69	Docs: Add A3 Docker image guidance for Atlas A3 machines (#5256 ) Fixes #3386 - Update Qwen3-30B-A3B.md to use A3-specific image tag - Update Qwen3-Dense.md to provide both A2 and A3 image options - Update Qwen3-Next.md to use A3-specific image for Atlas A3 environments Previously, documentation only mentioned A2 images (vllm-ascend:version) but Atlas A3 machines require A3-specific images (vllm-ascend:version-a3). This change ensures users select the correct image for their hardware. 🤖 Generated with [Claude Code](https://claude.com/claude-code) - vLLM version: release/v0.13.0 - vLLM main: `ad32e3e19c` Signed-off-by: hu-qi <huqi1024@gmail.com> Co-authored-by: Claude <noreply@anthropic.com>	2026-01-05 19:42:42 +08:00
huqi	9d8b4c8d9d	[Doc] Add NNAL installation guide and requirements (#5235 ) Fixes #2727 - Add NNAL to the software requirements table with version information - Add note explaining that prebuilt Docker images include NNAL - Add warning message for manual installation when encountering libatb.so errors - Improve visibility of NNAL installation instructions to prevent runtime errors This addresses the issue where users encounter 'libatb.so not found' errors due to missing NNAL installation in their environment. ### What this PR does / why we need it? ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM version: release/v0.13.0 - vLLM main: `ad32e3e19c` --------- Signed-off-by: menogrey <1299267905@qq.com> Signed-off-by: hu-qi <huqi1024@gmail.com> Co-authored-by: zhangyiming <34808445+menogrey@users.noreply.github.com>	2026-01-05 19:40:26 +08:00
frankie	ec3563334b	Add the requirement of arctic-inference which speculative decoding with suffix_decode (#5045 ) ### Does this PR introduce _any_ user-facing change? suffix spec decode method rely on `arctic-inference` library. This PR add it into requirements to make sure the function works by default ### How was this patch tested? - vLLM version: v0.12.0 - vLLM main: `ad32e3e19c` --------- Signed-off-by: frankie-ys <yongshengwang@cmbchina.com> Signed-off-by: frankie <wangyongsheng686@gmail.com>	2026-01-05 19:15:49 +08:00
Icey	e7b623b363	[BugFix][Fusion] Fix graph fusion failure problem (#5253 ) Currently, the vllm pull request (https://github.com/vllm-project/vllm/pull/24252) is causing operator fusion to fail. This issue was previously fixed by patching the backend. The root cause has been identified, and the problem can be resolved with this pull request. - vLLM version: release/v0.13.0 - vLLM main: `ad32e3e19c` --------- Signed-off-by: wxsIcey <1790571317@qq.com>	2026-01-05 17:49:09 +08:00

1 2 3 4 5 ...

1985 Commits