xc-llm-ascend

Author	SHA1	Message	Date
Cao Yi	3953dcf784	[Feature][Quant] Auto-detect quantization format from model files (#6645 ) ## Summary - Add automatic quantization format detection, eliminating the need to manually specify `--quantization` when serving quantized models. - The detection inspects only lightweight JSON files (`quant_model_description.json` and `config.json`) at engine initialization time, with no `.safetensors` reads. - User-explicit `--quantization` flags are always respected; auto-detection only applies when the flag is omitted. ## Details Detection priority: 1. `quant_model_description.json` exists → `quantization="ascend"` (ModelSlim) 2. `config.json` contains `"quant_method": "compressed-tensors"` → `quantization="compressed-tensors"` (LLM-Compressor) 3. Neither → default float behavior Technical approach: Hooked into `NPUPlatform.check_and_update_config()` to run detection after `VllmConfig.__post_init__`. Since `quant_config` is already `None` at that point, we explicitly recreate it via `VllmConfig._get_quantization_config()` to trigger the full quantization initialization pipeline. ## Files Changed \| File \| Description \| \|------\|-------------\| \| `vllm_ascend/quantization/utils.py` \| Added `detect_quantization_method()` and `maybe_auto_detect_quantization()` \| \| `vllm_ascend/platform.py` \| Integrated auto-detection in `check_and_update_config()` \| \| `vllm_ascend/quantization/modelslim_config.py` \| Improved error handling for weight loading \| - vLLM version: v0.15.0 - vLLM main: `d7e17aaacd` --------- Signed-off-by: SlightwindSec <slightwindsec@gmail.com>	2026-02-26 10:59:25 +08:00
starmountain1997	bc1622338c	[CI] Add long and short prompt tests for DeepSeek-V3.2 (#6536 ) ### What this PR does / why we need it? This version has no divisibility constraint between tp and mtp+1. However, cudagraph_capture_sizes must be a common multiple of tp and mtp+1, with a maximum of tp * (mtp+1). Therefore, we fixed cudagraph_capture_sizes. We added a long-sequence test (64k input, 3k output) for the two-node mixed deployment scenario. Due to the excessive time required for performance benchmarking, we are only verifying functionality. The single-node scenario is skipped because VRAM limitations prevent launching the model with a max-model-len of 68,000. and we also add aime2025 test for dual-node deepseek 3.2 nightly test. ### How was this patch tested? test at nightly environment. - vLLM version: v0.15.0 - vLLM main: https://github.com/vllm-project/vllm/commit/v0.15.0 Signed-off-by: guozr <guozr1997@hotmail.com> Co-authored-by: guozr <guozr1997@hotmail.com>	2026-02-26 10:58:50 +08:00
Dijurido	169e434f78	[CI] Fix EAGLE CI problems (#6702 ) ### What this PR does / why we need it? New FIA operator requires queryT equal to the last element of actualSequenceLengthQ. ### Does this PR introduce _any_ user-facing change? No. ### How was this patch tested? Passed existing test (test_mtp_eagle_correctness.py). - vLLM version: v0.15.0 - vLLM main: `9562912cea` --------- Signed-off-by: Wangbingjie <wangbj1207@126.com> Signed-off-by: Wangbingjie <w30061490@china.huawei.com> Co-authored-by: Wangbingjie <w30061490@china.huawei.com>	2026-02-26 10:26:01 +08:00
Li-Yongwen	2870f7c8ad	[Feat] Support routing replay (#6696 ) ### What this PR does / why we need it? [Feat] Support routing replay same as https://github.com/vllm-project/vllm-ascend/pull/6666 resubmit because of DOC failure ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM version: v0.15.0 - vLLM main: `9562912cea` --------- Signed-off-by: liyongwen <1310439159@qq.com> Signed-off-by: Li-Yongwen <63399187+Li-Yongwen@users.noreply.github.com> Co-authored-by: wangxiyuan <wangxiyuan1007@gmail.com>	2026-02-26 10:22:47 +08:00
Rozwel-dx	a9cca0c5c4	[Refactor] Modify the binding logic, added memory migration and interrupt core binding functions. (#6785 ) [Refactor] Modify the binding logic, added memory migration and interrupt core binding functions. ### What this PR does / why we need it? Controls the use of memory on a closer NUMA node to achieve a lower memory access latency, while binding interrupts to different CPU cores to prevent them form interrupting the inference process. ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? `b8eaaa073b` Signed-off-by: rowzwel_dx <1392851715@qq.com> Signed-off-by: Rozwel-dx <1392851715@qq.com> - vLLM version: v0.15.0 - vLLM main: `9562912cea` Signed-off-by: Rozwel-dx <1392851715@qq.com>	2026-02-26 08:49:50 +08:00
Shanshan Shen	3a4292e5b7	[MM][Perf] Use `seq_lens` CPU cache to avoid frequent d2h copy for better performance (#6448 ) ### What this PR does / why we need it? Currently, the performance of multi-modal encoding (i.e., `AscendMMEncoderAttention` forward) is considerably bounded by the heavy host pre-process operations. We can see from the profiling results below, before the real computation of Attention, there are long free time in the device, which will lead to extremely low NPU utilization. <img width="2264" height="1398" alt="iShot_2026-01-23_16 26 39" src="https://github.com/user-attachments/assets/37f21d06-e526-4f28-82fe-005746cf13bd" /> --- To opitimize this, this PR has proposed four changes: 1. Use `seq_lens` CPU cache to avoid frequent d2h copy. Before this PR, `AscendMMEncoderAttention` will copy the `cu_seqlens` from NPU to CPU in every forward, since the op `_npu_flash_attention_unpad()` requires CPU `cu_seqlens` (otherwise it will crash). Thus, we use `seq_lens_cpu_cache` to cache this tensor, since it's shared between all layers, but may change in different forward step. When the current `layer_index` is `0`, we update the cache, otherwise we directly use the cache to avoid frequent `diff` and `copy` operations, which are costful. 2. Pre-compute the scale value to avoid calculating it in every forward. 3. Move the judgment of `enable_pad` from forward to the `__init__` method. 4. Revert https://github.com/vllm-project/vllm-ascend/pull/6204. Performance after these optimizations: - TTFT has been reduced by 7.43% ⬇️. - Throughput has been increased by 1.23% ⬆️. --- > [!NOTE] > This PR requires https://github.com/vllm-project/vllm/pull/33674 be merged. --- ### Does this PR introduce _any_ user-facing change? No. ### How was this patch tested? Launch the server: ```bash vllm serve /root/.cache/modelscope/hub/models/Qwen/Qwen3-VL-8B-Instruct \ --dtype bfloat16 \ --limit-mm-per-prompt '{"image": 1}' \ --max-model-len 16384 \ --max-num-batched-tokens 16384 \ --no-async-scheduling ``` Run benchmark: ```bash vllm bench serve \ --model /root/.cache/modelscope/hub/models/Qwen/Qwen3-VL-8B-Instruct \ --backend openai-chat \ --endpoint /v1/chat/completions \ --dataset-name hf \ --hf-split train \ --dataset-path lmarena-ai/vision-arena-bench-v0.1 \ --num-prompts 500 \ --request-rate 10 \ --burstiness 5 \ --no-stream ``` Before this PR: ``` ============ Serving Benchmark Result ============ Successful requests: 500 Failed requests: 0 Request rate configured (RPS): 10.00 Benchmark duration (s): 82.23 Total input tokens: 33418 Total generated tokens: 61543 Request throughput (req/s): 6.08 Output token throughput (tok/s): 748.45 Peak output token throughput (tok/s): 3203.00 Peak concurrent requests: 402.00 Total token throughput (tok/s): 1154.86 ---------------Time to First Token---------------- Mean TTFT (ms): 10275.37 Median TTFT (ms): 6297.88 P99 TTFT (ms): 22918.26 -----Time per Output Token (excl. 1st token)------ Mean TPOT (ms): 263.02 Median TPOT (ms): 277.61 P99 TPOT (ms): 483.56 ---------------Inter-token Latency---------------- Mean ITL (ms): 257.31 Median ITL (ms): 94.83 P99 ITL (ms): 1773.90 ================================================== ``` After this PR: ``` ============ Serving Benchmark Result ============ Successful requests: 500 Failed requests: 0 Request rate configured (RPS): 10.00 Benchmark duration (s): 81.20 Total input tokens: 33418 Total generated tokens: 61509 Request throughput (req/s): 6.16 Output token throughput (tok/s): 757.54 Peak output token throughput (tok/s): 2562.00 Peak concurrent requests: 395.00 Total token throughput (tok/s): 1169.11 ---------------Time to First Token---------------- Mean TTFT (ms): 9511.91 Median TTFT (ms): 5479.78 P99 TTFT (ms): 21427.21 -----Time per Output Token (excl. 1st token)------ Mean TPOT (ms): 261.12 Median TPOT (ms): 276.03 P99 TPOT (ms): 446.99 ---------------Inter-token Latency---------------- Mean ITL (ms): 254.04 Median ITL (ms): 97.71 P99 ITL (ms): 1516.67 ================================================== ``` - vLLM version: v0.15.0 - vLLM main: `dc917cceb8` Signed-off-by: shen-shanshan <467638484@qq.com>	2026-02-26 08:49:36 +08:00
jack	29e3cdde20	[Doc][Skill] Introduce AI-assisted model-adaptation workflow for vllm-ascend (#6731 ) ### What this PR does / why we need it This PR introduces the first AI-assisted model-adaptation skill package for `vllm-ascend`. The goal is to make model adaptation work (especially for recurring feature-request issues) repeatable, auditable, and easier to hand off. ### Scope in this PR This PR adds only skill/workflow assets under: - `.agents/skills/vllm-ascend-model-adapter/SKILL.md` - `.agents/skills/vllm-ascend-model-adapter/references/workflow-checklist.md` - `.agents/skills/vllm-ascend-model-adapter/references/troubleshooting.md` - `.agents/skills/vllm-ascend-model-adapter/references/multimodal-ep-aclgraph-lessons.md` - `.agents/skills/vllm-ascend-model-adapter/references/fp8-on-npu-lessons.md` - `.agents/skills/vllm-ascend-model-adapter/references/deliverables.md` ### Workflow improvements The skill standardizes: 1. Environment assumptions used in our Docker setup - implementation roots: `/vllm-workspace/vllm` and `/vllm-workspace/vllm-ascend` - serving root: `/workspace` - model path convention: `/models/<model-name>` 2. Validation strategy - Stage A: fast `--load-format dummy` gate - Stage B: mandatory real-weight gate before sign-off - avoid false-ready by requiring request-level checks (not startup log only) 3. Feature-first verification checklist - ACLGraph / EP / flashcomm1 / MTP / multimodal - explicit `supported / unsupported / not-applicable / checkpoint-missing` outcomes 4. Delivery contract - minimal scoped code changes - required artifacts (Chinese report + runbook, e2e config YAML, tutorial doc) - one signed commit in delivery repo ### What this PR does NOT do - No runtime/kernel/model patch is included in this PR. - No direct model support claim is made by this PR alone. - Model-specific adaptation/fix work should be submitted in follow-up PRs using this skill as the workflow baseline. ### Why this matters for maintainers This gives the repo a shared, explicit AI-assistance protocol, so future model-adaptation PRs are easier to review, compare, and reproduce. --------- Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com> Co-authored-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>	2026-02-26 08:48:15 +08:00
wangxiyuan	3b59d0ebe9	[Doc][Feature] Add vLLM Ascend development guidelines AGETNS.md (#6797 ) ### What this PR does / why we need it? This PR adds a new document, `AGENTS.md`, which provides detailed development guidelines for contributors to the vLLM Ascend project. These guidelines cover code style, testing, NPU-specific considerations, and the contribution process to ensure code quality and consistency. ### Does this PR introduce _any_ user-facing change? No, this is a documentation-only update for developers. ### How was this patch tested? This is a documentation change and does not require testing. - vLLM version: v0.15.0 - vLLM main: `83b47f67b1` Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>	2026-02-26 08:47:46 +08:00
Zhu Yi Lin	aa7fb5d707	[Bugfix] Fix DeepseekV3.1 Accuracy issue (#6805 ) ### What this PR does / why we need it? In order to adapt to the GLM model, logits were passed in the sample, which can cause accuracy issues in version 0.15.0. - vLLM version: v0.15.0 - vLLM main: `83b47f67b1` Signed-off-by: GDzhu01 <809721801@qq.com>	2026-02-25 23:02:00 +08:00
bowenli	e3927cc8f5	[Bugfix] fix bug for mtp (#6514 ) ### What this PR does / why we need it? fix(mtp): resolve MTP core bugs and enhance eager mode test cases 1. Resolved critical issues in eager mode MTP core execution logic; 2. Fixed functional bugs in the _update_states_after_model_execute function; 3. Updated and released test_mtp_qwen3_next.py to validate eager mode acceptance rate. ### Does this PR introduce _any_ user-facing change? None ### How was this patch tested? - vLLM version: v0.15.0 - vLLM main: https://github.com/vllm-project/vllm/commit/v0.15.0 Signed-off-by: Bowen-Leee <caoshankuangren@gmail.com>	2026-02-25 17:50:57 +08:00
LoganJane	ed051737e9	[Bugfix] Support Kimi-K2.5 models (#6755 ) ### What this PR does / why we need it? This PR supports the Kimi-K2.5 models on the NPU of bf16 and w4a8 weights. The corresponding PR in the vllm community has been merged: https://github.com/vllm-project/vllm/pull/34501 ### Does this PR introduce _any_ user-facing change? - No. ### How was this patch tested? We test the Kimi-K2.5 weights. The weights path: https://modelscope.cn/models/Eco-Tech/Kimi-K2.5-W4A8 Successfully ran on 910B NPU using vllm-ascend by the w4a8 weights. - vLLM version: v0.15.0 - vLLM main: `9562912cea` --------- Signed-off-by: LoganJane <LoganJane73@hotmail.com>	2026-02-25 14:51:46 +08:00
kx	4efd362bac	[fix]change num_commmon_tokens to num_common_tokens (#6792 ) ### What this PR does / why we need it? change num_commmon_tokens to num_common_tokens in vllm_ascend/_310p/model_runner_310p.py，which caused CI test failure - vLLM version: v0.15.0 - vLLM main: `9562912cea` Signed-off-by: 01267596 <xiongkai123@cmbchina.com> Co-authored-by: 01267596 <xiongkai123@cmbchina.com>	2026-02-25 14:48:54 +08:00
starmountain1997	2260af405f	[DOC] add request forwarding (#6780 ) ### What this PR does / why we need it? - New section: "Request Forwarding" documentation in docs/source/tutorials/models/DeepSeek-V3.2.md - Environment fix: Changed VLLM_ASCEND_ENABLE_FLASHCOMM1 from 0 to 1 in the DeepSeek-V3 configuration examples ### Does this PR introduce _any_ user-facing change? Documentation update only - provides new configuration guidance for request forwarding setups ### How was this patch tested? - vLLM version: v0.15.0 - vLLM main: `9562912cea` --------- Signed-off-by: guozr <guozr1997@hotmail.com> Co-authored-by: guozr <guozr1997@hotmail.com>	2026-02-25 14:43:51 +08:00
Canlin Guo	ad9d9569ea	[Bugfix] Add the missing parentheses to @torch.inference_mode (#6757 ) ### What this PR does / why we need it? This PR fixes a bug in `vllm_ascend/worker/model_runner_v1.py` where the `@torch.inference_mode` decorator was used without parentheses. Using the decorator without instantiation is deprecated and may not correctly disable gradient calculations, leading to performance degradation and increased memory usage during inference. This change adds the required parentheses to ensure `torch.inference_mode` is applied correctly. ### Does this PR introduce _any_ user-facing change? No. ### How was this patch tested? The change is a minor syntax correction. Existing CI tests should cover this. - vLLM version: v0.15.0 - vLLM main: `9562912cea` Signed-off-by: gcanlin <canlinguosdu@gmail.com>	2026-02-25 14:37:53 +08:00
Shanshan Shen	957804df56	[Refactor][Bugfix] Use upstream `mem_utils` for profiling and correct non-torch memory recorded during profiling (#6625 ) ### What this PR does / why we need it? 1. Following https://github.com/vllm-project/vllm/pull/32322, use the `memory_profiling` context manager from vllm for profiling. 2. Fix wrong non-torch memory value recorded during profiling, which is not its peak during inference. --- More details about point 2: After profling, the non-torch memory value we recorded is lower than that in real inference. This is mainly because of the different memory management behaviour between `torch.cuda.empty_cache()` and `torch.npu.empty_cache()`. With regard to `torch.cuda.empty_cache()`, it only recycle the unused memory in pytorch memory pool (i.e., memory managed by pytorch caching allocator), with no affect to non-torch memory. However, as for `torch.npu.empty_cache()`, it has a totally different memory management mechanism, i.e., it may call `aclrtSynchronize` and enable Ascend runtime to free up non-torch memory. Thus, the non-torch memory value we recorded after `torch.npu.empty_cache()` is much lower than its peak during profling. Resolution: We record the peak non-torch memory value (`non_torch_memory_before_empty_cache`) after profiling, but before `torch.npu.empty_cache()`. Then, we add the diff (`non_torch_memory_cleared_by_empty_cache = non_torch_memory_before_empty_cache - self.non_torch_memory`) to non-torch memory when calculating available KV cache memory, which will lead to less KV cache memory (i.e., it's safer to avoid OOM issues). --- > [!NOTE] > This PR needs to wait for main2main aligning to latest vllm commit before merging. ### Does this PR introduce _any_ user-facing change? no. ### How was this patch tested? Before this PR, the non-torch memory we used to calculate available KV cache memory is 0.90 G, whereas its peak during real inference is 1.08 G, diff: 182.00 M. After this PR, we add this diff to non-torch memory after profiling and thus make the profiling results more accurate. - vLLM version: v0.15.0 - vLLM main: `d7e17aaacd` --------- Signed-off-by: shen-shanshan <467638484@qq.com>	2026-02-25 14:28:08 +08:00
DreamerLeader	812c722cfb	[KVPool][BugFix] Correctly initialize head_or_tp_rank for mooncake backend (#6498 ) ### What this PR does / why we need it? The problem that the local priority is not used in the A2 environment on the Mooncake node is resolved. - vLLM version: v0.15.0 - vLLM main: https://github.com/vllm-project/vllm/commit/v0.15.0 --------- Signed-off-by: 房建伟 <fangjianwei@fangjianweideMacBook-Air.local> Co-authored-by: Pz1116 <zpbzpb123123@gmail.com>	2026-02-25 14:22:00 +08:00
Frank Chen	3da2ba22eb	[Platform] Enable ARM-only CPU binding with NUMA-balanced A3 policy and update docs/tests (#6686 ) ### What this PR does / why we need it? - Keeps enable_cpu_binding default on, but skips binding on non‑ARM CPUs inside bind_cpus, with a clear log. - Uses a table-driven binding policy: A3 uses NUMA‑balanced binding; other device types use NUMA‑affinity binding. - Updates docs to reflect the exact behavior and adds/updates unit tests for the new logic. ### Does this PR introduce _any_ user-facing change? - Yes. CPU binding is now enabled by default via additional_config, and documented in the user guide. - CPU binding behavior differs by device type (A3 vs. others). ### How was this patch tested? Added/updated unit tests: test_cpu_binding.py 1. test_binding_mode_table covers A2 vs A3 binding mode mapping. 2. test_build_cpu_pools_fallback_to_numa_balanced covers fallback when affinity info is missing. 3. TestBindingSwitch.test_is_arm_cpu covers ARM/x86/unknown arch detection. 4. test_bind_cpus_skip_non_arm covers non‑ARM skip path in bind_cpus. test_worker_v1.py 1. Updated mocks for enable_cpu_binding default True to align with new config default. - vLLM version: v0.14.1 - vLLM main: d7de043 --------- Signed-off-by: chenchuw886 <chenchuw@huawei.com> Co-authored-by: chenchuw886 <chenchuw@huawei.com>	2026-02-25 11:15:14 +08:00
Li Wang	ac9a7d1301	[Nightly] Increase VLLM_ENGINE_READY_TIMEOUT_S to avoid nightly failure (#6778 ) ### What this PR does / why we need it? After some observation, I found some cases failed for timeout, just like https://github.com/vllm-project/vllm-ascend/actions/runs/22280996034/job/64487867977#step:9:921 and https://github.com/vllm-project/vllm-ascend/actions/runs/22315540111/job/64574590762#step:9:1809, this may caused by the excessively long model loading time (currently we are still loading weights from network storage), it is necessary to adjust the timeout seconds 600s -> 1800s ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM version: v0.15.0 - vLLM main: `9562912cea` Signed-off-by: wangli <wangli858794774@gmail.com>	2026-02-25 10:14:51 +08:00
weiguihua2	db51a1b9b6	[Feat]ds3.2 support pcp (#6733 ) ### What this PR does / why we need it? The ds3.2 model adaptation supports the PCP feature. The solution is as follows: When saving the KV cache, first perform an allgather operation on the KVs, and then each node saves its own copy. When the attention or indexer performs calculations, they all gather the KV cache and then perform the calculations. ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? 02/12 23:05:10 - AISBench - INFO - Running 1-th replica of evaluation 02/12 23:05:10 - AISBench - INFO - Task [vllm-api-general-chat/gsm8k]: {'accuracy': 96.35416666666667, 'type': 'GEN'} 02/12 23:05:10 - AISBench - INFO - time elapsed: 2.87s 02/12 23:05:12 - AISBench - INFO - Evaluation tasks completed. 02/12 23:05:12 - AISBench - INFO - Summarizing evaluation results... dataset version metric mode vllm-api-general-chat gsm8kdataset - accuracy gen 96.35 - vLLM version: v0.15.0 - vLLM main: `9562912cea` --------- Signed-off-by: weiguihua2 <weiguihua2@huawei.com> Co-authored-by: wangxiyuan <wangxiyuan1007@gmail.com>	2026-02-25 09:46:57 +08:00
Icey	ee59429015	upgrade main to 0212 (#6712 ) ### What this PR does / why we need it? Fixes `transformers_utils/processors/__init__` import error, due to https://github.com/vllm-project/vllm/pull/33247 Fixes Fused MoE break introduced by `MoERunner abstraction,` due to https://github.com/vllm-project/vllm/pull/32344 > delete AscendMoERunnere when https://github.com/vllm-project/vllm/pull/35178 is merged Fixes `Make Qwen3VL compatible with Transformers v5`, due to https://github.com/vllm-project/vllm/pull/34262 ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM version: v0.15.0 - vLLM main: `9562912cea` --------- Signed-off-by: wxsIcey <1790571317@qq.com>	2026-02-25 09:17:29 +08:00
LI SHENGYONG	0331f16a50	[EPLB] Reduce the memory used for heat aggregation (#6729 ) ### What this PR does / why we need it? If dist.all_gather is used directly, 2 x HCCL_BUFFSIZE memory will be consumed, but the actual memory required for hotspot aggregation is less than 1 MB. Therefore, a separate small communication domain is created for it. ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? Original： ![1](https://github.com/user-attachments/assets/8880b461-c26f-497c-9a05-2ca60cc46aa4) Current： ![2](https://github.com/user-attachments/assets/c9da32b5-9200-4fa2-aff9-d8c4978ac602) - vLLM version: v0.15.0 - vLLM main: `9562912cea` Signed-off-by: shenchuxiaofugui <1311027364@qq.com>	2026-02-24 18:02:24 +08:00
zzzzwwjj	5c8ab7af39	[main]update release note & support matrix (#6759 ) ### What this PR does / why we need it? Update release note & support matrix to add experimental tag for features and models. ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM version: v0.15.0 - vLLM main: `9562912cea` 0.13.0 branch: https://github.com/vllm-project/vllm-ascend/pull/6751 Signed-off-by: zzzzwwjj <1183291235@qq.com>	2026-02-24 17:39:35 +08:00
pu-zhe	a8e951e6f5	[Feat] 310p supports PrefillCacheHit State (#6756 ) ### What this PR does / why we need it? This PR extends the Ascend 310P attention backend to support the `PrefillCacheHit` state. Previously, only `PrefillNoCache`, `DecodeOnly`, and `ChunkedPrefill` were supported. This PR handles this state by routing it to the existing `forward_chunked_prefill_310` implementation, which is suitable for this scenario. The changes also include refactoring the main `forward_impl` dispatch method for better clarity and updating unit tests to cover the new state and ensure correctness. ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? Accuracy test when chunked prefill is disabled. - vLLM version: v0.15.0 - vLLM main: `9562912cea` --------- Signed-off-by: pu-zhe <zpuaa@outlook.com>	2026-02-24 16:48:05 +08:00
SILONG ZENG	62ea664aa7	[Lint]Style: Convert `test/` to ruff format(Batch #5 ) (#6747 ) ### What this PR does / why we need it? \| File Path \| \| :--- \| \| `tests/e2e/singlecard/compile/backend.py` \| \| `tests/e2e/singlecard/compile/test_graphex_norm_quant_fusion.py` \| \| `tests/e2e/singlecard/compile/test_graphex_qknorm_rope_fusion.py` \| \| `tests/e2e/singlecard/compile/test_norm_quant_fusion.py` \| \| `tests/e2e/singlecard/model_runner_v2/test_basic.py` \| \| `tests/e2e/singlecard/test_aclgraph_accuracy.py` \| \| `tests/e2e/singlecard/test_aclgraph_batch_invariant.py` \| \| `tests/e2e/singlecard/test_aclgraph_mem.py` \| \| `tests/e2e/singlecard/test_async_scheduling.py` \| \| `tests/e2e/singlecard/test_auto_fit_max_mode_len.py` \| \| `tests/e2e/singlecard/test_batch_invariant.py` \| \| `tests/e2e/singlecard/test_camem.py` \| \| `tests/e2e/singlecard/test_completion_with_prompt_embeds.py` \| \| `tests/e2e/singlecard/test_cpu_offloading.py` \| \| `tests/e2e/singlecard/test_guided_decoding.py` \| \| `tests/e2e/singlecard/test_ilama_lora.py` \| \| `tests/e2e/singlecard/test_llama32_lora.py` \| \| `tests/e2e/singlecard/test_models.py` \| \| `tests/e2e/singlecard/test_multistream_overlap_shared_expert.py` \| \| `tests/e2e/singlecard/test_quantization.py` \| \| `tests/e2e/singlecard/test_qwen3_multi_loras.py` \| \| `tests/e2e/singlecard/test_sampler.py` \| \| `tests/e2e/singlecard/test_vlm.py` \| \| `tests/e2e/singlecard/test_xlite.py` \| \| `tests/e2e/singlecard/utils.py` \| ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM version: v0.15.0 - vLLM main: `9562912cea` --------- Signed-off-by: MrZ20 <2609716663@qq.com>	2026-02-24 15:50:00 +08:00
xleoken	747484cb64	[Bugfix] Fix wrong computed_tokens when meet exception. (#6522 ) <!-- Thanks for sending a pull request! BEFORE SUBMITTING, PLEASE READ https://docs.vllm.ai/en/latest/contributing/overview.html --> ### What this PR does / why we need it? <!-- - Please clarify what changes you are proposing. The purpose of this section is to outline the changes and how this PR fixes the issue. If possible, please consider writing useful notes for better and faster reviews in your PR. - Please clarify why the changes are needed. For instance, the use case and bug description. - Fixes # --> Fix wrong computed_tokens when meet exception. This pull request addresses a bug in the KV transfer mechanism where an exception during token lookup operations could lead to an incorrect count of computed_tokens. By modifying the exception handling in both the lookup and lookup_scheduler functions to return 0 instead of the start index, the system now correctly indicates that no tokens were successfully processed when a remote connection failure occurs. This enhancement improves the robustness and accuracy of token management within the vllm_ascend distributed KV pool. ### Does this PR introduce _any_ user-facing change? <!-- Note that it means any user-facing change including all aspects such as API, interface or other behavior changes. Documentation-only updates are not considered user-facing changes. --> NO. ### How was this patch tested? <!-- CI passed with new added/existing test. If it was tested in a way different from regular unit tests, please clarify how you tested step by step, ideally copy and paste-able, so that other reviewers can test and check, and descendants can verify in the future. If tests were not added, please describe why they were not added and/or why it was difficult to add. --> Signed-off-by: xleoken <xleoken@163.com>	2026-02-24 15:29:30 +08:00
LI SHENGYONG	ff29e029de	[EPLB][Bugfix] Bugfix for ineffective dynamic eplb (#6653 ) ### What this PR does / why we need it? #6043 deleted the forward_before phase of the dynamic eplb. Currently, the end-to-end precision is monitored in the UT, and the log is not printed in the key place. As a result, the eplb does not take effect and is not intercepted. 1. The forward_before function is added back. 2. Delete unnecessary logs and add key logs. 3. Warm-up of algorithm 3 is added. ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? ![Snipaste_2026-02-10_15-57-31](https://github.com/user-attachments/assets/03813e5f-3d19-42d8-8118-76223afe8298) #### The conversation is normal. Okay, the user is asking, \"What is deep learning?\" I need to explain this in a clear and concise way. Let me start by recalling what I know about deep learning. It's a subset of machine learning, right? So first, I should mention that it's part of machine learning, which itself is a branch of AI. Then, the key aspect of deep learning is the use of neural networks with multiple layers. These are called deep neural networks.\n\nWait, I should define neural networks first. Maybe start with the basics. A neural network is inspired by the human brain, with layers of nodes (neurons) that process data. But deep learning specifically refers to networks with many layers—hence \"deep.\" So the term \"deep\" comes from the number of layers. \n\nI should explain how deep learning works. It involves training these networks on large datasets, allowing them to automatically learn features from the data. Unlike traditional machine learning, where you might have to manually extract features, deep learning models can do this automatically. That's a key point. For example, in image recognition, a deep learning model can learn to detect edges, shapes, and then more complex patterns without human intervention.\n\nApplications are important too. The user might want to know where deep learning is used. Common examples include image and speech recognition, natural language processing, autonomous vehicles, and recommendation systems. Maybe mention specific technologies like self-driving cars using computer vision or virtual assistants like Siri or Alexa - vLLM version: v0.15.0 - vLLM main: `13397841ab` Signed-off-by: shenchuxiaofugui <1311027364@qq.com>	2026-02-24 14:43:04 +08:00
luomin2005	f41eeeb11e	Refactor the ops PyTorch adapter，cleanup for csrc/torch_binding.cpp (#6732 ) ### What this PR does / why we need it? Refactor the ops PyTorch adapter，cleanup for csrc/torch_binding.cpp, more details see https://github.com/vllm-project/vllm-ascend/issues/6486 ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? install the new package to test the new modification, here is the result: - vLLM version: v0.15.0 - vLLM main: `9562912cea` --------- Signed-off-by: liziyu <liziyu16@huawei.com> Signed-off-by: wangxiaoteng <wangxiaoteng@huawei.com> Signed-off-by: luomin2005 <luomin2005@huawei.com> Co-authored-by: liziyu <56102866+liziyu179@users.noreply.github.com> Co-authored-by: wangxiaoteng <wangxiaoteng@huawei.com>	2026-02-24 09:12:43 +08:00
Nengjun Ma	f0caeeadcb	[CI] unlock when load model (#6771 ) ### What this PR does / why we need it? ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM version: v0.15.0 - vLLM main: `9562912cea` Signed-off-by: leo-pony <nengjunma@outlook.com>	2026-02-14 18:54:04 +08:00
yydyzr	70e26551cf	[Doc] modify glm doc (#6770 ) ### What this PR does / why we need it? 1. add description of another version of glm5-w4a8 weight 2. update the introduction of installation 3. introduce a script to enable bf16 MTP ### Does this PR introduce _any_ user-facing change? N/A ### How was this patch tested? N/A - vLLM version: v0.15.0 - vLLM main: `9562912cea` --------- Signed-off-by: yydyzr <liuyuncong1@huawei.com>	2026-02-14 16:47:23 +08:00
SILONG ZENG	e2237819a9	[CI]Fixed the spell check function in `typos.toml` (#6753 ) ### What this PR does / why we need it? The incorrect regular expression syntax `.[UE4M3\|ue4m3].` actually ignores all words containing any of the following characters: `u, e, 4, m, 3, \|` ```yaml extend-ignore-identifiers-re = [".Unc.", "._thw", ".UE8M0.", ".[UE4M3\|ue4m3].", ".eles.", ".fo.", ".ba.", ".ot.", ".[Tt]h[rR]."] ``` ===fix===> ```yaml extend-ignore-identifiers-re = [".Unc.", "._thw", ".UE8M0.", ".(UE4M3\|ue4m3]).", ".eles.", ".fo.", ".ba.", ".ot.", ".[Tt]h[rR]."] ``` ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM version: v0.15.0 - vLLM main: `9562912cea` Signed-off-by: MrZ20 <2609716663@qq.com>	2026-02-14 11:57:26 +08:00
JIACHENG XU	64aea60f2e	[EPLB][Nightly] Refactor UT (#6543 ) ### What this PR does / why we need it? The basic configs are extracted and reused for eplb UT. This is done so that if the basic configs are changed later, eplb UT does not need to be modified repeatedly. ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM version: v0.15.0 - vLLM main: https://github.com/vllm-project/vllm/commit/v0.15.0 Signed-off-by: bigsir007 <xujiacheng12@huawei.com> Co-authored-by: bigsir007 <xujiacheng12@huawei.com>	2026-02-14 10:56:29 +08:00
xulei	1e77077788	[Bugfix][DispatchFFNCombine] resolve vec error caused by unaligned UB access (#6707 ) ### What this PR does / why we need it? 1. Fix a vec error caused by unaligned UB accesss in the DispatchFFNCombine; 2. Fix expert_token_nums tensor defined on host instead of NPU in moe_comm_method.py 3. Fix multi-core copy issue of expert_token_nums in dispatchffnCombine op (single aiv copy is sufficient) ### Does this PR introduce _any_ user-facing change? No, this PR does not introduce any user-facing changes. The fix only addresses internal memory access logic and does not modify any public APIs, interfaces, or user-visible behaviors. ### How was this patch tested? `export VLLM_ASCEND_ENABLE_FUSED_MC2=1` vLLM version: v0.15.0 - vLLM version: v0.15.0 - vLLM main: `9562912cea` Signed-off-by: xulei_ict <xulei292@huawei.com> Co-authored-by: xulei_ict <xulei292@huawei.com>	2026-02-14 10:32:50 +08:00
whx	e2175d9c7e	[Lint] Adapt lint tools for windows (#6727 ) ### What this PR does / why we need it? If users run bash format.sh with `git bash` on windows system, there exists `Executable /bin/bash not found` error. This is because in Windows Git Bash environment, the Bash executable is actually located at `/usr/bin/bash`, while the `/bin` directory may not exist, or may be just an empty directory or a broken symlink that does not contain bash. ### Does this PR introduce _any_ user-facing change? None ### How was this patch tested? With this PR and `pre-commit` installed, windows coders can directly run `bash format.sh` to clean lint issues. - vLLM version: v0.15.0 - vLLM main: `9562912cea` Signed-off-by: whx-sjtu <2952154980@qq.com>	2026-02-13 15:53:16 +08:00
Cao Yi	6de207de88	[main][Docs] Fix typos across documentation (#6728 ) ## Summary Fix typos and improve grammar consistency across 50 documentation files. ### Changes include: - Spelling corrections (e.g., "Facotory" → "Factory", "certainty" → "determinism") - Grammar improvements (e.g., "multi-thread" → "multi-threaded", "re-routed" → "re-run") - Punctuation fixes (semicolon consistency in filter parameters) - Code style fixes (correct flag name `--num-prompts` instead of `--num-prompt`) - Capitalization consistency (e.g., "python" → "Python", "ascend" → "Ascend") - vLLM version: v0.15.0 - vLLM main: `9562912cea` --------- Signed-off-by: SlightwindSec <slightwindsec@gmail.com>	2026-02-13 15:50:05 +08:00
Shaoxu Cheng	b6bc3d2f9d	[Feat.][310P]: weightNZ feature with quant or unquant. (#6705 ) NZ Format Support for Linear Layers: Implemented support for the NZ (N-dimensional Z-order) format for linear layer weights on Ascend 310P, enhancing performance for both quantized and unquantized layers. Unquantized Linear Method for Ascend 310P: Introduced AscendUnquantizedLinearMethod310 to specifically handle and apply NZ format casting to unquantized linear layer weights during the loading process. MRotaryEmbedding Integration: Extended Rotary Embedding support by adding AscendMRotaryEmbedding310 to provide an Ascend-specific implementation for MRotaryEmbedding. Quantization Method Updates: Updated the w8a8_static quantization method to directly transpose weights and apply NZ format casting, ensuring consistency with the new format. - vLLM version: v0.15.0 - vLLM main: `9562912cea` --------- Signed-off-by: Tflowers-0129 <2906339855@qq.com>	2026-02-13 15:41:02 +08:00
Shaoxu Cheng	f40256b697	[Feat.][310P] addrmsnorm for 300I DUO (#6704 ) ### What this PR does / why we need it? This PR integrates the `npu_add_rms_norm` fused kernel for RMSNorm operations with residual connections on 310P devices. This change optimizes the computation by replacing a two-step process (manual residual addition followed by RMSNorm) with a single, more efficient fused operation. This is needed to improve the performance of models utilizing RMSNorm with residual connections on the 310P architecture. Fixes # ### Does this PR introduce _any_ user-facing change? No, this PR introduces an internal optimization and does not change any user-facing APIs or behaviors. ### How was this patch tested? This patch was tested with updated unit tests (`test_RMSNorm_forward_310p`) that mock the `npu_add_rms_norm` operation to verify the correctness of the fused kernel integration. --------- Signed-off-by: Tflowers-0129 <2906339855@qq.com>	2026-02-13 15:40:49 +08:00
Icey	7164990904	[Graph][Fusion] Integrating inductor pass and npugraph ex pass (#6354 ) ### What this PR does / why we need it? Integrating inductor pass and npugraph ex pass, see RFC: https://github.com/vllm-project/vllm-ascend/issues/6347 ### Does this PR introduce _any_ user-facing change? N/A ### How was this patch tested? all tests passed. - vLLM version: v0.14.1 - vLLM main: `dc917cceb8` --------- Signed-off-by: wxsIcey <1790571317@qq.com>	2026-02-13 15:34:55 +08:00
iiiklw	87a0b7b7c7	[bugfix] adapt bugfix for norm_quant_fusion_pass to npugraph_ex (#6726 ) ### What this PR does / why we need it? This PR adapts bugfixes from `norm_quant_fusion_pass` to `graphex_norm_quant_fusion_pass` for the `npugraph_ex` backend. The main changes are: - Replaced `torch.ops.npu.npu_add_rms_norm` with `torch.ops._C_ascend.npu_add_rms_norm_bias`. - For patterns without bias, `None` is passed as the bias argument. - For patterns with bias, the separate `add` operation for bias is removed and the bias is passed directly to `npu_add_rms_norm_bias`. This improves fusion. These changes ensure consistency and correctness for RMSNorm and quantization fusion patterns when using `npugraph_ex`. ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM version: v0.15.0 - vLLM main: `9562912cea` Signed-off-by: huyuanquan1 <huyuanquan1@huawei.com> Co-authored-by: huyuanquan1 <huyuanquan1@huawei.com>	2026-02-13 10:10:39 +08:00
taoyao1221	41d056f947	[doc] add A2 series doc for GLM5.md (#6717 ) ### What this PR does / why we need it? Added support for A2 in the GLM-5 doc. ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? vLLM version: v0.15.0 vLLM main: `9562912cea` - vLLM version: v0.15.0 - vLLM main: `9562912cea`	2026-02-12 16:08:17 +08:00
wangxiaoteng888	b881fab416	[P/D][PCP] mooncake layerwise support pcp function (#6627 ) ### What this PR does / why we need it? mooncake layerwise support pcp function PCP (Prefill Context Parallelism) Support: Introduced explicit support for Prefill Context Parallelism (PCP) and Decode Context Parallelism (DCP) in the Mooncake layerwise KV cache transfer mechanism, allowing for more granular control and awareness of parallel configurations during data transfer. ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? By ci - vLLM version: v0.15.0 - vLLM main: `d7e17aaacd` --------- Signed-off-by: wangxiaoteng <wangxiaoteng@huawei.com> Signed-off-by: liziyu <liziyu16@huawei.com> Co-authored-by: liziyu <liziyu16@huawei.com>	2026-02-12 11:02:25 +08:00
yejj	8b23554741	[Misc] gen kv events in ascendconnector (#6593 ) ### What this PR does / why we need it? refer to https://github.com/vllm-project/vllm-ascend/issues/6391, Currently adapted the complete process of event publishing in vllm: * `kv_connector_model_runner_mixin` invoke kv-connector `get_kv_connector_kv_cache_events` func to collect kvevents * in `scheduler.py` , it's `update_from_output` func will invoke `_update_from_kv_xfer_finished` which invoke `connector.update_connector_output` to collect kv-events from all kv-worker, and then scheduler will invoke `connector.take_events` api to collect all kv-events and add it to the events which from `kv_cache_manager` ### Does this PR introduce _any_ user-facing change? no ### How was this patch tested? You can add `--kv-events-config` parameter to the `vllm server` command to enable this feature. - vLLM version: v0.15.0 - vLLM main: `d7e17aaacd` --------- Signed-off-by: yejj710 <abyss1999@163.com> Co-authored-by: fems14 <1804143737@qq.com>	2026-02-12 11:01:09 +08:00
jiahao.quan	7221045777	[Attention] add gpt-oss support (#5901 ) ### What this PR does / why we need it? Please refer to the following link for the historical conversation https://github.com/vllm-project/vllm-ascend/pull/4467. We have made updates in light of the comments from the prior PR review. Given the refactoring of the attention_v1 component, we have carried out necessary adjustments to fit the newly revised code. ### Does this PR introduce _any_ user-facing change? 1. Modified the code in the Attention section to adapt to the SWA and Sink features required by gpt-oss. 2. Modified the code in the MoE section to add support for bias and swigluoai. ### How was this patch tested? Please refer to the https://github.com/vllm-project/vllm-ascend/pull/4467 for performance tests, on the basis of which the accuracy tests from AIME2024 have been newly added. ![img_v3_02tu_501e88e3-2217-4565-8edf-b9acf4f43f2g](https://github.com/user-attachments/assets/024f8283-18ab-4d4d-ab12-27917b5d7d06) - vLLM version: v0.13.0 - vLLM main: `bde38c11df` --------- Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com> Signed-off-by: mikequan0425 <mikequan0425@foxmail.com> Signed-off-by: hfadzxy <starmoon_zhang@163.com> Signed-off-by: shenchuxiaofugui <1311027364@qq.com> Signed-off-by: jiangyunfan1 <jiangyunfan1@h-partners.com> Signed-off-by: pu-zhe <zpuaa@outlook.com> Signed-off-by: liziyu <liziyu16@huawei.com> Signed-off-by: wangxiaoteng <wangxiaoteng@huawei.com> Signed-off-by: luomin2005 <luomin2005@huawei.com> Signed-off-by: whx-sjtu <2952154980@qq.com> Signed-off-by: SlightwindSec <slightwindsec@gmail.com> Signed-off-by: wxsIcey <1790571317@qq.com> Signed-off-by: MrZ20 <2609716663@qq.com> Co-authored-by: wangxiyuan <wangxiyuan1007@gmail.com> Co-authored-by: leon_tao <taoyao2@huawei.com> Co-authored-by: nurxat <738457498@qq.com> Co-authored-by: hfadzxy <starmoon_zhang@163.com> Co-authored-by: mikequan <199741451@qq.com> Co-authored-by: LI SHENGYONG <49200266+shenchuxiaofugui@users.noreply.github.com> Co-authored-by: jiangyunfan1 <jiangyunfan1@h-partners.com> Co-authored-by: pu-zhe <zpuaa@outlook.com> Co-authored-by: luomin2005 <luomin2005@huawei.com> Co-authored-by: liziyu <56102866+liziyu179@users.noreply.github.com> Co-authored-by: wangxiaoteng <wangxiaoteng@huawei.com> Co-authored-by: whx <56632993+whx-sjtu@users.noreply.github.com> Co-authored-by: Cao Yi <slightwindsec@gmail.com> Co-authored-by: Icey <1790571317@qq.com> Co-authored-by: SILONG ZENG <2609716663@qq.com>	2026-02-12 10:55:34 +08:00
lih827	f71812011d	[Feature] DispatchGmmCombineDecode support bf16/float16 gmm1/gmm2 weight and support gmm weight with ND format (#6393 ) ### What this PR does / why we need it? 1. support ND format gmm weight input. Before this pr, gmm1_weight and gmm2_weight could only be passed as input to the DispatchGmmCombineDecode operator in NZ data format. After the modification, they are allowed to be passed in ND data format. 2. support bf16/float16 gmm weight The current PR modification enables the DispatchGmmCombineDecode operator to support non-W8A8 scenarios, allowing gmm1_weight and gmm2_weight to be passed as float16/bfloat16 which is correspond with input token data type. ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? - vLLM version: v0.14.1 - vLLM main: `dc917cceb8` Signed-off-by: lih827 <383084552@qq.com>	2026-02-12 10:37:41 +08:00
Ronald	f1ffb5fb19	[Feature] adapt to uva buffer and main2main (#6657 ) ### What this PR does / why we need it? vllm model runner v2 use uva buffer to prepare input data, but npu doesn't support uva yet, this pr implement a uvawrapper class to mimic gpu's uva backend. what's more, this pr make some modifications to adapt to the newer main branch. ### Does this PR introduce _any_ user-facing change? no ### How was this patch tested? - vLLM main: `13397841ab` --------- Signed-off-by: Ronald1995 <ronaldautomobile@163.com>	2026-02-12 10:36:31 +08:00
ZYang6263	56269eae0e	[BugFix] Fix AddRMSNormQuant not taking effect (#6620 ) ### What this PR does / why we need it? Fix the issue where, in graph mode, the fused `AddRMSNormQuant` operator does not take effect when there is no bias. ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM version: v0.15.0 - vLLM main: `d7e17aaacd` --------- Signed-off-by: ZYang6263 <zy626375@gmail.com>	2026-02-12 09:26:05 +08:00
Canlin Guo	052cc4e61b	[Docs] Fix GLM-5 deploy command (#6711 ) This pull request refines the GLM-5 deployment documentation by updating the Docker run command to include a more comprehensive set of device mappings and by removing an extraneous quantization flag from the `vllm serve` commands. These changes aim to correct and clarify the deployment instructions, ensuring users can successfully set up and run the GLM-5 model as intended. - vLLM version: v0.15.0 - vLLM main: `9562912cea` Signed-off-by: Canlin Guo <961750412@qq.com>	2026-02-12 08:55:48 +08:00
iiiklw	a0315f6697	[npugraph_ex]enable npugraph_ex by default (#6664 ) ### What this PR does / why we need it? This pull request enables the `npugraph_ex` backend by default to improve performance on Ascend NPUs, as proposed in the [RFC](https://github.com/vllm-project/vllm-ascend/issues/6214). ### Does this PR introduce _any_ user-facing change? Yes. `npugraph_ex` is now enabled by default. Users can disable it by setting `enable: false` in the `npugraph_ex_config` section of the `additional_config`. ### How was this patch tested? CI passed. The changes are covered by existing and new E2E tests (`test_aclgraph_accuracy.py`) and unit tests (`test_ascend_config.py`) that have been updated to reflect the new default behavior. The tests verify correctness and consistency with `npugraph_ex` enabled and disabled, as well as with the new static kernel option. Signed-off-by: huyuanquan1 <huyuanquan1@huawei.com> Co-authored-by: huyuanquan1 <huyuanquan1@huawei.com>	2026-02-12 08:44:06 +08:00
rika	b86ea66b0a	[doc]add GLM5.md (#6709 ) ### What this PR does / why we need it? Add GLM5 doc ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? - vLLM version: v0.15.0 - vLLM main: `9562912cea` Signed-off-by: nakairika <982275964@qq.com>	2026-02-12 04:00:40 +08:00
yydyzr	ff3a50d011	[Model] GLM5 adaptation (#6642 ) ### What this PR does / why we need it? GLM5 adaptation 1. use torch_npu.npu_lightning_indexer for GLM5 2. forbid eagle proposer when fullgraph mode is enabled because of bugs 3. add quatization config for GLM5 ### Does this PR introduce _any_ user-facing change? N/A ### How was this patch tested? by ci - vLLM main: `978a37c823` --------- Signed-off-by: yydyzr <liuyuncong1@huawei.com> Signed-off-by: shenchuxiaofugui <1311027364@qq.com> Co-authored-by: shenchuxiaofugui <1311027364@qq.com>	2026-02-11 22:22:22 +08:00
Zetong Li	140fcaffc3	[Bugfix] Update target probs to target logits in rejection sample (#6685 ) ### What this PR does / why we need it? This PR aims to update `target_probs` to `target_logits` in `rejection_sample`, for catching up with https://github.com/vllm-project/vllm/pull/32852. Otherwise, sampling with temperature will incur accuracy problem where tokens can be accepted or rejected unreasonably. ### Does this PR introduce _any_ user-facing change? N/A ### How was this patch tested? by ci - vLLM version: v0.15.0 - vLLM main: `13397841ab` Signed-off-by: Zetong Li <slippersss@126.com>	2026-02-11 21:31:40 +08:00

1 2 3 4 5 ...

2428 Commits