Files
enginex-ascend-910-vllm/vllm_ascend/patch/__init__.py
Sun Ruoxi 7f8a1b1f7a init v0.23.0
Signed-off-by: Sun Ruoxi <sunruoxi@4paradigm.com>
2026-08-27 15:11:51 +08:00

1103 lines
56 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

#
# Copyright (c) 2025 Huawei Technologies Co., Ltd. All Rights Reserved.
# This file is a part of the vllm-ascend project.
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
# ----------------------------------------------------------------------------------
# This module manage the patch for vllm. There are two folders in this module:
# - platform: contains the patches applied before worker starts. It's called by
# `vllm_ascend.utils.adapt_patch(is_global_patch=True)` in
# `vllm_ascend.platform.NPUPlatform.pre_register_and_update()` function.
# - worker: contains the patches applied when worker starts. It's called by
# `vllm_ascend.utils.adapt_patch(is_global_patch=False)` in
# each worker's `__init__` function.
#
# Once a new patch is added in vllm-ascend, please add the patch description into this file as well.
# ----------------------------------------------------------------------------------
# What's Patched and how it works:
# --------------------------------
# * Platform Patch:
# =================
# ** 1. File: platform/patch_distributed.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `torch.distributed.all_reduce`, `torch.distributed.broadcast`
# Why:
# tensor alignment for 310p
# How
# rewrite all_reduce and broadcast in torch.distributed
# Related PR (if no, explain why):
# No, not ready yet.
# Future Plan:
# Find a better way to support tensor alignment for 310p without this patch.
#
# ** 2. File: platform/patch_mamba_config.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.model_executor.models.config.HybridAttentionMambaModelConfig.verify_and_update_config`
# Why:
# block size is set to 16 in vLLM which is not supported by Ascend.
# How
# Set block size to 128 on npu.
# Related PR (if no, explain why):
# we'll fix this in vLLM soon.
# Future Plan:
# Remove this patch when vLLM merges the PR.
#
# ** 3. File: platform/patch_multiproc_executor.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.v1.executor.multiproc_executor.MultiprocExecutor`
# Why:
# vLLM create child process with daemon=True, which doesn't work with EPLB case, since EPLB will create
# a new process which is not allowed by daemon=True.
# How
# Set daemon=False in MultiprocExecutor.
# Related PR (if no, explain why):
# Find a way to support daemon=False in vLLM
# Future Plan:
# Remove this patch when vLLM fix the issue.
#
# ** 4. File: platform/patch_shm_broadcast.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.distributed.device_communicators.shm_broadcast.MessageQueue`
# Why:
# vLLM 0.23.0 local readers can wait indefinitely after a best-effort ZMQ
# notification is lost. A caller exception can also leave a read slot
# unreleased and block the writer.
# How:
# Replace timeout_ms and acquire_read with the implementations from the
# upstream fix. Idle waits are capped at five seconds, and read-slot cleanup
# runs in a finally block.
# Related PR (if no, explain why):
# https://github.com/vllm-project/vllm/pull/45224
# Future Plan:
# Remove this patch when the supported vLLM release includes PR #45224.
#
# ** 5. File: platform/patch_balance_schedule.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.v1.engine.core.EngineCoreProc.run_engine_core`
# `vllm.v1.core.sched.scheduler.Scheduler`
# Why:
# vLLM v1 scheduling currently enables chunkedprefill by default, which processes prefill and decode
# requests simultaneously in a single scheduling session. This can impact the overall system throughput
# and performance in some scenarios.
# How
# Set --additional-config '{"enable_balance_scheduling": true}' or
# set environmental variable VLLM_ASCEND_BALANCE_SCHEDULING=1 (deprecated).
# Related PR (if no, explain why):
# https://github.com/vllm-project/vllm/pull/29721
# Future Plan:
# Remove this patch when vLLM merge the PR.
#
# 2. Disable automatic scheduler preemption on vLLM 0.23.0 PD-disaggregated
# prefill nodes.
# Why:
# Async scheduling can finish the one-token prefill request while memory
# pressure preempts it, releasing delayed KV blocks before their block IDs
# are sent to the decode node.
# How:
# When a pure KV producer cannot allocate slots, the default, async, and
# profiling-chunk schedulers stop the current step without preempting a
# running request; forced prefix-cache reset is rejected until requests
# drain.
# Related PR (if no, explain why):
# No, this is a temporary vLLM 0.23.0 compatibility fix.
# Future Plan:
# Remove this patch after the upstream scheduler and KV connector handle
# async prefill completion and delayed block release atomically.
#
# ** 6. File: platform/patch_minimax_m2_config.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.config.model.ModelConfig._verify_quantization`
# Why:
# MiniMax-M2 fp8 checkpoints on NPU may fail upstream quantization validation.
# vllm-ascend needs to disable fp8 quantization and load bf16 dequantized
# weights in worker-side patches instead.
# How
# Monkey-patch `_verify_quantization` and intercept platform quantization
# verification to force `cfg.quantization=None` for MiniMax-M2 fp8 on NPU.
# Related PR (if no, explain why):
# No, upstream behavior differs across versions and needs discussion.
# Future Plan:
# Remove this patch once upstream supports MiniMax-M2 fp8 on NPU or provides
# a backend-safe validation / override mechanism.
#
# 2. `vllm.config.model.ModelConfig._verify_cuda_graph`
# Why:
# For MiniMax-M2 on NPU with ACL graph capture enabled, HCCL op expansion
# mode affects graph shape coverage. Users may forget to set it.
# How
# If user doesn't set it, set `HCCL_OP_EXPANSION_MODE=AIV` for this model
# and log a warning when a different value is detected.
# Related PR (if no, explain why):
# No, this is an environment-specific tuning knob.
# Future Plan:
# Remove this patch if upstream provides an official NPU graph-capture
# guidance / auto-configuration path for HCCL.
#
# 3. `vllm.config.speculative.SpeculativeConfig._verify_args`
# Why:
# Upstream vLLM's eagle3/extract_hidden_states restricts target model types
# via a whitelist. MiniMax-M2 should be allowed once the worker-side model
# can emit auxiliary hidden states.
# How
# Monkey-patch `_verify_args` to bypass only the whitelist ValueError for
# MiniMax model_type when method is eagle3/extract_hidden_states.
# SpeculativeConfig is a Pydantic dataclass (`@config`); init validation calls
# `__pydantic_decorators__.model_validators["_verify_args"].func`, so that
# `Decorator.func` must be replaced (not only `SpeculativeConfig._verify_args`),
# then `rebuild_dataclass(SpeculativeConfig, force=True)`.
# If `VllmConfig` was imported earlier, also `rebuild_dataclass(VllmConfig, ...)`
# so nested `speculative_config` validation does not use a stale schema.
# Related PR (if no, explain why):
# https://github.com/vllm-project/vllm/pull/37512
# Future Plan:
# Remove this patch once upstream whitelist includes MiniMax.
#
# 4. `vllm.model_executor.models.registry` (spec decode aliases)
# Why:
# Some Eagle3 draft checkpoints may declare a MiniMax-specific architecture
# string while reusing the shared Eagle3 implementation.
# How
# Register `Eagle3MiniMaxM2ForCausalLM` as an alias pointing to the
# existing Eagle3 implementation in the speculative decoding registry.
# Related PR (if no, explain why):
# https://github.com/vllm-project/vllm/pull/37512
# Future Plan:
# Drop the alias once upstream registry includes it or the checkpoint
# standardizes architecture strings.
#
# ** 7. File: platform/patch_minimax_usage_accounting.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.entrypoints.openai.chat_completion.serving.OpenAIServingChat`
# `vllm.reasoning.minimax_m2_reasoning_parser`
# Why:
# MiniMax-M2 chat usage accounting needs to report
# `completion_tokens_details.reasoning_tokens` for both streaming and
# non-streaming chat completions without slowing other reasoning models.
# How
# Monkey-patch MiniMax reasoning token counters and bind usage-accounting
# wrappers only on MiniMax chat-serving instances.
# Related PR (if no, explain why):
# https://github.com/vllm-project/vllm/pull/45701
# https://github.com/vllm-project/vllm/pull/45802
# Future Plan:
# Remove this patch after both upstream vLLM PRs are merged and the
# supported vLLM revision used by vLLM Ascend includes them through the
# regular main-to-main sync.
#
# ** 7a. File: platform/patch_glm_tool_call_streaming.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.entrypoints.openai.chat_completion.serving.OpenAIServingChat`
# Why:
# GLM tool-call streaming can emit final remaining-argument chunks with
# repeated tool-call metadata, and can combine terminal argument bytes with
# `finish_reason="tool_calls"` in the same SSE chunk.
# How
# Monkey-patch remaining-argument delta construction to emit only argument
# fragments by default, and split terminal argument chunks into an argument
# chunk followed by an empty finish chunk.
# Related PR (if no, explain why):
# https://github.com/vllm-project/vllm/issues/44098
# https://github.com/vllm-project/vllm/pull/44099
# https://github.com/vllm-project/vllm-ascend/issues/8327
# https://github.com/vllm-project/vllm-ascend/pull/8178
# Future Plan:
# Remove this patch once the supported vLLM version contains the upstream
# GLM tool-call final chunk fixes.
#
# ** 7b. File: platform/patch_glm47_tool_call_parser.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.tool_parsers.glm47_moe_tool_parser.Glm47MoeModelToolParser`
# Why:
# vLLM's GLM47 streaming parser can drop complete inline zero-argument
# tool calls such as `<tool_call>get_current_time</tool_call>`, while
# non-streaming parses the same output correctly.
# How
# Monkey-patch GLM47 tool-call region extraction so complete inline
# zero-argument regions are normalized for the existing streaming name
# extractor without emitting partial names for incomplete regions.
# Related PR (if no, explain why):
# https://github.com/vllm-project/vllm/issues/44326
# https://github.com/vllm-project/vllm/pull/44327
# Future Plan:
# Remove this patch once the supported vLLM version contains the upstream
# GLM47 inline zero-argument streaming parser fix.
#
# ** 10a. File: platform/patch_kv_cache_utils.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.v1.core.kv_cache_utils.resolve_kv_cache_block_sizes`
# `vllm.v1.engine.core.resolve_kv_cache_block_sizes`
# Why:
# vLLM PR #40860 added a restriction that hybrid KV cache groups with
# multiple block sizes do not support context parallelism (dcp/pcp > 1).
# This restriction is correct for CUDA but not for Ascend, which
# implements context parallelism for MLA and SWA-MLA layers separately.
# How
# Monkey-patch resolve_kv_cache_block_sizes to handle the multiple-groups
# + CP case by returning lcm(block_sizes) * dcp * pcp as scheduler_block_size
# instead of raising ValueError.
# Related PR (if no, explain why):
# vLLM PR #40860 ([Feat] DeepSeek V4 Rebased).
# Future Plan:
# Remove this patch once upstream vLLM supports hybrid KV cache + CP for
# non-CUDA backends, or exposes a platform hook for this behavior.
#
# ** 10ab. File: worker/patch_v2/patch_attn_utils.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.v1.worker.gpu.attn_utils.get_kv_cache_spec`
# Why:
# The current v2 worker still goes through the shared upstream v1 helper
# to build KV cache specs. For Ascend MLA layers that helper returns the
# generic `MLAAttentionSpec`, but NPU-side cache allocation and reshape
# logic expects `AscendMLAAttentionSpec`.
# How
# Monkey-patch `get_kv_cache_spec` so regular attention layers keep the
# upstream behavior while MLA layers are rewritten to
# `AscendMLAAttentionSpec`, including the FA-quant head-size adjustment.
# Related PR (if no, explain why):
# No. This is a plugin-side compatibility patch for the current upstream
# helper path.
# Future Plan:
# Remove this patch once upstream adds a backend hook for KV cache spec
# construction or v2 worker no longer depends on the shared v1 helper.
#
# ** 10. File: platform/patch_profiling_chunk.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.v1.engine.core.EngineCore.__init__`
# 2. `vllm.v1.engine.core.EngineCoreProc.run_engine_core`
# 3. `Scheduler.update_from_output` (scheduler class, wrapped when profiling chunk is enabled)
# Why:
# Profiling-based dynamic chunk sizing needs to run a one-shot profiling pass
# after `model_executor` is ready, and to feed per-step execution latency back
# into `ProfilingChunkManager` so the history-aware chunk predictor can refine
# online. In multiprocessing `spawn` mode the child process starts a fresh
# interpreter, so monkey-patches applied in the parent are lost unless the
# subprocess entry point re-applies them before any `EngineCore` is created.
# How
# Replace `EngineCore.__init__` to call `scheduler.run_profiling_chunk_init`
# when present, then wrap `scheduler.update_from_output` once per process to
# read `model_output.execution_time_ms` and `scheduler_output` token/chunk
# metadata and call `ProfilingChunkManager.record_batch_execution_time` (and
# bootstrap target latency for the first chunk when needed). Replace
# `EngineCoreProc.run_engine_core` so importing this module in the child
# re-runs the idempotent patch helper before delegating to the original
# implementation.
# Related PR (if no, explain why):
# No, vllm-ascend-specific profiling / scheduling integration.
# Future Plan:
# Remove or narrow this patch if upstream exposes stable hooks for backend
# profiling startup and per-step timing callbacks without monkey-patching
# `EngineCore` and the multiprocess entry point.
#
# ** 10b. File: platform/patch_pp_mtp.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.config.model.ModelConfig.verify_with_parallel_config`
# Why:
# Local Eagle/MTP drafters are loaded on the last PP stage rather than
# partitioned across all PP ranks. Upstream `ModelConfig.verify_with_parallel_config`
# validates against `pipeline_parallel_size`, which fails for these drafters
# since they run locally with effective PP=1.
# How
# Monkey-patch `verify_with_parallel_config` to detect Eagle/MTP drafter
# models (by `model_type` and `architectures`) when `runner="draft"` and
# `pipeline_parallel_size > 1`. For such configs, call the original verify
# with a patched `pipeline_parallel_size=1` copy, preserving normal target-model
# validation for non-drafter models.
# Related PR (if no, explain why):
# Backport of local vLLM PP+MTP branch changes.
# Future Plan:
# Remove this patch once upstream vLLM's `ModelConfig.verify_with_parallel_config`
# supports local drafter models with PP > 1, or moves the PP validation to a
# separate hook that can be overridden per-model-type.
#
# ** 11. File: platform/patch_tool_choice_none_content.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.entrypoints.openai.chat_completion.protocol.ChatCompletionResponse`
# `vllm.entrypoints.openai.chat_completion.protocol.ChatCompletionStreamResponse`
# Why:
# vLLM v0.23.0 can serialize empty `tool_calls: []` fields for content-only
# OpenAI chat responses / streaming deltas, while OpenAI-compatible SDKs
# expect those empty fields to be omitted so clients see `tool_calls=None`.
# How
# Wrap `model_dump` / `model_dump_json` for chat response payloads and drop
# empty `tool_calls` lists from `message` / `delta` objects.
# Related PR (if no, explain why):
# https://github.com/vllm-project/vllm/pull/44105
# Future Plan:
# Remove this patch once the supported vLLM version contains PR #44105.
#
# ** 12. File: platform/patch_deepseek_v4_tool_call_parser.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.tool_parsers.deepseekv4_tool_parser.DeepSeekV4ToolParser`
# Why:
# Upstream vLLM now includes DeepSeek V4 tokenizer/renderer/reasoning
# registration, but its streaming tool-call delta parsing does not guarantee
# incremental `arguments` emission for long argument payloads.
# How:
# Monkey-patch `DeepSeekV4ToolParser` stream parsing to emit tool-call
# metadata in the first delta and stream argument fragments incrementally.
# Related PR (if no, explain why):
# Upstream vLLM main behavior as of current runtime.
# Future Plan:
# Remove this patch if upstream streaming behavior is updated to satisfy the
# same DeepSeek DSML incrementality contract.
#
# ** 12a. File: platform/patch_minimax_m2_tool_call_parser.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.tool_parsers.minimax_m2_tool_parser.MinimaxM2ToolParser`
# Why:
# vLLM 0.21.0 only emits MiniMax-M2 tool-call arguments after a complete
# `<invoke>...</invoke>` block, so long arguments are buffered instead of
# streamed incrementally.
# How:
# Monkey-patch the MiniMax-M2 parser to emit the tool name once the
# `<invoke name=...>` header is available and then stream partial
# `<parameter>` values as JSON argument fragments.
# Related PR (if no, explain why):
# https://github.com/vllm-project/vllm/pull/40253
# https://github.com/vllm-project/vllm/pull/40298
# Future Plan:
# Remove this patch once the supported vLLM version contains the upstream
# MiniMax-M2 incremental tool-call streaming fix.
#
# ** 12b. File: platform/patch_structured_output.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.sampling_params.SamplingParams._validate_structured_outputs`
# `vllm.v1.structured_output.StructuredOutputManager.grammar_init`
# Why:
# V1 structured outputs use one engine-level backend, while `backend=auto`
# resolves the backend per request. After one request initializes
# `xgrammar`, a later request that resolves to `guidance` can still reach
# the initialized `xgrammar` backend and crash during grammar compilation.
# How:
# Record the first resolved backend on the structured-output config and
# reject later requests that resolve to a different backend. Also guard
# `grammar_init` so requests that bypass API-side validation fail before
# backend grammar compilation.
# Related PR (if no, explain why):
# https://github.com/vllm-project/vllm/issues/43920
# https://github.com/vllm-project/vllm/pull/44401
# Future Plan:
# Remove this patch once upstream vLLM either enforces backend consistency
# before grammar compilation or safely handles mixed-backend grammar
# failures without killing the engine.
#
# ** 13. File: platform/patch_camem_allocator.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.config.model.is_cumem_allocator_available`
# Why:
# Upstream vLLM main enables and validates the CUDA/ROCm CuMem allocator
# when `enable_sleep_mode=True`. Ascend implements sleep mode with its own
# CaMem allocator, so the upstream CuMem-only availability check fails
# during `ModelConfig` validation before Ascend worker code can run.
# How:
# Treat Ascend's platform sleep allocator as satisfying the allocator
# availability check, while preserving the original vLLM CuMem check as
# fallback.
# Related PR (if no, explain why):
# No, this maps an upstream CUDA/ROCm allocator validation to Ascend's
# backend-specific CaMem implementation.
# Future Plan:
# Remove this patch if upstream exposes a platform allocator capability hook
# for sleep mode validation.
#
# ** 15. File: platform/patch_weight_transfer_engine.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.distributed.weight_transfer.factory.WeightTransferEngineFactory._registry["nccl"]`
# Why:
# Upstream vLLM's WeightTransferConfig.backend is a pydantic Literal["nccl", "ipc"]
# which does not accept "hccl". On Ascend NPU, NCCL is unavailable and HCCL must
# be used for trainer-to-worker weight broadcasting.
# How
# Replace the "nccl" factory entry with a lambda that returns
# HCCLWeightTransferEngine. Users pass the already-accepted "nccl" string
# (e.g. --weight-transfer-config '{"backend": "nccl"}') and the factory
# resolves it to the HCCL engine at runtime.
# Related PR (if no, explain why):
# No. Adding "hccl" to the Literal requires modifying pydantic core schemas,
# which is fragile across pydantic versions.
# Future Plan:
# Remove this patch when upstream vLLM relaxes the Literal type to str or
# provides an extension point for out-of-tree weight transfer backends.
# 2. `vllm.distributed.weight_transfer.factory.WeightTransferEngineFactory._registry["ipc"]`
# Why:
# The "ipc" backend must resolve to NPUIPCWeightTransferEngine on Ascend NPU.
# However, this patch runs during global plugin patching - extremely early in
# startup, before any weight transfer backend is selected. Importing the IPC
# engine eagerly pulls in vllm.distributed.weight_transfer.ipc_engine, which
# does `import ray` at module top level. Since ray is an optional dependency,
# its absence aborts the whole vllm_ascend plugin load and crashes every
# `vllm serve` invocation - even workloads that never use weight transfer.
# How
# Register a lazy loader function (instead of an eager import) that imports
# and returns NPUIPCWeightTransferEngine only when create_engine() is invoked
# for the "ipc" backend. This matches the factory's zero-arg-callable
# lazy-loading contract, so the ray-importing module is loaded only when ipc
# is actually requested. (HCCL keeps its eager import - it never imports ray.)
# Related PR (if no, explain why):
# No. The eager `import ray` lives in upstream vLLM's ipc_engine module; the
# lazy loader is a local workaround until upstream defers that import.
# Future Plan:
# Remove this workaround once upstream vLLM stops importing ray at module top
# level in vllm.distributed.weight_transfer.ipc_engine (e.g. defers it into
# the code path that actually needs ray), so importing the IPC engine no
# longer requires the optional ray dependency.
#
# ** 15. File: platform/patch_kv_cache_coordinator.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.v1.core.kv_cache_coordinator.HybridKVCacheCoordinator.find_longest_cache_hit_per_group`
# Why:
# In PD disaggregation with hybrid Mamba models, the D side receives
# FullAttention KV blocks from the P side but has no local prefix-cache
# hit for Mamba groups. Upstream's min-reduction across all KV groups
# collapses the FullAttention hit length to 0, preventing partial
# FullAttention-only prefix cache reuse on the D side.
# How:
# For Mamba hybrid models,
# num_new_local_computed_tokens should be the FA hit
# length. This value is passed to the connector's
# get_num_new_matched_tokens which computes:
# external = total - local_computed.
# Using the FA hit skips re-transferring FA blocks
# already cached on D-side.
# Related PR (if no, explain why):
# https://github.com/vllm-project/vllm/pull/42524
# https://github.com/vllm-project/vllm/pull/44243
# Future Plan:
# Remove this patch when vLLM PR #42524 and #44243 is included in the supported
# upstream vLLM version.
#
# * Worker Patch:
# ===============
#
# ** 1. File: worker/patch_distributed.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.distributed.parallel_state.GroupCoordinator`
# Why:
# vllm doesn't support all_to_all for GroupCoordinator.
# How
# Add all_to_all implementation for GroupCoordinator.
# Related PR (if no, explain why):
# No, we should use vlLM all2all manager to support all_to_all for npu.
# Future Plan:
# Remove this patch when the refactor of all2all manager is done.
#
# ** 3. File: worker/patch_triton.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.model_executor.layers.mamba.ops`, `vllm.model_executor.layers.fla.ops`,
# `vllm.v1.worker.gpu.sample.gumbel.gumbel_sample`
# Why:
# triton ops in vLLM perform not good on NPU. And there is no dispatch mechanism for triton ops.
# How
# override triton ops in vLLM with ascend implementation
# Related PR (if no, explain why):
# Let vLLM support triton ops dispatch.
# Future Plan:
# Remove this patch when vLLM support the dispatch function.
#
# 2. `triton.next_power_of_2`
# Why:
# The Triton version bundled with torch_npu on Ascend NPU
# does not include `next_power_of_2`, which is called by
# upstream vLLM and vLLM-Ascend code in 94+ places.
# Additionally, when Triton is not available (HAS_TRITON=False),
# vLLM uses TritonPlaceholder which also lacks this function.
# How
# Import `triton` from vllm.triton_utils (which handles both
# real Triton and TritonPlaceholder) and inject `next_power_of_2`
# onto the module, reusing `vllm.utils.math_utils.next_power_of_2`.
# Related PR (if no, explain why):
# No, torch_npu Triton compatibility issue.
# Future Plan:
# Remove this patch when torch_npu's Triton includes
# next_power_of_2 or when vLLM no longer calls triton.next_power_of_2.
#
# ** 4. File: worker/patch_qwen3_next_mtp.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.v1.worker.utils.bind_kv_cache`
# Why:
# 'bind_kv_cache' func will raise an exception when current_platform is npu.
# How
# Replace with a new bind_kv_cache.
# Skip the raise.
# Related PR (if no, explain why):
# It need discuss.
# Future Plan:
# Remove this patch after discussing with vllm community and adapting bind_kv_cache to npu.
#
# ** 5. File: worker/patch_rejection_sampler.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.v1.sample.rejection_sampler`
# Why:
# - some functions from `rejection_sampler` are not supported or slow on npu.
# How
# - add npu_top_k_top_p to 'apply_sampling_constraints' func
# - add custom triton kernel to `expand_batch_to_tokens` and `rejection_sample`
# Related PR (if no, explain why):
# Let vLLM support triton ops dispatch.
# Future Plan:
# 1. make these functions as class func of RejectionSampler, create AscendRejectionSampler
# to override them, then delete the patch file `worker/patch_rejection_sampler.py`.
# 2. make these functions as costom op, then remove AscendRejectionSampler
#
## ** 6. File: worker/patch_module.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.v1.attention.backends.gdn_attn.torch.argsort`
# Why:
# 1. 'torch.argsort' func of npu does not support bool.
# 2. Without `stable=True`, the output will have a lot of redundant tokens.
# How
# Replace with a new torch.argsort that will cast the input to torch.int32
# and do stable sort.
# Related PR (if no, explain why):
# 1. It depends on torch_npu.
# 2. https://github.com/vllm-project/vllm/pull/30632
# Future Plan:
# Remove this patch when bool is supported in 'torch.argsort' func of npu.
# Make 'torch.argsort' in `vllm.v1.attention.backends.gdn_attn` be stable.
# 2. `vllm_ascend.ops.gdn_attn_builder.AscendGDNAttentionMetadataBuilder.build`
# Why:
# Qwen3.5/Qwen3Next GDN Decode/Specific Decode on NPU needs prebuilt varlen chunk metadata
# to avoid forward-time host round-trips that break async scheduling.
# How
# Override the GDN attention metadata builder for Ascend backend and attach
# prebuilt device metadata bundle onto the returned attention metadata object.
# Future Plan:
# Remove this patch when upstream exposes a backend hook for extending GDN
# metadata or when the optimization is accepted upstream directly.
#
# ** 8. File: worker/patch_qwen3_next.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.model_executor.models.qwen3_next.Qwen3NextGatedDeltaNet.forward`
# Why:
# The Qwen3Next GatedDeltaNet forward cannot directly add custom operators.
# How
# Add a branch in Qwen3NextGatedDeltaNet.forward to adapt to fused_qkvzba_split_reshape_cat.
# Related PR (if no, explain why):
# https://github.com/vllm-project/vllm/pull/30863
# Future Plan:
# Remove this patch when vLLM support these operators.
#
# 2. `vllm.model_executor.models.qwen3_next.Qwen3NextGatedDeltaNet._forward_core`
# Why:
# triton ops fused_recurrent_gated_delta_rule and fused_gdn_gating in vLLM perform not good on NPU.
# How
# add a new fused triton ops in vLLM with ascend implementation.
# Related PR (if no, explain why):
# https://github.com/vllm-project/vllm/pull/30860
# Future Plan:
# Remove this patch when vLLM support these operators.
#
# 3. `vllm.model_executor.models.qwen3_next.Qwen3NextGatedDeltaNet._forward_core`
# Why:
# The Qwen3Next GatedDeltaNet _forward_core cannot directly add custom operators.
# How
# Add a branch in Qwen3NextGatedDeltaNet._forward_core to adapt to fused_gdn_gating_patch.
# Related PR (if no, explain why):
# https://github.com/vllm-project/vllm/pull/31002
# Future Plan:
# Remove this patch when vLLM support these operators.
#
# ** 10. File: worker/patch_qwen3vl.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.model_executor.models.qwen3_vl.Qwen3VLForConditionalGeneration._get_deepstack_input_embeds`
# Why:
# support flash comm v1 for qwen3vl.
# How
# override _get_deepstack_input_embeds method with the flash comm v1 implementation.
# Future Plan:
# Remove this patch when https://github.com/vllm-project/vllm-ascend/issues/5712 is completed.
# 2. `vllm.model_executor.models.qwen3_vl_moe.Qwen3MoeLLMForCausalLM.start_layer`,
# `vllm.model_executor.models.qwen3_vl_moe.Qwen3MoeLLMForCausalLM.end_layer`
# Why:
# Qwen3-VL-MoE checks the language-model pipeline boundary on non-first
# PP ranks, but Qwen3MoeLLMForCausalLM keeps start_layer/end_layer only
# on the inner model object.
# How:
# Expose start_layer/end_layer properties on Qwen3MoeLLMForCausalLM and
# forward them to the inner model.
# Future Plan:
# Remove this patch when upstream vLLM exposes these PP layer boundaries
# on the Qwen3-VL-MoE language-model wrapper.
#
# ** 11. File: worker/patch_npugraph_ex_triton.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `npugraph_ex.core._concrete_graph.ValuePack`,
# `npugraph_ex.npu_fx_compiler._unpack_meta`,
# `npugraph_ex.npu_fx_compiler._NpuGraphConverter._unpack_npu`
# Why:
# In the Triton scenario, npugraph_ex backend needs to process the value pack of the input parameters.
# How
# Supplement the relevant processing logic through patches.
# Related PR (if no, explain why):
# https://gitcode.com/Ascend/torchair/pull/2575
# Future Plan:
# Remove this patch when the PTA version used by vllm-ascend has been upgraded.
#
# ** 12. File: worker/patch_v2/patch_uva.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.v1.worker.gpu.states.UvaBuffer`
# Why:
# ASCEND NPUs do not support UVA yet, so we need to wrap it in vLLM.
# How
# make UvaBuffer a dummy class, mimic the interface of vllm UvaBuffer.
# Future Plan:
# Remove this patch when NPU support UVA.
#
# ** 13. File: worker/patch_kimi_k25.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.model_executor.models.kimi_k25_vit.Learnable2DInterpPosEmbDivided_fixed.forward`
# Why:
# The forward method uses interpolate with ops not supported on NPU.
# How
# Replace with a new forward that uses CPU for interpolate when shape mismatch,
# and use get_rope_shape to handle the rope shape interpolation.
# Future Plan:
# Remove this patch when vLLM aligns with the latest main.
#
# ** 14. File: worker/patch_draft_quarot.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.model_executor.models.llama_eagle3.Eagle3LlamaForCausalLM.load_weights`
# Why:
# vllm-ascend reused the loading logic of drafter model from vllm,
# but vllm doesn't need to apply to Ascend quantization.
# How
# Dynamically replace the `load_weights` function at runtime,
# and fix `target_config` into the new implementation with a closure.
# Related PR (if no, explain why):
# https://github.com/vllm-project/vllm/pull/36225
# Future Plan:
# Remove this patch when vLLM merges the PR.
#
# ** 15. File: worker/patch_minimax_m2.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.model_executor.models.minimax_m2.MiniMaxM2MoE.forward`
# Why:
# MiniMax-M2 routing should keep router logits in fp32 on NPU.
# How
# Replace the forward to cast hidden states to fp32 before the gate.
# Related PR (if no, explain why):
# No, model-specific behavior.
# Future Plan:
# Remove this patch once upstream behavior is sufficient for Ascend.
#
# 2. `vllm.model_executor.models.minimax_m2.MiniMaxM2Attention.forward`
# Why:
# MiniMax-M2 attention benefits from the NPU fused split-qkv + RMSNorm + rope
# kernel path.
# How
# Replace `forward` to call `torch.ops.vllm.split_qkv_tp_rmsnorm_rope` before
# the upstream attention and output projection steps.
# Related PR (if no, explain why):
# No, backend-specific fused kernel path.
# Future Plan:
# Remove this patch when upstream exposes a backend dispatch path for this
# fused attention preparation.
#
# 3. `vllm.model_executor.models.minimax_m2.MiniMaxM2Model.load_weights`
# Why:
# MiniMax-M2 fp8 checkpoints may store fp8 weights with per-block inverse
# scales. On NPU we load bf16 weights by dequantizing at load time.
# How
# Inject fp8 dequant helpers and wrap `load_weights` to convert fp8 weight +
# `weight_scale_inv` pairs into bf16 blocks before delegating to upstream.
# Related PR (if no, explain why):
# No, fp8 load format and backend constraints are model/backend specific.
# Future Plan:
# Remove this patch when upstream supports MiniMax-M2 fp8 loading on NPU.
#
# ** 16. File: worker/patch_minimax_m2_linear_attn.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.model_executor.layers.mamba.linear_attn.MiniMaxText01RMSNormTP.__init__`
# `vllm.model_executor.layers.mamba.linear_attn.MiniMaxText01RMSNormTP.weight_loader`
# Why:
# MiniMax-M2 linear attention RMSNorm needs weight sharding that can follow
# TP layout (and sometimes kv-head replication) on NPU.
# How
# Override `__init__` to parameterize weight shard world/rank and install a
# sharded `weight_loader` implementation.
# Related PR (if no, explain why):
# No, upstream API surface differs across versions.
# Future Plan:
# Remove this patch when upstream exposes stable sharding hooks for this layer.
#
# 2. `vllm.model_executor.layers.mamba.linear_attn.MiniMaxText01RMSNormTP.forward_qk`
# (or older `_normalize_qk`)
# Why:
# q/k norm for linear attention is performance-sensitive. On NPU, a fused
# rms_norm kernel is faster and TP needs a global rstd correction.
# How
# Replace q/k normalization with NPU rms_norm fast path and TP-global rstd
# correction; fall back to upstream implementation on non-NPU.
# Related PR (if no, explain why):
# No, backend-specific optimization.
# Future Plan:
# Remove this patch when upstream adds a backend dispatch path for q/k norm.
#
# ** 17. File: worker/patch_qwen3_5.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.model_executor.models.qwen3_5.Qwen3_5GatedDeltaNet._forward_core`
# Why:
# The class Qwen3_5GatedDeltaNet reuse the `_forward_core` method of Qwen3NextGatedDeltaNet,
# but the ascendC ops of Qwen3NextGatedDeltaNet do not support ssm_state with float32 format.
# How
# patch Qwen3_5GatedDeltaNet._forward_core to use triton ops like `fused_recurrent_gated_delta_rule`.
# Future Plan:
# Remove this patch when all ops in _forward_core support both Qwen3_5 and Qwen3Next.
#
# ** 17a. File: worker/patch_idex_310.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.model_executor.layers.fla.ops.index.prepare_chunk_indices`
# `vllm.model_executor.layers.fla.ops.index.prepare_chunk_offsets`
# Why:
# 310P uses Ascend-friendly chunk index helpers for Qwen GDN prefill.
# How:
# Replace upstream FLA chunk index helper functions with 310P implementations.
#
# 2. `vllm_ascend.spec_decode.llm_base_proposer.AscendSpecDecodeBaseProposer.set_inputs_first_pass`
# Why:
# 310P needs to protect the tail slot during MTP input_ids shift to avoid
# GatherV2 corruption from persistent drafter input buffers.
# How:
# Reuse the 310P proposer implementation for the first-pass input shift.
#
# 3. `vllm.model_executor.layers.mamba.gdn.qwen_gdn_linear_attn.QwenGatedDeltaNetAttention`
# Why:
# Qwen GDN needs 310P-specific state helpers, forward core, state dtype,
# and attention backend/builder wiring.
# How:
# Patch Qwen GDN methods to use Ascend GDN implementations and the 310P
# GDN attention backend. RC devices also route upstream GDNAttentionBackend
# to the 310P metadata builder.
# Related PR (if no, explain why):
# No, 310P custom operator and backend behavior are vllm-ascend specific.
# Future Plan:
# Remove this patch when upstream exposes stable hooks for 310P GDN
# chunk metadata, spec-decode input layout, and backend selection.
#
# ** 18. File: worker/patch_cudagraph.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.v1.cudagraph_dispatcher.CudagraphDispatcher._create_padded_batch_descriptor`
# Why:
# vllm's FULL mode will cause error, we use a patch to avoid it.
# After that, FULL can be enable now.
# How
# Dynamically replace the `_create_padded_batch_descriptor` function at runtime,
# and change the condition of if.
# Related PR (if no, explain why):
# https://github.com/vllm-project/vllm/pull/34880
# Future Plan:
# Remove this patch when vLLM merges the PR.
#
# ** 19. File: worker/patch_deepseek_mtp.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.model_executor.models.deepseek_v2.get_spec_layer_idx_from_weight_name` and
# `vllm.model_executor.models.deepseek_mtp.get_spec_layer_idx_from_weight_name`
# Why:
# When GLM5 uses rotary quant in vllm-ascend, the MTP layer needs to load an extra weight
# named `rot.weight`.
# How
# If weight name starts with `rot`, return `layer_id + i` like other tensors in MTP layer.
# Related PR (if no, explain why):
# Rotary quant is a unique feature of vllm-ascend.
# Future Plan:
# Remove this patch when vllm supports rotary quant or pluggable `MultiTokenPredictorLayer`.
# 2. `vllm.model_executor.models.deepseek_mtp.DeepSeekMultiTokenPredictorLayer`
# Why:
# When GLM5 uses rotary quant in vllm-ascend, the `previous_hidden_states` does not .
# How
# If the target model uses rotary quant, a new linear operation is added before `ehnorm`.
# Related PR (if no, explain why):
# Rotary quant is a unique feature of vllm-ascend.
# Future Plan:
# Remove this patch when vllm supports rotary quant or pluggable `MultiTokenPredictorLayer`.
# 3. `vllm.model_executor.models.deepseek_mtp.DeepSeekMTP._rewrite_spec_layer_name`
# Why:
# Rename `rot.weight` to match the format of weights in `DeepSeekMTP`.
# How
# If the weight name is `rot`, rename it to `model.layers.{spec_layer}.rot.weight`.
# Related PR (if no, explain why):
# Rotary quant is a unique feature of vllm-ascend.
# Future Plan:
# Remove this patch when vllm supports rotary quant or pluggable `MultiTokenPredictorLayer`.
# 4. `vllm.model_executor.models.deepseek_v2.GlmMoeDsaForCausalLM.load_weights`
# Why:
# After vllm PR #41706, GlmMoeDsaForCausalLM.load_weights uses `AutoWeightsLoader` which
# does not skip `rot.weight`, and will cause ValueError while loading weights.
# How
# Use the `skip_prefixes` parameter to skip certain weight tensors.
# Related PR (if no, explain why):
# https://github.com/vllm-project/vllm/pull/41706
# Future Plan:
# Remove this patch when vllm supports rotary quant or pluggable `MultiTokenPredictorLayer`.
# ** 19a. File: worker/patch_deepseek_v2.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.model_executor.models.deepseek_v2.DeepseekV2MLAAttention.__init__`
# Why:
# GLM-5.2 checkpoints omit `Indexer` weights on shared-indexer layers,
# while GLM-5.1 IndexCache overrides only skip top-k computation and keep
# per-layer `Indexer` weights. Treating both layouts alike breaks GLM-5.1
# weight loading.
# How:
# Skip `Indexer` construction only when the layer both skips top-k and is
# explicitly marked `shared` in `indexer_types`. MTP layers always retain
# a complete `Indexer`.
# Related PR (if no, explain why):
# https://github.com/vllm-project/vllm/pull/45895
# Future Plan:
# Remove this patch when vLLM Ascend depends on a vLLM version that includes
# PR #45895.
#
# ** 19b. File: worker/model_runner_v1.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `NPUModelRunner._check_and_update_cudagraph_mode`
# Why:
# The upstream `GPUModelRunner._check_and_update_cudagraph_mode` initializes
# drafter cudagraph keys unconditionally, but in PP mode only the last rank
# loads the drafter. The previous hacky workaround temporarily set
# `self.speculative_config = None` to bypass super()'s drafter init, then
# restored it and called a separate `_maybe_initialize_drafter_cudagraph_keys`
# helper. This state-mutation pattern is fragile and hard to maintain.
# How
# Directly inline the upstream cudagraph mode resolution logic with Ascend-specific
# additions: wrap `resolve_cudagraph_mode_and_sizes` with `update_pass_config` for
# `enable_sp`, add PP last-rank guard for drafter initialization, and call
# `set_graph_params`/`set_draft_graph_params` for ACL graph params. Remove the
# `_maybe_initialize_drafter_cudagraph_keys` helper entirely.
# Related PR (if no, explain why):
# No, cleaner PP+MTP support without speculative_config state mutation.
# Future Plan:
# Remove this override once upstream exposes a hook for drafter cudagraph key
# initialization that respects PP rank boundaries.
# ** 20. File: worker/patch_mamba_utils.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.v1.worker.mamba_utils.batch_memcpy_kernel = batch_memcpy_kernel`
# Why:
# Oringnal batch_memcpy_kernel implemented in vLLM might encounter bugs when running on
# Ascend hardwares.
# How
# patch to fix related bugs.
# Future Plan:
# Remove this patch when:
# (1) oringnal batch_memcpy_kernel can run on Ascend hardware.
# or
# (2) design a dispatch mechanism for batch_memcpy_kernel.
# 2. `vllm.v1.worker.mamba_utils.batch_memcpy = batch_memcpy`
# Why:
# vLLM use BLOCK_SIZE 1024 for batch_memcpy_kernel. This results in suboptimal performance
# on Ascend hardwares.
# How
# patch to change BLOCK_SIZE to 8192.
# Future Plan:
# Remove this patch when:
# design a dispatch mechanism for batch_memcpy_kernel.
# 3. `mamba_utils.preprocess_mamba = preprocess_mamba`
# Why:
# 1. preprocess_mamba has a assert logic, cause kv transfer call fails
# 2. preprocess_mamba copy the state of previous step to the last block before kv transfer load
# How:
# 1. patch to remove assert
# 2. path to only collect copy metadata in preprocess_mamba(and do actual copy after kv transfer load).
# Future Plan:
# Remove this patch when:
# vLLM itself supports kv transfer for mamba
# ** 21. File: worker/patch_weight_utils.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.model_executor.models.deepseek_v2.DeepseekV2ForCausalLM.load_weights`
# Why:
# The C8 weight quantized by modelslim will modify the model structure,
# and the scale and offset required for kvcache quantization will increase.
# In addition, the names of the quantization parameters are different from
# those in the community.
# How
# we have enhanced the maybe_remap_kv_scale_name function.
# Future Plan:
# The maybe_remap_kv_scale_name function of the community is reconstructed to support
# multiple backends.
# ** 21b. File: worker/patch_process_weights_after_loading.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.model_executor.model_loader.utils.process_weights_after_loading`
# `vllm.model_executor.model_loader.base_loader.process_weights_after_loading`
# and imported references in vllm-ascend model loaders
# Why:
# DSA attention is implemented in vllm-ascend as the plugin layer
# `DSAAttention`. Upstream vLLM only runs post-load attention weight
# processing for built-in attention classes, so
# `DSAAttention.process_weights_after_loading()` is skipped in the
# original loader flow. DSV4 DSA-CP o-proj TP initialization must run in
# this post-load phase rather than being initialized lazily in forward.
# How:
# Rebind the upstream `process_weights_after_loading` helper, including
# already-imported loader references, so `DSAAttention` participates in
# the same post-load traversal while preserving the original quant-method
# and torchao reload behavior.
# Related PR (if no, explain why):
# https://github.com/vllm-project/vllm-ascend/pull/10694
# https://github.com/vllm-project/vllm/pull/46828
# Future Plan:
# Remove this patch once the supported vLLM version includes PR #46828.
# Then register `DSAAttention` through vLLM's post-load weight-processing
# registry instead of monkey-patching model-loader helpers.
# ** 22. File: worker/patch_v2/patch_input_batch.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.v1.worker.gpu.input_batch.InputBatch`
# Why:
# vllm use InputBatch to make dummy tensors. in `model_runner.py` and `cudagraph_utils.py`
# which make it difficult to inherit from vllm methods.
# How
# replace InputBatch with AscendInputBatch.
# Future Plan:
# remove this patch when vLLM-ascend's make_dummy behavior aligns with vLLM.
# ** 23. File: worker/patch_v2/patch_block_table.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.v1.worker.gpu.block_table.BlockTables`
# Why:
## vllm-ascend need to initialize slot mapping as torch.int32 dtype,
# but vllm default is torch.int64 dtype.
# How
# replace BlockTables with AscendBlockTables which initialize slot mapping
# as torch.int32 dtype.
# Future Plan:
# remove this patch when vLLM-ascend's BlockTables can initialize
# slot mapping as torch.int64 dtype.
# ** 24. File: worker/patch_v2/patch_model_state.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.v1.worker.gpu.model_states.default.init_model_state`
# Why:
## vllm's prepare_attn in ModelState is different from vllm,
# we need to override init_model_state.
# How
# Define AscendModelState and initialize it in init_model_state.
# Future Plan:
# remove this when vllm-ascend's attention metadata is align with vllm.
# ** 25. File: worker/patch_v2/patch_triton.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.v1.worker.gpu.sample.logprob`, `vllm.v1.worker.gpu.sample.penalties.apply_penalties`,
# `vllm.v1.worker.gpu.sample.gumbel.gumbel_sample`
# Why:
# triton ops in vLLM perform not good on NPU. And there is no dispatch mechanism for triton ops.
# How
# override triton ops in vLLM with ascend implementation
# Related PR (if no, explain why):
# Let vLLM support triton ops dispatch.
# Future Plan:
# Remove this patch when vLLM support the dispatch function.
#
# ** 26. File: worker/patch_gqa_c8.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.model_executor.models.qwen3.Qwen3ForCausalLM.load_weights`
# Why:
# The GQA W8A8C8 model stores per-channel KV cache scales and offsets
# (k_cache_scale, k_cache_offset, v_cache_scale, v_cache_offset) under
# weight names that AutoWeightsLoader does not recognise and would
# silently discard. Without these scales the INT8 KV cache cannot be
# dequantised correctly at inference time.
# How:
# Wrap load_weights to intercept the C8 scale/offset tensors before they
# reach the base loader. Each intercepted tensor is routed to the
# corresponding nn.Parameter via its weight_loader, then excluded from
# the remaining weight stream so the base loader never sees it.
# Related PR (if no, explain why):
# This PR (Qwen3-32B and GLM4.7 W8A8C8 support). Upstream vLLM's weight-loading
# pipeline does not yet have a generic hook for hardware-plugin-defined
# KV cache parameters.
# Future Plan:
# Remove this patch when vLLM provides a first-class extension point
# for loading extra KV cache quantisation parameters in model load_weights,
# or when the GQA model's weight names are aligned with the parameter
# names expected by the quantisation backend.
# ** 27. File: worker/patch_qwen3vl.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.model_executor.models.qwen3.Qwen3Attention.forward` and
# `vllm.model_executor.models.qwen3_moe.Qwen3MoeAttention.forward`
# Why:
# support triton_split_qkv_rmsnorm_mrope fused kernel for Qwen3Attention and Qwen3MoeAttention.
# How
# override forward method with the triton_split_qkv_rmsnorm_mrope fused kernel,
# when using mrope.
# Future Plan:
# Remove this patch when vllm-ascend supports pattern matching for this fused kernel.
# ** 28. File: worker/patch_qwen3_dflash.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.model_executor.models.qwen3_dflash.DFlashQwen3Model.precompute_and_store_context_kv`
# Why:
# The function directly calls the ops.rms_norm and ops.rotary_imbedding operators,
# but NPU does not have a corresponding implementation.
# How
# Replace ops.* with the internal implementation of vllm-ascend.
# Future Plan:
# Remove this patch when vllm-ascend supports pattern matching for ops.*.
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.model_executor.layers.fused_moe.routed_experts_capturer.RoutedExpertsCapturer.capture`
# Why:
# The upstream implementation doesn't support vllm-ascend specific MoE communication types
# (ALLTOALL and MC2). In the SP + modular-kernel path, the original code cannot correctly
# handle tensor splitting and all-gather operations on NPU, especially when tokens are
# unevenly distributed across TP ranks or padded to max_tokens in MC2 mode.
# How
# Override the capture method to add support for vllm-ascend's MoECommType:
# - Check `_EXTRA_CTX.moe_comm_type` to determine if ALLTOALL or MC2 mode is active
# - Calculate correct gather_topk_ids_shape based on communication type:
# * ALLTOALL: uses actual token_num_per_dp for shape calculation
# * MC2: uses padded max_tokens * tp_size for shape calculation
# - Properly handle tensor_split and all_gather operations for NPU distributed communication
# Future Plan:
# Remove this patch when upstream vLLM supports MoE communication type abstraction that
# can be extended by hardware plugins like vllm-ascend.
#
# ** 29. File: platform/patch_mamba_manager.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.v1.core.single_type_kv_cache_manager.MambaManager`
# Why:
# 1. Upstream hybrid prefix cache lookup does not support PCP/DCP.
# 2. Upstream MambaManager#get_num_blocks_to_allocate give the
# wrong number of blocks when an external cache hit occurred
# How:
# 1. Replace MambaManager with AscendMambaManager for prefix cache hit lookup
# on hybrid Mamba paths (logical mamba block_size when caching is enabled).
# 2. Override the get_num_blocks_to_allocate method to fix the number of blocks
# when hitting the external cache and loading synchronously
# Related PR (if no, explain why):
# 1. https://github.com/vllm-project/vllm/pull/40996
# 2. https://github.com/vllm-project/vllm/pull/46892
# Future Plan:
# 1. Upstream PR #40996 adds hybrid prefix cache lookup for DCP only; PCP is
# not supported yet. Remove this patch once upstream supports both PCP and DCP.
# 2. Remove this patch once upstream accept 46892 pr or fixed the bug by other pr.
#
# ** 30. File: platform/patch_use_v2_model_runner.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
# 1. `vllm.config.vllm.VllmConfig.use_v2_model_runner`
# Why:
# Upstream vLLM enables the v2 model runner not only via the
# VLLM_USE_V2_MODEL_RUNNER env var but also based on model
# architecture whitelists, Triton availability, and feature
# compatibility checks. On Ascend the NPU v2 runner is not yet
# compatible with all upstream-defaulted models and features, so
# enabling by model architecture can crash. We override the
# property to read only VLLM_USE_V2_MODEL_RUNNER, deferring
# model/framework checks to the NPU runner itself.
# How:
# Monkey-patch VllmConfig.use_v2_model_runner to return
# envs.VLLM_USE_V2_MODEL_RUNNER (defaulting to False when unset).
# worker/patch_v2/patch_use_v2_model_runner.py reuses this platform
# patch so EngineCore and worker processes share the same behavior.
# Related PR (if no, explain why):
# 1. https://github.com/vllm-project/vllm-ascend/pull/11389
# Future Plan:
# Remove this patch once vllm-ascend fully supports the v2 model
# runner and can rely on upstream's default enablement heuristics
# (model architecture, Triton, feature checks) without crashes or
# degraded functionality.