1103 lines
56 KiB
Python
1103 lines
56 KiB
Python
#
|
||
# Copyright (c) 2025 Huawei Technologies Co., Ltd. All Rights Reserved.
|
||
# This file is a part of the vllm-ascend project.
|
||
#
|
||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||
# you may not use this file except in compliance with the License.
|
||
# You may obtain a copy of the License at
|
||
#
|
||
# http://www.apache.org/licenses/LICENSE-2.0
|
||
#
|
||
# Unless required by applicable law or agreed to in writing, software
|
||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||
# See the License for the specific language governing permissions and
|
||
# limitations under the License.
|
||
|
||
# ----------------------------------------------------------------------------------
|
||
# This module manage the patch for vllm. There are two folders in this module:
|
||
# - platform: contains the patches applied before worker starts. It's called by
|
||
# `vllm_ascend.utils.adapt_patch(is_global_patch=True)` in
|
||
# `vllm_ascend.platform.NPUPlatform.pre_register_and_update()` function.
|
||
# - worker: contains the patches applied when worker starts. It's called by
|
||
# `vllm_ascend.utils.adapt_patch(is_global_patch=False)` in
|
||
# each worker's `__init__` function.
|
||
#
|
||
# Once a new patch is added in vllm-ascend, please add the patch description into this file as well.
|
||
# ----------------------------------------------------------------------------------
|
||
|
||
# What's Patched and how it works:
|
||
# --------------------------------
|
||
# * Platform Patch:
|
||
# =================
|
||
# ** 1. File: platform/patch_distributed.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `torch.distributed.all_reduce`, `torch.distributed.broadcast`
|
||
# Why:
|
||
# tensor alignment for 310p
|
||
# How:
|
||
# rewrite all_reduce and broadcast in torch.distributed
|
||
# Related PR (if no, explain why):
|
||
# No, not ready yet.
|
||
# Future Plan:
|
||
# Find a better way to support tensor alignment for 310p without this patch.
|
||
#
|
||
# ** 2. File: platform/patch_mamba_config.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.model_executor.models.config.HybridAttentionMambaModelConfig.verify_and_update_config`
|
||
# Why:
|
||
# block size is set to 16 in vLLM which is not supported by Ascend.
|
||
# How:
|
||
# Set block size to 128 on npu.
|
||
# Related PR (if no, explain why):
|
||
# we'll fix this in vLLM soon.
|
||
# Future Plan:
|
||
# Remove this patch when vLLM merges the PR.
|
||
#
|
||
# ** 3. File: platform/patch_multiproc_executor.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.v1.executor.multiproc_executor.MultiprocExecutor`
|
||
# Why:
|
||
# vLLM create child process with daemon=True, which doesn't work with EPLB case, since EPLB will create
|
||
# a new process which is not allowed by daemon=True.
|
||
# How:
|
||
# Set daemon=False in MultiprocExecutor.
|
||
# Related PR (if no, explain why):
|
||
# Find a way to support daemon=False in vLLM
|
||
# Future Plan:
|
||
# Remove this patch when vLLM fix the issue.
|
||
#
|
||
# ** 4. File: platform/patch_shm_broadcast.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.distributed.device_communicators.shm_broadcast.MessageQueue`
|
||
# Why:
|
||
# vLLM 0.23.0 local readers can wait indefinitely after a best-effort ZMQ
|
||
# notification is lost. A caller exception can also leave a read slot
|
||
# unreleased and block the writer.
|
||
# How:
|
||
# Replace timeout_ms and acquire_read with the implementations from the
|
||
# upstream fix. Idle waits are capped at five seconds, and read-slot cleanup
|
||
# runs in a finally block.
|
||
# Related PR (if no, explain why):
|
||
# https://github.com/vllm-project/vllm/pull/45224
|
||
# Future Plan:
|
||
# Remove this patch when the supported vLLM release includes PR #45224.
|
||
#
|
||
# ** 5. File: platform/patch_balance_schedule.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.v1.engine.core.EngineCoreProc.run_engine_core`
|
||
# `vllm.v1.core.sched.scheduler.Scheduler`
|
||
# Why:
|
||
# vLLM v1 scheduling currently enables chunkedprefill by default, which processes prefill and decode
|
||
# requests simultaneously in a single scheduling session. This can impact the overall system throughput
|
||
# and performance in some scenarios.
|
||
# How:
|
||
# Set --additional-config '{"enable_balance_scheduling": true}' or
|
||
# set environmental variable VLLM_ASCEND_BALANCE_SCHEDULING=1 (deprecated).
|
||
# Related PR (if no, explain why):
|
||
# https://github.com/vllm-project/vllm/pull/29721
|
||
# Future Plan:
|
||
# Remove this patch when vLLM merge the PR.
|
||
#
|
||
# 2. Disable automatic scheduler preemption on vLLM 0.23.0 PD-disaggregated
|
||
# prefill nodes.
|
||
# Why:
|
||
# Async scheduling can finish the one-token prefill request while memory
|
||
# pressure preempts it, releasing delayed KV blocks before their block IDs
|
||
# are sent to the decode node.
|
||
# How:
|
||
# When a pure KV producer cannot allocate slots, the default, async, and
|
||
# profiling-chunk schedulers stop the current step without preempting a
|
||
# running request; forced prefix-cache reset is rejected until requests
|
||
# drain.
|
||
# Related PR (if no, explain why):
|
||
# No, this is a temporary vLLM 0.23.0 compatibility fix.
|
||
# Future Plan:
|
||
# Remove this patch after the upstream scheduler and KV connector handle
|
||
# async prefill completion and delayed block release atomically.
|
||
#
|
||
# ** 6. File: platform/patch_minimax_m2_config.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.config.model.ModelConfig._verify_quantization`
|
||
# Why:
|
||
# MiniMax-M2 fp8 checkpoints on NPU may fail upstream quantization validation.
|
||
# vllm-ascend needs to disable fp8 quantization and load bf16 dequantized
|
||
# weights in worker-side patches instead.
|
||
# How:
|
||
# Monkey-patch `_verify_quantization` and intercept platform quantization
|
||
# verification to force `cfg.quantization=None` for MiniMax-M2 fp8 on NPU.
|
||
# Related PR (if no, explain why):
|
||
# No, upstream behavior differs across versions and needs discussion.
|
||
# Future Plan:
|
||
# Remove this patch once upstream supports MiniMax-M2 fp8 on NPU or provides
|
||
# a backend-safe validation / override mechanism.
|
||
#
|
||
# 2. `vllm.config.model.ModelConfig._verify_cuda_graph`
|
||
# Why:
|
||
# For MiniMax-M2 on NPU with ACL graph capture enabled, HCCL op expansion
|
||
# mode affects graph shape coverage. Users may forget to set it.
|
||
# How:
|
||
# If user doesn't set it, set `HCCL_OP_EXPANSION_MODE=AIV` for this model
|
||
# and log a warning when a different value is detected.
|
||
# Related PR (if no, explain why):
|
||
# No, this is an environment-specific tuning knob.
|
||
# Future Plan:
|
||
# Remove this patch if upstream provides an official NPU graph-capture
|
||
# guidance / auto-configuration path for HCCL.
|
||
#
|
||
# 3. `vllm.config.speculative.SpeculativeConfig._verify_args`
|
||
# Why:
|
||
# Upstream vLLM's eagle3/extract_hidden_states restricts target model types
|
||
# via a whitelist. MiniMax-M2 should be allowed once the worker-side model
|
||
# can emit auxiliary hidden states.
|
||
# How:
|
||
# Monkey-patch `_verify_args` to bypass only the whitelist ValueError for
|
||
# MiniMax model_type when method is eagle3/extract_hidden_states.
|
||
# SpeculativeConfig is a Pydantic dataclass (`@config`); init validation calls
|
||
# `__pydantic_decorators__.model_validators["_verify_args"].func`, so that
|
||
# `Decorator.func` must be replaced (not only `SpeculativeConfig._verify_args`),
|
||
# then `rebuild_dataclass(SpeculativeConfig, force=True)`.
|
||
# If `VllmConfig` was imported earlier, also `rebuild_dataclass(VllmConfig, ...)`
|
||
# so nested `speculative_config` validation does not use a stale schema.
|
||
# Related PR (if no, explain why):
|
||
# https://github.com/vllm-project/vllm/pull/37512
|
||
# Future Plan:
|
||
# Remove this patch once upstream whitelist includes MiniMax.
|
||
#
|
||
# 4. `vllm.model_executor.models.registry` (spec decode aliases)
|
||
# Why:
|
||
# Some Eagle3 draft checkpoints may declare a MiniMax-specific architecture
|
||
# string while reusing the shared Eagle3 implementation.
|
||
# How:
|
||
# Register `Eagle3MiniMaxM2ForCausalLM` as an alias pointing to the
|
||
# existing Eagle3 implementation in the speculative decoding registry.
|
||
# Related PR (if no, explain why):
|
||
# https://github.com/vllm-project/vllm/pull/37512
|
||
# Future Plan:
|
||
# Drop the alias once upstream registry includes it or the checkpoint
|
||
# standardizes architecture strings.
|
||
#
|
||
# ** 7. File: platform/patch_minimax_usage_accounting.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.entrypoints.openai.chat_completion.serving.OpenAIServingChat`
|
||
# `vllm.reasoning.minimax_m2_reasoning_parser`
|
||
# Why:
|
||
# MiniMax-M2 chat usage accounting needs to report
|
||
# `completion_tokens_details.reasoning_tokens` for both streaming and
|
||
# non-streaming chat completions without slowing other reasoning models.
|
||
# How:
|
||
# Monkey-patch MiniMax reasoning token counters and bind usage-accounting
|
||
# wrappers only on MiniMax chat-serving instances.
|
||
# Related PR (if no, explain why):
|
||
# https://github.com/vllm-project/vllm/pull/45701
|
||
# https://github.com/vllm-project/vllm/pull/45802
|
||
# Future Plan:
|
||
# Remove this patch after both upstream vLLM PRs are merged and the
|
||
# supported vLLM revision used by vLLM Ascend includes them through the
|
||
# regular main-to-main sync.
|
||
#
|
||
# ** 7a. File: platform/patch_glm_tool_call_streaming.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.entrypoints.openai.chat_completion.serving.OpenAIServingChat`
|
||
# Why:
|
||
# GLM tool-call streaming can emit final remaining-argument chunks with
|
||
# repeated tool-call metadata, and can combine terminal argument bytes with
|
||
# `finish_reason="tool_calls"` in the same SSE chunk.
|
||
# How:
|
||
# Monkey-patch remaining-argument delta construction to emit only argument
|
||
# fragments by default, and split terminal argument chunks into an argument
|
||
# chunk followed by an empty finish chunk.
|
||
# Related PR (if no, explain why):
|
||
# https://github.com/vllm-project/vllm/issues/44098
|
||
# https://github.com/vllm-project/vllm/pull/44099
|
||
# https://github.com/vllm-project/vllm-ascend/issues/8327
|
||
# https://github.com/vllm-project/vllm-ascend/pull/8178
|
||
# Future Plan:
|
||
# Remove this patch once the supported vLLM version contains the upstream
|
||
# GLM tool-call final chunk fixes.
|
||
#
|
||
# ** 7b. File: platform/patch_glm47_tool_call_parser.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.tool_parsers.glm47_moe_tool_parser.Glm47MoeModelToolParser`
|
||
# Why:
|
||
# vLLM's GLM47 streaming parser can drop complete inline zero-argument
|
||
# tool calls such as `<tool_call>get_current_time</tool_call>`, while
|
||
# non-streaming parses the same output correctly.
|
||
# How:
|
||
# Monkey-patch GLM47 tool-call region extraction so complete inline
|
||
# zero-argument regions are normalized for the existing streaming name
|
||
# extractor without emitting partial names for incomplete regions.
|
||
# Related PR (if no, explain why):
|
||
# https://github.com/vllm-project/vllm/issues/44326
|
||
# https://github.com/vllm-project/vllm/pull/44327
|
||
# Future Plan:
|
||
# Remove this patch once the supported vLLM version contains the upstream
|
||
# GLM47 inline zero-argument streaming parser fix.
|
||
#
|
||
# ** 10a. File: platform/patch_kv_cache_utils.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.v1.core.kv_cache_utils.resolve_kv_cache_block_sizes`
|
||
# `vllm.v1.engine.core.resolve_kv_cache_block_sizes`
|
||
# Why:
|
||
# vLLM PR #40860 added a restriction that hybrid KV cache groups with
|
||
# multiple block sizes do not support context parallelism (dcp/pcp > 1).
|
||
# This restriction is correct for CUDA but not for Ascend, which
|
||
# implements context parallelism for MLA and SWA-MLA layers separately.
|
||
# How:
|
||
# Monkey-patch resolve_kv_cache_block_sizes to handle the multiple-groups
|
||
# + CP case by returning lcm(block_sizes) * dcp * pcp as scheduler_block_size
|
||
# instead of raising ValueError.
|
||
# Related PR (if no, explain why):
|
||
# vLLM PR #40860 ([Feat] DeepSeek V4 Rebased).
|
||
# Future Plan:
|
||
# Remove this patch once upstream vLLM supports hybrid KV cache + CP for
|
||
# non-CUDA backends, or exposes a platform hook for this behavior.
|
||
#
|
||
# ** 10ab. File: worker/patch_v2/patch_attn_utils.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.v1.worker.gpu.attn_utils.get_kv_cache_spec`
|
||
# Why:
|
||
# The current v2 worker still goes through the shared upstream v1 helper
|
||
# to build KV cache specs. For Ascend MLA layers that helper returns the
|
||
# generic `MLAAttentionSpec`, but NPU-side cache allocation and reshape
|
||
# logic expects `AscendMLAAttentionSpec`.
|
||
# How:
|
||
# Monkey-patch `get_kv_cache_spec` so regular attention layers keep the
|
||
# upstream behavior while MLA layers are rewritten to
|
||
# `AscendMLAAttentionSpec`, including the FA-quant head-size adjustment.
|
||
# Related PR (if no, explain why):
|
||
# No. This is a plugin-side compatibility patch for the current upstream
|
||
# helper path.
|
||
# Future Plan:
|
||
# Remove this patch once upstream adds a backend hook for KV cache spec
|
||
# construction or v2 worker no longer depends on the shared v1 helper.
|
||
#
|
||
# ** 10. File: platform/patch_profiling_chunk.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.v1.engine.core.EngineCore.__init__`
|
||
# 2. `vllm.v1.engine.core.EngineCoreProc.run_engine_core`
|
||
# 3. `Scheduler.update_from_output` (scheduler class, wrapped when profiling chunk is enabled)
|
||
# Why:
|
||
# Profiling-based dynamic chunk sizing needs to run a one-shot profiling pass
|
||
# after `model_executor` is ready, and to feed per-step execution latency back
|
||
# into `ProfilingChunkManager` so the history-aware chunk predictor can refine
|
||
# online. In multiprocessing `spawn` mode the child process starts a fresh
|
||
# interpreter, so monkey-patches applied in the parent are lost unless the
|
||
# subprocess entry point re-applies them before any `EngineCore` is created.
|
||
# How:
|
||
# Replace `EngineCore.__init__` to call `scheduler.run_profiling_chunk_init`
|
||
# when present, then wrap `scheduler.update_from_output` once per process to
|
||
# read `model_output.execution_time_ms` and `scheduler_output` token/chunk
|
||
# metadata and call `ProfilingChunkManager.record_batch_execution_time` (and
|
||
# bootstrap target latency for the first chunk when needed). Replace
|
||
# `EngineCoreProc.run_engine_core` so importing this module in the child
|
||
# re-runs the idempotent patch helper before delegating to the original
|
||
# implementation.
|
||
# Related PR (if no, explain why):
|
||
# No, vllm-ascend-specific profiling / scheduling integration.
|
||
# Future Plan:
|
||
# Remove or narrow this patch if upstream exposes stable hooks for backend
|
||
# profiling startup and per-step timing callbacks without monkey-patching
|
||
# `EngineCore` and the multiprocess entry point.
|
||
#
|
||
# ** 10b. File: platform/patch_pp_mtp.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.config.model.ModelConfig.verify_with_parallel_config`
|
||
# Why:
|
||
# Local Eagle/MTP drafters are loaded on the last PP stage rather than
|
||
# partitioned across all PP ranks. Upstream `ModelConfig.verify_with_parallel_config`
|
||
# validates against `pipeline_parallel_size`, which fails for these drafters
|
||
# since they run locally with effective PP=1.
|
||
# How:
|
||
# Monkey-patch `verify_with_parallel_config` to detect Eagle/MTP drafter
|
||
# models (by `model_type` and `architectures`) when `runner="draft"` and
|
||
# `pipeline_parallel_size > 1`. For such configs, call the original verify
|
||
# with a patched `pipeline_parallel_size=1` copy, preserving normal target-model
|
||
# validation for non-drafter models.
|
||
# Related PR (if no, explain why):
|
||
# Backport of local vLLM PP+MTP branch changes.
|
||
# Future Plan:
|
||
# Remove this patch once upstream vLLM's `ModelConfig.verify_with_parallel_config`
|
||
# supports local drafter models with PP > 1, or moves the PP validation to a
|
||
# separate hook that can be overridden per-model-type.
|
||
#
|
||
# ** 11. File: platform/patch_tool_choice_none_content.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.entrypoints.openai.chat_completion.protocol.ChatCompletionResponse`
|
||
# `vllm.entrypoints.openai.chat_completion.protocol.ChatCompletionStreamResponse`
|
||
# Why:
|
||
# vLLM v0.23.0 can serialize empty `tool_calls: []` fields for content-only
|
||
# OpenAI chat responses / streaming deltas, while OpenAI-compatible SDKs
|
||
# expect those empty fields to be omitted so clients see `tool_calls=None`.
|
||
# How:
|
||
# Wrap `model_dump` / `model_dump_json` for chat response payloads and drop
|
||
# empty `tool_calls` lists from `message` / `delta` objects.
|
||
# Related PR (if no, explain why):
|
||
# https://github.com/vllm-project/vllm/pull/44105
|
||
# Future Plan:
|
||
# Remove this patch once the supported vLLM version contains PR #44105.
|
||
#
|
||
# ** 12. File: platform/patch_deepseek_v4_tool_call_parser.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.tool_parsers.deepseekv4_tool_parser.DeepSeekV4ToolParser`
|
||
# Why:
|
||
# Upstream vLLM now includes DeepSeek V4 tokenizer/renderer/reasoning
|
||
# registration, but its streaming tool-call delta parsing does not guarantee
|
||
# incremental `arguments` emission for long argument payloads.
|
||
# How:
|
||
# Monkey-patch `DeepSeekV4ToolParser` stream parsing to emit tool-call
|
||
# metadata in the first delta and stream argument fragments incrementally.
|
||
# Related PR (if no, explain why):
|
||
# Upstream vLLM main behavior as of current runtime.
|
||
# Future Plan:
|
||
# Remove this patch if upstream streaming behavior is updated to satisfy the
|
||
# same DeepSeek DSML incrementality contract.
|
||
#
|
||
# ** 12a. File: platform/patch_minimax_m2_tool_call_parser.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.tool_parsers.minimax_m2_tool_parser.MinimaxM2ToolParser`
|
||
# Why:
|
||
# vLLM 0.21.0 only emits MiniMax-M2 tool-call arguments after a complete
|
||
# `<invoke>...</invoke>` block, so long arguments are buffered instead of
|
||
# streamed incrementally.
|
||
# How:
|
||
# Monkey-patch the MiniMax-M2 parser to emit the tool name once the
|
||
# `<invoke name=...>` header is available and then stream partial
|
||
# `<parameter>` values as JSON argument fragments.
|
||
# Related PR (if no, explain why):
|
||
# https://github.com/vllm-project/vllm/pull/40253
|
||
# https://github.com/vllm-project/vllm/pull/40298
|
||
# Future Plan:
|
||
# Remove this patch once the supported vLLM version contains the upstream
|
||
# MiniMax-M2 incremental tool-call streaming fix.
|
||
#
|
||
# ** 12b. File: platform/patch_structured_output.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.sampling_params.SamplingParams._validate_structured_outputs`
|
||
# `vllm.v1.structured_output.StructuredOutputManager.grammar_init`
|
||
# Why:
|
||
# V1 structured outputs use one engine-level backend, while `backend=auto`
|
||
# resolves the backend per request. After one request initializes
|
||
# `xgrammar`, a later request that resolves to `guidance` can still reach
|
||
# the initialized `xgrammar` backend and crash during grammar compilation.
|
||
# How:
|
||
# Record the first resolved backend on the structured-output config and
|
||
# reject later requests that resolve to a different backend. Also guard
|
||
# `grammar_init` so requests that bypass API-side validation fail before
|
||
# backend grammar compilation.
|
||
# Related PR (if no, explain why):
|
||
# https://github.com/vllm-project/vllm/issues/43920
|
||
# https://github.com/vllm-project/vllm/pull/44401
|
||
# Future Plan:
|
||
# Remove this patch once upstream vLLM either enforces backend consistency
|
||
# before grammar compilation or safely handles mixed-backend grammar
|
||
# failures without killing the engine.
|
||
#
|
||
# ** 13. File: platform/patch_camem_allocator.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.config.model.is_cumem_allocator_available`
|
||
# Why:
|
||
# Upstream vLLM main enables and validates the CUDA/ROCm CuMem allocator
|
||
# when `enable_sleep_mode=True`. Ascend implements sleep mode with its own
|
||
# CaMem allocator, so the upstream CuMem-only availability check fails
|
||
# during `ModelConfig` validation before Ascend worker code can run.
|
||
# How:
|
||
# Treat Ascend's platform sleep allocator as satisfying the allocator
|
||
# availability check, while preserving the original vLLM CuMem check as
|
||
# fallback.
|
||
# Related PR (if no, explain why):
|
||
# No, this maps an upstream CUDA/ROCm allocator validation to Ascend's
|
||
# backend-specific CaMem implementation.
|
||
# Future Plan:
|
||
# Remove this patch if upstream exposes a platform allocator capability hook
|
||
# for sleep mode validation.
|
||
#
|
||
# ** 15. File: platform/patch_weight_transfer_engine.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.distributed.weight_transfer.factory.WeightTransferEngineFactory._registry["nccl"]`
|
||
# Why:
|
||
# Upstream vLLM's WeightTransferConfig.backend is a pydantic Literal["nccl", "ipc"]
|
||
# which does not accept "hccl". On Ascend NPU, NCCL is unavailable and HCCL must
|
||
# be used for trainer-to-worker weight broadcasting.
|
||
# How:
|
||
# Replace the "nccl" factory entry with a lambda that returns
|
||
# HCCLWeightTransferEngine. Users pass the already-accepted "nccl" string
|
||
# (e.g. --weight-transfer-config '{"backend": "nccl"}') and the factory
|
||
# resolves it to the HCCL engine at runtime.
|
||
# Related PR (if no, explain why):
|
||
# No. Adding "hccl" to the Literal requires modifying pydantic core schemas,
|
||
# which is fragile across pydantic versions.
|
||
# Future Plan:
|
||
# Remove this patch when upstream vLLM relaxes the Literal type to str or
|
||
# provides an extension point for out-of-tree weight transfer backends.
|
||
# 2. `vllm.distributed.weight_transfer.factory.WeightTransferEngineFactory._registry["ipc"]`
|
||
# Why:
|
||
# The "ipc" backend must resolve to NPUIPCWeightTransferEngine on Ascend NPU.
|
||
# However, this patch runs during global plugin patching - extremely early in
|
||
# startup, before any weight transfer backend is selected. Importing the IPC
|
||
# engine eagerly pulls in vllm.distributed.weight_transfer.ipc_engine, which
|
||
# does `import ray` at module top level. Since ray is an optional dependency,
|
||
# its absence aborts the whole vllm_ascend plugin load and crashes every
|
||
# `vllm serve` invocation - even workloads that never use weight transfer.
|
||
# How:
|
||
# Register a lazy loader function (instead of an eager import) that imports
|
||
# and returns NPUIPCWeightTransferEngine only when create_engine() is invoked
|
||
# for the "ipc" backend. This matches the factory's zero-arg-callable
|
||
# lazy-loading contract, so the ray-importing module is loaded only when ipc
|
||
# is actually requested. (HCCL keeps its eager import - it never imports ray.)
|
||
# Related PR (if no, explain why):
|
||
# No. The eager `import ray` lives in upstream vLLM's ipc_engine module; the
|
||
# lazy loader is a local workaround until upstream defers that import.
|
||
# Future Plan:
|
||
# Remove this workaround once upstream vLLM stops importing ray at module top
|
||
# level in vllm.distributed.weight_transfer.ipc_engine (e.g. defers it into
|
||
# the code path that actually needs ray), so importing the IPC engine no
|
||
# longer requires the optional ray dependency.
|
||
#
|
||
# ** 15. File: platform/patch_kv_cache_coordinator.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.v1.core.kv_cache_coordinator.HybridKVCacheCoordinator.find_longest_cache_hit_per_group`
|
||
# Why:
|
||
# In PD disaggregation with hybrid Mamba models, the D side receives
|
||
# FullAttention KV blocks from the P side but has no local prefix-cache
|
||
# hit for Mamba groups. Upstream's min-reduction across all KV groups
|
||
# collapses the FullAttention hit length to 0, preventing partial
|
||
# FullAttention-only prefix cache reuse on the D side.
|
||
# How:
|
||
# For Mamba hybrid models,
|
||
# num_new_local_computed_tokens should be the FA hit
|
||
# length. This value is passed to the connector's
|
||
# get_num_new_matched_tokens which computes:
|
||
# external = total - local_computed.
|
||
# Using the FA hit skips re-transferring FA blocks
|
||
# already cached on D-side.
|
||
# Related PR (if no, explain why):
|
||
# https://github.com/vllm-project/vllm/pull/42524
|
||
# https://github.com/vllm-project/vllm/pull/44243
|
||
# Future Plan:
|
||
# Remove this patch when vLLM PR #42524 and #44243 is included in the supported
|
||
# upstream vLLM version.
|
||
#
|
||
# * Worker Patch:
|
||
# ===============
|
||
#
|
||
# ** 1. File: worker/patch_distributed.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.distributed.parallel_state.GroupCoordinator`
|
||
# Why:
|
||
# vllm doesn't support all_to_all for GroupCoordinator.
|
||
# How:
|
||
# Add all_to_all implementation for GroupCoordinator.
|
||
# Related PR (if no, explain why):
|
||
# No, we should use vlLM all2all manager to support all_to_all for npu.
|
||
# Future Plan:
|
||
# Remove this patch when the refactor of all2all manager is done.
|
||
#
|
||
# ** 3. File: worker/patch_triton.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.model_executor.layers.mamba.ops`, `vllm.model_executor.layers.fla.ops`,
|
||
# `vllm.v1.worker.gpu.sample.gumbel.gumbel_sample`
|
||
# Why:
|
||
# triton ops in vLLM perform not good on NPU. And there is no dispatch mechanism for triton ops.
|
||
# How:
|
||
# override triton ops in vLLM with ascend implementation
|
||
# Related PR (if no, explain why):
|
||
# Let vLLM support triton ops dispatch.
|
||
# Future Plan:
|
||
# Remove this patch when vLLM support the dispatch function.
|
||
#
|
||
# 2. `triton.next_power_of_2`
|
||
# Why:
|
||
# The Triton version bundled with torch_npu on Ascend NPU
|
||
# does not include `next_power_of_2`, which is called by
|
||
# upstream vLLM and vLLM-Ascend code in 94+ places.
|
||
# Additionally, when Triton is not available (HAS_TRITON=False),
|
||
# vLLM uses TritonPlaceholder which also lacks this function.
|
||
# How:
|
||
# Import `triton` from vllm.triton_utils (which handles both
|
||
# real Triton and TritonPlaceholder) and inject `next_power_of_2`
|
||
# onto the module, reusing `vllm.utils.math_utils.next_power_of_2`.
|
||
# Related PR (if no, explain why):
|
||
# No, torch_npu Triton compatibility issue.
|
||
# Future Plan:
|
||
# Remove this patch when torch_npu's Triton includes
|
||
# next_power_of_2 or when vLLM no longer calls triton.next_power_of_2.
|
||
#
|
||
# ** 4. File: worker/patch_qwen3_next_mtp.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.v1.worker.utils.bind_kv_cache`
|
||
# Why:
|
||
# 'bind_kv_cache' func will raise an exception when current_platform is npu.
|
||
# How:
|
||
# Replace with a new bind_kv_cache.
|
||
# Skip the raise.
|
||
# Related PR (if no, explain why):
|
||
# It need discuss.
|
||
# Future Plan:
|
||
# Remove this patch after discussing with vllm community and adapting bind_kv_cache to npu.
|
||
#
|
||
# ** 5. File: worker/patch_rejection_sampler.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.v1.sample.rejection_sampler`
|
||
# Why:
|
||
# - some functions from `rejection_sampler` are not supported or slow on npu.
|
||
# How:
|
||
# - add npu_top_k_top_p to 'apply_sampling_constraints' func
|
||
# - add custom triton kernel to `expand_batch_to_tokens` and `rejection_sample`
|
||
# Related PR (if no, explain why):
|
||
# Let vLLM support triton ops dispatch.
|
||
# Future Plan:
|
||
# 1. make these functions as class func of RejectionSampler, create AscendRejectionSampler
|
||
# to override them, then delete the patch file `worker/patch_rejection_sampler.py`.
|
||
# 2. make these functions as costom op, then remove AscendRejectionSampler
|
||
#
|
||
## ** 6. File: worker/patch_module.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.v1.attention.backends.gdn_attn.torch.argsort`
|
||
# Why:
|
||
# 1. 'torch.argsort' func of npu does not support bool.
|
||
# 2. Without `stable=True`, the output will have a lot of redundant tokens.
|
||
# How:
|
||
# Replace with a new torch.argsort that will cast the input to torch.int32
|
||
# and do stable sort.
|
||
# Related PR (if no, explain why):
|
||
# 1. It depends on torch_npu.
|
||
# 2. https://github.com/vllm-project/vllm/pull/30632
|
||
# Future Plan:
|
||
# Remove this patch when bool is supported in 'torch.argsort' func of npu.
|
||
# Make 'torch.argsort' in `vllm.v1.attention.backends.gdn_attn` be stable.
|
||
# 2. `vllm_ascend.ops.gdn_attn_builder.AscendGDNAttentionMetadataBuilder.build`
|
||
# Why:
|
||
# Qwen3.5/Qwen3Next GDN Decode/Specific Decode on NPU needs prebuilt varlen chunk metadata
|
||
# to avoid forward-time host round-trips that break async scheduling.
|
||
# How:
|
||
# Override the GDN attention metadata builder for Ascend backend and attach
|
||
# prebuilt device metadata bundle onto the returned attention metadata object.
|
||
# Future Plan:
|
||
# Remove this patch when upstream exposes a backend hook for extending GDN
|
||
# metadata or when the optimization is accepted upstream directly.
|
||
#
|
||
# ** 8. File: worker/patch_qwen3_next.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.model_executor.models.qwen3_next.Qwen3NextGatedDeltaNet.forward`
|
||
# Why:
|
||
# The Qwen3Next GatedDeltaNet forward cannot directly add custom operators.
|
||
# How:
|
||
# Add a branch in Qwen3NextGatedDeltaNet.forward to adapt to fused_qkvzba_split_reshape_cat.
|
||
# Related PR (if no, explain why):
|
||
# https://github.com/vllm-project/vllm/pull/30863
|
||
# Future Plan:
|
||
# Remove this patch when vLLM support these operators.
|
||
#
|
||
# 2. `vllm.model_executor.models.qwen3_next.Qwen3NextGatedDeltaNet._forward_core`
|
||
# Why:
|
||
# triton ops fused_recurrent_gated_delta_rule and fused_gdn_gating in vLLM perform not good on NPU.
|
||
# How:
|
||
# add a new fused triton ops in vLLM with ascend implementation.
|
||
# Related PR (if no, explain why):
|
||
# https://github.com/vllm-project/vllm/pull/30860
|
||
# Future Plan:
|
||
# Remove this patch when vLLM support these operators.
|
||
#
|
||
# 3. `vllm.model_executor.models.qwen3_next.Qwen3NextGatedDeltaNet._forward_core`
|
||
# Why:
|
||
# The Qwen3Next GatedDeltaNet _forward_core cannot directly add custom operators.
|
||
# How:
|
||
# Add a branch in Qwen3NextGatedDeltaNet._forward_core to adapt to fused_gdn_gating_patch.
|
||
# Related PR (if no, explain why):
|
||
# https://github.com/vllm-project/vllm/pull/31002
|
||
# Future Plan:
|
||
# Remove this patch when vLLM support these operators.
|
||
#
|
||
# ** 10. File: worker/patch_qwen3vl.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.model_executor.models.qwen3_vl.Qwen3VLForConditionalGeneration._get_deepstack_input_embeds`
|
||
# Why:
|
||
# support flash comm v1 for qwen3vl.
|
||
# How:
|
||
# override _get_deepstack_input_embeds method with the flash comm v1 implementation.
|
||
# Future Plan:
|
||
# Remove this patch when https://github.com/vllm-project/vllm-ascend/issues/5712 is completed.
|
||
# 2. `vllm.model_executor.models.qwen3_vl_moe.Qwen3MoeLLMForCausalLM.start_layer`,
|
||
# `vllm.model_executor.models.qwen3_vl_moe.Qwen3MoeLLMForCausalLM.end_layer`
|
||
# Why:
|
||
# Qwen3-VL-MoE checks the language-model pipeline boundary on non-first
|
||
# PP ranks, but Qwen3MoeLLMForCausalLM keeps start_layer/end_layer only
|
||
# on the inner model object.
|
||
# How:
|
||
# Expose start_layer/end_layer properties on Qwen3MoeLLMForCausalLM and
|
||
# forward them to the inner model.
|
||
# Future Plan:
|
||
# Remove this patch when upstream vLLM exposes these PP layer boundaries
|
||
# on the Qwen3-VL-MoE language-model wrapper.
|
||
#
|
||
# ** 11. File: worker/patch_npugraph_ex_triton.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `npugraph_ex.core._concrete_graph.ValuePack`,
|
||
# `npugraph_ex.npu_fx_compiler._unpack_meta`,
|
||
# `npugraph_ex.npu_fx_compiler._NpuGraphConverter._unpack_npu`
|
||
# Why:
|
||
# In the Triton scenario, npugraph_ex backend needs to process the value pack of the input parameters.
|
||
# How:
|
||
# Supplement the relevant processing logic through patches.
|
||
# Related PR (if no, explain why):
|
||
# https://gitcode.com/Ascend/torchair/pull/2575
|
||
# Future Plan:
|
||
# Remove this patch when the PTA version used by vllm-ascend has been upgraded.
|
||
#
|
||
# ** 12. File: worker/patch_v2/patch_uva.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.v1.worker.gpu.states.UvaBuffer`
|
||
# Why:
|
||
# ASCEND NPUs do not support UVA yet, so we need to wrap it in vLLM.
|
||
# How:
|
||
# make UvaBuffer a dummy class, mimic the interface of vllm UvaBuffer.
|
||
# Future Plan:
|
||
# Remove this patch when NPU support UVA.
|
||
#
|
||
# ** 13. File: worker/patch_kimi_k25.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.model_executor.models.kimi_k25_vit.Learnable2DInterpPosEmbDivided_fixed.forward`
|
||
# Why:
|
||
# The forward method uses interpolate with ops not supported on NPU.
|
||
# How:
|
||
# Replace with a new forward that uses CPU for interpolate when shape mismatch,
|
||
# and use get_rope_shape to handle the rope shape interpolation.
|
||
# Future Plan:
|
||
# Remove this patch when vLLM aligns with the latest main.
|
||
#
|
||
# ** 14. File: worker/patch_draft_quarot.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.model_executor.models.llama_eagle3.Eagle3LlamaForCausalLM.load_weights`
|
||
# Why:
|
||
# vllm-ascend reused the loading logic of drafter model from vllm,
|
||
# but vllm doesn't need to apply to Ascend quantization.
|
||
# How:
|
||
# Dynamically replace the `load_weights` function at runtime,
|
||
# and fix `target_config` into the new implementation with a closure.
|
||
# Related PR (if no, explain why):
|
||
# https://github.com/vllm-project/vllm/pull/36225
|
||
# Future Plan:
|
||
# Remove this patch when vLLM merges the PR.
|
||
#
|
||
# ** 15. File: worker/patch_minimax_m2.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.model_executor.models.minimax_m2.MiniMaxM2MoE.forward`
|
||
# Why:
|
||
# MiniMax-M2 routing should keep router logits in fp32 on NPU.
|
||
# How:
|
||
# Replace the forward to cast hidden states to fp32 before the gate.
|
||
# Related PR (if no, explain why):
|
||
# No, model-specific behavior.
|
||
# Future Plan:
|
||
# Remove this patch once upstream behavior is sufficient for Ascend.
|
||
#
|
||
# 2. `vllm.model_executor.models.minimax_m2.MiniMaxM2Attention.forward`
|
||
# Why:
|
||
# MiniMax-M2 attention benefits from the NPU fused split-qkv + RMSNorm + rope
|
||
# kernel path.
|
||
# How:
|
||
# Replace `forward` to call `torch.ops.vllm.split_qkv_tp_rmsnorm_rope` before
|
||
# the upstream attention and output projection steps.
|
||
# Related PR (if no, explain why):
|
||
# No, backend-specific fused kernel path.
|
||
# Future Plan:
|
||
# Remove this patch when upstream exposes a backend dispatch path for this
|
||
# fused attention preparation.
|
||
#
|
||
# 3. `vllm.model_executor.models.minimax_m2.MiniMaxM2Model.load_weights`
|
||
# Why:
|
||
# MiniMax-M2 fp8 checkpoints may store fp8 weights with per-block inverse
|
||
# scales. On NPU we load bf16 weights by dequantizing at load time.
|
||
# How:
|
||
# Inject fp8 dequant helpers and wrap `load_weights` to convert fp8 weight +
|
||
# `weight_scale_inv` pairs into bf16 blocks before delegating to upstream.
|
||
# Related PR (if no, explain why):
|
||
# No, fp8 load format and backend constraints are model/backend specific.
|
||
# Future Plan:
|
||
# Remove this patch when upstream supports MiniMax-M2 fp8 loading on NPU.
|
||
#
|
||
# ** 16. File: worker/patch_minimax_m2_linear_attn.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.model_executor.layers.mamba.linear_attn.MiniMaxText01RMSNormTP.__init__`
|
||
# `vllm.model_executor.layers.mamba.linear_attn.MiniMaxText01RMSNormTP.weight_loader`
|
||
# Why:
|
||
# MiniMax-M2 linear attention RMSNorm needs weight sharding that can follow
|
||
# TP layout (and sometimes kv-head replication) on NPU.
|
||
# How:
|
||
# Override `__init__` to parameterize weight shard world/rank and install a
|
||
# sharded `weight_loader` implementation.
|
||
# Related PR (if no, explain why):
|
||
# No, upstream API surface differs across versions.
|
||
# Future Plan:
|
||
# Remove this patch when upstream exposes stable sharding hooks for this layer.
|
||
#
|
||
# 2. `vllm.model_executor.layers.mamba.linear_attn.MiniMaxText01RMSNormTP.forward_qk`
|
||
# (or older `_normalize_qk`)
|
||
# Why:
|
||
# q/k norm for linear attention is performance-sensitive. On NPU, a fused
|
||
# rms_norm kernel is faster and TP needs a global rstd correction.
|
||
# How:
|
||
# Replace q/k normalization with NPU rms_norm fast path and TP-global rstd
|
||
# correction; fall back to upstream implementation on non-NPU.
|
||
# Related PR (if no, explain why):
|
||
# No, backend-specific optimization.
|
||
# Future Plan:
|
||
# Remove this patch when upstream adds a backend dispatch path for q/k norm.
|
||
#
|
||
# ** 17. File: worker/patch_qwen3_5.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.model_executor.models.qwen3_5.Qwen3_5GatedDeltaNet._forward_core`
|
||
# Why:
|
||
# The class Qwen3_5GatedDeltaNet reuse the `_forward_core` method of Qwen3NextGatedDeltaNet,
|
||
# but the ascendC ops of Qwen3NextGatedDeltaNet do not support ssm_state with float32 format.
|
||
# How:
|
||
# patch Qwen3_5GatedDeltaNet._forward_core to use triton ops like `fused_recurrent_gated_delta_rule`.
|
||
# Future Plan:
|
||
# Remove this patch when all ops in _forward_core support both Qwen3_5 and Qwen3Next.
|
||
#
|
||
# ** 17a. File: worker/patch_idex_310.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.model_executor.layers.fla.ops.index.prepare_chunk_indices`
|
||
# `vllm.model_executor.layers.fla.ops.index.prepare_chunk_offsets`
|
||
# Why:
|
||
# 310P uses Ascend-friendly chunk index helpers for Qwen GDN prefill.
|
||
# How:
|
||
# Replace upstream FLA chunk index helper functions with 310P implementations.
|
||
#
|
||
# 2. `vllm_ascend.spec_decode.llm_base_proposer.AscendSpecDecodeBaseProposer.set_inputs_first_pass`
|
||
# Why:
|
||
# 310P needs to protect the tail slot during MTP input_ids shift to avoid
|
||
# GatherV2 corruption from persistent drafter input buffers.
|
||
# How:
|
||
# Reuse the 310P proposer implementation for the first-pass input shift.
|
||
#
|
||
# 3. `vllm.model_executor.layers.mamba.gdn.qwen_gdn_linear_attn.QwenGatedDeltaNetAttention`
|
||
# Why:
|
||
# Qwen GDN needs 310P-specific state helpers, forward core, state dtype,
|
||
# and attention backend/builder wiring.
|
||
# How:
|
||
# Patch Qwen GDN methods to use Ascend GDN implementations and the 310P
|
||
# GDN attention backend. RC devices also route upstream GDNAttentionBackend
|
||
# to the 310P metadata builder.
|
||
# Related PR (if no, explain why):
|
||
# No, 310P custom operator and backend behavior are vllm-ascend specific.
|
||
# Future Plan:
|
||
# Remove this patch when upstream exposes stable hooks for 310P GDN
|
||
# chunk metadata, spec-decode input layout, and backend selection.
|
||
#
|
||
# ** 18. File: worker/patch_cudagraph.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.v1.cudagraph_dispatcher.CudagraphDispatcher._create_padded_batch_descriptor`
|
||
# Why:
|
||
# vllm's FULL mode will cause error, we use a patch to avoid it.
|
||
# After that, FULL can be enable now.
|
||
# How:
|
||
# Dynamically replace the `_create_padded_batch_descriptor` function at runtime,
|
||
# and change the condition of if.
|
||
# Related PR (if no, explain why):
|
||
# https://github.com/vllm-project/vllm/pull/34880
|
||
# Future Plan:
|
||
# Remove this patch when vLLM merges the PR.
|
||
#
|
||
# ** 19. File: worker/patch_deepseek_mtp.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.model_executor.models.deepseek_v2.get_spec_layer_idx_from_weight_name` and
|
||
# `vllm.model_executor.models.deepseek_mtp.get_spec_layer_idx_from_weight_name`
|
||
# Why:
|
||
# When GLM5 uses rotary quant in vllm-ascend, the MTP layer needs to load an extra weight
|
||
# named `rot.weight`.
|
||
# How:
|
||
# If weight name starts with `rot`, return `layer_id + i` like other tensors in MTP layer.
|
||
# Related PR (if no, explain why):
|
||
# Rotary quant is a unique feature of vllm-ascend.
|
||
# Future Plan:
|
||
# Remove this patch when vllm supports rotary quant or pluggable `MultiTokenPredictorLayer`.
|
||
# 2. `vllm.model_executor.models.deepseek_mtp.DeepSeekMultiTokenPredictorLayer`
|
||
# Why:
|
||
# When GLM5 uses rotary quant in vllm-ascend, the `previous_hidden_states` does not .
|
||
# How:
|
||
# If the target model uses rotary quant, a new linear operation is added before `ehnorm`.
|
||
# Related PR (if no, explain why):
|
||
# Rotary quant is a unique feature of vllm-ascend.
|
||
# Future Plan:
|
||
# Remove this patch when vllm supports rotary quant or pluggable `MultiTokenPredictorLayer`.
|
||
# 3. `vllm.model_executor.models.deepseek_mtp.DeepSeekMTP._rewrite_spec_layer_name`
|
||
# Why:
|
||
# Rename `rot.weight` to match the format of weights in `DeepSeekMTP`.
|
||
# How:
|
||
# If the weight name is `rot`, rename it to `model.layers.{spec_layer}.rot.weight`.
|
||
# Related PR (if no, explain why):
|
||
# Rotary quant is a unique feature of vllm-ascend.
|
||
# Future Plan:
|
||
# Remove this patch when vllm supports rotary quant or pluggable `MultiTokenPredictorLayer`.
|
||
# 4. `vllm.model_executor.models.deepseek_v2.GlmMoeDsaForCausalLM.load_weights`
|
||
# Why:
|
||
# After vllm PR #41706, GlmMoeDsaForCausalLM.load_weights uses `AutoWeightsLoader` which
|
||
# does not skip `rot.weight`, and will cause ValueError while loading weights.
|
||
# How:
|
||
# Use the `skip_prefixes` parameter to skip certain weight tensors.
|
||
# Related PR (if no, explain why):
|
||
# https://github.com/vllm-project/vllm/pull/41706
|
||
# Future Plan:
|
||
# Remove this patch when vllm supports rotary quant or pluggable `MultiTokenPredictorLayer`.
|
||
# ** 19a. File: worker/patch_deepseek_v2.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.model_executor.models.deepseek_v2.DeepseekV2MLAAttention.__init__`
|
||
# Why:
|
||
# GLM-5.2 checkpoints omit `Indexer` weights on shared-indexer layers,
|
||
# while GLM-5.1 IndexCache overrides only skip top-k computation and keep
|
||
# per-layer `Indexer` weights. Treating both layouts alike breaks GLM-5.1
|
||
# weight loading.
|
||
# How:
|
||
# Skip `Indexer` construction only when the layer both skips top-k and is
|
||
# explicitly marked `shared` in `indexer_types`. MTP layers always retain
|
||
# a complete `Indexer`.
|
||
# Related PR (if no, explain why):
|
||
# https://github.com/vllm-project/vllm/pull/45895
|
||
# Future Plan:
|
||
# Remove this patch when vLLM Ascend depends on a vLLM version that includes
|
||
# PR #45895.
|
||
#
|
||
# ** 19b. File: worker/model_runner_v1.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `NPUModelRunner._check_and_update_cudagraph_mode`
|
||
# Why:
|
||
# The upstream `GPUModelRunner._check_and_update_cudagraph_mode` initializes
|
||
# drafter cudagraph keys unconditionally, but in PP mode only the last rank
|
||
# loads the drafter. The previous hacky workaround temporarily set
|
||
# `self.speculative_config = None` to bypass super()'s drafter init, then
|
||
# restored it and called a separate `_maybe_initialize_drafter_cudagraph_keys`
|
||
# helper. This state-mutation pattern is fragile and hard to maintain.
|
||
# How:
|
||
# Directly inline the upstream cudagraph mode resolution logic with Ascend-specific
|
||
# additions: wrap `resolve_cudagraph_mode_and_sizes` with `update_pass_config` for
|
||
# `enable_sp`, add PP last-rank guard for drafter initialization, and call
|
||
# `set_graph_params`/`set_draft_graph_params` for ACL graph params. Remove the
|
||
# `_maybe_initialize_drafter_cudagraph_keys` helper entirely.
|
||
# Related PR (if no, explain why):
|
||
# No, cleaner PP+MTP support without speculative_config state mutation.
|
||
# Future Plan:
|
||
# Remove this override once upstream exposes a hook for drafter cudagraph key
|
||
# initialization that respects PP rank boundaries.
|
||
# ** 20. File: worker/patch_mamba_utils.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.v1.worker.mamba_utils.batch_memcpy_kernel = batch_memcpy_kernel`
|
||
# Why:
|
||
# Oringnal batch_memcpy_kernel implemented in vLLM might encounter bugs when running on
|
||
# Ascend hardwares.
|
||
# How:
|
||
# patch to fix related bugs.
|
||
# Future Plan:
|
||
# Remove this patch when:
|
||
# (1) oringnal batch_memcpy_kernel can run on Ascend hardware.
|
||
# or
|
||
# (2) design a dispatch mechanism for batch_memcpy_kernel.
|
||
# 2. `vllm.v1.worker.mamba_utils.batch_memcpy = batch_memcpy`
|
||
# Why:
|
||
# vLLM use BLOCK_SIZE 1024 for batch_memcpy_kernel. This results in suboptimal performance
|
||
# on Ascend hardwares.
|
||
# How:
|
||
# patch to change BLOCK_SIZE to 8192.
|
||
# Future Plan:
|
||
# Remove this patch when:
|
||
# design a dispatch mechanism for batch_memcpy_kernel.
|
||
# 3. `mamba_utils.preprocess_mamba = preprocess_mamba`
|
||
# Why:
|
||
# 1. preprocess_mamba has a assert logic, cause kv transfer call fails
|
||
# 2. preprocess_mamba copy the state of previous step to the last block before kv transfer load
|
||
# How:
|
||
# 1. patch to remove assert
|
||
# 2. path to only collect copy metadata in preprocess_mamba(and do actual copy after kv transfer load).
|
||
# Future Plan:
|
||
# Remove this patch when:
|
||
# vLLM itself supports kv transfer for mamba
|
||
# ** 21. File: worker/patch_weight_utils.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.model_executor.models.deepseek_v2.DeepseekV2ForCausalLM.load_weights`
|
||
# Why:
|
||
# The C8 weight quantized by modelslim will modify the model structure,
|
||
# and the scale and offset required for kvcache quantization will increase.
|
||
# In addition, the names of the quantization parameters are different from
|
||
# those in the community.
|
||
# How:
|
||
# we have enhanced the maybe_remap_kv_scale_name function.
|
||
# Future Plan:
|
||
# The maybe_remap_kv_scale_name function of the community is reconstructed to support
|
||
# multiple backends.
|
||
# ** 21b. File: worker/patch_process_weights_after_loading.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.model_executor.model_loader.utils.process_weights_after_loading`
|
||
# `vllm.model_executor.model_loader.base_loader.process_weights_after_loading`
|
||
# and imported references in vllm-ascend model loaders
|
||
# Why:
|
||
# DSA attention is implemented in vllm-ascend as the plugin layer
|
||
# `DSAAttention`. Upstream vLLM only runs post-load attention weight
|
||
# processing for built-in attention classes, so
|
||
# `DSAAttention.process_weights_after_loading()` is skipped in the
|
||
# original loader flow. DSV4 DSA-CP o-proj TP initialization must run in
|
||
# this post-load phase rather than being initialized lazily in forward.
|
||
# How:
|
||
# Rebind the upstream `process_weights_after_loading` helper, including
|
||
# already-imported loader references, so `DSAAttention` participates in
|
||
# the same post-load traversal while preserving the original quant-method
|
||
# and torchao reload behavior.
|
||
# Related PR (if no, explain why):
|
||
# https://github.com/vllm-project/vllm-ascend/pull/10694
|
||
# https://github.com/vllm-project/vllm/pull/46828
|
||
# Future Plan:
|
||
# Remove this patch once the supported vLLM version includes PR #46828.
|
||
# Then register `DSAAttention` through vLLM's post-load weight-processing
|
||
# registry instead of monkey-patching model-loader helpers.
|
||
# ** 22. File: worker/patch_v2/patch_input_batch.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.v1.worker.gpu.input_batch.InputBatch`
|
||
# Why:
|
||
# vllm use InputBatch to make dummy tensors. in `model_runner.py` and `cudagraph_utils.py`
|
||
# which make it difficult to inherit from vllm methods.
|
||
# How:
|
||
# replace InputBatch with AscendInputBatch.
|
||
# Future Plan:
|
||
# remove this patch when vLLM-ascend's make_dummy behavior aligns with vLLM.
|
||
# ** 23. File: worker/patch_v2/patch_block_table.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.v1.worker.gpu.block_table.BlockTables`
|
||
# Why:
|
||
## vllm-ascend need to initialize slot mapping as torch.int32 dtype,
|
||
# but vllm default is torch.int64 dtype.
|
||
# How:
|
||
# replace BlockTables with AscendBlockTables which initialize slot mapping
|
||
# as torch.int32 dtype.
|
||
# Future Plan:
|
||
# remove this patch when vLLM-ascend's BlockTables can initialize
|
||
# slot mapping as torch.int64 dtype.
|
||
# ** 24. File: worker/patch_v2/patch_model_state.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.v1.worker.gpu.model_states.default.init_model_state`
|
||
# Why:
|
||
## vllm's prepare_attn in ModelState is different from vllm,
|
||
# we need to override init_model_state.
|
||
# How:
|
||
# Define AscendModelState and initialize it in init_model_state.
|
||
# Future Plan:
|
||
# remove this when vllm-ascend's attention metadata is align with vllm.
|
||
# ** 25. File: worker/patch_v2/patch_triton.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.v1.worker.gpu.sample.logprob`, `vllm.v1.worker.gpu.sample.penalties.apply_penalties`,
|
||
# `vllm.v1.worker.gpu.sample.gumbel.gumbel_sample`
|
||
# Why:
|
||
# triton ops in vLLM perform not good on NPU. And there is no dispatch mechanism for triton ops.
|
||
# How:
|
||
# override triton ops in vLLM with ascend implementation
|
||
# Related PR (if no, explain why):
|
||
# Let vLLM support triton ops dispatch.
|
||
# Future Plan:
|
||
# Remove this patch when vLLM support the dispatch function.
|
||
#
|
||
# ** 26. File: worker/patch_gqa_c8.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.model_executor.models.qwen3.Qwen3ForCausalLM.load_weights`
|
||
# Why:
|
||
# The GQA W8A8C8 model stores per-channel KV cache scales and offsets
|
||
# (k_cache_scale, k_cache_offset, v_cache_scale, v_cache_offset) under
|
||
# weight names that AutoWeightsLoader does not recognise and would
|
||
# silently discard. Without these scales the INT8 KV cache cannot be
|
||
# dequantised correctly at inference time.
|
||
# How:
|
||
# Wrap load_weights to intercept the C8 scale/offset tensors before they
|
||
# reach the base loader. Each intercepted tensor is routed to the
|
||
# corresponding nn.Parameter via its weight_loader, then excluded from
|
||
# the remaining weight stream so the base loader never sees it.
|
||
# Related PR (if no, explain why):
|
||
# This PR (Qwen3-32B and GLM4.7 W8A8C8 support). Upstream vLLM's weight-loading
|
||
# pipeline does not yet have a generic hook for hardware-plugin-defined
|
||
# KV cache parameters.
|
||
# Future Plan:
|
||
# Remove this patch when vLLM provides a first-class extension point
|
||
# for loading extra KV cache quantisation parameters in model load_weights,
|
||
# or when the GQA model's weight names are aligned with the parameter
|
||
# names expected by the quantisation backend.
|
||
# ** 27. File: worker/patch_qwen3vl.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.model_executor.models.qwen3.Qwen3Attention.forward` and
|
||
# `vllm.model_executor.models.qwen3_moe.Qwen3MoeAttention.forward`
|
||
# Why:
|
||
# support triton_split_qkv_rmsnorm_mrope fused kernel for Qwen3Attention and Qwen3MoeAttention.
|
||
# How:
|
||
# override forward method with the triton_split_qkv_rmsnorm_mrope fused kernel,
|
||
# when using mrope.
|
||
# Future Plan:
|
||
# Remove this patch when vllm-ascend supports pattern matching for this fused kernel.
|
||
# ** 28. File: worker/patch_qwen3_dflash.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.model_executor.models.qwen3_dflash.DFlashQwen3Model.precompute_and_store_context_kv`
|
||
# Why:
|
||
# The function directly calls the ops.rms_norm and ops.rotary_imbedding operators,
|
||
# but NPU does not have a corresponding implementation.
|
||
# How:
|
||
# Replace ops.* with the internal implementation of vllm-ascend.
|
||
# Future Plan:
|
||
# Remove this patch when vllm-ascend supports pattern matching for ops.*.
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.model_executor.layers.fused_moe.routed_experts_capturer.RoutedExpertsCapturer.capture`
|
||
# Why:
|
||
# The upstream implementation doesn't support vllm-ascend specific MoE communication types
|
||
# (ALLTOALL and MC2). In the SP + modular-kernel path, the original code cannot correctly
|
||
# handle tensor splitting and all-gather operations on NPU, especially when tokens are
|
||
# unevenly distributed across TP ranks or padded to max_tokens in MC2 mode.
|
||
# How:
|
||
# Override the capture method to add support for vllm-ascend's MoECommType:
|
||
# - Check `_EXTRA_CTX.moe_comm_type` to determine if ALLTOALL or MC2 mode is active
|
||
# - Calculate correct gather_topk_ids_shape based on communication type:
|
||
# * ALLTOALL: uses actual token_num_per_dp for shape calculation
|
||
# * MC2: uses padded max_tokens * tp_size for shape calculation
|
||
# - Properly handle tensor_split and all_gather operations for NPU distributed communication
|
||
# Future Plan:
|
||
# Remove this patch when upstream vLLM supports MoE communication type abstraction that
|
||
# can be extended by hardware plugins like vllm-ascend.
|
||
#
|
||
# ** 29. File: platform/patch_mamba_manager.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.v1.core.single_type_kv_cache_manager.MambaManager`
|
||
# Why:
|
||
# 1. Upstream hybrid prefix cache lookup does not support PCP/DCP.
|
||
# 2. Upstream MambaManager#get_num_blocks_to_allocate give the
|
||
# wrong number of blocks when an external cache hit occurred
|
||
# How:
|
||
# 1. Replace MambaManager with AscendMambaManager for prefix cache hit lookup
|
||
# on hybrid Mamba paths (logical mamba block_size when caching is enabled).
|
||
# 2. Override the get_num_blocks_to_allocate method to fix the number of blocks
|
||
# when hitting the external cache and loading synchronously
|
||
# Related PR (if no, explain why):
|
||
# 1. https://github.com/vllm-project/vllm/pull/40996
|
||
# 2. https://github.com/vllm-project/vllm/pull/46892
|
||
# Future Plan:
|
||
# 1. Upstream PR #40996 adds hybrid prefix cache lookup for DCP only; PCP is
|
||
# not supported yet. Remove this patch once upstream supports both PCP and DCP.
|
||
# 2. Remove this patch once upstream accept 46892 pr or fixed the bug by other pr.
|
||
#
|
||
# ** 30. File: platform/patch_use_v2_model_runner.py**
|
||
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||
# 1. `vllm.config.vllm.VllmConfig.use_v2_model_runner`
|
||
# Why:
|
||
# Upstream vLLM enables the v2 model runner not only via the
|
||
# VLLM_USE_V2_MODEL_RUNNER env var but also based on model
|
||
# architecture whitelists, Triton availability, and feature
|
||
# compatibility checks. On Ascend the NPU v2 runner is not yet
|
||
# compatible with all upstream-defaulted models and features, so
|
||
# enabling by model architecture can crash. We override the
|
||
# property to read only VLLM_USE_V2_MODEL_RUNNER, deferring
|
||
# model/framework checks to the NPU runner itself.
|
||
# How:
|
||
# Monkey-patch VllmConfig.use_v2_model_runner to return
|
||
# envs.VLLM_USE_V2_MODEL_RUNNER (defaulting to False when unset).
|
||
# worker/patch_v2/patch_use_v2_model_runner.py reuses this platform
|
||
# patch so EngineCore and worker processes share the same behavior.
|
||
# Related PR (if no, explain why):
|
||
# 1. https://github.com/vllm-project/vllm-ascend/pull/11389
|
||
# Future Plan:
|
||
# Remove this patch once vllm-ascend fully supports the v2 model
|
||
# runner and can rely on upstream's default enablement heuristics
|
||
# (model architecture, Triton, feature checks) without crashes or
|
||
# degraded functionality.
|