xc-llm-ascend/vllm_ascend/patch/__init__.py

#
# Copyright (c) 2025 Huawei Technologies Co., Ltd. All Rights Reserved.
# This file is a part of the vllm-ascend project.
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
#     http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.

# ----------------------------------------------------------------------------------
# This module manage the patch for vllm. There are two folders in this module:
# - platform: contains the patches applied before worker starts. It's called by
#             `vllm_ascend.utils.adapt_patch(is_global_patch=True)` in
#             `vllm_ascend.platform.NPUPlatform.pre_register_and_update()` function.
# - worker: contains the patches applied when worker starts. It's called by
#           `vllm_ascend.utils.adapt_patch(is_global_patch=False)` in
#           each worker's `__init__` function.
#
# Then in each kind of patch, there are three folders:
# - patch_0_8_5: contains the patches applied when vllm version is 0.8.5.
# - patch_main: contains the patches applied when vllm version is main branch.
# - patch_common: contains the patches applied in both 0.8.5 and main branch.
#
# In the future, with the vllm version upgrade, the new patch folder such as
# patch_0_8_5, patch_0_8_6, etc. will be added to manage the patch for different
# vllm version. And the patch_common will contain the patches applied in all the
# vllm version.
# Once the vllm version is too old that vllm-ascend will not support, the related
# patch folder will be removed as well.
#
# Once a new patch is added in vllm-ascend, please add the patch description into this file as well.
# ----------------------------------------------------------------------------------

# What's Patched and how it works:
# --------------------------------
# * Platform Patch:
# =================
# ** File: platform/patch_common/patch_distributed.py**
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
#   1. `vllm.distributed.parallel_state.destroy_model_parallel()`
#    Why:
#       vllm dose not support outside platform maintain its own `CoordinatorGroup`, vllm-ascend maintain EP and ETP
#       inside of the repo, and needs a common interface to destroy them, this patch add the interface of destroy
#       platform owned `CoordinatorGroup` to make sure all the CoordinateGroup can be properly destroyed
#    How：
#       Call `vllm_ascend.distributed.parallel_state method `destroy_platform_model_parallel` to destroy all the `CoordinateGroup`
#    Related PR (if no, explain why): no related PR, we want add this ability into vllm
#    Future Plan:
#       Remove those patch when vllm merged them
#   2. `vllm.distributed.stateless_init_torch_distributed_process_group()`
#    Why:
#       The stateless process group can not be initialized except from gloo and nccl backend, vllm-ascend
#       needs to initialize its own stateless process group for communication, so we add the platform related
#       call to the `stateless_init_torch_distributed_process_group`, to enable other platform which may support
#       stateless process group initialize method
#    How：
#       rewrite stateless_init_torch_distributed_process_group to judge if there is a stateless process group initialize
#       method and call platform method `platform_register_backend` to initialize them
#    Related PR (if no, explain why): no related PR, we want add this ability into vllm
#    Future Plan:
#       Remove those patch when vllm merged them
#   3. `ParallelConfig.get_next_dp_init_port`
#    Why:
#       We want to get dp port from env variable, so the multi-node inference can be properly initialized and run.
#    How：
#       Get the dp port from env variable enable multi-mode dp inference
#    Related PR (if no, explain why): no related PR, we want add this ability into vllm
#    Future Plan:
#       Its a workaround in vllm-ascend to enable multi-node dp inference, maybe removed if vllm have better plan
#       on multi-node dp inference implementation
#   4. `ParallelConfig.stateless_init_dp_group`
#    Why:
#       vLLM use gloo backend by default to initialize stateless dp process gourp, but we want to use hccl here to
#       get better performance
#    How：
#       adopt nccl backend to init process group
#    Related PR (if no, explain why): no related PR, we want add this ability into vllm
#    Future Plan:
#       Remove those patch when vllm merged them
#
#
# * Worker Patch:
# ===============
# ** File: worker/patch_common/patch_metrics.py **
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
#   1. `vllm.spec_decode.metrics.AsyncMetricsCollector.maybe_collect_rejsample_metrics`
#    Why:
#       There are cuda hard code (current_platform.is_cuda_alike()) in
#       `AsyncMetricsCollector.maybe_collect_rejsample_metrics`
#    How：
#       Change to use `current_platform.Event` to determine whether to return None
#    Related PR (if no, explain why): 1. refused by vllm. 2. vllm doesn't support 3. prepare to submit....
#       https://github.com/vllm-project/vllm/pull/14411
#    Future Plan:
#       Revert it when the related pr is merged in vllm.
#
# ** File: worker/patch_common/patch_minicpm.py **
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
#   1. `vllm.model_executor.models.minicpm.MiniCPMAttention.forward`
#    Why:
#       The forward func of MiniCPMAttention in vllm do a datatype convert
#       (original datatype --> float32) to ensure the precision on cuda.
#       However float32 is not supported in cann rope op, thus we keep this patch
#    How：
#       Removed the dtype convert operations in forward
#    Related PR (if no, explain why): 1. refused by vllm. 2. vllm doesn't support 3. prepare to submit....
#       NO, only for npu due to rope op.
#    Future Plan:
#       Keep this patch in vllm-ascend.
#
# ** File: worker/patch_common/patch_multi_step_worker.py **
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
#   1. `vllm.spec_decode.multi_step_worker.MultiStepWorker.sampler_output`
#    Why:
#       There are cuda hard code (current_platform.is_cuda_alike()) in
#       `MultiStepWorker.sampler_output`, and we need to use the patched `TP1DraftModelRunner` in it.
#    How：
#       Make speculative decoding extensible to different backends.
#       - support attention metadata register to the set supported spec decode
#       - offer a api in platform to determine whether spec decode is supported,
#         and deprecate is_cuda_alike in it.
#    Related PR (if no, explain why): 1. refused by vllm. 2. vllm doesn't support 3. prepare to submit....
#       - https://github.com/vllm-project/vllm/pull/15195
#       - https://github.com/vllm-project/vllm-ascend/pull/395
#    Future Plan:
#       Revert it when the related pr is merged in vllm and vllm-ascend.
#
#   2. `vllm.spec_decode.multi_step_worker.MultiStepWorker.set_include_gpu_probs_tensor` and
#       `vllm.spec_decode.multi_step_worker.MultiStepWorker.set_should_modify_greedy_probs_inplace`
#    Why:
#       vLLM `Remove Sampler from Model Code` so vllm-ascend needs adapt to this change.
#    How：
#       Use vLLM 0.8.4 method to patch it.
#    Related PR (if no, explain why): 1. refused by vllm. 2. vllm doesn't support 3. prepare to submit....
#       - https://github.com/vllm-project/vllm/pull/15195
#       - https://github.com/vllm-project/vllm-ascend/pull/395
#    Future Plan:
#       Remove it when we identify the reasons clearly.
#
# ** File: worker/patch_common/patch_spec_decode_worker.py **
# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
#   1. `vllm.spec_decode.spec_decode_worker.SpecDecodeWorker.create_worker`
#    Why:
#       We need to use the patched `TP1DraftModelRunner` in `SpecDecodeWorker.create_worker`.
#       The mainly reason to overwrite `TP1DraftModelRunner`is the hard code of
#           `FlashAttentionMetadata`
#    How：
#       ditto
#    Related PR (if no, explain why): 1. refused by vllm. 2. vllm doesn't support 3. prepare to submit....
#       - https://github.com/vllm-project/vllm/pull/15195
#       - https://github.com/vllm-project/vllm-ascend/pull/395
#    Future Plan:
#       Revert it when the related pr is merged in vllm and vllm-ascend.
#
-												port deepseekv2 and mtp to main branch (#429)

### What this PR does / why we need it?
This PR ports all the deepseek graph mode code and mtp code from v0.7.3
to the main branch
---------

Signed-off-by: SidaoY <1024863041@qq.com>
Signed-off-by: linfeng-yuan <1102311262@qq.com>
Signed-off-by: Yizhou Liu <liuyizhou5@h-partners.com>
Signed-off-by: mengwei805 <mengwei25@huawei.com>
Signed-off-by: libaokui <libaokui@huawei.com>
Signed-off-by: q00832892 <qiaoyang19@huawei.com>
Signed-off-by: ganyi <pleaplusone.gy@gmail.com>
Co-authored-by: SidaoY <1024863041@qq.com>
Co-authored-by: linfeng-yuan <1102311262@qq.com>
Co-authored-by: Yizhou Liu <liuyizhou5@h-partners.com>
Co-authored-by: mengwei805 <mengwei25@huawei.com>
Co-authored-by: libaokui <libaokui@huawei.com>
											
										
										
											2025-04-19 17:38:18 +08:00
+								#
 								# Copyright (c) 2025 Huawei Technologies Co., Ltd. All Rights Reserved.
 								# This file is a part of the vllm-ascend project.
 								#
 								# Licensed under the Apache License, Version 2.0 (the "License");
 								# you may not use this file except in compliance with the License.
 								# You may obtain a copy of the License at
 								#
 								#     http://www.apache.org/licenses/LICENSE-2.0
 								#
 								# Unless required by applicable law or agreed to in writing, software
 								# distributed under the License is distributed on an "AS IS" BASIS,
 								# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
 								# See the License for the specific language governing permissions and
 								# limitations under the License.
-												[Patch] format patch module to make it more clear (#601)

Format patch module to make it more clear. 
Add the patch doc description, the new patch must follow this guide.

Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>
											
										
										
											2025-04-22 14:13:00 +08:00
 								# ----------------------------------------------------------------------------------
 								# This module manage the patch for vllm. There are two folders in this module:
 								# - platform: contains the patches applied before worker starts. It's called by
 								#             `vllm_ascend.utils.adapt_patch(is_global_patch=True)` in
 								#             `vllm_ascend.platform.NPUPlatform.pre_register_and_update()` function.
 								# - worker: contains the patches applied when worker starts. It's called by
 								#           `vllm_ascend.utils.adapt_patch(is_global_patch=False)` in
 								#           each worker's `__init__` function.
 								#
 								# Then in each kind of patch, there are three folders:
-												[CI] upgrade vllm to 0.8.5 (#715)

1. Upgrade vllm to 0.8.5
2. Drop 0.8.4 support
3. Keep doc to 0.8.4rc2 until we release 0.8.5

Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>
											
										
										
											2025-04-30 09:15:50 +08:00
+								# - patch_0_8_5: contains the patches applied when vllm version is 0.8.5.
-												[Patch] format patch module to make it more clear (#601)

Format patch module to make it more clear. 
Add the patch doc description, the new patch must follow this guide.

Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>
											
										
										
											2025-04-22 14:13:00 +08:00
+								# - patch_main: contains the patches applied when vllm version is main branch.
-												[CI] upgrade vllm to 0.8.5 (#715)

1. Upgrade vllm to 0.8.5
2. Drop 0.8.4 support
3. Keep doc to 0.8.4rc2 until we release 0.8.5

Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>
											
										
										
											2025-04-30 09:15:50 +08:00
+								# - patch_common: contains the patches applied in both 0.8.5 and main branch.
-												[Patch] format patch module to make it more clear (#601)

Format patch module to make it more clear. 
Add the patch doc description, the new patch must follow this guide.

Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>
											
										
										
											2025-04-22 14:13:00 +08:00
+								#
 								# In the future, with the vllm version upgrade, the new patch folder such as
 								# patch_0_8_5, patch_0_8_6, etc. will be added to manage the patch for different
 								# vllm version. And the patch_common will contain the patches applied in all the
 								# vllm version.
 								# Once the vllm version is too old that vllm-ascend will not support, the related
 								# patch folder will be removed as well.
 								#
 								# Once a new patch is added in vllm-ascend, please add the patch description into this file as well.
 								# ----------------------------------------------------------------------------------
 								# What's Patched and how it works:
 								# --------------------------------
 								# * Platform Patch:
 								# =================
 								# ** File: platform/patch_common/patch_distributed.py**
 								# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
 								#   1. `vllm.distributed.parallel_state.destroy_model_parallel()`
 								#    Why:
 								#       vllm dose not support outside platform maintain its own `CoordinatorGroup`, vllm-ascend maintain EP and ETP
 								#       inside of the repo, and needs a common interface to destroy them, this patch add the interface of destroy
 								#       platform owned `CoordinatorGroup` to make sure all the CoordinateGroup can be properly destroyed
 								#    How：
-												[Platform] format platform to make it more clear (#610)

Platform should only contain the function that based from vllm. This PR
move the unrelated function to the right place to make platform more
clear.

Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>
											
										
										
											2025-04-30 09:03:10 +08:00
+								#       Call `vllm_ascend.distributed.parallel_state method `destroy_platform_model_parallel` to destroy all the `CoordinateGroup`
-												[Patch] format patch module to make it more clear (#601)

Format patch module to make it more clear. 
Add the patch doc description, the new patch must follow this guide.

Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>
											
										
										
											2025-04-22 14:13:00 +08:00
+								#    Related PR (if no, explain why): no related PR, we want add this ability into vllm
 								#    Future Plan:
 								#       Remove those patch when vllm merged them
 								#   2. `vllm.distributed.stateless_init_torch_distributed_process_group()`
 								#    Why:
 								#       The stateless process group can not be initialized except from gloo and nccl backend, vllm-ascend
 								#       needs to initialize its own stateless process group for communication, so we add the platform related
 								#       call to the `stateless_init_torch_distributed_process_group`, to enable other platform which may support
 								#       stateless process group initialize method
 								#    How：
-												[Platform] format platform to make it more clear (#610)

Platform should only contain the function that based from vllm. This PR
move the unrelated function to the right place to make platform more
clear.

Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>
											
										
										
											2025-04-30 09:03:10 +08:00
+								#       rewrite stateless_init_torch_distributed_process_group to judge if there is a stateless process group initialize
-												[Patch] format patch module to make it more clear (#601)

Format patch module to make it more clear. 
Add the patch doc description, the new patch must follow this guide.

Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>
											
										
										
											2025-04-22 14:13:00 +08:00
+								#       method and call platform method `platform_register_backend` to initialize them
 								#    Related PR (if no, explain why): no related PR, we want add this ability into vllm
 								#    Future Plan:
 								#       Remove those patch when vllm merged them
 								#   3. `ParallelConfig.get_next_dp_init_port`
 								#    Why:
 								#       We want to get dp port from env variable, so the multi-node inference can be properly initialized and run.
 								#    How：
 								#       Get the dp port from env variable enable multi-mode dp inference
 								#    Related PR (if no, explain why): no related PR, we want add this ability into vllm
 								#    Future Plan:
 								#       Its a workaround in vllm-ascend to enable multi-node dp inference, maybe removed if vllm have better plan
 								#       on multi-node dp inference implementation
-												Add dp initialize patch with hccl backend (#626)

<!--  Thanks for sending a pull request!

BEFORE SUBMITTING, PLEASE READ
https://docs.vllm.ai/en/latest/contributing/overview.html

-->
### What this PR does / why we need it?
<!--
- Please clarify what changes you are proposing. The purpose of this
section is to outline the changes and how this PR fixes the issue.
If possible, please consider writing useful notes for better and faster
reviews in your PR.

- Please clarify why the changes are needed. For instance, the use case
and bug description.

- Fixes #
-->
Add dp stateless process group initialization path with hccl backend as
vllm-ascend patch.
### Does this PR introduce _any_ user-facing change?
<!--
Note that it means *any* user-facing change including all aspects such
as API, interface or other behavior changes.
Documentation-only updates are not considered user-facing changes.
-->

### How was this patch tested?
<!--
CI passed with new added/existing test.
If it was tested in a way different from regular unit tests, please
clarify how you tested step by step, ideally copy and paste-able, so
that other reviewers can test and check, and descendants can verify in
the future.
If tests were not added, please describe why they were not added and/or
why it was difficult to add.
-->

---------

Signed-off-by: ganyi <pleaplusone.gy@gmail.com>
											
										
										
											2025-04-23 15:47:51 +08:00
+								#   4. `ParallelConfig.stateless_init_dp_group`
 								#    Why:
 								#       vLLM use gloo backend by default to initialize stateless dp process gourp, but we want to use hccl here to
 								#       get better performance
 								#    How：
 								#       adopt nccl backend to init process group
 								#    Related PR (if no, explain why): no related PR, we want add this ability into vllm
 								#    Future Plan:
 								#       Remove those patch when vllm merged them
-												[Model][MiniCPM] support MiniCPM (#645)

### What this PR does / why we need it?
This pr support minicpm in branch main. see
https://github.com/vllm-project/vllm-ascend/pull/164


### How was this patch tested?
test locally with minicpm

---------

Signed-off-by: MengqingCao <cmq0113@163.com>
											
										
										
											2025-04-27 11:27:24 +08:00
+								#
 								#
-												[Patch] format patch module to make it more clear (#601)

Format patch module to make it more clear. 
Add the patch doc description, the new patch must follow this guide.

Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>
											
										
										
											2025-04-22 14:13:00 +08:00
+								# * Worker Patch:
 								# ===============
-												[Misc] Remove some parts of metrics patch (#603)

### What this PR does / why we need it?
Remove some parts of metrics patch, since the `cuda` hard code has been
fixed by https://github.com/vllm-project/vllm/pull/14411.

Signed-off-by: shen-shanshan <467638484@qq.com>
											
										
										
											2025-04-22 18:45:21 +08:00
+								# ** File: worker/patch_common/patch_metrics.py **
 								# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
 								#   1. `vllm.spec_decode.metrics.AsyncMetricsCollector.maybe_collect_rejsample_metrics`
-												[Patch] format patch module to make it more clear (#601)

Format patch module to make it more clear. 
Add the patch doc description, the new patch must follow this guide.

Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>
											
										
										
											2025-04-22 14:13:00 +08:00
+								#    Why:
 								#       There are cuda hard code (current_platform.is_cuda_alike()) in
 								#       `AsyncMetricsCollector.maybe_collect_rejsample_metrics`
 								#    How：
 								#       Change to use `current_platform.Event` to determine whether to return None
 								#    Related PR (if no, explain why): 1. refused by vllm. 2. vllm doesn't support 3. prepare to submit....
 								#       https://github.com/vllm-project/vllm/pull/14411
 								#    Future Plan:
 								#       Revert it when the related pr is merged in vllm.
 								#
-												[Model][MiniCPM] support MiniCPM (#645)

### What this PR does / why we need it?
This pr support minicpm in branch main. see
https://github.com/vllm-project/vllm-ascend/pull/164


### How was this patch tested?
test locally with minicpm

---------

Signed-off-by: MengqingCao <cmq0113@163.com>
											
										
										
											2025-04-27 11:27:24 +08:00
+								# ** File: worker/patch_common/patch_minicpm.py **
 								# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
 								#   1. `vllm.model_executor.models.minicpm.MiniCPMAttention.forward`
 								#    Why:
 								#       The forward func of MiniCPMAttention in vllm do a datatype convert
 								#       (original datatype --> float32) to ensure the precision on cuda.
 								#       However float32 is not supported in cann rope op, thus we keep this patch
 								#    How：
 								#       Removed the dtype convert operations in forward
 								#    Related PR (if no, explain why): 1. refused by vllm. 2. vllm doesn't support 3. prepare to submit....
 								#       NO, only for npu due to rope op.
 								#    Future Plan:
 								#       Keep this patch in vllm-ascend.
 								#
-												[Patch] format patch module to make it more clear (#601)

Format patch module to make it more clear. 
Add the patch doc description, the new patch must follow this guide.

Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>
											
										
										
											2025-04-22 14:13:00 +08:00
+								# ** File: worker/patch_common/patch_multi_step_worker.py **
 								# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
 								#   1. `vllm.spec_decode.multi_step_worker.MultiStepWorker.sampler_output`
 								#    Why:
 								#       There are cuda hard code (current_platform.is_cuda_alike()) in
 								#       `MultiStepWorker.sampler_output`, and we need to use the patched `TP1DraftModelRunner` in it.
 								#    How：
 								#       Make speculative decoding extensible to different backends.
 								#       - support attention metadata register to the set supported spec decode
 								#       - offer a api in platform to determine whether spec decode is supported,
 								#         and deprecate is_cuda_alike in it.
 								#    Related PR (if no, explain why): 1. refused by vllm. 2. vllm doesn't support 3. prepare to submit....
 								#       - https://github.com/vllm-project/vllm/pull/15195
 								#       - https://github.com/vllm-project/vllm-ascend/pull/395
 								#    Future Plan:
 								#       Revert it when the related pr is merged in vllm and vllm-ascend.
 								#
-												[MTP] follow custom deepseek modeling changes to support graph mode (#636)

<!--  Thanks for sending a pull request!

BEFORE SUBMITTING, PLEASE READ
https://docs.vllm.ai/en/latest/contributing/overview.html

-->
### What this PR does / why we need it?

As custom deepseek modeling do some changes to support graph mode in
https://github.com/vllm-project/vllm-ascend/pull/585, so i follow it to
change custom deepseek_mtp modeling.

And some modifications for k>1 were not carried over by the
https://github.com/vllm-project/vllm-ascend/pull/429, now i add it.

In order to better take care of the MTP feature in the vllm-ascend
repository, I added cases related to graph mode(torchair), but i skip it
since torchair can not correctly clean up memory in vllmrunner.

Also i add some case for MTP quantization weights, but test weight is
not ready, so i skip it and i will open it when test quant weights is
ready.

https://github.com/vllm-project/vllm-ascend/pull/648 did not completely
fix the sample
change(https://github.com/vllm-project/vllm-ascend/issues/660) issue, I
added the relevant changes.

### Does this PR introduce _any_ user-facing change?
now, u can use following method to use mtp in deepseek v3/r1 float or
quant weights with eager mode.
```python
llm = LLM(
    model="wemaster/deepseek_mtp_main_random_bf16",
    tensor_parallel_size=2,
    speculative_config={
        "num_speculative_tokens": 1,
    },
    enforce_eager=True,
    trust_remote_code=True,
    disable_log_stats=False,
    gpu_memory_utilization=0.8,
    max_model_len=64,
)
```

or use mtp in deepseek v3/r1 float or quant weights with graph
mode（torchair）
```python
llm = LLM(
    model="wemaster/deepseek_mtp_main_random_bf16",
    tensor_parallel_size=2,
    speculative_config={
        "num_speculative_tokens": 1,
    },
    trust_remote_code=True,
    additional_config={
        'enable_graph_mode': True,
    },
    disable_log_stats=False,
    gpu_memory_utilization=0.8,
    max_model_len=64,
)
```

add notes:
1. now, we support k>1, so u can set num_speculative_tokens > 1 if there
is sufficient redundant computing power;
2. MTP is not supported in V1, we will support it when vLLM does it in
https://github.com/vllm-project/vllm/issues/13500.
3. if u run MTP failed by `segmentation fault`, u can follow v0.7.3
patch https://github.com/vllm-project/vllm-ascend/pull/236 file
`vllm_ascend/patch/patch_metrics.py` method
`__npu_async_metrics_collector_init__`

### How was this patch tested?
local tested passed and test by CI

Signed-off-by: mengwei805 <mengwei25@huawei.com>
											
										
										
											2025-04-28 21:18:53 +08:00
+								#   2. `vllm.spec_decode.multi_step_worker.MultiStepWorker.set_include_gpu_probs_tensor` and
 								#       `vllm.spec_decode.multi_step_worker.MultiStepWorker.set_should_modify_greedy_probs_inplace`
 								#    Why:
 								#       vLLM `Remove Sampler from Model Code` so vllm-ascend needs adapt to this change.
 								#    How：
 								#       Use vLLM 0.8.4 method to patch it.
 								#    Related PR (if no, explain why): 1. refused by vllm. 2. vllm doesn't support 3. prepare to submit....
 								#       - https://github.com/vllm-project/vllm/pull/15195
 								#       - https://github.com/vllm-project/vllm-ascend/pull/395
 								#    Future Plan:
 								#       Remove it when we identify the reasons clearly.
 								#
 								# ** File: worker/patch_common/patch_spec_decode_worker.py **
-												[Patch] format patch module to make it more clear (#601)

Format patch module to make it more clear. 
Add the patch doc description, the new patch must follow this guide.

Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>
											
										
										
											2025-04-22 14:13:00 +08:00
+								# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
 								#   1. `vllm.spec_decode.spec_decode_worker.SpecDecodeWorker.create_worker`
 								#    Why:
 								#       We need to use the patched `TP1DraftModelRunner` in `SpecDecodeWorker.create_worker`.
 								#       The mainly reason to overwrite `TP1DraftModelRunner`is the hard code of
 								#           `FlashAttentionMetadata`
 								#    How：
 								#       ditto
 								#    Related PR (if no, explain why): 1. refused by vllm. 2. vllm doesn't support 3. prepare to submit....
 								#       - https://github.com/vllm-project/vllm/pull/15195
 								#       - https://github.com/vllm-project/vllm-ascend/pull/395
 								#    Future Plan:
 								#       Revert it when the related pr is merged in vllm and vllm-ascend.
-												[Core] Cleanup triton patch which has been fixed in vllm (#764)

### What this PR does / why we need it?
- Revert "Re-patch TritonPlaceholder on main to make CI happy (#753)"
because upstream main CI already merged:
https://github.com/vllm-project/vllm/pull/17446
- Keep 0.8.5.post1 compatible

### Does this PR introduce _any_ user-facing change?
No

### How was this patch tested?
CI passed

---------

Signed-off-by: Yikun Jiang <yikunkero@gmail.com>
											
										
										
											2025-05-06 18:52:15 +08:00
+								#