xc-llm-ascend/vllm_ascend/spec_decode/interface.py

import enum
from typing import Optional

import torch
from vllm.config import CUDAGraphMode, VllmConfig
from vllm.v1.core.sched.output import SchedulerOutput
from vllm.v1.sample.metadata import SamplingMetadata
from vllm.v1.spec_decode.metadata import SpecDecodeMetadata


class SpecDcodeType(enum.Enum):
    NGRAM = 0
    EAGLE = 1
    EAGLE3 = 2
    MTP = 4


class Proposer:

    def __init__(self,
                 vllm_config: VllmConfig,
                 device: torch.device = None,
                 runner=None):
        pass

    def load_model(self, model):
        """Called by load_model in model_runner"""
        raise NotImplementedError

    @torch.inference_mode()
    def dummy_run(self,
                  num_tokens: int,
                  with_prefill: bool = False,
                  skip_attn: bool = False,
                  num_reqs: int = 0,
                  num_tokens_across_dp: Optional[torch.Tensor] = None,
                  aclgraph_runtime_mode: CUDAGraphMode = CUDAGraphMode.NONE,
                  batch_descriptor=None,
                  dummy_compute_logits=lambda hidden_states: None):
        """Called by dummy_run in modle_runner"""
        raise NotImplementedError

    def generate_token_ids(self,
                           valid_sampled_token_ids: list[list[int]],
                           sampling_metadata: SamplingMetadata = None,
                           scheduler_output: SchedulerOutput = None,
                           spec_decode_metadata: SpecDecodeMetadata = None,
                           positions: torch.Tensor = None,
                           num_scheduled_tokens: int = 0,
                           hidden_states: torch.Tensor = None,
                           attn_metadata=None,
                           aux_hidden_states: torch.Tensor = None):
        """Called by execute_model in model_runner"""
        raise NotImplementedError
[Refactor] Refactor Spec Decode (#2668) ### What this PR does / why we need it? Refactor spec decode ### Does this PR introduce _any_ user-facing change? N/A ### How was this patch tested? CI passed with new added/existing test. - vLLM version: v0.10.1.1 - vLLM main: https://github.com/vllm-project/vllm/commit/6997a25ac65ed6cc3c2be6d09ca45f633a345f63 --------- Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com> Signed-off-by: Icey <1790571317@qq.com> Co-authored-by: wangxiyuan <wangxiyuan1007@gmail.com> 2025-09-04 11:34:47 +08:00			`import enum`
			`from typing import Optional`

			`import torch`
[Feat]mtp aclgraph support (#3244) ### What this PR does / why we need it? Currently, MTP Model in deepseek can not be capture in ACLGraph. This PR is use to allow MTP to be captured in ACLGraph mode. ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM version: v0.11.0rc3 - vLLM main: https://github.com/vllm-project/vllm/commit/v0.11.0 Signed-off-by: anon189Ty <Stari_Falcon@outlook.com> 2025-10-17 18:14:49 +08:00			`from vllm.config import CUDAGraphMode, VllmConfig`
[Refactor] Refactor Spec Decode (#2668) ### What this PR does / why we need it? Refactor spec decode ### Does this PR introduce _any_ user-facing change? N/A ### How was this patch tested? CI passed with new added/existing test. - vLLM version: v0.10.1.1 - vLLM main: https://github.com/vllm-project/vllm/commit/6997a25ac65ed6cc3c2be6d09ca45f633a345f63 --------- Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com> Signed-off-by: Icey <1790571317@qq.com> Co-authored-by: wangxiyuan <wangxiyuan1007@gmail.com> 2025-09-04 11:34:47 +08:00			`from vllm.v1.core.sched.output import SchedulerOutput`
			`from vllm.v1.sample.metadata import SamplingMetadata`
			`from vllm.v1.spec_decode.metadata import SpecDecodeMetadata`


			`class SpecDcodeType(enum.Enum):`
			`NGRAM = 0`
			`EAGLE = 1`
			`EAGLE3 = 2`
			`MTP = 4`


			`class Proposer:`

			`def __init__(self,`
			`vllm_config: VllmConfig,`
			`device: torch.device = None,`
			`runner=None):`
			`pass`

			`def load_model(self, model):`
			`"""Called by load_model in model_runner"""`
			`raise NotImplementedError`

			`@torch.inference_mode()`
			`def dummy_run(self,`
			`num_tokens: int,`
			`with_prefill: bool = False,`
			`skip_attn: bool = False,`
			`num_reqs: int = 0,`
[Feat]mtp aclgraph support (#3244) ### What this PR does / why we need it? Currently, MTP Model in deepseek can not be capture in ACLGraph. This PR is use to allow MTP to be captured in ACLGraph mode. ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? - vLLM version: v0.11.0rc3 - vLLM main: https://github.com/vllm-project/vllm/commit/v0.11.0 Signed-off-by: anon189Ty <Stari_Falcon@outlook.com> 2025-10-17 18:14:49 +08:00			`num_tokens_across_dp: Optional[torch.Tensor] = None,`
			`aclgraph_runtime_mode: CUDAGraphMode = CUDAGraphMode.NONE,`
[v0.11.0-dev][CI] Fix ngram lacking of input arg `dummy_compute_logits` error (#4648) ### What this PR does / why we need it? Fix ngram lacking of input arg `dummy_compute_logits` error ### How was this patch tested? CI passed with existing test. --------- Signed-off-by: MengqingCao <cmq0113@163.com> 2025-12-03 09:22:07 +08:00			`batch_descriptor=None,`
			`dummy_compute_logits=lambda hidden_states: None):`
[Refactor] Refactor Spec Decode (#2668) ### What this PR does / why we need it? Refactor spec decode ### Does this PR introduce _any_ user-facing change? N/A ### How was this patch tested? CI passed with new added/existing test. - vLLM version: v0.10.1.1 - vLLM main: https://github.com/vllm-project/vllm/commit/6997a25ac65ed6cc3c2be6d09ca45f633a345f63 --------- Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com> Signed-off-by: Icey <1790571317@qq.com> Co-authored-by: wangxiyuan <wangxiyuan1007@gmail.com> 2025-09-04 11:34:47 +08:00			`"""Called by dummy_run in modle_runner"""`
			`raise NotImplementedError`

			`def generate_token_ids(self,`
			`valid_sampled_token_ids: list[list[int]],`
			`sampling_metadata: SamplingMetadata = None,`
			`scheduler_output: SchedulerOutput = None,`
			`spec_decode_metadata: SpecDecodeMetadata = None,`
			`positions: torch.Tensor = None,`
			`num_scheduled_tokens: int = 0,`
			`hidden_states: torch.Tensor = None,`
			`attn_metadata=None,`
			`aux_hidden_states: torch.Tensor = None):`
			`"""Called by execute_model in model_runner"""`
			`raise NotImplementedError`