xc-llm-ascend

Files

Mengqing Cao 6c65dd891f [ModelRunner][Qwen3-Next] Fix attn_group initialization timing (#3477 )

### What this PR does / why we need it?
Fix attn_group initialization timing so that fix qwen3-next model

### Does this PR introduce _any_ user-facing change?

### How was this patch tested?

- vLLM version: v0.11.0rc3
- vLLM main: https://github.com/vllm-project/vllm/commit/v0.11.0

---------

Signed-off-by: MengqingCao <cmq0113@163.com>

2025-10-20 09:39:40 +08:00

__init__.py

[Misc][V0 Deprecation] Remove Cache Engine Used for V0 Worker (#1878 )

2025-07-19 09:42:32 +08:00

block_table.py

[HybridKV] Fix prefill disaggregation kvcache addr alignment & use hybrid kv cache only when running qwen3_next (#3007 )

2025-09-18 21:43:22 +08:00

model_runner_v1.py

[ModelRunner][Qwen3-Next] Fix attn_group initialization timing (#3477 )

2025-10-20 09:39:40 +08:00

npu_input_batch.py

Drop 0.10.2 (#3284 )

2025-10-09 10:28:38 +08:00

worker_v1.py

[Feat]Make full graph mode compalible with MTP (#3276 )

2025-10-17 20:19:56 +08:00