xc-llm-ascend

Files

Nicholas Tao 7bec1a9b9c qwen3_moe/qwen25 support torchair graph (#2403 )

### What this PR does / why we need it?
Added support for the TorchAir graph mode in qwen3_moe and qwen2.5
### Does this PR introduce _any_ user-facing change?
No
### How was this patch tested?
```bash
llm = LLM(
    model=model,
    tensor_parallel_size=GPUs_per_dp_rank,
    enforce_eager=False,
    enable_expert_parallel=True,
    max_model_len=4096,
    max_num_seqs=16,
    trust_remote_code=trust_remote_code,
    gpu_memory_utilization=0.4,
    additional_config={
             "torchair_graph_config": {
                 "enabled": True,
                 "use_cached_graph": False,
                 "graph_batch_sizes_init": False,
                 "graph_batch_sizes": [16]
             },
             "ascend_scheduler_config": {
                 "enabled": True,
                 "chunked_prefill_enabled":True,
             },
             "refresh": True,
    },
)
```

- vLLM version: v0.10.0
- vLLM main:
b87cb97a53

Signed-off-by: taoyuxiang <oui.nicholas.tao@gmail.com>

2025-08-20 11:23:50 +08:00

__init__.py

ut:add ut for qwen2_5_vl_without_padding.py (#1988 )

2025-07-25 14:12:44 +08:00

test_deepseek_mtp.py

[V1] MTP supports torchair (#2145 )

2025-08-06 19:37:43 +08:00

test_deepseek_v2.py

[main][refactor] Refactoring forward_context and model_runner_v1 (#1979 )

2025-07-28 14:06:20 +08:00

test_qwen2_5_vl_without_padding.py

[Bugfix] fix tensor not same device in qwen2_5_vl_without_padding (#2051 )