xc-llm-ascend

Files

panchao-hub 8069442b41 enable npugraph_ex (#5120 )

### What this PR does / why we need it?
We will expose the enabling switch for npugraph_ex to better facilitate
subsequent optimization.

### Does this PR introduce _any_ user-facing change?
Previously, the enable_npugraph_ex switch would trigger an error; now we
have removed the error reporting mechanism to better facilitate
subsequent optimization efforts.
Basic functionalities are available in CANN and torch_npu for Q3, while
advanced optimizations will depend on the Q4 release.

### How was this patch tested?
llm =LLM(
    model=model,
    enforce_eager=False ,
        additional_config={
        "enable_npugraph_ex":  True
        },
        compilation_config={
            "cudagraph_mode": "FULL_DECODE_ONLY",
            "cudagraph_capture_sizes": [16],
        },
}


- vLLM version: v0.12.0
- vLLM main:
ad32e3e19c

---------

Signed-off-by: p00465316 <panchao13@huawei.com>
Co-authored-by: p00465316 <panchao13@huawei.com>
Co-authored-by: weijinqian0 <1184188277@qq.com>

2025-12-18 09:08:40 +08:00

npugraph_ex_passes

[Feature] Support npuhraph_ex backend (#4700 )

2025-12-10 20:48:05 +08:00

passes

[Fusion] [Graph] Add qknorm rope fusion operator (#4711 )

2025-12-17 08:53:44 +08:00

__init__.py

[Bugfix] add compilation/__init__.py to fix import error (#1152 )

2025-06-10 17:14:25 +08:00

acl_graph.py

[Attention] Temporarily add back pa for small batch sizes. (#4765 )