xc-llm-ascend

Files

LHXuuu 45a573cff1 [Quantization][Feature] Support compressed tensors moe w4a8 dynamic weight (#5889 )

### What this PR does / why we need it?

While using the LLM Compressor quantization tool from the VLLM community
to generate quantized weights, the VLLM Ascend engine needs to be
adapted to support the compressed tensors quantization format.

1. Support Moe model W4A8 dynamic weight.

- vLLM version: v0.13.0
- vLLM main:
bde38c11df

---------

Signed-off-by: LHXuuu <scut_xlh@163.com>
Signed-off-by: menogrey <1299267905@qq.com>
Co-authored-by: menogrey <1299267905@qq.com>

2026-02-02 16:39:32 +08:00

w4a8_dynamic_moe.py

[Quantization][Feature] Support compressed tensors moe w4a8 dynamic weight (#5889 )

2026-02-02 16:39:32 +08:00

w8a8_int8_dynamic_moe.py

[CI] Fix lint CI (#5880 )

2026-01-14 11:23:38 +08:00

w8a8_int8_dynamic.py

[Lint]Style: Convert example to ruff format (#5863 )

2026-01-13 20:46:50 +08:00

w8a8_int8.py

[Lint]Style: Convert example to ruff format (#5863 )

2026-01-13 20:46:50 +08:00