xc-llm-ascend

Files

Shaoxu Cheng 39e77fb9e4 [Feat.]: support 310p w8a8 (#6454 )

### What this PR does / why we need it?
Introduced 310P W8A8 Quantization Support: New modules and methods have
been added to enable W8A8 static quantization specifically for the
Ascend 310P platform.
Platform-Specific Quantization Configuration Loading: The system now
dynamically loads the appropriate quantization configurations
(AscendCompressedTensorsConfig, AscendModelSlimConfig) based on whether
the current hardware is an Ascend 310P device.
Implemented AscendW8A8LinearMethod310P: A dedicated linear quantization
method for 310P is provided, handling the specifics of weight and
activation quantization, including input parameter broadcasting and
weight data manipulation.
Extended AscendModelSlimConfig for 310P: A specialized configuration
class for 310P integrates the new W8A8 linear method for both standard
linear layers and vocabulary parallel embeddings, ensuring proper
quantization application.

- vLLM version: v0.14.1
- vLLM main:
dc917cceb8

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: Shaoxu Cheng <2906339855@qq.com>

2026-02-03 14:13:06 +08:00

__init__.py

[Refactor] Quantization Module Refactor (#5738 )

2026-01-23 14:13:47 +08:00

base.py

[Refactor] Quantization Module Refactor (#5738 )

2026-01-23 14:13:47 +08:00

registry.py

[Refactor] Quantization Module Refactor (#5738 )

2026-01-23 14:13:47 +08:00

w4a4_flatquant.py

[Refactor] Quantization Module Refactor (#5738 )