@@ -1,71 +1,112 @@
|
||||
# Quantization Guide
|
||||
|
||||
Model quantization is a technique that reduces the size and computational requirements of a model by lowering the data precision of the weights and activation values in the model, thereby saving the memory and improving the inference speed.
|
||||
Model quantization is a technique that reduces model size and computational overhead by lowering the numerical precision of weights and activations, thereby saving memory and improving inference speed.
|
||||
|
||||
Since 0.9.0rc2 version, quantization feature is experimentally supported in vLLM Ascend. Users can enable quantization feature by specifying `--quantization ascend`. Currently, only Qwen, DeepSeek series models are well tested. We’ll support more quantization algorithm and models in the future.
|
||||
`vLLM Ascend` supports multiple quantization methods. This guide provides instructions for using different quantization tools and running quantized models on vLLM Ascend.
|
||||
|
||||
## Install modelslim
|
||||
> **Note**
|
||||
>
|
||||
> You can choose to convert the model yourself or use the quantized model we uploaded.
|
||||
> See <https://www.modelscope.cn/models/vllm-ascend/Kimi-K2-Instruct-W8A8>.
|
||||
> Before you quantize a model, ensure sufficient RAM is available.
|
||||
|
||||
To quantize a model, users should install [ModelSlim](https://gitee.com/ascend/msit/blob/master/msmodelslim/README.md) which is the Ascend compression and acceleration tool. It is an affinity-based compression tool designed for acceleration, using compression as its core technology and built upon the Ascend platform.
|
||||
## Quantization Tools
|
||||
|
||||
Install modelslim:
|
||||
vLLM Ascend supports models quantized by two main tools: `ModelSlim` and `LLM-Compressor`.
|
||||
|
||||
### 1. ModelSlim (Recommended)
|
||||
|
||||
[ModelSlim](https://gitcode.com/Ascend/msmodelslim/blob/master/README.md) is an Ascend-friendly compression tool focused on acceleration, using compression techniques, and built for Ascend hardware. It includes a series of inference optimization technologies such as quantization and compression, aiming to accelerate large language dense models, MoE models, multimodal understanding models, multimodal generation models, etc.
|
||||
|
||||
#### Installation
|
||||
|
||||
To use ModelSlim for model quantization, install it from its [Git repository](https://gitcode.com/Ascend/msmodelslim):
|
||||
|
||||
```bash
|
||||
# The branch(br_release_MindStudio_8.1.RC2_TR5_20260624) has been verified
|
||||
git clone -b br_release_MindStudio_8.1.RC2_TR5_20260624 https://gitee.com/ascend/msit
|
||||
# Install 26.0.0 version, this is currently the latest stable branch
|
||||
git clone https://gitcode.com/Ascend/msmodelslim.git -b 26.0.0
|
||||
|
||||
cd msit/msmodelslim
|
||||
cd msmodelslim
|
||||
|
||||
bash install.sh
|
||||
pip install accelerate
|
||||
```
|
||||
|
||||
## Quantize model
|
||||
#### Model Quantization
|
||||
|
||||
:::{note}
|
||||
You can choose to convert the model yourself or use the quantized model we uploaded,
|
||||
see https://www.modelscope.cn/models/vllm-ascend/Kimi-K2-Instruct-W8A8
|
||||
This conversion process will require a larger CPU memory, please ensure that the RAM size is greater than 2TB
|
||||
:::
|
||||
The following example shows how to generate W8A8 quantized weights for the [Qwen3-MoE model](https://gitcode.com/Ascend/msmodelslim/blob/master/example/Qwen3-MOE/README.md).
|
||||
|
||||
### Adapts and change
|
||||
1. Ascend does not support the `flash_attn` library. To run the model, you need to follow the [guide](https://gitee.com/ascend/msit/blob/master/msmodelslim/example/DeepSeek/README.md#deepseek-v3r1) and comment out certain parts of the code in `modeling_deepseek.py` located in the weights folder.
|
||||
2. The current version of transformers does not support loading weights in FP8 quantization format. you need to follow the [guide](https://gitee.com/ascend/msit/blob/master/msmodelslim/example/DeepSeek/README.md#deepseek-v3r1) and delete the quantization related fields from `config.json` in the weights folder
|
||||
|
||||
### Generate the w8a8 weights
|
||||
**Quantization Script:**
|
||||
|
||||
```bash
|
||||
cd example/DeepSeek
|
||||
cd example/Qwen3-MOE
|
||||
|
||||
# Support multi-card quantization
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3,4,5,6,7
|
||||
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:False
|
||||
export MODEL_PATH="/root/.cache/Kimi-K2-Instruct"
|
||||
export SAVE_PATH="/root/.cache/Kimi-K2-Instruct-W8A8"
|
||||
|
||||
python3 quant_deepseek_w8a8.py --model_path $MODEL_PATH --save_path $SAVE_PATH --batch_size 4
|
||||
# Set model and save paths
|
||||
export MODEL_PATH="/path/to/your/model"
|
||||
export SAVE_PATH="/path/to/your/quantized_model"
|
||||
|
||||
# Run quantization script
|
||||
python3 quant_qwen_moe_w8a8.py --model_path $MODEL_PATH \
|
||||
--save_path $SAVE_PATH \
|
||||
--anti_dataset ../common/qwen3-moe_anti_prompt_50.json \
|
||||
--calib_dataset ../common/qwen3-moe_calib_prompt_50.json \
|
||||
--trust_remote_code True
|
||||
```
|
||||
|
||||
Here is the full converted model files except safetensors:
|
||||
After quantization completes, the output directory will contain the quantized model files.
|
||||
|
||||
For more examples, refer to the [official examples](https://gitcode.com/Ascend/msmodelslim/tree/master/example).
|
||||
|
||||
### 2. LLM-Compressor
|
||||
|
||||
[LLM-Compressor](https://github.com/vllm-project/llm-compressor) is a unified compressed model library for faster vLLM inference.
|
||||
|
||||
#### Installation
|
||||
|
||||
```bash
|
||||
.
|
||||
|-- config.json
|
||||
|-- configuration.json
|
||||
|-- configuration_deepseek.py
|
||||
|-- generation_config.json
|
||||
|-- modeling_deepseek.py
|
||||
|-- quant_model_description.json
|
||||
|-- quant_model_weight_w8a8_dynamic.safetensors.index.json
|
||||
|-- tiktoken.model
|
||||
|-- tokenization_kimi.py
|
||||
`-- tokenizer_config.json
|
||||
pip install llmcompressor
|
||||
```
|
||||
|
||||
## Run the model
|
||||
#### Model Quantization
|
||||
|
||||
Now, you can run the quantized models with vLLM Ascend. Here is the example for online and offline inference.
|
||||
`LLM-Compressor` provides various quantization scheme examples.
|
||||
|
||||
### Offline inference
|
||||
##### Dense Quantization
|
||||
|
||||
An example to generate W8A8 dynamic quantized weights for dense model:
|
||||
|
||||
```bash
|
||||
# Navigate to LLM-Compressor examples directory
|
||||
cd examples/quantization/llm-compressor
|
||||
|
||||
# Run quantization script
|
||||
python3 w8a8_int8_dynamic.py
|
||||
```
|
||||
|
||||
##### MoE Quantization
|
||||
|
||||
An example to generate W8A8 dynamic quantized weights for MoE model:
|
||||
|
||||
```bash
|
||||
# Navigate to LLM-Compressor examples directory
|
||||
cd examples/quantization/llm-compressor
|
||||
|
||||
# Run quantization script
|
||||
python3 w8a8_int8_dynamic_moe.py
|
||||
```
|
||||
|
||||
For more content, refer to the [official examples](https://github.com/vllm-project/llm-compressor/tree/main/examples).
|
||||
|
||||
The quantization types currently supported by LLM-Compressor can be viewed in the `vllm_ascend/quantization/compressed_tensors_config.py` file.
|
||||
|
||||
## Running Quantized Models
|
||||
|
||||
Once you have a quantized model which is generated by **ModelSlim**, you can use vLLM Ascend for inference by specifying the `--quantization ascend` parameter to enable quantization features, while for models quantized by **LLM-Compressor**, it is not necessary to add this parameter.
|
||||
|
||||
### Offline Inference
|
||||
|
||||
```python
|
||||
import torch
|
||||
@@ -76,12 +117,20 @@ prompts = [
|
||||
"Hello, my name is",
|
||||
"The future of AI is",
|
||||
]
|
||||
# Set sampling parameters
|
||||
sampling_params = SamplingParams(temperature=0.6, top_p=0.95, top_k=40)
|
||||
|
||||
llm = LLM(model="{quantized_model_save_path}",
|
||||
max_model_len=2048,
|
||||
llm = LLM(model="/path/to/your/quantized_model",
|
||||
max_model_len=4096,
|
||||
trust_remote_code=True,
|
||||
# Enable quantization by specifying `quantization="ascend"`
|
||||
# Set appropriate TP and DP values
|
||||
tensor_parallel_size=2,
|
||||
data_parallel_size=1,
|
||||
# Set an unused port
|
||||
port=8000,
|
||||
# Set serving model name
|
||||
served_model_name="quantized_model",
|
||||
# Specify `quantization="ascend"` to enable quantization for models quantized by ModelSlim
|
||||
quantization="ascend")
|
||||
|
||||
outputs = llm.generate(prompts, sampling_params)
|
||||
@@ -91,36 +140,22 @@ for output in outputs:
|
||||
print(f"Prompt: {prompt!r}, Generated text: {generated_text!r}")
|
||||
```
|
||||
|
||||
### Online inference
|
||||
|
||||
Enable quantization by specifying `--quantization ascend`, for more details, see DeepSeek-V3-W8A8 [tutorial](https://vllm-ascend.readthedocs.io/en/latest/tutorials/multi_node.html)
|
||||
|
||||
## FAQs
|
||||
|
||||
### 1. How to solve the KeyError: 'xxx.layers.0.self_attn.q_proj.weight' problem?
|
||||
|
||||
First, make sure you specify `ascend` quantization method. Second, check if your model is converted by this `br_release_MindStudio_8.1.RC2_TR5_20260624` modelslim version. Finally, if it still doesn't work, please
|
||||
submit a issue, maybe some new models need to be adapted.
|
||||
|
||||
### 2. How to solve the error "Could not locate the configuration_deepseek.py"?
|
||||
|
||||
Please convert DeepSeek series models using `br_release_MindStudio_8.1.RC2_TR5_20260624` modelslim, this version has fixed the missing configuration_deepseek.py error.
|
||||
|
||||
### 3. When converting deepseek series models with modelslim, what should you pay attention?
|
||||
|
||||
When the mla portion of the weights used `W8A8_DYNAMIC` quantization, if torchair graph mode is enabled, please modify the configuration file in the CANN package to prevent incorrect inference results.
|
||||
|
||||
The operation steps are as follows:
|
||||
|
||||
1. Search in the CANN package directory used, for example:
|
||||
find /usr/local/Ascend/ -name fusion_config.json
|
||||
|
||||
2. Add `"AddRmsNormDynamicQuantFusionPass":"off",` and `"MultiAddRmsNormDynamicQuantFusionPass":"off",` to the fusion_config.json you find, the location is as follows:
|
||||
### Online Inference
|
||||
|
||||
```bash
|
||||
{
|
||||
"Switch":{
|
||||
"GraphFusion":{
|
||||
"AddRmsNormDynamicQuantFusionPass":"off",
|
||||
"MultiAddRmsNormDynamicQuantFusionPass":"off",
|
||||
# Corresponding to offline inference
|
||||
python -m vllm.entrypoints.api_server \
|
||||
--model /path/to/your/quantized_model \
|
||||
--max-model-len 4096 \
|
||||
--port 8000 \
|
||||
--tensor-parallel-size 2 \
|
||||
--data-parallel-size 1 \
|
||||
--served-model-name quantized_model \
|
||||
--trust-remote-code
|
||||
```
|
||||
|
||||
## References
|
||||
|
||||
- [ModelSlim GitCode](https://gitcode.com/Ascend/msmodelslim)
|
||||
- [LLM-Compressor GitHub](https://github.com/vllm-project/llm-compressor)
|
||||
- [vLLM Quantization Guide](https://docs.vllm.ai/en/latest/features/quantization/)
|
||||
|
||||
Reference in New Issue
Block a user