Model: swift/Qwen3-32B-AWQ Source: Original Platform
license, tasks, base_model
| license | tasks | base_model | ||
|---|---|---|---|---|
| Apache License 2.0 |
|
|
Inference
import torch
from modelscope import AutoModelForCausalLM, AutoTokenizer
model_name = "swift/Qwen3-32B-AWQ"
# load the tokenizer and the model
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype=torch.float16,
device_map="auto"
)
# prepare the model input
prompt = "Give me a short introduction to large language model."
messages = [
{"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
enable_thinking=True # Switches between thinking and non-thinking modes. Default is True.
)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)
# conduct text completion
generated_ids = model.generate(
**model_inputs,
max_new_tokens=32768
)
output_ids = generated_ids[0][len(model_inputs.input_ids[0]):].tolist()
# parsing thinking content
try:
# rindex finding 151668 (</think>)
index = len(output_ids) - output_ids[::-1].index(151668)
except ValueError:
index = 0
thinking_content = tokenizer.decode(output_ids[:index], skip_special_tokens=True).strip("\n")
content = tokenizer.decode(output_ids[index:], skip_special_tokens=True).strip("\n")
print("thinking content:", thinking_content)
print("content:", content)
Quantization
The model has undergone AWQ int4 quantization using the ms-swift framework.
If you have fine-tuned the model and wish to quantize the fine-tuned version, you can refer to the following quantization scripts:
With these scripts, you can easily complete the quantization process for the model.
Evaluation
We evaluate the quality of this AWQ quantization with EvalScope. For the best practice for evaluating Qwen3 models, one may refer to the following:
Performance of Qwen3-32B-AWQ is evaluated on our mixed-benchmark of Qwen3 Evaluation Collection, with the results listed below:
The performance comparison of Qwen3-32B-AWQ and Qwen3-32B
| task_type | dataset_name | metric | average_score(AWQ) | average_score(without AWQ) | count |
|---|---|---|---|---|---|
| exam | MMLU-Pro | AverageAccuracy | 0.7906 | 0.8018 | 12032 |
| exam | MMLU-Redux | AverageAccuracy | 0.8918 | 0.893 | 5700 |
| exam | C-Eval | AverageAccuracy | 0.8834 | 0.8915 | 1346 |
| instruction | IFEval | inst_level_strict_acc | 0.8842 | 0.8802 | 541 |
| instruction | IFEval | inst_level_loose_acc | 0.915 | 0.9125 | 541 |
| instruction | IFEval | prompt_level_loose_acc | 0.8725 | 0.8595 | 541 |
| instruction | IFEval | prompt_level_strict_acc | 0.8336 | 0.8189 | 541 |
| math | MATH-500 | AveragePass@1 | 0.932 | 0.942 | 500 |
| knowledge | GPQA | AveragePass@1 | 0.6566 | 0.6465 | 198 |
| code | LiveCodeBench | Pass@1 | 0.5 | 0.544 | 182 |
| exam | iQuiz | AverageAccuracy | 0.8 | 0.775 | 120 |
| math | AIME 2024 | AveragePass@1 | 0.7333 | 0.7667 | 30 |
| math | AIME 2025 | AveragePass@1 | 0.5667 | 0.6667 | 30 |
Conclusion
As we can see from the comparison above, evaluatoin results across different tasks and datasets suggest that our quantized-version with AWQ exihibit minimum fluctuation on model performance. In fact, for most benchmarks, AWQ version performs mostly on-par with the original version, except for math-related benchmarks (such as AIME2024/AIME2025) where performance degration is a bit more noticable.
Please also note that the result for the Qwen3-32B model is not Qwen offical. It is done in the same setup that we used to evaluate Qwen3-32B-AWQ. You may reproduce this evaluation result by following our best-practice guides above.