Files
Qwen3-32B-AWQ/README.md
ModelHub XC 64835c6916 初始化项目,由ModelHub XC社区提供模型
Model: swift/Qwen3-32B-AWQ
Source: Original Platform
2026-06-14 15:10:13 +08:00

5.0 KiB

license, tasks, base_model
license tasks base_model
Apache License 2.0
text-generation
Qwen/Qwen3-32B

Inference

import torch
from modelscope import AutoModelForCausalLM, AutoTokenizer

model_name = "swift/Qwen3-32B-AWQ"

# load the tokenizer and the model
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype=torch.float16,
    device_map="auto"
)

# prepare the model input
prompt = "Give me a short introduction to large language model."
messages = [
    {"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
    enable_thinking=True # Switches between thinking and non-thinking modes. Default is True.
)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)

# conduct text completion
generated_ids = model.generate(
    **model_inputs,
    max_new_tokens=32768
)
output_ids = generated_ids[0][len(model_inputs.input_ids[0]):].tolist() 

# parsing thinking content
try:
    # rindex finding 151668 (</think>)
    index = len(output_ids) - output_ids[::-1].index(151668)
except ValueError:
    index = 0

thinking_content = tokenizer.decode(output_ids[:index], skip_special_tokens=True).strip("\n")
content = tokenizer.decode(output_ids[index:], skip_special_tokens=True).strip("\n")

print("thinking content:", thinking_content)
print("content:", content)

Quantization

The model has undergone AWQ int4 quantization using the ms-swift framework.

If you have fine-tuned the model and wish to quantize the fine-tuned version, you can refer to the following quantization scripts:

With these scripts, you can easily complete the quantization process for the model.

Evaluation

We evaluate the quality of this AWQ quantization with EvalScope. For the best practice for evaluating Qwen3 models, one may refer to the following:

Performance of Qwen3-32B-AWQ is evaluated on our mixed-benchmark of Qwen3 Evaluation Collection, with the results listed below:

The performance comparison of Qwen3-32B-AWQ and Qwen3-32B

task_type dataset_name metric average_score(AWQ) average_score(without AWQ) count
exam MMLU-Pro AverageAccuracy 0.7906 0.8018 12032
exam MMLU-Redux AverageAccuracy 0.8918 0.893 5700
exam C-Eval AverageAccuracy 0.8834 0.8915 1346
instruction IFEval inst_level_strict_acc 0.8842 0.8802 541
instruction IFEval inst_level_loose_acc 0.915 0.9125 541
instruction IFEval prompt_level_loose_acc 0.8725 0.8595 541
instruction IFEval prompt_level_strict_acc 0.8336 0.8189 541
math MATH-500 AveragePass@1 0.932 0.942 500
knowledge GPQA AveragePass@1 0.6566 0.6465 198
code LiveCodeBench Pass@1 0.5 0.544 182
exam iQuiz AverageAccuracy 0.8 0.775 120
math AIME 2024 AveragePass@1 0.7333 0.7667 30
math AIME 2025 AveragePass@1 0.5667 0.6667 30

Conclusion

As we can see from the comparison above, evaluatoin results across different tasks and datasets suggest that our quantized-version with AWQ exihibit minimum fluctuation on model performance. In fact, for most benchmarks, AWQ version performs mostly on-par with the original version, except for math-related benchmarks (such as AIME2024/AIME2025) where performance degration is a bit more noticable.

Please also note that the result for the Qwen3-32B model is not Qwen offical. It is done in the same setup that we used to evaluate Qwen3-32B-AWQ. You may reproduce this evaluation result by following our best-practice guides above.