9.4 KiB
Qwen3-8B Quantization Comparison Summary
F16 Baseline Reference
| Metric | Value |
|---|---|
| F16 Perplexity | 10.1114 |
| File Size | 15.26 GiB (16.00 BPW) |
All precision loss percentages below are calculated relative to this F16 baseline.
Q4_K_HIFI (INT8 Residuals + Per-Block Scale)
Pros:
- 🎯 Uses intelligent outlier preservation with INT8 residuals on critical tensors
- 🔬 10 tensors use Q6_K_HIFI_RES8 format for maximum precision on sensitive weights
- ✅ Best perplexity at 10.4133 (beats Q4_K_M's 10.4239)
- ✅ Best imatrix perplexity at 10.2225 (beats Q4_K_M's 10.2355)
- 📊 +3.0% PPL vs F16 without imatrix, +1.1% with imatrix ✅
Cons:
- 💾 Larger file size at 4.93 GiB (+5.3% vs Q4_K_M)
- 🐢 Slower than Q4_K_M (117.53 vs 118.70 TPS, 1.0% slower)
- 🐢 Slower than Q4_K_S (117.53 vs 123.04 TPS, 4.5% slower)
Best for: Maximum quality at Q4-level efficiency. Ideal when perplexity matters most.
Performance Comparison (Q4_K_HIFI vs the others)
Q4_K_M
| Metric | Q4_K_HIFI | Q4_K_M | Difference |
|---|---|---|---|
| Speed (TPS) | 117.53 | 118.70 | -1.17 (1.0% slower) |
| Perplexity | 10.4133 | 10.4239 | -0.01 (0.1% better) ✅ |
| PPL vs F16 | +3.0% | +3.1% | 0.1% less precision loss |
| File Size | 4.93 GiB | 4.68 GiB | +0.25 GiB (5.3% larger) |
| Bits Per Weight | 5.17 | 4.90 | +0.27 (5.5% more) |
Pros:
- ⚖️ Traditional "balanced" approach between speed and quality
- 📚 Well-documented, standard quantization method
- 💾 Smaller file size than Q4_K_HIFI
- ⚡ Slightly faster than Q4_K_HIFI (1.0%)
- 📊 +3.1% PPL vs F16 — similar precision loss to Q4_K_HIFI
Cons:
- ❌ Worse perplexity than Q4_K_HIFI (10.4239 vs 10.4133)
Best for: When speed and size are prioritized over marginal quality gains. Summary: Q4_K_HIFI delivers better quality (0.1% lower perplexity) at the cost of 5.3% larger file size and 1.0% slower inference.
Q4_K_S
| Metric | Q4_K_HIFI | Q4_K_S | Difference |
|---|---|---|---|
| Speed (TPS) | 117.53 | 123.04 | -5.51 (4.5% slower) |
| Perplexity | 10.4133 | 10.6837 | -0.27 (2.5% better) ✅ |
| PPL vs F16 | +3.0% | +5.7% | 2.7% less precision loss |
| File Size | 4.93 GiB | 4.47 GiB | +0.46 GiB (10.3% larger) |
| Bits Per Weight | 5.17 | 4.68 | +0.49 (10.5% more) |
Pros:
- ⚡ Fastest inference at 123.04 TPS (4.5% faster than Q4_K_HIFI)
- 💾 Smallest file size at 4.47 GiB
- ✅ Best choice when speed and storage are critical
Cons:
- ❌ Significantly worse quality with perplexity of 10.68 (2.5% higher than Q4_K_HIFI)
- 📊 +5.7% PPL vs F16 — highest precision loss of Q4_K variants
- ⚠️ Uses minimal q6_K enhancement — impacts quality on sensitive weights
Best for: Extreme resource constraints, quick prototyping, or bulk processing where quality is less important. Summary: Q4_K_HIFI trades a 4.5% speed reduction and 10.3% larger file size for a 2.5% improvement in quality over Q4_K_S.
Recommendation Matrix
| Priority | Recommended Model | Rationale |
|---|---|---|
| Quality First | Q4_K_HIFI | Best perplexity (10.4133) without imatrix, (10.2225) with imatrix |
| Speed First | Q4_K_S | 4.5% faster than Q4_K_HIFI, acceptable if quality degradation is tolerable |
| Best Balance | Q4_K_HIFI | Best quality with acceptable speed/size overhead |
| Smallest Size | Q4_K_S | 10.3% smaller than Q4_K_HIFI, 4.5% smaller than Q4_K_M |
Key Insight
At 8B scale, Q4_K_HIFI outperforms Q4_K_M with Q6_K_HIFI_RES8 enhancement:
- 0.1% lower perplexity than Q4_K_M (10.4133 vs 10.4239)
- 5.3% larger file size (4.93 GiB vs 4.68 GiB)
- 1.0% slower than Q4_K_M (117.53 vs 118.70 TPS)
- All variants lose only 3.0-5.7% precision vs F16 (10.11 baseline) — excellent retention!
The Q6_K_HIFI_RES8 format (upgraded from Q5_K_HIFI_RES8) provides the precision needed at 8B scale. The INT8 residuals with 6-bit base quantization capture weight distributions that Q4_K_M's standard q6_K enhancement misses.
💡 Scale Effect: At 8B scale, models benefit from higher-precision HIFI enhancement. The Q6_K base (vs Q5_K) in the HIFI format provides the extra headroom needed for accurate outlier representation.
Precision Loss Summary (vs F16 Baseline: PPL 10.1114)
| Model | PPL (no imatrix) | vs F16 | PPL (imatrix) | vs F16 |
|---|---|---|---|---|
| Q4_K_HIFI | 10.4133 | +3.0% | 10.2225 | +1.1% ✅ |
| Q4_K_M | 10.4239 | +3.1% | 10.2355 | +1.2% |
| Q4_K_S | 10.6837 | +5.7% | 10.2943 | +1.8% |
Key Observations:
- Without imatrix: Q4_K variants lose only 3.0-5.7% precision vs F16 — excellent retention at 8B scale
- With imatrix: Precision loss drops to just 1.1-1.8% — nearly F16 quality!
- Q4_K_HIFI (imatrix) achieves the lowest precision loss at only +1.1% vs F16
- The 8B model shows the best precision retention across all model sizes tested
Tensor Distribution
| Model | q4_K | q5_K | q6_K | Q6_K_HIFI_RES8 | f32 | Total |
|---|---|---|---|---|---|---|
| Q4_K_S | 245 | 8 | 1 | 0 | 145 | 399 |
| Q4_K_M | 217 | 0 | 37 | 0 | 145 | 399 |
| Q4_K_HIFI | 213 | 0 | 31 | 10 | 145 | 399 |
Q4_K_HIFI Enhancement: 10 critical tensors use Q6_K_HIFI_RES8 format with INT8 residuals + per-block scale (upgraded from Q5_K_HIFI_RES8 for 8B+ models).
Addendum: Impact of imatrix on Q4_K_M and Q4_K_S
When all models are quantized with an importance matrix (imatrix), quality improves significantly across all variants.
imatrix Perplexity Improvements
| Model | Without imatrix | With imatrix | Improvement | PPL vs F16 (imatrix) |
|---|---|---|---|---|
| Q4_K_HIFI | 10.4133 | 10.2225 | -0.191 (1.8% better) | +1.1% ✅ |
| Q4_K_M | 10.4239 | 10.2355 | -0.188 (1.8% better) | +1.2% |
| Q4_K_S | 10.6837 | 10.2943 | -0.389 (3.6% better) | +1.8% |
Revised Comparison (All with imatrix)
| Model | PPL (imatrix) | vs Q4_K_HIFI | vs F16 | Size |
|---|---|---|---|---|
| Q4_K_HIFI | 10.2225 | baseline | +1.1% ✅ | 4.93 GiB |
| Q4_K_M | 10.2355 | +0.013 (+0.1%) | +1.2% | 4.68 GiB |
| Q4_K_S | 10.2943 | +0.072 (+0.7%) | +1.8% | 4.47 GiB |
Key Findings
Q4_K_HIFI is the best choice with imatrix:
| Comparison | Without imatrix | With imatrix |
|---|---|---|
| Q4_K_HIFI vs Q4_K_M | -0.1% (Q4_K_HIFI better) ✅ | -0.1% (Q4_K_HIFI better) ✅ |
| Q4_K_HIFI vs Q4_K_S | -2.5% (Q4_K_HIFI better) ✅ | -0.7% (Q4_K_HIFI better) ✅ |
Revised Recommendations (When Using imatrix)
| Priority | Without imatrix | With imatrix |
|---|---|---|
| Quality First | Q4_K_HIFI ✅ | Q4_K_HIFI ✅ (best perplexity) |
| Best Balance | Q4_K_HIFI ✅ | Q4_K_HIFI ✅ |
| Size/Speed | Q4_K_S | Q4_K_S |
Conclusion
For both imatrix and non-imatrix quantization:
- Q4_K_HIFI consistently outperforms Q4_K_M in perplexity (10.2225 vs 10.2355 with imatrix, 10.4133 vs 10.4239 without)
- Q4_K_HIFI is 5.3% larger than Q4_K_M (4.93 GiB vs 4.68 GiB)
- Q4_K_HIFI + imatrix offers the best quality with acceptable size/speed tradeoffs
Q4_K_HIFI advantages:
- Best perplexity in all scenarios (with and without imatrix)
- Q6_K_HIFI_RES8 format captures weight outliers that standard quantization misses
- Significantly better quality than Q4_K_S (2.5% without imatrix, 0.7% with imatrix)
Appendix (Test Environment Details)
| Component | Specification |
|---|---|
| OS | Ubuntu 24.04.3 LTS |
| CPU | AMD EPYC 9254 24-Core Processor |
| CPU Cores | 96 cores (2 threads/core) |
| RAM | 1.0Ti |
| GPU | NVIDIA L40S × 2 |
| VRAM | 46068 MiB per GPU |
| CUDA | 12.9 |
| Test Data | wikitext-2-raw, 584 chunks |
| Context | 512 tokens |
| Samples | 100 per speed benchmark |
| imatrix | mixed-imatrix-dataset.txt, 4697 chunks |