初始化项目,由ModelHub XC社区提供模型

Model: geoffmunn/Qwen3-8B-f16
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-09-19 23:46:22 +08:00
commit b85152c86b
44 changed files with 44634 additions and 0 deletions

71
.gitattributes vendored Normal file
View File

@@ -0,0 +1,71 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16:Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16:Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16:Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16:Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16:Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16:Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16:Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16:Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16:Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16-imatrix-5000.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16-Q3_HIFI.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16-imatrix-4697.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16-imatrix-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16-imatrix-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16-Q4_HIFI.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16:Q3_HIFI.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16:Q4_HIFI.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16-imatrix:Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16-imatrix:Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16:Q4_K_HIFI.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16-imatrix:Q4_K_HIFI.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16:Q3_K_HIFI.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16-imatrix:Q3_K_HIFI.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16-imatrix:Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16-imatrix:Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16-imatrix-4697-coder.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16-imatrix-4697-generic.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16-imatrix-8843-coder.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16-imatrix-9343-generic.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16-imatrix:Q5_K_HIFI.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16-imatrix:Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16-imatrix:Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16:Q5_K_HIFI.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16-imatrix:Q2_K_HIFI.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16-imatrix:Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-8B-f16:Q2_K_HIFI.gguf filter=lfs diff=lfs merge=lfs -text

25
MODELFILE Normal file
View File

@@ -0,0 +1,25 @@
# MODELFILE for Qwen3-8B-GGUF
# Used by LM Studio, OpenWebUI, GPT4All, etc.
context_length: 32768
embedding: false
f16: cpu
# Chat template using ChatML (used by Qwen)
prompt_template: >-
<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
# Stop sequences help end generation cleanly
stop: "<|im_end|>"
stop: "<|im_start|>"
# Default sampling (optimized for thinking mode)
temperature: 0.6
top_p: 0.95
top_k: 20
min_p: 0.0
repeat_penalty: 1.1

View File

@@ -0,0 +1,92 @@
# Qwen3-8B Quantization Comparison Summary
## Q3_HIFI (Adaptive/Custom)
**Pros:**
- 🏆 **Best quality** with lowest perplexity of 10.56 (4.4% better than Q3_K_M, 7.2% better than Q3_K_S)
- 📦 **Smaller than Q3_K_M** (3.72 vs 3.84 GiB) while being significantly better quality
- 🎯 Uses intelligent layer-sensitive quantization (Q3_HIFI on sensitive layers, mixed q3_K/q4_K elsewhere)
- 📊 Most consistent results (lowest relative standard deviation in perplexity)
**Cons:**
- 🐢 **Slowest inference** at 143.98 TPS (6.3% slower than Q3_K_S)
- 🔧 Custom quantization may have less community support
**Best for:** Production deployments where output quality matters, tasks requiring accuracy (reasoning, coding, complex instructions), or when you want the best quality-to-size ratio.
## Performance Comparison (Q3_HIFI vs the others)
### Q3_K_M
| Metric | Q3_HIFI | Q3_K_M | Difference |
|---------------------|----------|----------|------------------------------|
| **Speed (TPS)** | 143.98 | 144.72 | -0.74 (0.5% slower) |
| **Perplexity** | 10.56 | 11.05 | **-0.49 (4.4% better)** |
| **File Size** | 3.72 GiB | 3.84 GiB | **-0.12 GiB (3.1% smaller)** |
| **Bits Per Weight** | 3.90 | 4.02 | -0.12 (3.0% less) |
**Pros:**
- ⚖️ Traditional "balanced" approach between speed and quality
- 📚 Well-documented, standard quantization method
**Cons:**
- 💾 **Largest file size** at 3.84 GiB despite not being the best quality
- 🐌 Middle-of-the-road speed (144.7 TPS)
-**Outclassed by Q3_HIFI** which is smaller AND better quality
**Best for:** Legacy compatibility or when you need a proven, standard quantization approach.
**Summary:** Q3_HIFI delivers significantly better quality (4.4% lower perplexity) in a smaller package (3.1% less storage) with virtually no speed penalty (0.5% slower).
### Q3_K_S
| Metric | Q3_HIFI | Q3_K_S | Difference |
|---------------------|----------|----------|-------------------------|
| **Speed (TPS)** | 143.98 | 153.74 | -9.76 (6.3% slower) |
| **Perplexity** | 10.56 | 11.38 | **-0.82 (7.2% better)** |
| **File Size** | 3.72 GiB | 3.51 GiB | +0.21 GiB (6.0% larger) |
| **Bits Per Weight** | 3.90 | 3.68 | +0.22 (6.0% more) |
**Pros:**
-**Fastest inference** at 153.74 TPS (~6% faster than Q3_K_M, ~7% faster than Q3_HIFI)
- 💾 **Smallest file size** at 3.51 GiB
- ✅ Best choice when speed and storage are critical
**Cons:**
-**Worst quality** with perplexity of 11.38 (7.2% higher than Q3_HIFI)
- Uses only q3_K quantization throughout (no mixed precision)
**Best for:** Resource-constrained environments, real-time applications where latency matters more than accuracy, or initial prototyping.
**Summary:** Q3_HIFI trades a 6.3% speed reduction and 6.0% larger file size for a substantial 7.2% improvement in quality (lower perplexity).
---
## Recommendation Matrix
| Priority | Recommended Model | Rationale |
|-------------------|-------------------|-----------------------------------------------------------------------------|
| **Quality First** | Q3_HIFI | 7.2% better perplexity than Q3_K_S with acceptable speed loss |
| **Speed First** | Q3_K_S | 6.3% faster inference, acceptable quality tradeoff for latency-sensitive apps |
| **Best Balance** | Q3_HIFI | Better quality AND smaller size than Q3_K_M, only 0.5% slower |
| **Smallest Size** | Q3_K_S | 6% smaller than alternatives |
---
## Key Insight
**Q3_HIFI represents a clear advancement** over the traditional Q3_K_M approach. It achieves:
- **4.4% lower perplexity** (better accuracy)
- **3.1% smaller file size** (3.72 vs 3.84 GiB)
- Only **0.5% slower** inference (144.0 vs 144.7 TPS)
The Q3_K_M quantization is essentially obsoleted by Q3_HIFI for most use cases. The only remaining choice is between **Q3_K_S** (maximum speed, acceptable quality) and **Q3_HIFI** (maximum quality, acceptable speed).
## Appendix (Test Environment Details)
| Component | Specification |
|---------------|---------------------------------|
| **OS** | Ubuntu 24.04.3 LTS |
| **CPU** | AMD EPYC 9254 24-Core Processor |
| **CPU Cores** | 96 cores (2 threads/core) |
| **RAM** | 1.0Ti |
| **GPU** | NVIDIA L40S × 2 |
| **VRAM** | 46068 MiB per GPU |
| **CUDA** | 12.9 |

View File

@@ -0,0 +1,196 @@
# Qwen3-8B Quantization Comparison Summary
## F16 Baseline Reference
| Metric | Value |
|--------|-------|
| **F16 Perplexity** | 10.1114 |
| **File Size** | 15.26 GiB (16.00 BPW) |
All precision loss percentages below are calculated relative to this F16 baseline.
---
## Q4_K_HIFI (INT8 Residuals + Per-Block Scale)
**Pros:**
- 🎯 Uses intelligent outlier preservation with INT8 residuals on critical tensors
- 🔬 10 tensors use Q6_K_HIFI_RES8 format for maximum precision on sensitive weights
-**Best perplexity** at 10.4133 (beats Q4_K_M's 10.4239)
-**Best imatrix perplexity** at 10.2225 (beats Q4_K_M's 10.2355)
- 📊 **+3.0% PPL vs F16** without imatrix, **+1.1% with imatrix** ✅
**Cons:**
- 💾 Larger file size at 4.93 GiB (+5.3% vs Q4_K_M)
- 🐢 Slower than Q4_K_M (117.53 vs 118.70 TPS, 1.0% slower)
- 🐢 Slower than Q4_K_S (117.53 vs 123.04 TPS, 4.5% slower)
**Best for:** Maximum quality at Q4-level efficiency. Ideal when perplexity matters most.
## Performance Comparison (Q4_K_HIFI vs the others)
### Q4_K_M
| Metric | Q4_K_HIFI | Q4_K_M | Difference |
|---------------------|------------|------------|-------------------------------|
| **Speed (TPS)** | 117.53 | 118.70 | -1.17 (1.0% slower) |
| **Perplexity** | 10.4133 | 10.4239 | **-0.01 (0.1% better)** ✅ |
| **PPL vs F16** | +3.0% | +3.1% | 0.1% less precision loss |
| **File Size** | 4.93 GiB | 4.68 GiB | +0.25 GiB (5.3% larger) |
| **Bits Per Weight** | 5.17 | 4.90 | +0.27 (5.5% more) |
**Pros:**
- ⚖️ Traditional "balanced" approach between speed and quality
- 📚 Well-documented, standard quantization method
- 💾 Smaller file size than Q4_K_HIFI
- ⚡ Slightly faster than Q4_K_HIFI (1.0%)
- 📊 **+3.1% PPL vs F16** — similar precision loss to Q4_K_HIFI
**Cons:**
-**Worse perplexity** than Q4_K_HIFI (10.4239 vs 10.4133)
**Best for:** When speed and size are prioritized over marginal quality gains.
**Summary:** Q4_K_HIFI delivers better quality (0.1% lower perplexity) at the cost of 5.3% larger file size and 1.0% slower inference.
### Q4_K_S
| Metric | Q4_K_HIFI | Q4_K_S | Difference |
|---------------------|------------|------------|-------------------------------|
| **Speed (TPS)** | 117.53 | 123.04 | -5.51 (4.5% slower) |
| **Perplexity** | 10.4133 | 10.6837 | **-0.27 (2.5% better)** ✅ |
| **PPL vs F16** | +3.0% | +5.7% | 2.7% less precision loss |
| **File Size** | 4.93 GiB | 4.47 GiB | +0.46 GiB (10.3% larger) |
| **Bits Per Weight** | 5.17 | 4.68 | +0.49 (10.5% more) |
**Pros:**
-**Fastest inference** at 123.04 TPS (4.5% faster than Q4_K_HIFI)
- 💾 **Smallest file size** at 4.47 GiB
- ✅ Best choice when speed and storage are critical
**Cons:**
-**Significantly worse quality** with perplexity of 10.68 (2.5% higher than Q4_K_HIFI)
- 📊 **+5.7% PPL vs F16** — highest precision loss of Q4_K variants
- ⚠️ Uses minimal q6_K enhancement — impacts quality on sensitive weights
**Best for:** Extreme resource constraints, quick prototyping, or bulk processing where quality is less important.
**Summary:** Q4_K_HIFI trades a 4.5% speed reduction and 10.3% larger file size for a 2.5% improvement in quality over Q4_K_S.
---
## Recommendation Matrix
| Priority | Recommended Model | Rationale |
|-------------------|-------------------|--------------------------------------------------------------------------------|
| **Quality First** | Q4_K_HIFI | Best perplexity (10.4133) without imatrix, (10.2225) with imatrix |
| **Speed First** | Q4_K_S | 4.5% faster than Q4_K_HIFI, acceptable if quality degradation is tolerable |
| **Best Balance** | Q4_K_HIFI | Best quality with acceptable speed/size overhead |
| **Smallest Size** | Q4_K_S | 10.3% smaller than Q4_K_HIFI, 4.5% smaller than Q4_K_M |
---
## Key Insight
**At 8B scale, Q4_K_HIFI outperforms Q4_K_M with Q6_K_HIFI_RES8 enhancement:**
- **0.1% lower perplexity** than Q4_K_M (10.4133 vs 10.4239)
- **5.3% larger** file size (4.93 GiB vs 4.68 GiB)
- **1.0% slower** than Q4_K_M (117.53 vs 118.70 TPS)
- **All variants lose only 3.0-5.7% precision vs F16** (10.11 baseline) — excellent retention!
The Q6_K_HIFI_RES8 format (upgraded from Q5_K_HIFI_RES8) provides the precision needed at 8B scale. The INT8 residuals with 6-bit base quantization capture weight distributions that Q4_K_M's standard q6_K enhancement misses.
💡 **Scale Effect:** At 8B scale, models benefit from higher-precision HIFI enhancement. The Q6_K base (vs Q5_K) in the HIFI format provides the extra headroom needed for accurate outlier representation.
---
## Precision Loss Summary (vs F16 Baseline: PPL 10.1114)
| Model | PPL (no imatrix) | vs F16 | PPL (imatrix) | vs F16 |
|-----------|------------------|--------|---------------|--------|
| Q4_K_HIFI | 10.4133 | **+3.0%** | 10.2225 | **+1.1%** ✅ |
| Q4_K_M | 10.4239 | **+3.1%** | 10.2355 | **+1.2%** |
| Q4_K_S | 10.6837 | **+5.7%** | 10.2943 | **+1.8%** |
**Key Observations:**
- Without imatrix: Q4_K variants lose only 3.0-5.7% precision vs F16 — excellent retention at 8B scale
- With imatrix: Precision loss drops to just 1.1-1.8% — nearly F16 quality!
- Q4_K_HIFI (imatrix) achieves the **lowest precision loss** at only +1.1% vs F16
- The 8B model shows the best precision retention across all model sizes tested
---
## Tensor Distribution
| Model | q4_K | q5_K | q6_K | Q6_K_HIFI_RES8 | f32 | Total |
|---------|------|------|------|----------------|-----|-------|
| Q4_K_S | 245 | 8 | 1 | 0 | 145 | 399 |
| Q4_K_M | 217 | 0 | 37 | 0 | 145 | 399 |
| Q4_K_HIFI | 213 | 0 | 31 | 10 | 145 | 399 |
**Q4_K_HIFI Enhancement:** 10 critical tensors use Q6_K_HIFI_RES8 format with INT8 residuals + per-block scale (upgraded from Q5_K_HIFI_RES8 for 8B+ models).
---
## Addendum: Impact of imatrix on Q4_K_M and Q4_K_S
When all models are quantized **with an importance matrix (imatrix)**, quality improves significantly across all variants.
### imatrix Perplexity Improvements
| Model | Without imatrix | With imatrix | Improvement | PPL vs F16 (imatrix) |
|---------|-----------------|--------------|-------------|----------------------|
| Q4_K_HIFI | 10.4133 | **10.2225** | **-0.191 (1.8% better)** | **+1.1%** ✅ |
| Q4_K_M | 10.4239 | **10.2355** | **-0.188 (1.8% better)** | **+1.2%** |
| Q4_K_S | 10.6837 | **10.2943** | **-0.389 (3.6% better)** | **+1.8%** |
### Revised Comparison (All with imatrix)
| Model | PPL (imatrix) | vs Q4_K_HIFI | vs F16 | Size |
|---------|---------------|------------|--------|------|
| Q4_K_HIFI | 10.2225 | baseline | **+1.1%** ✅ | 4.93 GiB |
| Q4_K_M | 10.2355 | +0.013 (+0.1%) | +1.2% | 4.68 GiB |
| Q4_K_S | 10.2943 | +0.072 (+0.7%) | +1.8% | 4.47 GiB |
### Key Findings
**Q4_K_HIFI is the best choice with imatrix:**
| Comparison | Without imatrix | With imatrix |
|------------|-----------------|--------------|
| Q4_K_HIFI vs Q4_K_M | **-0.1%** (Q4_K_HIFI better) ✅ | **-0.1%** (Q4_K_HIFI better) ✅ |
| Q4_K_HIFI vs Q4_K_S | **-2.5%** (Q4_K_HIFI better) ✅ | **-0.7%** (Q4_K_HIFI better) ✅ |
### Revised Recommendations (When Using imatrix)
| Priority | Without imatrix | With imatrix |
|-------------------|-----------------|--------------|
| **Quality First** | Q4_K_HIFI ✅ | **Q4_K_HIFI** ✅ (best perplexity) |
| **Best Balance** | Q4_K_HIFI ✅ | **Q4_K_HIFI** ✅ |
| **Size/Speed** | Q4_K_S | Q4_K_S |
### Conclusion
**For both imatrix and non-imatrix quantization:**
- Q4_K_HIFI consistently outperforms Q4_K_M in perplexity (10.2225 vs 10.2355 with imatrix, 10.4133 vs 10.4239 without)
- Q4_K_HIFI is 5.3% larger than Q4_K_M (4.93 GiB vs 4.68 GiB)
- **Q4_K_HIFI + imatrix** offers the best quality with acceptable size/speed tradeoffs
**Q4_K_HIFI advantages:**
- Best perplexity in all scenarios (with and without imatrix)
- Q6_K_HIFI_RES8 format captures weight outliers that standard quantization misses
- Significantly better quality than Q4_K_S (2.5% without imatrix, 0.7% with imatrix)
---
## Appendix (Test Environment Details)
| Component | Specification |
|---------------|----------------------------------------|
| **OS** | Ubuntu 24.04.3 LTS |
| **CPU** | AMD EPYC 9254 24-Core Processor |
| **CPU Cores** | 96 cores (2 threads/core) |
| **RAM** | 1.0Ti |
| **GPU** | NVIDIA L40S × 2 |
| **VRAM** | 46068 MiB per GPU |
| **CUDA** | 12.9 |
| **Test Data** | wikitext-2-raw, 584 chunks |
| **Context** | 512 tokens |
| **Samples** | 100 per speed benchmark |
| **imatrix** | mixed-imatrix-dataset.txt, 4697 chunks |

1989
Qwen3-8B-analysis.md Normal file

File diff suppressed because it is too large Load Diff

206
Qwen3-8B-f16-Q2_K/README.md Normal file
View File

@@ -0,0 +1,206 @@
---
license: apache-2.0
tags:
- gguf
- qwen
- qwen3-8b
- qwen3-8b-q2
- qwen3-8b-q2_k
- qwen3-8b-q2_k-gguf
- llama.cpp
- quantized
- text-generation
- reasoning
- agent
- chat
- multilingual
- imatrix
- q3_hifi
base_model: Qwen/Qwen3-8B
author: geoffmunn
---
# Qwen3-8B-f16:Q2_K
Quantized version of [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B) at **Q2_K** level, derived from **f16** base weights.
## Model Info
- **Format**: GGUF (for llama.cpp and compatible runtimes)
- **Size**: 3.28 GB
- **Precision**: Q2_K
- **Base Model**: [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B)
- **Conversion Tool**: [llama.cpp](https://github.com/ggerganov/llama.cpp)
## Quality & Performance
| Metric | Value |
|--------------------|-------------------------------------------------------------------------------|
| **Speed** | ⚡ Fast |
| **RAM Required** | ~3.0 GB |
| **Recommendation** | Not recommended. Came first in the bat & ball question, no other appearances. |
## Prompt Template (ChatML)
This model uses the **ChatML** format used by Qwen:
```text
<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
Set this in your app (LM Studio, OpenWebUI, etc.) for best results.
## Generation Parameters
### Thinking Mode (Recommended for Logic)
Use when solving math, coding, or logical problems.
| Parameter | Value |
|----------------|-------|
| Temperature | 0.6 |
| Top-P | 0.95 |
| Top-K | 20 |
| Min-P | 0.0 |
| Repeat Penalty | 1.1 |
> ❗ DO NOT use greedy decoding — it causes infinite loops.
Enable via:
- `enable_thinking=True` in tokenizer
- Or add `/think` in user input during conversation
### Non-Thinking Mode (Fast Dialogue)
For casual chat and quick replies.
| Parameter | Value |
|----------------|-------|
| Temperature | 0.7 |
| Top-P | 0.8 |
| Top-K | 20 |
| Min-P | 0.0 |
| Repeat Penalty | 1.1 |
Enable via:
- `enable_thinking=False`
- Or add `/no_think` in prompt
Stop sequences: `<|im_end|>`, `<|im_start|>`
## 💡 Usage Tips
> This model supports two operational modes:
>
> ### 🔍 Thinking Mode (Recommended for Logic)
> Activate with `enable_thinking=True` or append `/think` in prompt.
>
> - Ideal for: math, coding, planning, analysis
> - Use sampling: `temp=0.6`, `top_p=0.95`, `top_k=20`
> - Avoid greedy decoding
>
> ### ⚡ Non-Thinking Mode (Fast Chat)
> Use `enable_thinking=False` or `/no_think`.
>
> - Best for: casual conversation, quick answers
> - Sampling: `temp=0.7`, `top_p=0.8`
>
> ---
>
> 🔄 **Switch Dynamically**
> In multi-turn chats, the last `/think` or `/no_think` directive takes precedence.
>
> 🔁 **Avoid Repetition**
> Set `presence_penalty=1.5` if stuck in loops.
>
> 📏 **Use Full Context**
> Allow up to 32,768 output tokens for complex tasks.
>
> 🧰 **Agent Ready**
> Works with Qwen-Agent, MCP servers, and custom tools.
## Customisation & Troubleshooting
Importing directly into Ollama should work, but you might encounter this error: `Error: invalid character '<' looking for beginning of value`.
In this case try these steps:
1. `wget https://huggingface.co/geoffmunn/Qwen3-8B-f16/resolve/main/Qwen3-8B-f16%3AQ2_K.gguf`
2. `nano Modelfile` and enter these details:
```text
FROM ./Qwen3-8B-f16:Q2_K.gguf
# Chat template using ChatML (used by Qwen)
SYSTEM You are a helpful assistant
TEMPLATE "{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>{{ end }}<|im_start|>user
{{ .Prompt }}<|im_end|>
<|im_start|>assistant
"
PARAMETER stop <|im_start|>
PARAMETER stop <|im_end|>
# Default sampling
PARAMETER temperature 0.6
PARAMETER top_p 0.95
PARAMETER top_k 20
PARAMETER min_p 0.0
PARAMETER repeat_penalty 1.1
PARAMETER num_ctx 4096
```
The `num_ctx` value has been dropped to increase speed significantly.
3. Then run this command: `ollama create Qwen3-8B-f16:Q2_K -f Modelfile`
You will now see "Qwen3-8B-f16:Q2_K" in your Ollama model list.
These import steps are also useful if you want to customise the default parameters or system prompt.
## 🖥️ CLI Example Using Ollama or TGI Server
Heres how you can query this model via API using `curl` and `jq`. Replace the endpoint with your local server (e.g., Ollama, Text Generation Inference).
```bash
curl http://localhost:11434/api/generate -s -N -d '{
"model": "hf.co/geoffmunn/Qwen3-8B-f16:Q2_K",
"prompt": "Repeat the following instruction exactly as given: Summarize what a neural network is in one sentence.",
"temperature": 0.5,
"top_p": 0.95,
"top_k": 20,
"min_p": 0.0,
"repeat_penalty": 1.1,
"stream": false
}' | jq -r '.response'
```
🎯 **Why this works well**:
- The prompt is meaningful and demonstrates either **reasoning**, **creativity**, or **clarity** depending on quant level.
- Temperature is tuned appropriately: lower for factual responses (`0.4`), higher for creative ones (`0.7`).
- Uses `jq` to extract clean output.
> 💬 Tip: For interactive streaming, set `"stream": true` and process line-by-line.
## Verification
Check integrity:
```bash
sha256sum -c ../SHA256SUMS.txt
```
## Usage
Compatible with:
- [LM Studio](https://lmstudio.ai) local AI model runner with GPU acceleration
- [OpenWebUI](https://openwebui.com) self-hosted AI platform with RAG and tools
- [GPT4All](https://gpt4all.io) private, offline AI chatbot
- Directly via `llama.cpp`
Supports dynamic switching between thinking modes via `/think` and `/no_think` in multi-turn conversations.
## License
Apache 2.0 see base model for full terms.

View File

@@ -0,0 +1,204 @@
---
license: apache-2.0
tags:
- gguf
- qwen
- qwen3-8b
- qwen3-8b-q3
- qwen3-8b-q3_k_m
- qwen3-8b-q3_k_m-gguf
- llama.cpp
- quantized
- text-generation
- reasoning
- agent
- chat
- multilingual
base_model: Qwen/Qwen3-8B
author: geoffmunn
---
# Qwen3-8B-f16:Q3_K_M
Quantized version of [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B) at **Q3_K_M** level, derived from **f16** base weights.
## Model Info
- **Format**: GGUF (for llama.cpp and compatible runtimes)
- **Size**: 4.12 GB
- **Precision**: Q3_K_M
- **Base Model**: [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B)
- **Conversion Tool**: [llama.cpp](https://github.com/ggerganov/llama.cpp)
## Quality & Performance
| Metric | Value |
|--------------------|-------------------------------------------------------------------------------------|
| **Speed** | ⚡ Fast |
| **RAM Required** | ~3.6 GB |
| **Recommendation** | 🥇 **Best overall model.** Was a top 3 finisher for all questions except the haiku. |
## Prompt Template (ChatML)
This model uses the **ChatML** format used by Qwen:
```text
<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
Set this in your app (LM Studio, OpenWebUI, etc.) for best results.
## Generation Parameters
### Thinking Mode (Recommended for Logic)
Use when solving math, coding, or logical problems.
| Parameter | Value |
|----------------|-------|
| Temperature | 0.6 |
| Top-P | 0.95 |
| Top-K | 20 |
| Min-P | 0.0 |
| Repeat Penalty | 1.1 |
> ❗ DO NOT use greedy decoding — it causes infinite loops.
Enable via:
- `enable_thinking=True` in tokenizer
- Or add `/think` in user input during conversation
### Non-Thinking Mode (Fast Dialogue)
For casual chat and quick replies.
| Parameter | Value |
|----------------|-------|
| Temperature | 0.7 |
| Top-P | 0.8 |
| Top-K | 20 |
| Min-P | 0.0 |
| Repeat Penalty | 1.1 |
Enable via:
- `enable_thinking=False`
- Or add `/no_think` in prompt
Stop sequences: `<|im_end|>`, `<|im_start|>`
## 💡 Usage Tips
> This model supports two operational modes:
>
> ### 🔍 Thinking Mode (Recommended for Logic)
> Activate with `enable_thinking=True` or append `/think` in prompt.
>
> - Ideal for: math, coding, planning, analysis
> - Use sampling: `temp=0.6`, `top_p=0.95`, `top_k=20`
> - Avoid greedy decoding
>
> ### ⚡ Non-Thinking Mode (Fast Chat)
> Use `enable_thinking=False` or `/no_think`.
>
> - Best for: casual conversation, quick answers
> - Sampling: `temp=0.7`, `top_p=0.8`
>
> ---
>
> 🔄 **Switch Dynamically**
> In multi-turn chats, the last `/think` or `/no_think` directive takes precedence.
>
> 🔁 **Avoid Repetition**
> Set `presence_penalty=1.5` if stuck in loops.
>
> 📏 **Use Full Context**
> Allow up to 32,768 output tokens for complex tasks.
>
> 🧰 **Agent Ready**
> Works with Qwen-Agent, MCP servers, and custom tools.
## Customisation & Troubleshooting
Importing directly into Ollama should work, but you might encounter this error: `Error: invalid character '<' looking for beginning of value`.
In this case try these steps:
1. `wget https://huggingface.co/geoffmunn/Qwen3-8B-f16/resolve/main/Qwen3-8B-f16%3AQ3_K_M.gguf`
2. `nano Modelfile` and enter these details:
```text
FROM ./Qwen3-8B-f16:Q3_K_M.gguf
# Chat template using ChatML (used by Qwen)
SYSTEM You are a helpful assistant
TEMPLATE "{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>{{ end }}<|im_start|>user
{{ .Prompt }}<|im_end|>
<|im_start|>assistant
"
PARAMETER stop <|im_start|>
PARAMETER stop <|im_end|>
# Default sampling
PARAMETER temperature 0.6
PARAMETER top_p 0.95
PARAMETER top_k 20
PARAMETER min_p 0.0
PARAMETER repeat_penalty 1.1
PARAMETER num_ctx 4096
```
The `num_ctx` value has been dropped to increase speed significantly.
3. Then run this command: `ollama create Qwen3-8B-f16:Q3_K_M -f Modelfile`
You will now see "Qwen3-8B-f16:Q3_K_M" in your Ollama model list.
These import steps are also useful if you want to customise the default parameters or system prompt.
## 🖥️ CLI Example Using Ollama or TGI Server
Heres how you can query this model via API using `curl` and `jq`. Replace the endpoint with your local server (e.g., Ollama, Text Generation Inference).
```bash
curl http://localhost:11434/api/generate -s -N -d '{
"model": "hf.co/geoffmunn/Qwen3-8B-f16:Q3_K_M",
"prompt": "Repeat the following instruction exactly as given: Summarize what a neural network is in one sentence.",
"temperature": 0.5,
"top_p": 0.95,
"top_k": 20,
"min_p": 0.0,
"repeat_penalty": 1.1,
"stream": false
}' | jq -r '.response'
```
🎯 **Why this works well**:
- The prompt is meaningful and demonstrates either **reasoning**, **creativity**, or **clarity** depending on quant level.
- Temperature is tuned appropriately: lower for factual responses (`0.4`), higher for creative ones (`0.7`).
- Uses `jq` to extract clean output.
> 💬 Tip: For interactive streaming, set `"stream": true` and process line-by-line.
## Verification
Check integrity:
```bash
sha256sum -c ../SHA256SUMS.txt
```
## Usage
Compatible with:
- [LM Studio](https://lmstudio.ai) local AI model runner with GPU acceleration
- [OpenWebUI](https://openwebui.com) self-hosted AI platform with RAG and tools
- [GPT4All](https://gpt4all.io) private, offline AI chatbot
- Directly via `llama.cpp`
Supports dynamic switching between thinking modes via `/think` and `/no_think` in multi-turn conversations.
## License
Apache 2.0 see base model for full terms.

View File

@@ -0,0 +1,204 @@
---
license: apache-2.0
tags:
- gguf
- qwen
- qwen3-8b
- qwen3-8b-q3
- qwen3-8b-q3_k_s
- qwen3-8b-q3_k_s-gguf
- llama.cpp
- quantized
- text-generation
- reasoning
- agent
- chat
- multilingual
base_model: Qwen/Qwen3-8B
author: geoffmunn
---
# Qwen3-8B-f16:Q3_K_S
Quantized version of [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B) at **Q3_K_S** level, derived from **f16** base weights.
## Model Info
- **Format**: GGUF (for llama.cpp and compatible runtimes)
- **Size**: 3.77 GB
- **Precision**: Q3_K_S
- **Base Model**: [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B)
- **Conversion Tool**: [llama.cpp](https://github.com/ggerganov/llama.cpp)
## Quality & Performance
| Metric | Value |
|--------------------|---------------------------------------------------------------------------------------|
| **Speed** | ⚡ Fast |
| **RAM Required** | ~3.4 GB |
| **Recommendation** | 🥉 Came first and second in questions covering both ends of the temperature spectrum. |
## Prompt Template (ChatML)
This model uses the **ChatML** format used by Qwen:
```text
<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
Set this in your app (LM Studio, OpenWebUI, etc.) for best results.
## Generation Parameters
### Thinking Mode (Recommended for Logic)
Use when solving math, coding, or logical problems.
| Parameter | Value |
|----------------|-------|
| Temperature | 0.6 |
| Top-P | 0.95 |
| Top-K | 20 |
| Min-P | 0.0 |
| Repeat Penalty | 1.1 |
> ❗ DO NOT use greedy decoding — it causes infinite loops.
Enable via:
- `enable_thinking=True` in tokenizer
- Or add `/think` in user input during conversation
### Non-Thinking Mode (Fast Dialogue)
For casual chat and quick replies.
| Parameter | Value |
|----------------|-------|
| Temperature | 0.7 |
| Top-P | 0.8 |
| Top-K | 20 |
| Min-P | 0.0 |
| Repeat Penalty | 1.1 |
Enable via:
- `enable_thinking=False`
- Or add `/no_think` in prompt
Stop sequences: `<|im_end|>`, `<|im_start|>`
## 💡 Usage Tips
> This model supports two operational modes:
>
> ### 🔍 Thinking Mode (Recommended for Logic)
> Activate with `enable_thinking=True` or append `/think` in prompt.
>
> - Ideal for: math, coding, planning, analysis
> - Use sampling: `temp=0.6`, `top_p=0.95`, `top_k=20`
> - Avoid greedy decoding
>
> ### ⚡ Non-Thinking Mode (Fast Chat)
> Use `enable_thinking=False` or `/no_think`.
>
> - Best for: casual conversation, quick answers
> - Sampling: `temp=0.7`, `top_p=0.8`
>
> ---
>
> 🔄 **Switch Dynamically**
> In multi-turn chats, the last `/think` or `/no_think` directive takes precedence.
>
> 🔁 **Avoid Repetition**
> Set `presence_penalty=1.5` if stuck in loops.
>
> 📏 **Use Full Context**
> Allow up to 32,768 output tokens for complex tasks.
>
> 🧰 **Agent Ready**
> Works with Qwen-Agent, MCP servers, and custom tools.
## Customisation & Troubleshooting
Importing directly into Ollama should work, but you might encounter this error: `Error: invalid character '<' looking for beginning of value`.
In this case try these steps:
1. `wget https://huggingface.co/geoffmunn/Qwen3-8B-f16/resolve/main/Qwen3-8B-f16%3AQ3_K_S.gguf`
2. `nano Modelfile` and enter these details:
```text
FROM ./Qwen3-8B-f16:Q3_K_S.gguf
# Chat template using ChatML (used by Qwen)
SYSTEM You are a helpful assistant
TEMPLATE "{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>{{ end }}<|im_start|>user
{{ .Prompt }}<|im_end|>
<|im_start|>assistant
"
PARAMETER stop <|im_start|>
PARAMETER stop <|im_end|>
# Default sampling
PARAMETER temperature 0.6
PARAMETER top_p 0.95
PARAMETER top_k 20
PARAMETER min_p 0.0
PARAMETER repeat_penalty 1.1
PARAMETER num_ctx 4096
```
The `num_ctx` value has been dropped to increase speed significantly.
3. Then run this command: `ollama create Qwen3-8B-f16:Q3_K_S -f Modelfile`
You will now see "Qwen3-8B-f16:Q3_K_S" in your Ollama model list.
These import steps are also useful if you want to customise the default parameters or system prompt.
## 🖥️ CLI Example Using Ollama or TGI Server
Heres how you can query this model via API using `curl` and `jq`. Replace the endpoint with your local server (e.g., Ollama, Text Generation Inference).
```bash
curl http://localhost:11434/api/generate -s -N -d '{
"model": "hf.co/geoffmunn/Qwen3-8B-f16:Q3_K_S",
"prompt": "Repeat the following instruction exactly as given: Summarize what a neural network is in one sentence.",
"temperature": 0.5,
"top_p": 0.95,
"top_k": 20,
"min_p": 0.0,
"repeat_penalty": 1.1,
"stream": false
}' | jq -r '.response'
```
🎯 **Why this works well**:
- The prompt is meaningful and demonstrates either **reasoning**, **creativity**, or **clarity** depending on quant level.
- Temperature is tuned appropriately: lower for factual responses (`0.4`), higher for creative ones (`0.7`).
- Uses `jq` to extract clean output.
> 💬 Tip: For interactive streaming, set `"stream": true` and process line-by-line.
## Verification
Check integrity:
```bash
sha256sum -c ../SHA256SUMS.txt
```
## Usage
Compatible with:
- [LM Studio](https://lmstudio.ai) local AI model runner with GPU acceleration
- [OpenWebUI](https://openwebui.com) self-hosted AI platform with RAG and tools
- [GPT4All](https://gpt4all.io) private, offline AI chatbot
- Directly via `llama.cpp`
Supports dynamic switching between thinking modes via `/think` and `/no_think` in multi-turn conversations.
## License
Apache 2.0 see base model for full terms.

View File

@@ -0,0 +1,206 @@
---
license: apache-2.0
tags:
- gguf
- qwen
- qwen3-8b
- qwen3-8b-q4
- qwen3-8b-q4_k_m
- qwen3-8b-q4_k_m-gguf
- llama.cpp
- quantized
- text-generation
- reasoning
- agent
- chat
- multilingual
- imatrix
- q3_hifi
base_model: Qwen/Qwen3-8B
author: geoffmunn
---
# Qwen3-8B-f16:Q4_K_M
Quantized version of [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B) at **Q4_K_M** level, derived from **f16** base weights.
## Model Info
- **Format**: GGUF (for llama.cpp and compatible runtimes)
- **Size**: 5.85 GB
- **Precision**: Q4_K_M
- **Base Model**: [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B)
- **Conversion Tool**: [llama.cpp](https://github.com/ggerganov/llama.cpp)
## Quality & Performance
| Metric | Value |
|--------------------|-------------------------------------------------------------------------|
| **Speed** | 🚀 Fast |
| **RAM Required** | ~4.3 GB |
| **Recommendation** | Came first and second in questions covering high temperature questions. |
## Prompt Template (ChatML)
This model uses the **ChatML** format used by Qwen:
```text
<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
Set this in your app (LM Studio, OpenWebUI, etc.) for best results.
## Generation Parameters
### Thinking Mode (Recommended for Logic)
Use when solving math, coding, or logical problems.
| Parameter | Value |
|----------------|-------|
| Temperature | 0.6 |
| Top-P | 0.95 |
| Top-K | 20 |
| Min-P | 0.0 |
| Repeat Penalty | 1.1 |
> ❗ DO NOT use greedy decoding — it causes infinite loops.
Enable via:
- `enable_thinking=True` in tokenizer
- Or add `/think` in user input during conversation
### Non-Thinking Mode (Fast Dialogue)
For casual chat and quick replies.
| Parameter | Value |
|----------------|-------|
| Temperature | 0.7 |
| Top-P | 0.8 |
| Top-K | 20 |
| Min-P | 0.0 |
| Repeat Penalty | 1.1 |
Enable via:
- `enable_thinking=False`
- Or add `/no_think` in prompt
Stop sequences: `<|im_end|>`, `<|im_start|>`
## 💡 Usage Tips
> This model supports two operational modes:
>
> ### 🔍 Thinking Mode (Recommended for Logic)
> Activate with `enable_thinking=True` or append `/think` in prompt.
>
> - Ideal for: math, coding, planning, analysis
> - Use sampling: `temp=0.6`, `top_p=0.95`, `top_k=20`
> - Avoid greedy decoding
>
> ### ⚡ Non-Thinking Mode (Fast Chat)
> Use `enable_thinking=False` or `/no_think`.
>
> - Best for: casual conversation, quick answers
> - Sampling: `temp=0.7`, `top_p=0.8`
>
> ---
>
> 🔄 **Switch Dynamically**
> In multi-turn chats, the last `/think` or `/no_think` directive takes precedence.
>
> 🔁 **Avoid Repetition**
> Set `presence_penalty=1.5` if stuck in loops.
>
> 📏 **Use Full Context**
> Allow up to 32,768 output tokens for complex tasks.
>
> 🧰 **Agent Ready**
> Works with Qwen-Agent, MCP servers, and custom tools.
## Customisation & Troubleshooting
Importing directly into Ollama should work, but you might encounter this error: `Error: invalid character '<' looking for beginning of value`.
In this case try these steps:
1. `wget https://huggingface.co/geoffmunn/Qwen3-8B-f16/resolve/main/Qwen3-8B-f16%3AQ4_K_M.gguf`
2. `nano Modelfile` and enter these details:
```text
FROM ./Qwen3-8B-f16:Q4_K_M.gguf
# Chat template using ChatML (used by Qwen)
SYSTEM You are a helpful assistant
TEMPLATE "{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>{{ end }}<|im_start|>user
{{ .Prompt }}<|im_end|>
<|im_start|>assistant
"
PARAMETER stop <|im_start|>
PARAMETER stop <|im_end|>
# Default sampling
PARAMETER temperature 0.6
PARAMETER top_p 0.95
PARAMETER top_k 20
PARAMETER min_p 0.0
PARAMETER repeat_penalty 1.1
PARAMETER num_ctx 4096
```
The `num_ctx` value has been dropped to increase speed significantly.
3. Then run this command: `ollama create Qwen3-8B-f16:Q4_K_M -f Modelfile`
You will now see "Qwen3-8B-f16:Q4_K_M" in your Ollama model list.
These import steps are also useful if you want to customise the default parameters or system prompt.
## 🖥️ CLI Example Using Ollama or TGI Server
Heres how you can query this model via API using `curl` and `jq`. Replace the endpoint with your local server (e.g., Ollama, Text Generation Inference).
```bash
curl http://localhost:11434/api/generate -s -N -d '{
"model": "hf.co/geoffmunn/Qwen3-8B-f16:Q4_K_M",
"prompt": "Repeat the following instruction exactly as given: Write a short haiku about autumn leaves falling gently in a quiet forest.",
"temperature": 0.7,
"top_p": 0.95,
"top_k": 20,
"min_p": 0.0,
"repeat_penalty": 1.1,
"stream": false
}' | jq -r '.response'
```
🎯 **Why this works well**:
- The prompt is meaningful and demonstrates either **reasoning**, **creativity**, or **clarity** depending on quant level.
- Temperature is tuned appropriately: lower for factual responses (`0.4`), higher for creative ones (`0.7`).
- Uses `jq` to extract clean output.
> 💬 Tip: For interactive streaming, set `"stream": true` and process line-by-line.
## Verification
Check integrity:
```bash
sha256sum -c ../SHA256SUMS.txt
```
## Usage
Compatible with:
- [LM Studio](https://lmstudio.ai) local AI model runner with GPU acceleration
- [OpenWebUI](https://openwebui.com) self-hosted AI platform with RAG and tools
- [GPT4All](https://gpt4all.io) private, offline AI chatbot
- Directly via `llama.cpp`
Supports dynamic switching between thinking modes via `/think` and `/no_think` in multi-turn conversations.
## License
Apache 2.0 see base model for full terms.

View File

@@ -0,0 +1,206 @@
---
license: apache-2.0
tags:
- gguf
- qwen
- qwen3-8b
- qwen3-8b-q4
- qwen3-8b-q4_k_s
- qwen3-8b-q4_k_s-gguf
- llama.cpp
- quantized
- text-generation
- reasoning
- agent
- chat
- multilingual
- imatrix
- q3_hifi
base_model: Qwen/Qwen3-8B
author: geoffmunn
---
# Qwen3-8B-f16:Q4_K_S
Quantized version of [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B) at **Q4_K_S** level, derived from **f16** base weights.
## Model Info
- **Format**: GGUF (for llama.cpp and compatible runtimes)
- **Size**: 4.8 GB
- **Precision**: Q4_K_S
- **Base Model**: [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B)
- **Conversion Tool**: [llama.cpp](https://github.com/ggerganov/llama.cpp)
## Quality & Performance
| Metric | Value |
|--------------------|---------------------------------------------------------------------------------------|
| **Speed** | 🚀 Fast |
| **RAM Required** | ~4.1 GB |
| **Recommendation** | 🥉 Came first and second in questions covering both ends of the temperature spectrum. |
## Prompt Template (ChatML)
This model uses the **ChatML** format used by Qwen:
```text
<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
Set this in your app (LM Studio, OpenWebUI, etc.) for best results.
## Generation Parameters
### Thinking Mode (Recommended for Logic)
Use when solving math, coding, or logical problems.
| Parameter | Value |
|----------------|-------|
| Temperature | 0.6 |
| Top-P | 0.95 |
| Top-K | 20 |
| Min-P | 0.0 |
| Repeat Penalty | 1.1 |
> ❗ DO NOT use greedy decoding — it causes infinite loops.
Enable via:
- `enable_thinking=True` in tokenizer
- Or add `/think` in user input during conversation
### Non-Thinking Mode (Fast Dialogue)
For casual chat and quick replies.
| Parameter | Value |
|----------------|-------|
| Temperature | 0.7 |
| Top-P | 0.8 |
| Top-K | 20 |
| Min-P | 0.0 |
| Repeat Penalty | 1.1 |
Enable via:
- `enable_thinking=False`
- Or add `/no_think` in prompt
Stop sequences: `<|im_end|>`, `<|im_start|>`
## 💡 Usage Tips
> This model supports two operational modes:
>
> ### 🔍 Thinking Mode (Recommended for Logic)
> Activate with `enable_thinking=True` or append `/think` in prompt.
>
> - Ideal for: math, coding, planning, analysis
> - Use sampling: `temp=0.6`, `top_p=0.95`, `top_k=20`
> - Avoid greedy decoding
>
> ### ⚡ Non-Thinking Mode (Fast Chat)
> Use `enable_thinking=False` or `/no_think`.
>
> - Best for: casual conversation, quick answers
> - Sampling: `temp=0.7`, `top_p=0.8`
>
> ---
>
> 🔄 **Switch Dynamically**
> In multi-turn chats, the last `/think` or `/no_think` directive takes precedence.
>
> 🔁 **Avoid Repetition**
> Set `presence_penalty=1.5` if stuck in loops.
>
> 📏 **Use Full Context**
> Allow up to 32,768 output tokens for complex tasks.
>
> 🧰 **Agent Ready**
> Works with Qwen-Agent, MCP servers, and custom tools.
## Customisation & Troubleshooting
Importing directly into Ollama should work, but you might encounter this error: `Error: invalid character '<' looking for beginning of value`.
In this case try these steps:
1. `wget https://huggingface.co/geoffmunn/Qwen3-8B-f16/resolve/main/Qwen3-8B-f16%3AQ4_K_S.gguf`
2. `nano Modelfile` and enter these details:
```text
FROM ./Qwen3-8B-f16:Q4_K_S.gguf
# Chat template using ChatML (used by Qwen)
SYSTEM You are a helpful assistant
TEMPLATE "{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>{{ end }}<|im_start|>user
{{ .Prompt }}<|im_end|>
<|im_start|>assistant
"
PARAMETER stop <|im_start|>
PARAMETER stop <|im_end|>
# Default sampling
PARAMETER temperature 0.6
PARAMETER top_p 0.95
PARAMETER top_k 20
PARAMETER min_p 0.0
PARAMETER repeat_penalty 1.1
PARAMETER num_ctx 4096
```
The `num_ctx` value has been dropped to increase speed significantly.
3. Then run this command: `ollama create Qwen3-8B-f16:Q4_K_S -f Modelfile`
You will now see "Qwen3-8B-f16:Q4_K_S" in your Ollama model list.
These import steps are also useful if you want to customise the default parameters or system prompt.
## 🖥️ CLI Example Using Ollama or TGI Server
Heres how you can query this model via API using `curl` and `jq`. Replace the endpoint with your local server (e.g., Ollama, Text Generation Inference).
```bash
curl http://localhost:11434/api/generate -s -N -d '{
"model": "hf.co/geoffmunn/Qwen3-8B-f16:Q4_K_S",
"prompt": "Repeat the following instruction exactly as given: Summarize what a neural network is in one sentence.",
"temperature": 0.5,
"top_p": 0.95,
"top_k": 20,
"min_p": 0.0,
"repeat_penalty": 1.1,
"stream": false
}' | jq -r '.response'
```
🎯 **Why this works well**:
- The prompt is meaningful and demonstrates either **reasoning**, **creativity**, or **clarity** depending on quant level.
- Temperature is tuned appropriately: lower for factual responses (`0.4`), higher for creative ones (`0.7`).
- Uses `jq` to extract clean output.
> 💬 Tip: For interactive streaming, set `"stream": true` and process line-by-line.
## Verification
Check integrity:
```bash
sha256sum -c ../SHA256SUMS.txt
```
## Usage
Compatible with:
- [LM Studio](https://lmstudio.ai) local AI model runner with GPU acceleration
- [OpenWebUI](https://openwebui.com) self-hosted AI platform with RAG and tools
- [GPT4All](https://gpt4all.io) private, offline AI chatbot
- Directly via `llama.cpp`
Supports dynamic switching between thinking modes via `/think` and `/no_think` in multi-turn conversations.
## License
Apache 2.0 see base model for full terms.

View File

@@ -0,0 +1,206 @@
---
license: apache-2.0
tags:
- gguf
- qwen
- qwen3-8b
- qwen3-8b-q5
- qwen3-8b-q5_k_m
- qwen3-8b-q5_k_m-gguf
- llama.cpp
- quantized
- text-generation
- reasoning
- agent
- chat
- multilingual
- imatrix
- q3_hifi
base_model: Qwen/Qwen3-8B
author: geoffmunn
---
# Qwen3-8B-f16:Q5_K_M
Quantized version of [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B) at **Q5_K_M** level, derived from **f16** base weights.
## Model Info
- **Format**: GGUF (for llama.cpp and compatible runtimes)
- **Size**: 5.85 GB
- **Precision**: Q5_K_M
- **Base Model**: [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B)
- **Conversion Tool**: [llama.cpp](https://github.com/ggerganov/llama.cpp)
## Quality & Performance
| Metric | Value |
|--------------------|-----------------------------------------------------------------|
| **Speed** | 🐢 Medium |
| **RAM Required** | ~4.9 GB |
| **Recommendation** | Not recommended, no appeareances in the top 3 for any question. |
## Prompt Template (ChatML)
This model uses the **ChatML** format used by Qwen:
```text
<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
Set this in your app (LM Studio, OpenWebUI, etc.) for best results.
## Generation Parameters
### Thinking Mode (Recommended for Logic)
Use when solving math, coding, or logical problems.
| Parameter | Value |
|----------------|-------|
| Temperature | 0.6 |
| Top-P | 0.95 |
| Top-K | 20 |
| Min-P | 0.0 |
| Repeat Penalty | 1.1 |
> ❗ DO NOT use greedy decoding — it causes infinite loops.
Enable via:
- `enable_thinking=True` in tokenizer
- Or add `/think` in user input during conversation
### Non-Thinking Mode (Fast Dialogue)
For casual chat and quick replies.
| Parameter | Value |
|----------------|-------|
| Temperature | 0.7 |
| Top-P | 0.8 |
| Top-K | 20 |
| Min-P | 0.0 |
| Repeat Penalty | 1.1 |
Enable via:
- `enable_thinking=False`
- Or add `/no_think` in prompt
Stop sequences: `<|im_end|>`, `<|im_start|>`
## 💡 Usage Tips
> This model supports two operational modes:
>
> ### 🔍 Thinking Mode (Recommended for Logic)
> Activate with `enable_thinking=True` or append `/think` in prompt.
>
> - Ideal for: math, coding, planning, analysis
> - Use sampling: `temp=0.6`, `top_p=0.95`, `top_k=20`
> - Avoid greedy decoding
>
> ### ⚡ Non-Thinking Mode (Fast Chat)
> Use `enable_thinking=False` or `/no_think`.
>
> - Best for: casual conversation, quick answers
> - Sampling: `temp=0.7`, `top_p=0.8`
>
> ---
>
> 🔄 **Switch Dynamically**
> In multi-turn chats, the last `/think` or `/no_think` directive takes precedence.
>
> 🔁 **Avoid Repetition**
> Set `presence_penalty=1.5` if stuck in loops.
>
> 📏 **Use Full Context**
> Allow up to 32,768 output tokens for complex tasks.
>
> 🧰 **Agent Ready**
> Works with Qwen-Agent, MCP servers, and custom tools.
## Customisation & Troubleshooting
Importing directly into Ollama should work, but you might encounter this error: `Error: invalid character '<' looking for beginning of value`.
In this case try these steps:
1. `wget https://huggingface.co/geoffmunn/Qwen3-8B-f16/resolve/main/Qwen3-8B-f16%3AQ5_K_M.gguf`
2. `nano Modelfile` and enter these details:
```text
FROM ./Qwen3-8B-f16:Q5_K_M.gguf
# Chat template using ChatML (used by Qwen)
SYSTEM You are a helpful assistant
TEMPLATE "{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>{{ end }}<|im_start|>user
{{ .Prompt }}<|im_end|>
<|im_start|>assistant
"
PARAMETER stop <|im_start|>
PARAMETER stop <|im_end|>
# Default sampling
PARAMETER temperature 0.6
PARAMETER top_p 0.95
PARAMETER top_k 20
PARAMETER min_p 0.0
PARAMETER repeat_penalty 1.1
PARAMETER num_ctx 4096
```
The `num_ctx` value has been dropped to increase speed significantly.
3. Then run this command: `ollama create Qwen3-8B-f16:Q5_K_M -f Modelfile`
You will now see "Qwen3-8B-f16:Q5_K_M" in your Ollama model list.
These import steps are also useful if you want to customise the default parameters or system prompt.
## 🖥️ CLI Example Using Ollama or TGI Server
Heres how you can query this model via API using `curl` and `jq`. Replace the endpoint with your local server (e.g., Ollama, Text Generation Inference).
```bash
curl http://localhost:11434/api/generate -s -N -d '{
"model": "hf.co/geoffmunn/Qwen3-8B-f16:Q5_K_M",
"prompt": "Repeat the following instruction exactly as given: Explain why the sky appears blue during the day but red at sunrise and sunset, using physics principles like Rayleigh scattering.",
"temperature": 0.4,
"top_p": 0.95,
"top_k": 20,
"min_p": 0.0,
"repeat_penalty": 1.1,
"stream": false
}' | jq -r '.response'
```
🎯 **Why this works well**:
- The prompt is meaningful and demonstrates either **reasoning**, **creativity**, or **clarity** depending on quant level.
- Temperature is tuned appropriately: lower for factual responses (`0.4`), higher for creative ones (`0.7`).
- Uses `jq` to extract clean output.
> 💬 Tip: For interactive streaming, set `"stream": true` and process line-by-line.
## Verification
Check integrity:
```bash
sha256sum -c ../SHA256SUMS.txt
```
## Usage
Compatible with:
- [LM Studio](https://lmstudio.ai) local AI model runner with GPU acceleration
- [OpenWebUI](https://openwebui.com) self-hosted AI platform with RAG and tools
- [GPT4All](https://gpt4all.io) private, offline AI chatbot
- Directly via `llama.cpp`
Supports dynamic switching between thinking modes via `/think` and `/no_think` in multi-turn conversations.
## License
Apache 2.0 see base model for full terms.

View File

@@ -0,0 +1,206 @@
---
license: apache-2.0
tags:
- gguf
- qwen
- qwen3-8b
- qwen3-8b-q5
- qwen3-8b-q5_k_s
- qwen3-8b-q5_k_s-gguf
- llama.cpp
- quantized
- text-generation
- reasoning
- agent
- chat
- multilingual
- imatrix
- q3_hifi
base_model: Qwen/Qwen3-8B
author: geoffmunn
---
# Qwen3-8B-f16:Q5_K_S
Quantized version of [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B) at **Q5_K_S** level, derived from **f16** base weights.
## Model Info
- **Format**: GGUF (for llama.cpp and compatible runtimes)
- **Size**: 5.72 GB
- **Precision**: Q5_K_S
- **Base Model**: [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B)
- **Conversion Tool**: [llama.cpp](https://github.com/ggerganov/llama.cpp)
## Quality & Performance
| Metric | Value |
|--------------------|---------------------------------------------------|
| **Speed** | 🐢 Medium |
| **RAM Required** | ~4.8 GB |
| **Recommendation** | 🥈 A good second place. Good for all query types. |
## Prompt Template (ChatML)
This model uses the **ChatML** format used by Qwen:
```text
<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
Set this in your app (LM Studio, OpenWebUI, etc.) for best results.
## Generation Parameters
### Thinking Mode (Recommended for Logic)
Use when solving math, coding, or logical problems.
| Parameter | Value |
|----------------|-------|
| Temperature | 0.6 |
| Top-P | 0.95 |
| Top-K | 20 |
| Min-P | 0.0 |
| Repeat Penalty | 1.1 |
> ❗ DO NOT use greedy decoding — it causes infinite loops.
Enable via:
- `enable_thinking=True` in tokenizer
- Or add `/think` in user input during conversation
### Non-Thinking Mode (Fast Dialogue)
For casual chat and quick replies.
| Parameter | Value |
|----------------|-------|
| Temperature | 0.7 |
| Top-P | 0.8 |
| Top-K | 20 |
| Min-P | 0.0 |
| Repeat Penalty | 1.1 |
Enable via:
- `enable_thinking=False`
- Or add `/no_think` in prompt
Stop sequences: `<|im_end|>`, `<|im_start|>`
## 💡 Usage Tips
> This model supports two operational modes:
>
> ### 🔍 Thinking Mode (Recommended for Logic)
> Activate with `enable_thinking=True` or append `/think` in prompt.
>
> - Ideal for: math, coding, planning, analysis
> - Use sampling: `temp=0.6`, `top_p=0.95`, `top_k=20`
> - Avoid greedy decoding
>
> ### ⚡ Non-Thinking Mode (Fast Chat)
> Use `enable_thinking=False` or `/no_think`.
>
> - Best for: casual conversation, quick answers
> - Sampling: `temp=0.7`, `top_p=0.8`
>
> ---
>
> 🔄 **Switch Dynamically**
> In multi-turn chats, the last `/think` or `/no_think` directive takes precedence.
>
> 🔁 **Avoid Repetition**
> Set `presence_penalty=1.5` if stuck in loops.
>
> 📏 **Use Full Context**
> Allow up to 32,768 output tokens for complex tasks.
>
> 🧰 **Agent Ready**
> Works with Qwen-Agent, MCP servers, and custom tools.
## Customisation & Troubleshooting
Importing directly into Ollama should work, but you might encounter this error: `Error: invalid character '<' looking for beginning of value`.
In this case try these steps:
1. `wget https://huggingface.co/geoffmunn/Qwen3-8B-f16/resolve/main/Qwen3-8B-f16%3AQ5_K_S.gguf`
2. `nano Modelfile` and enter these details:
```text
FROM ./Qwen3-8B-f16:Q5_K_S.gguf
# Chat template using ChatML (used by Qwen)
SYSTEM You are a helpful assistant
TEMPLATE "{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>{{ end }}<|im_start|>user
{{ .Prompt }}<|im_end|>
<|im_start|>assistant
"
PARAMETER stop <|im_start|>
PARAMETER stop <|im_end|>
# Default sampling
PARAMETER temperature 0.6
PARAMETER top_p 0.95
PARAMETER top_k 20
PARAMETER min_p 0.0
PARAMETER repeat_penalty 1.1
PARAMETER num_ctx 4096
```
The `num_ctx` value has been dropped to increase speed significantly.
3. Then run this command: `ollama create Qwen3-8B-f16:Q5_K_S -f Modelfile`
You will now see "Qwen3-8B-f16:Q5_K_S" in your Ollama model list.
These import steps are also useful if you want to customise the default parameters or system prompt.
## 🖥️ CLI Example Using Ollama or TGI Server
Heres how you can query this model via API using `curl` and `jq`. Replace the endpoint with your local server (e.g., Ollama, Text Generation Inference).
```bash
curl http://localhost:11434/api/generate -s -N -d '{
"model": "hf.co/geoffmunn/Qwen3-8B-f16:Q5_K_S",
"prompt": "Repeat the following instruction exactly as given: Write a short haiku about autumn leaves falling gently in a quiet forest.",
"temperature": 0.7,
"top_p": 0.95,
"top_k": 20,
"min_p": 0.0,
"repeat_penalty": 1.1,
"stream": false
}' | jq -r '.response'
```
🎯 **Why this works well**:
- The prompt is meaningful and demonstrates either **reasoning**, **creativity**, or **clarity** depending on quant level.
- Temperature is tuned appropriately: lower for factual responses (`0.4`), higher for creative ones (`0.7`).
- Uses `jq` to extract clean output.
> 💬 Tip: For interactive streaming, set `"stream": true` and process line-by-line.
## Verification
Check integrity:
```bash
sha256sum -c ../SHA256SUMS.txt
```
## Usage
Compatible with:
- [LM Studio](https://lmstudio.ai) local AI model runner with GPU acceleration
- [OpenWebUI](https://openwebui.com) self-hosted AI platform with RAG and tools
- [GPT4All](https://gpt4all.io) private, offline AI chatbot
- Directly via `llama.cpp`
Supports dynamic switching between thinking modes via `/think` and `/no_think` in multi-turn conversations.
## License
Apache 2.0 see base model for full terms.

206
Qwen3-8B-f16-Q6_K/README.md Normal file
View File

@@ -0,0 +1,206 @@
---
license: apache-2.0
tags:
- gguf
- qwen
- qwen3-8b
- qwen3-8b-q6
- qwen3-8b-q6_k
- qwen3-8b-q6_k-gguf
- llama.cpp
- quantized
- text-generation
- reasoning
- agent
- chat
- multilingual
- imatrix
- q3_hifi
base_model: Qwen/Qwen3-8B
author: geoffmunn
---
# Qwen3-8B-f16:Q6_K
Quantized version of [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B) at **Q6_K** level, derived from **f16** base weights.
## Model Info
- **Format**: GGUF (for llama.cpp and compatible runtimes)
- **Size**: 6.73 GB
- **Precision**: Q6_K
- **Base Model**: [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B)
- **Conversion Tool**: [llama.cpp](https://github.com/ggerganov/llama.cpp)
## Quality & Performance
| Metric | Value |
|--------------------|---------------------------------------------------|
| **Speed** | 🐌 Slow |
| **RAM Required** | ~5.5 GB |
| **Recommendation** | Showed up in a few results, but not recommended. |
## Prompt Template (ChatML)
This model uses the **ChatML** format used by Qwen:
```text
<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
Set this in your app (LM Studio, OpenWebUI, etc.) for best results.
## Generation Parameters
### Thinking Mode (Recommended for Logic)
Use when solving math, coding, or logical problems.
| Parameter | Value |
|----------------|-------|
| Temperature | 0.6 |
| Top-P | 0.95 |
| Top-K | 20 |
| Min-P | 0.0 |
| Repeat Penalty | 1.1 |
> ❗ DO NOT use greedy decoding — it causes infinite loops.
Enable via:
- `enable_thinking=True` in tokenizer
- Or add `/think` in user input during conversation
### Non-Thinking Mode (Fast Dialogue)
For casual chat and quick replies.
| Parameter | Value |
|----------------|-------|
| Temperature | 0.7 |
| Top-P | 0.8 |
| Top-K | 20 |
| Min-P | 0.0 |
| Repeat Penalty | 1.1 |
Enable via:
- `enable_thinking=False`
- Or add `/no_think` in prompt
Stop sequences: `<|im_end|>`, `<|im_start|>`
## 💡 Usage Tips
> This model supports two operational modes:
>
> ### 🔍 Thinking Mode (Recommended for Logic)
> Activate with `enable_thinking=True` or append `/think` in prompt.
>
> - Ideal for: math, coding, planning, analysis
> - Use sampling: `temp=0.6`, `top_p=0.95`, `top_k=20`
> - Avoid greedy decoding
>
> ### ⚡ Non-Thinking Mode (Fast Chat)
> Use `enable_thinking=False` or `/no_think`.
>
> - Best for: casual conversation, quick answers
> - Sampling: `temp=0.7`, `top_p=0.8`
>
> ---
>
> 🔄 **Switch Dynamically**
> In multi-turn chats, the last `/think` or `/no_think` directive takes precedence.
>
> 🔁 **Avoid Repetition**
> Set `presence_penalty=1.5` if stuck in loops.
>
> 📏 **Use Full Context**
> Allow up to 32,768 output tokens for complex tasks.
>
> 🧰 **Agent Ready**
> Works with Qwen-Agent, MCP servers, and custom tools.
## Customisation & Troubleshooting
Importing directly into Ollama should work, but you might encounter this error: `Error: invalid character '<' looking for beginning of value`.
In this case try these steps:
1. `wget https://huggingface.co/geoffmunn/Qwen3-8B-f16/resolve/main/Qwen3-8B-f16%3AQ6_K.gguf`
2. `nano Modelfile` and enter these details:
```text
FROM ./Qwen3-8B-f16:Q6_K.gguf
# Chat template using ChatML (used by Qwen)
SYSTEM You are a helpful assistant
TEMPLATE "{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>{{ end }}<|im_start|>user
{{ .Prompt }}<|im_end|>
<|im_start|>assistant
"
PARAMETER stop <|im_start|>
PARAMETER stop <|im_end|>
# Default sampling
PARAMETER temperature 0.6
PARAMETER top_p 0.95
PARAMETER top_k 20
PARAMETER min_p 0.0
PARAMETER repeat_penalty 1.1
PARAMETER num_ctx 4096
```
The `num_ctx` value has been dropped to increase speed significantly.
3. Then run this command: `ollama create Qwen3-8B-f16:Q6_K -f Modelfile`
You will now see "Qwen3-8B-f16:Q6_K" in your Ollama model list.
These import steps are also useful if you want to customise the default parameters or system prompt.
## 🖥️ CLI Example Using Ollama or TGI Server
Heres how you can query this model via API using `curl` and `jq`. Replace the endpoint with your local server (e.g., Ollama, Text Generation Inference).
```bash
curl http://localhost:11434/api/generate -s -N -d '{
"model": "hf.co/geoffmunn/Qwen3-8B-f16:Q6_K",
"prompt": "Repeat the following instruction exactly as given: Explain why the sky appears blue during the day but red at sunrise and sunset, using physics principles like Rayleigh scattering.",
"temperature": 0.4,
"top_p": 0.95,
"top_k": 20,
"min_p": 0.0,
"repeat_penalty": 1.1,
"stream": false
}' | jq -r '.response'
```
🎯 **Why this works well**:
- The prompt is meaningful and demonstrates either **reasoning**, **creativity**, or **clarity** depending on quant level.
- Temperature is tuned appropriately: lower for factual responses (`0.4`), higher for creative ones (`0.7`).
- Uses `jq` to extract clean output.
> 💬 Tip: For interactive streaming, set `"stream": true` and process line-by-line.
## Verification
Check integrity:
```bash
sha256sum -c ../SHA256SUMS.txt
```
## Usage
Compatible with:
- [LM Studio](https://lmstudio.ai) local AI model runner with GPU acceleration
- [OpenWebUI](https://openwebui.com) self-hosted AI platform with RAG and tools
- [GPT4All](https://gpt4all.io) private, offline AI chatbot
- Directly via `llama.cpp`
Supports dynamic switching between thinking modes via `/think` and `/no_think` in multi-turn conversations.
## License
Apache 2.0 see base model for full terms.

206
Qwen3-8B-f16-Q8_0/README.md Normal file
View File

@@ -0,0 +1,206 @@
---
license: apache-2.0
tags:
- gguf
- qwen
- qwen3-8b
- qwen3-8b-q8
- qwen3-8b-q8_0
- qwen3-8b-q8_0-gguf
- llama.cpp
- quantized
- text-generation
- reasoning
- agent
- chat
- multilingual
- imatrix
- q3_hifi
base_model: Qwen/Qwen3-8B
author: geoffmunn
---
# Qwen3-8B-f16:Q8_0
Quantized version of [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B) at **Q8_0** level, derived from **f16** base weights.
## Model Info
- **Format**: GGUF (for llama.cpp and compatible runtimes)
- **Size**: 8.71 GB
- **Precision**: Q8_0
- **Base Model**: [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B)
- **Conversion Tool**: [llama.cpp](https://github.com/ggerganov/llama.cpp)
## Quality & Performance
| Metric | Value |
|--------------------|-----------------------------------------|
| **Speed** | 🐌 Slow |
| **RAM Required** | ~7.1 GB |
| **Recommendation** | Not recommended, Only one top 3 finish. |
## Prompt Template (ChatML)
This model uses the **ChatML** format used by Qwen:
```text
<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
Set this in your app (LM Studio, OpenWebUI, etc.) for best results.
## Generation Parameters
### Thinking Mode (Recommended for Logic)
Use when solving math, coding, or logical problems.
| Parameter | Value |
|----------------|-------|
| Temperature | 0.6 |
| Top-P | 0.95 |
| Top-K | 20 |
| Min-P | 0.0 |
| Repeat Penalty | 1.1 |
> ❗ DO NOT use greedy decoding — it causes infinite loops.
Enable via:
- `enable_thinking=True` in tokenizer
- Or add `/think` in user input during conversation
### Non-Thinking Mode (Fast Dialogue)
For casual chat and quick replies.
| Parameter | Value |
|----------------|-------|
| Temperature | 0.7 |
| Top-P | 0.8 |
| Top-K | 20 |
| Min-P | 0.0 |
| Repeat Penalty | 1.1 |
Enable via:
- `enable_thinking=False`
- Or add `/no_think` in prompt
Stop sequences: `<|im_end|>`, `<|im_start|>`
## 💡 Usage Tips
> This model supports two operational modes:
>
> ### 🔍 Thinking Mode (Recommended for Logic)
> Activate with `enable_thinking=True` or append `/think` in prompt.
>
> - Ideal for: math, coding, planning, analysis
> - Use sampling: `temp=0.6`, `top_p=0.95`, `top_k=20`
> - Avoid greedy decoding
>
> ### ⚡ Non-Thinking Mode (Fast Chat)
> Use `enable_thinking=False` or `/no_think`.
>
> - Best for: casual conversation, quick answers
> - Sampling: `temp=0.7`, `top_p=0.8`
>
> ---
>
> 🔄 **Switch Dynamically**
> In multi-turn chats, the last `/think` or `/no_think` directive takes precedence.
>
> 🔁 **Avoid Repetition**
> Set `presence_penalty=1.5` if stuck in loops.
>
> 📏 **Use Full Context**
> Allow up to 32,768 output tokens for complex tasks.
>
> 🧰 **Agent Ready**
> Works with Qwen-Agent, MCP servers, and custom tools.
## Customisation & Troubleshooting
Importing directly into Ollama should work, but you might encounter this error: `Error: invalid character '<' looking for beginning of value`.
In this case try these steps:
1. `wget https://huggingface.co/geoffmunn/Qwen3-8B-f16/resolve/main/Qwen3-8B-f16%3AQ8_0.gguf`
2. `nano Modelfile` and enter these details:
```text
FROM ./Qwen3-8B-f16:Q8_0.gguf
# Chat template using ChatML (used by Qwen)
SYSTEM You are a helpful assistant
TEMPLATE "{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>{{ end }}<|im_start|>user
{{ .Prompt }}<|im_end|>
<|im_start|>assistant
"
PARAMETER stop <|im_start|>
PARAMETER stop <|im_end|>
# Default sampling
PARAMETER temperature 0.6
PARAMETER top_p 0.95
PARAMETER top_k 20
PARAMETER min_p 0.0
PARAMETER repeat_penalty 1.1
PARAMETER num_ctx 4096
```
The `num_ctx` value has been dropped to increase speed significantly.
3. Then run this command: `ollama create Qwen3-8B-f16:Q8_0 -f Modelfile`
You will now see "Qwen3-8B-f16:Q8_0" in your Ollama model list.
These import steps are also useful if you want to customise the default parameters or system prompt.
## 🖥️ CLI Example Using Ollama or TGI Server
Heres how you can query this model via API using `curl` and `jq`. Replace the endpoint with your local server (e.g., Ollama, Text Generation Inference).
```bash
curl http://localhost:11434/api/generate -s -N -d '{
"model": "hf.co/geoffmunn/Qwen3-8B-f16:Q8_0",
"prompt": "Repeat the following instruction exactly as given: Explain why the sky appears blue during the day but red at sunrise and sunset, using physics principles like Rayleigh scattering.",
"temperature": 0.4,
"top_p": 0.95,
"top_k": 20,
"min_p": 0.0,
"repeat_penalty": 1.1,
"stream": false
}' | jq -r '.response'
```
🎯 **Why this works well**:
- The prompt is meaningful and demonstrates either **reasoning**, **creativity**, or **clarity** depending on quant level.
- Temperature is tuned appropriately: lower for factual responses (`0.4`), higher for creative ones (`0.7`).
- Uses `jq` to extract clean output.
> 💬 Tip: For interactive streaming, set `"stream": true` and process line-by-line.
## Verification
Check integrity:
```bash
sha256sum -c ../SHA256SUMS.txt
```
## Usage
Compatible with:
- [LM Studio](https://lmstudio.ai) local AI model runner with GPU acceleration
- [OpenWebUI](https://openwebui.com) self-hosted AI platform with RAG and tools
- [GPT4All](https://gpt4all.io) private, offline AI chatbot
- Directly via `llama.cpp`
Supports dynamic switching between thinking modes via `/think` and `/no_think` in multi-turn conversations.
## License
Apache 2.0 see base model for full terms.

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d2ac1c7c3bdf55f85355b01c144f6c5a1a65c9ffd28a9355453d95c444d099df
size 5347200

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7e59d35c1c4d3114f8ed305e2ab19139d2eaa615d5badab9a978025829d3daa2
size 5347200

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d2ac1c7c3bdf55f85355b01c144f6c5a1a65c9ffd28a9355453d95c444d099df
size 5347200

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7e59d35c1c4d3114f8ed305e2ab19139d2eaa615d5badab9a978025829d3daa2
size 5347200

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:85db3fdb95a5155b04fe5297f07ba20b2eb1714fd4b73eebde67e56b3ae130f8
size 3281733216

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:02cf559368fea369c3442ef58d6c8436b1175d343e805600476f3836a6c7278f
size 3442927200

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e85a7c1984ab12122261702e3d5ee18f55c251920ac11c5902d90bf513ec253a
size 4828370528

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3b0dfac8d304aede9580a307cf42abc9f1eee0599a7115d75ba9dee40bed0dc0
size 4124161632

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d6bcf405ed302a736c96d2ac5193a94111e07577dc187ebf3e7f876bc0011c50
size 3769611872

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:57b73385284b268a88568ef54d8dd4a10f3868a02da799c2000a2c1963017254
size 5703579264

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c09d414416f245c8d20759d7629806aeb6250b5de3b2c159f37f9ac168d4b2b5
size 5027784288

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:21cd101368068873d9b6ce431aba451e47c4b6b0cbafcd957ad34f673ac30601
size 4802012768

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f2ddf0729977105a55cdda882e9d799bff3bac36d1e7fef2a1982f6da00cec0f
size 6041810560

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e5625e30f36988e76e3886ac5cfc97e8e746d50dd1f413f08677f70b099e831c
size 5851113056

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a085d45efe804436e1d078957c4abbc96fe799622a544232ef0b7069a88f4186
size 5720761952

3
Qwen3-8B-f16:Q2_K.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:93551e45c3c4dd15ca31420a824662c2e5ec0f80310786bcb4364c2c54077d11
size 3281732960

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8a3013d2518eb264ec3e98fbfcf64b3d8e61786f4845cdd13c702a2900a09091
size 3442926944

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5caa61a0b2c4f2e008a3a700cc4d2c7958623677ba0415a4e0652f931bcd20f4
size 4811986272

3
Qwen3-8B-f16:Q3_K_M.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a85cd12d9887f0de21c0df6dda0601061e569a580b8121dd94096ea8fbacab2f
size 4124161376

3
Qwen3-8B-f16:Q3_K_S.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:34e25be87d35fb3f00779710b22b2b69e6210f67eefcf269f1b3758274a9b2ee
size 3769611616

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b5cd60d8a390afcee5a5f443ef25235b2592ed72487405d9666d88c6cfb33806
size 5703579040

3
Qwen3-8B-f16:Q4_K_M.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:10ded9291f250596ef7149438dd5da80bf3780b0eebcd2a1922c2b4a08c36e54
size 5027784032

3
Qwen3-8B-f16:Q4_K_S.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:07567ef46246e0f7bca5dad8fccc7c383e592d5912edd362b3ade8351ad30aab
size 4802012512

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:36d20cea46aaa9bcce80ea48bc0da74bac483218179bba82967829cc734ce1b8
size 6041810336

3
Qwen3-8B-f16:Q5_K_M.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:560737e3d4f9182d664858aea9c0a164f751154f78e32abc3276d1b4e606f17d
size 5851112800

3
Qwen3-8B-f16:Q5_K_S.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3c7625faba5fee4a39d627c1e6b2331d21423437a1240ca336a676973a3d8824
size 5720761696

3
Qwen3-8B-f16:Q6_K.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:863da90f35b7c22b1cd184a20d04cc0eaf0df67f8f52ab0a6d4f68d192600898
size 6725899552

3
Qwen3-8B-f16:Q8_0.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:21962d706e584d0058a3f078ba42a40b7c3c82a1aa1a25588d372d41c99a8b6e
size 8709518624

327
README.md Normal file
View File

@@ -0,0 +1,327 @@
---
license: apache-2.0
tags:
- gguf
- qwen
- qwen3
- qwen3-8b
- qwen3-8b-gguf
- llama.cpp
- quantized
- text-generation
- reasoning
- agent
- chat
- multilingual
- matrix
- q3_hifi
- q4_hifi
- q5_hifi
base_model: Qwen/Qwen3-8B
author: geoffmunn
pipeline_tag: text-generation
language:
- en
- zh
- es
- fr
- de
- ru
- ar
- ja
- ko
- hi
---
# Qwen3-8B-f16-GGUF
This is a **GGUF-quantized version** of the **[Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B)** language model - an **8-billion-parameter** LLM from Alibaba's Qwen series, designed for **advanced reasoning, agentic behavior, and multilingual tasks**.
Converted for use with `llama.cpp` and compatible tools like OpenWebUI, LM Studio, GPT4All, and more.
## Why Use an 8B Model?
The **Qwen3-8B** model represents a significant leap in capability while remaining remarkably accessible for local and edge deployment. It offers:
- **Near-state-of-the-art reasoning, coding, and multilingual performance** among open 8B-class models
- **Smooth inference on a single consumer GPU** (e.g., 1624 GB VRAM) or fast CPU runtime with quantization
- **Quantized versions (e.g., GGUF Q4_K_M, AWQ) that fit within ~68 GB of memory**, enabling use on mid-range hardware
- **Strong performance on complex tasks** like document summarization, structured output generation, and agentic workflows
Its ideal for:
- Local AI assistants that handle nuanced, multi-turn conversations
- Self-hosted RAG pipelines with deep document understanding
- Developers building production-grade on-prem AI features without cloud dependencies
- Researchers and tinkerers seeking a capable yet manageable open-weight foundation
Choose Qwen3-8B when you need high-quality output and robust general intelligence - but still value efficiency, privacy, and full control over your deployment environment.
# Qwen3 8B Quantization Guide: Cross-Bit Summary & Recommendations
## Executive Summary
At 8B scale, **quantization achieves exceptional resilience**—all bit widths deliver production-ready quality with imatrix, and even Q2_K becomes viable (+13.4% loss). The model's parameter redundancy provides a "sweet spot" where aggressive compression meets robust architecture. Q5_K_HIFI + imatrix achieves near-lossless fidelity (+0.27% vs F16), while Q4_K_M + imatrix offers the best balance of quality, speed, and compatibility:
| Bit Width | Best Variant (+ imatrix) | Quality vs F16 | File Size | Speed | Memory | Viability |
|-----------|--------------------------|----------------|-----------|-------|--------|-----------|
| **Q5_K** | Q5_K_HIFI + imatrix | **+0.27%** ✅✅✅ | 5.62 GiB | 109.7 TPS | 5,754 MiB | Exceptional |
| **Q4_K** | Q4_K_M + imatrix | **+1.3%** ✅✅ | 4.68 GiB | 125.5 TPS | 4,792 MiB | Excellent |
| **Q3_K** | Q3_K_HIFI + imatrix | **+3.5%** ✅ | 2.15 GiB | 151.3 TPS | 2,202 MiB | Very Good |
| **Q2_K** | Q2_K + imatrix | **+13.4%** ⚠️ | 3.05 GiB | 169.9 TPS | 3,134 MiB | Fair (viable) |
💡 **Critical insight**: 8B represents the **inflection point** where Q2_K becomes genuinely viable with imatrix (+13.4% loss vs +35% at 1.7B). Q5_K_HIFI + imatrix achieves near-lossless quality (+0.27%), while Q4_K_M + imatrix provides the best practical balance. All variants are production-ready with imatrix.
---
## Bit-Width Recommendations by Use Case
### ✅ Quality-Critical Applications
**→ Q5_K_HIFI + imatrix**
- Best perplexity at **10.1377 PPL (+0.27% vs F16)** — near-lossless fidelity
- Only 0.27% precision loss represents the closest approach to F16 quality across all quantization levels
- Requires custom llama.cpp build with `Q6_K_HIFI_RES8` support
- ⚠️ **Never use Q5_K_S without imatrix** — quality degrades to +1.62% vs F16
### ⚖️ Best Overall Balance (Recommended Default)
**→ Q4_K_M + imatrix**
- Excellent +1.3% precision loss vs F16 (PPL 10.2384)
- Strong 125.5 TPS speed (+171% vs F16)
- Compact 4.68 GiB file size (69.3% smaller than F16)
- **Standard llama.cpp compatibility** — no custom build required
- Ideal for most development and production scenarios
### 🚀 Maximum Speed
**→ Q2_K + imatrix**
- Fastest variant at **169.9 TPS** (+267% vs F16)
- Surprisingly viable quality at +13.4% loss with imatrix
- ⚠️ **Never use without imatrix** — quality degrades catastrophically to +57.9% loss
### 💎 Near-Lossless 3-Bit Option
**→ Q3_K_HIFI + imatrix**
- **Remarkable +3.5% precision loss** — exceptional for 3-bit quantization
- 71.2% memory reduction (2,202 MiB vs 7,670 MiB)
- Unique value: When you need maximum compression but cannot accept Q3_K_S quality
- ⚠️ **2738% slower than Q3_K_M** — significant speed trade-off
### 📱 Extreme Memory Constraints (< 2.0 GiB)
**→ Q3_K_S + imatrix**
- Absolute smallest footprint (1.75 GiB file, 1,792 MiB runtime)
- Acceptable +9.0% precision loss with imatrix
- Only viable option under 2.0 GiB budget
---
## Critical Warnings for 8B Scale
⚠️ **Q5_K quality ranking reversal with imatrix** — Q5_K_S + imatrix (10.1538 PPL) actually beats Q5_K_M + imatrix (10.1612 PPL) by 0.07 PPL points. This makes Q5_K_S + imatrix viable for speed-constrained deployments where the 3.2% speed advantage matters.
⚠️ **Q4_K_S without imatrix is unusable** — Suffers +5.7% precision loss (10.6893 PPL) — the highest degradation of any Q4 variant at 8B scale. **Always pair Q4_K_S with imatrix** (reduces loss to +1.9%).
⚠️ **Q2_K requires imatrix** — Without it, Q2_K suffers +57.9% precision loss (completely unusable). With imatrix, quality improves to +13.4% — viable for non-critical tasks.
⚠️ **Q2_K_HIFI is strictly worse than Q2_K** — At 8B scale, Q2_K_HIFI loses to Q2_K on every metric (quality, speed, size, memory). Always prefer standard Q2_K over Q2_K_HIFI.
⚠️ **Q3_K_HIFI requires no special handling** — Unlike at 0.6B/1.7B scales, Q3_K_HIFI at 8B delivers substantial quality gains (+3.5% vs F16 with imatrix) that justify its 13.5% memory premium over Q3_K_M.
⚠️ **All Q3 variants are production-ready** — Even Q3_K_S with imatrix (+9.0% loss) remains usable for non-critical tasks — a dramatic improvement over smaller scales where Q3 quantization often fails.
---
## Memory Budget Guide
| Available VRAM | Recommended Variant | Expected Quality | Why |
|----------------|---------------------|------------------|-----|
| **< 2.0 GiB** | Q3_K_S + imatrix | PPL 11.02, +9.0% loss | Only option that fits; quality acceptable for non-critical tasks |
| **2.0 2.5 GiB** | Q3_K_M + imatrix | PPL 10.62, +5.1% loss | Best Q3 balance; production-ready quality |
| **2.5 3.5 GiB** | Q2_K + imatrix | PPL 11.46, +13.4% loss | Maximum speed at 169.9 TPS; quality acceptable for simple tasks |
| **3.5 5.0 GiB** | Q4_K_M + imatrix | PPL 10.24, +1.3% loss | Best balance of quality/speed/size; standard compatibility |
| **5.0 6.5 GiB** | Q5_K_HIFI + imatrix | PPL 10.14, +0.27% loss | Near-lossless quality; requires custom build |
| **> 15.3 GiB** | F16 | Best quality (baseline) | Only if absolute precision required |
---
## Cross-Bit Performance Comparison
| Priority | Q2_K Best | Q3_K Best | Q4_K Best | Q5_K Best | Winner |
|----------|-----------|-----------|-----------|-----------|--------|
| **Quality (with imat)** | Q2_K (+13.4%) | Q3_K_HIFI (+3.5%) | Q4_K_M (+1.3%) | **Q5_K_HIFI (+0.27%)** ✅ | **Q5_K_HIFI** |
| **Speed** | **Q2_K (169.9 TPS)** ✅ | Q3_K_S (223.5 TPS) | Q4_K_S (131.0 TPS) | Q5_K_S (113.3 TPS) | **Q2_K** |
| **Smallest Size** | Q2_K (3.05 GiB) | **Q3_K_S (1.75 GiB)** ✅ | Q4_K_S (4.47 GiB) | Q5_K_S (5.32 GiB) | **Q3_K_S** |
| **Best Balance** | Q2_K + imat | Q3_K_M + imat | **Q4_K_M + imat** ✅ | Q5_K_HIFI + imat | **Q4_K_M** |
✅ = Recommended for general use
⚠️ = Context-dependent (see warnings above)
---
## Scale-Specific Insights: Why 8B Quantizes So Well
1. **Model redundancy threshold**: 8B represents the point where parameter count provides sufficient redundancy that quantization errors average out rather than accumulating catastrophically (unlike 0.6B/1.7B)
2. **Q2_K viability inflection**: 8B is the smallest scale where Q2_K becomes genuinely viable with imatrix (+13.4% loss). At 4B, Q2_K + imatrix is +18.7%; at 1.7B, +35.0%. This demonstrates a clear scale-dependent improvement curve.
3. **imatrix effectiveness plateau**: imatrix recovers 6276% of precision loss at 8B — less dramatic than at 1.7B (7078%) but more consistent across bit widths. Q5_K_S benefits most (74.1% recovery), making it competitive with Q5_K_M when imatrix is used.
4. **Residual quantization sweet spot**: Q5_K_HIFI's `Q6_K_HIFI_RES8` tensors provide maximal benefit at 8B scale — the 5 residual tensors capture precisely the right amount of quantization error without overhead.
5. **Q4_K_HIFI behavior shift**: Unlike at 14B where imatrix *harms* Q4_K_HIFI, at 8B imatrix *helps* it (-1.1% PPL improvement) — demonstrating non-linear scale effects.
6. **Q3_K viability threshold**: 8B is the smallest scale where Q3_K_HIFI achieves truly production-ready quality (+3.5% with imatrix) — below this, Q3 quantization requires careful validation.
---
## Decision Flowchart
```mermaid
Need best quality?
├─ Yes → Q5_K_HIFI + imatrix (+0.27% loss)
└─ No → Need max speed?
├─ Yes → Q2_K + imatrix (169.9 TPS, +13.4% loss)
└─ No → Need smallest size?
├─ Yes → Memory < 2.0 GiB?
│ ├─ Yes → Q3_K_S + imatrix (1,792 MiB, +9.0% loss)
│ └─ No → Q2_K + imatrix (3,134 MiB, +13.4% loss, fastest)
└─ No → Q4_K_M + imatrix (best balance, +1.3% loss, standard build)
```
---
## Practical Deployment Recommendations
### For Most Users
**→ Q4_K_M + imatrix**
Delivers excellent quality (+1.3% vs F16), strong speed (125.5 TPS), compact size (4.68 GiB), and universal llama.cpp compatibility. The safe, practical choice for 95% of deployments.
### For Quality-Critical Work
**→ Q5_K_HIFI + imatrix**
Achieves near-lossless quantization (+0.27% vs F16) with 64% memory reduction and 2.4× speedup. Requires custom build but worth it for research, content generation, or any task where output fidelity is non-negotiable.
### For Edge/Mobile Deployment
**→ Q3_K_HIFI + imatrix**
Best Q3 quality (+3.5% vs F16) with smallest viable footprint (2.15 GiB). Production-ready even without imatrix (+8.6% loss) — valuable for environments where imatrix generation isn't feasible.
### For High-Throughput Serving
**→ Q5_K_S + imatrix**
Fastest Q5 variant (113.3 TPS) with surprisingly good quality (+0.42% vs F16) that actually beats Q5_K_M with imatrix. Ideal when every TPS matters and marginal quality differences are acceptable.
### For Maximum Compression
**→ Q2_K + imatrix**
Only consider when memory/speed are absolutely critical and quality degradation is acceptable. At 8B scale, Q2_K + imatrix achieves +13.4% loss — viable for simple chatbots or non-critical inference.
---
## Bottom Line Recommendations
| Scenario | Recommended Variant | Rationale |
|----------|---------------------|-----------|
| **Default / General Purpose** | Q4_K_M + imatrix | Best balance of quality (+1.3%), speed (125.5 TPS), size (4.68 GiB), and compatibility |
| **Maximum Quality** | Q5_K_HIFI + imatrix | Near-lossless (+0.27% vs F16) with 64% memory reduction and 2.4× speedup |
| **Maximum Speed** | Q2_K + imatrix | Fastest (169.9 TPS, +267% vs F16) with acceptable quality (+13.4% loss) |
| **Minimum Size** | Q3_K_S + imatrix | Smallest footprint (1.75 GiB) with acceptable quality (+9.0% loss) |
| **No imatrix available** | Q5_K_HIFI (no imat) | Still excellent (+1.11% vs F16); all variants usable but quality reduced |
| **Extreme constraints** | Q3_K_S + imatrix | Only if memory < 2.0 GiB; +9.0% loss acceptable for non-critical tasks |
**Golden rules for 8B**:
1. **Always use imatrix** provides 6276% precision recovery across all bit widths
2. **Never use Q2_K without imatrix** completely unusable (+57.9% loss)
3. **Prefer Q2_K over Q2_K_HIFI** HIFI is strictly worse on all metrics at 8B
4. **Q5_K_S + imatrix beats Q5_K_M + imatrix** unexpected quality ranking reversal
5. **All four bit widths are viable** choose based on constraints, not quality cliffs
**8B is the quantization sweet spot**: Large enough for robustness across all bit widths (even Q2_K), small enough for dramatic efficiency gains. This scale demonstrates that intelligent quantization can deliver near-F16 quality at 1/3 the memory with 2.43.5× speed a compelling value proposition for nearly all deployments.
## Non-technical model anaysis and rankings
**NOTE:** This analysis does not include the HIFI models.
There are numerous good candidates - lots of different models showed up in the top 3 across all the quesionts. However, **Qwen3-8B-f16:Q3_K_M** was a finalist in all but one question so is the recommended model (or Qwen3-8B-f16:Q3_HIFI). **Qwen3-8B-f16:Q5_K_S** did nearly as well and is worth considering,
The 'hello' question is the first time that all models got it exactly right. All models in the 8B range did well and it's mainly a question of what one works best on your hardware.
You can read the results here: [Qwen3-8B-analysis.md](Qwen3-8B-analysis.md)
If you find this useful, please give the project a like.
## Non-HIFI recommentation table based on output
| Level | Speed | Size | Recommendation |
|-----------|-----------|-------------|---------------------------------------------------------------------------------------|
| Q2_K | Fastest | 3.28 GB | Not recommended. Came first in the bat & ball question, no other appearances. |
| 🥉Q3_K_S | Fast | 3.77 GB | 🥉 Came first and second in questions covering both ends of the temperature spectrum. |
| 🥇 Q3_K_M | Fast | 4.12 GB | 🥇 **Best overall model.** Was a top 3 finisher for all questions except the haiku. |
| 🥉Q4_K_S | 🚀 Fast | 4.8 GB | 🥉 Came first and second in questions covering both ends of the temperature spectrum. |
| Q4_K_M | 🚀 Fast | 5.85 GB | Came first and second in questions covering high temperature questions. |
| 🥈 Q5_K_S | 🐢 Medium | 5.72 GB | 🥈 A good second place. Good for all query types. |
| Q5_K_M | 🐢 Medium | 5.85 GB | Not recommended, no appeareances in the top 3 for any question. |
| Q6_K | 🐌 Slow | 6.73 GB | Showed up in a few results, but not recommended. |
| Q8_0 | 🐌 Slow | 8.71 GB | Not recommended, Only one top 3 finish. |
## Build notes
You can read the guide for building llama.cpp here: [HIFI_BUILD_GUIDE.md](https://github.com/geoffmunn/llama.cpp/blob/master/HIFI_BUILD_GUIDE.md).
The HIFI quantization also used a very large 4697 chunk imatrix file for extra precision. You can re-use it here: [Qwen3-8B-f16-imatrix-4697-generic.gguf](https://huggingface.co/geoffmunn/Qwen3-8B-f16/blob/main/Qwen3-8B-f16-imatrix-4697-generic.gguf)
The imatrix was created as a generic mix of Wikipedia, mathmatics, and coding examples.
### Source code
You can use the HIFI GitHub repository to build it from source if you're interested: [https://github.com/geoffmunn/llama.cpp](https://github.com/geoffmunn/llama.cpp).
Build notes: [HIFI_BUILD_GUIDE.md](https://github.com/geoffmunn/llama.cpp/blob/master/HIFI_BUILD_GUIDE.md)
Improvements and feedback are welcome.
## Usage
Load this model using:
- [OpenWebUI](https://openwebui.com) self-hosted AI interface with RAG & tools
- [LM Studio](https://lmstudio.ai) desktop app with GPU support and chat templates
- [GPT4All](https://gpt4all.io) private, local AI chatbot (offline-first)
- Or directly via `llama.cpp`
Each quantized model includes its own `README.md` and shares a common `MODELFILE` for optimal configuration.
Importing directly into Ollama should work, but you might encounter this error: `Error: invalid character '<' looking for beginning of value`.
In this case try these steps:
1. `wget https://huggingface.co/geoffmunn/Qwen3-8B-f16/resolve/main/Qwen3-8B-f16%3AQ3_K_M.gguf` (replace the quantised version with the one you want)
2. `nano Modelfile` and enter these details (again, replacing Q3_K_M with the version you want):
```text
FROM ./Qwen3-8B-f16:Q3_K_M.gguf
# Chat template using ChatML (used by Qwen)
SYSTEM You are a helpful assistant
TEMPLATE "{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>{{ end }}<|im_start|>user
{{ .Prompt }}<|im_end|>
<|im_start|>assistant
"
PARAMETER stop <|im_start|>
PARAMETER stop <|im_end|>
# Default sampling
PARAMETER temperature 0.6
PARAMETER top_p 0.95
PARAMETER top_k 20
PARAMETER min_p 0.0
PARAMETER repeat_penalty 1.1
PARAMETER num_ctx 4096
```
The `num_ctx` value has been dropped to increase speed significantly.
3. Then run this command: `ollama create Qwen3-8B-f16:Q3_K_M -f Modelfile`
You will now see "Qwen3-8B-f16:Q3_K_M" in your Ollama model list.
These import steps are also useful if you want to customise the default parameters or system prompt.
## Author
👤 Geoff Munn (@geoffmunn)
🔗 [Hugging Face Profile](https://huggingface.co/geoffmunn)
## Disclaimer
This is a community conversion for local inference. Not affiliated with Alibaba Cloud or the Qwen team.

40000
mixed-imatrix-dataset.txt Normal file

File diff suppressed because it is too large Load Diff