初始化项目,由ModelHub XC社区提供模型
Model: geoffmunn/Qwen3-8B-f16 Source: Original Platform
This commit is contained in:
71
.gitattributes
vendored
Normal file
71
.gitattributes
vendored
Normal file
@@ -0,0 +1,71 @@
|
||||
*.7z filter=lfs diff=lfs merge=lfs -text
|
||||
*.arrow filter=lfs diff=lfs merge=lfs -text
|
||||
*.bin filter=lfs diff=lfs merge=lfs -text
|
||||
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
||||
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
||||
*.ftz filter=lfs diff=lfs merge=lfs -text
|
||||
*.gz filter=lfs diff=lfs merge=lfs -text
|
||||
*.h5 filter=lfs diff=lfs merge=lfs -text
|
||||
*.joblib filter=lfs diff=lfs merge=lfs -text
|
||||
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
||||
*.model filter=lfs diff=lfs merge=lfs -text
|
||||
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
||||
*.npy filter=lfs diff=lfs merge=lfs -text
|
||||
*.npz filter=lfs diff=lfs merge=lfs -text
|
||||
*.onnx filter=lfs diff=lfs merge=lfs -text
|
||||
*.ot filter=lfs diff=lfs merge=lfs -text
|
||||
*.parquet filter=lfs diff=lfs merge=lfs -text
|
||||
*.pb filter=lfs diff=lfs merge=lfs -text
|
||||
*.pickle filter=lfs diff=lfs merge=lfs -text
|
||||
*.pkl filter=lfs diff=lfs merge=lfs -text
|
||||
*.pt filter=lfs diff=lfs merge=lfs -text
|
||||
*.pth filter=lfs diff=lfs merge=lfs -text
|
||||
*.rar filter=lfs diff=lfs merge=lfs -text
|
||||
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
||||
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar filter=lfs diff=lfs merge=lfs -text
|
||||
*.tflite filter=lfs diff=lfs merge=lfs -text
|
||||
*.tgz filter=lfs diff=lfs merge=lfs -text
|
||||
*.wasm filter=lfs diff=lfs merge=lfs -text
|
||||
*.xz filter=lfs diff=lfs merge=lfs -text
|
||||
*.zip filter=lfs diff=lfs merge=lfs -text
|
||||
*.zst filter=lfs diff=lfs merge=lfs -text
|
||||
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16:Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16:Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16:Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16:Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16:Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16:Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16:Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16:Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16:Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16-imatrix-5000.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16-Q3_HIFI.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16-imatrix-4697.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16-imatrix-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16-imatrix-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16-Q4_HIFI.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16:Q3_HIFI.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16:Q4_HIFI.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16-imatrix:Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16-imatrix:Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16:Q4_K_HIFI.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16-imatrix:Q4_K_HIFI.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16:Q3_K_HIFI.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16-imatrix:Q3_K_HIFI.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16-imatrix:Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16-imatrix:Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16-imatrix-4697-coder.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16-imatrix-4697-generic.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16-imatrix-8843-coder.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16-imatrix-9343-generic.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16-imatrix:Q5_K_HIFI.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16-imatrix:Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16-imatrix:Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16:Q5_K_HIFI.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16-imatrix:Q2_K_HIFI.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16-imatrix:Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Qwen3-8B-f16:Q2_K_HIFI.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
25
MODELFILE
Normal file
25
MODELFILE
Normal file
@@ -0,0 +1,25 @@
|
||||
# MODELFILE for Qwen3-8B-GGUF
|
||||
# Used by LM Studio, OpenWebUI, GPT4All, etc.
|
||||
|
||||
context_length: 32768
|
||||
embedding: false
|
||||
f16: cpu
|
||||
|
||||
# Chat template using ChatML (used by Qwen)
|
||||
prompt_template: >-
|
||||
<|im_start|>system
|
||||
You are a helpful assistant.<|im_end|>
|
||||
<|im_start|>user
|
||||
{prompt}<|im_end|>
|
||||
<|im_start|>assistant
|
||||
|
||||
# Stop sequences help end generation cleanly
|
||||
stop: "<|im_end|>"
|
||||
stop: "<|im_start|>"
|
||||
|
||||
# Default sampling (optimized for thinking mode)
|
||||
temperature: 0.6
|
||||
top_p: 0.95
|
||||
top_k: 20
|
||||
min_p: 0.0
|
||||
repeat_penalty: 1.1
|
||||
92
Q3_Quantization_Comparison.md
Normal file
92
Q3_Quantization_Comparison.md
Normal file
@@ -0,0 +1,92 @@
|
||||
# Qwen3-8B Quantization Comparison Summary
|
||||
|
||||
## Q3_HIFI (Adaptive/Custom)
|
||||
**Pros:**
|
||||
- 🏆 **Best quality** with lowest perplexity of 10.56 (4.4% better than Q3_K_M, 7.2% better than Q3_K_S)
|
||||
- 📦 **Smaller than Q3_K_M** (3.72 vs 3.84 GiB) while being significantly better quality
|
||||
- 🎯 Uses intelligent layer-sensitive quantization (Q3_HIFI on sensitive layers, mixed q3_K/q4_K elsewhere)
|
||||
- 📊 Most consistent results (lowest relative standard deviation in perplexity)
|
||||
|
||||
**Cons:**
|
||||
- 🐢 **Slowest inference** at 143.98 TPS (6.3% slower than Q3_K_S)
|
||||
- 🔧 Custom quantization may have less community support
|
||||
|
||||
**Best for:** Production deployments where output quality matters, tasks requiring accuracy (reasoning, coding, complex instructions), or when you want the best quality-to-size ratio.
|
||||
|
||||
## Performance Comparison (Q3_HIFI vs the others)
|
||||
|
||||
### Q3_K_M
|
||||
|
||||
| Metric | Q3_HIFI | Q3_K_M | Difference |
|
||||
|---------------------|----------|----------|------------------------------|
|
||||
| **Speed (TPS)** | 143.98 | 144.72 | -0.74 (0.5% slower) |
|
||||
| **Perplexity** | 10.56 | 11.05 | **-0.49 (4.4% better)** |
|
||||
| **File Size** | 3.72 GiB | 3.84 GiB | **-0.12 GiB (3.1% smaller)** |
|
||||
| **Bits Per Weight** | 3.90 | 4.02 | -0.12 (3.0% less) |
|
||||
|
||||
**Pros:**
|
||||
- ⚖️ Traditional "balanced" approach between speed and quality
|
||||
- 📚 Well-documented, standard quantization method
|
||||
|
||||
**Cons:**
|
||||
- 💾 **Largest file size** at 3.84 GiB despite not being the best quality
|
||||
- 🐌 Middle-of-the-road speed (144.7 TPS)
|
||||
- ❌ **Outclassed by Q3_HIFI** which is smaller AND better quality
|
||||
|
||||
**Best for:** Legacy compatibility or when you need a proven, standard quantization approach.
|
||||
**Summary:** Q3_HIFI delivers significantly better quality (4.4% lower perplexity) in a smaller package (3.1% less storage) with virtually no speed penalty (0.5% slower).
|
||||
|
||||
### Q3_K_S
|
||||
|
||||
| Metric | Q3_HIFI | Q3_K_S | Difference |
|
||||
|---------------------|----------|----------|-------------------------|
|
||||
| **Speed (TPS)** | 143.98 | 153.74 | -9.76 (6.3% slower) |
|
||||
| **Perplexity** | 10.56 | 11.38 | **-0.82 (7.2% better)** |
|
||||
| **File Size** | 3.72 GiB | 3.51 GiB | +0.21 GiB (6.0% larger) |
|
||||
| **Bits Per Weight** | 3.90 | 3.68 | +0.22 (6.0% more) |
|
||||
|
||||
**Pros:**
|
||||
- ⚡ **Fastest inference** at 153.74 TPS (~6% faster than Q3_K_M, ~7% faster than Q3_HIFI)
|
||||
- 💾 **Smallest file size** at 3.51 GiB
|
||||
- ✅ Best choice when speed and storage are critical
|
||||
|
||||
**Cons:**
|
||||
- ❌ **Worst quality** with perplexity of 11.38 (7.2% higher than Q3_HIFI)
|
||||
- Uses only q3_K quantization throughout (no mixed precision)
|
||||
|
||||
**Best for:** Resource-constrained environments, real-time applications where latency matters more than accuracy, or initial prototyping.
|
||||
**Summary:** Q3_HIFI trades a 6.3% speed reduction and 6.0% larger file size for a substantial 7.2% improvement in quality (lower perplexity).
|
||||
|
||||
---
|
||||
|
||||
## Recommendation Matrix
|
||||
|
||||
| Priority | Recommended Model | Rationale |
|
||||
|-------------------|-------------------|-----------------------------------------------------------------------------|
|
||||
| **Quality First** | Q3_HIFI | 7.2% better perplexity than Q3_K_S with acceptable speed loss |
|
||||
| **Speed First** | Q3_K_S | 6.3% faster inference, acceptable quality tradeoff for latency-sensitive apps |
|
||||
| **Best Balance** | Q3_HIFI | Better quality AND smaller size than Q3_K_M, only 0.5% slower |
|
||||
| **Smallest Size** | Q3_K_S | 6% smaller than alternatives |
|
||||
|
||||
---
|
||||
|
||||
## Key Insight
|
||||
|
||||
**Q3_HIFI represents a clear advancement** over the traditional Q3_K_M approach. It achieves:
|
||||
- **4.4% lower perplexity** (better accuracy)
|
||||
- **3.1% smaller file size** (3.72 vs 3.84 GiB)
|
||||
- Only **0.5% slower** inference (144.0 vs 144.7 TPS)
|
||||
|
||||
The Q3_K_M quantization is essentially obsoleted by Q3_HIFI for most use cases. The only remaining choice is between **Q3_K_S** (maximum speed, acceptable quality) and **Q3_HIFI** (maximum quality, acceptable speed).
|
||||
|
||||
## Appendix (Test Environment Details)
|
||||
|
||||
| Component | Specification |
|
||||
|---------------|---------------------------------|
|
||||
| **OS** | Ubuntu 24.04.3 LTS |
|
||||
| **CPU** | AMD EPYC 9254 24-Core Processor |
|
||||
| **CPU Cores** | 96 cores (2 threads/core) |
|
||||
| **RAM** | 1.0Ti |
|
||||
| **GPU** | NVIDIA L40S × 2 |
|
||||
| **VRAM** | 46068 MiB per GPU |
|
||||
| **CUDA** | 12.9 |
|
||||
196
Q4_Quantization_Comparison.md
Normal file
196
Q4_Quantization_Comparison.md
Normal file
@@ -0,0 +1,196 @@
|
||||
# Qwen3-8B Quantization Comparison Summary
|
||||
|
||||
## F16 Baseline Reference
|
||||
| Metric | Value |
|
||||
|--------|-------|
|
||||
| **F16 Perplexity** | 10.1114 |
|
||||
| **File Size** | 15.26 GiB (16.00 BPW) |
|
||||
|
||||
All precision loss percentages below are calculated relative to this F16 baseline.
|
||||
|
||||
---
|
||||
|
||||
## Q4_K_HIFI (INT8 Residuals + Per-Block Scale)
|
||||
**Pros:**
|
||||
- 🎯 Uses intelligent outlier preservation with INT8 residuals on critical tensors
|
||||
- 🔬 10 tensors use Q6_K_HIFI_RES8 format for maximum precision on sensitive weights
|
||||
- ✅ **Best perplexity** at 10.4133 (beats Q4_K_M's 10.4239)
|
||||
- ✅ **Best imatrix perplexity** at 10.2225 (beats Q4_K_M's 10.2355)
|
||||
- 📊 **+3.0% PPL vs F16** without imatrix, **+1.1% with imatrix** ✅
|
||||
|
||||
**Cons:**
|
||||
- 💾 Larger file size at 4.93 GiB (+5.3% vs Q4_K_M)
|
||||
- 🐢 Slower than Q4_K_M (117.53 vs 118.70 TPS, 1.0% slower)
|
||||
- 🐢 Slower than Q4_K_S (117.53 vs 123.04 TPS, 4.5% slower)
|
||||
|
||||
**Best for:** Maximum quality at Q4-level efficiency. Ideal when perplexity matters most.
|
||||
|
||||
## Performance Comparison (Q4_K_HIFI vs the others)
|
||||
|
||||
### Q4_K_M
|
||||
|
||||
| Metric | Q4_K_HIFI | Q4_K_M | Difference |
|
||||
|---------------------|------------|------------|-------------------------------|
|
||||
| **Speed (TPS)** | 117.53 | 118.70 | -1.17 (1.0% slower) |
|
||||
| **Perplexity** | 10.4133 | 10.4239 | **-0.01 (0.1% better)** ✅ |
|
||||
| **PPL vs F16** | +3.0% | +3.1% | 0.1% less precision loss |
|
||||
| **File Size** | 4.93 GiB | 4.68 GiB | +0.25 GiB (5.3% larger) |
|
||||
| **Bits Per Weight** | 5.17 | 4.90 | +0.27 (5.5% more) |
|
||||
|
||||
**Pros:**
|
||||
- ⚖️ Traditional "balanced" approach between speed and quality
|
||||
- 📚 Well-documented, standard quantization method
|
||||
- 💾 Smaller file size than Q4_K_HIFI
|
||||
- ⚡ Slightly faster than Q4_K_HIFI (1.0%)
|
||||
- 📊 **+3.1% PPL vs F16** — similar precision loss to Q4_K_HIFI
|
||||
|
||||
**Cons:**
|
||||
- ❌ **Worse perplexity** than Q4_K_HIFI (10.4239 vs 10.4133)
|
||||
|
||||
**Best for:** When speed and size are prioritized over marginal quality gains.
|
||||
**Summary:** Q4_K_HIFI delivers better quality (0.1% lower perplexity) at the cost of 5.3% larger file size and 1.0% slower inference.
|
||||
|
||||
### Q4_K_S
|
||||
|
||||
| Metric | Q4_K_HIFI | Q4_K_S | Difference |
|
||||
|---------------------|------------|------------|-------------------------------|
|
||||
| **Speed (TPS)** | 117.53 | 123.04 | -5.51 (4.5% slower) |
|
||||
| **Perplexity** | 10.4133 | 10.6837 | **-0.27 (2.5% better)** ✅ |
|
||||
| **PPL vs F16** | +3.0% | +5.7% | 2.7% less precision loss |
|
||||
| **File Size** | 4.93 GiB | 4.47 GiB | +0.46 GiB (10.3% larger) |
|
||||
| **Bits Per Weight** | 5.17 | 4.68 | +0.49 (10.5% more) |
|
||||
|
||||
**Pros:**
|
||||
- ⚡ **Fastest inference** at 123.04 TPS (4.5% faster than Q4_K_HIFI)
|
||||
- 💾 **Smallest file size** at 4.47 GiB
|
||||
- ✅ Best choice when speed and storage are critical
|
||||
|
||||
**Cons:**
|
||||
- ❌ **Significantly worse quality** with perplexity of 10.68 (2.5% higher than Q4_K_HIFI)
|
||||
- 📊 **+5.7% PPL vs F16** — highest precision loss of Q4_K variants
|
||||
- ⚠️ Uses minimal q6_K enhancement — impacts quality on sensitive weights
|
||||
|
||||
**Best for:** Extreme resource constraints, quick prototyping, or bulk processing where quality is less important.
|
||||
**Summary:** Q4_K_HIFI trades a 4.5% speed reduction and 10.3% larger file size for a 2.5% improvement in quality over Q4_K_S.
|
||||
|
||||
---
|
||||
|
||||
## Recommendation Matrix
|
||||
|
||||
| Priority | Recommended Model | Rationale |
|
||||
|-------------------|-------------------|--------------------------------------------------------------------------------|
|
||||
| **Quality First** | Q4_K_HIFI | Best perplexity (10.4133) without imatrix, (10.2225) with imatrix |
|
||||
| **Speed First** | Q4_K_S | 4.5% faster than Q4_K_HIFI, acceptable if quality degradation is tolerable |
|
||||
| **Best Balance** | Q4_K_HIFI | Best quality with acceptable speed/size overhead |
|
||||
| **Smallest Size** | Q4_K_S | 10.3% smaller than Q4_K_HIFI, 4.5% smaller than Q4_K_M |
|
||||
|
||||
---
|
||||
|
||||
## Key Insight
|
||||
|
||||
**At 8B scale, Q4_K_HIFI outperforms Q4_K_M with Q6_K_HIFI_RES8 enhancement:**
|
||||
- **0.1% lower perplexity** than Q4_K_M (10.4133 vs 10.4239)
|
||||
- **5.3% larger** file size (4.93 GiB vs 4.68 GiB)
|
||||
- **1.0% slower** than Q4_K_M (117.53 vs 118.70 TPS)
|
||||
- **All variants lose only 3.0-5.7% precision vs F16** (10.11 baseline) — excellent retention!
|
||||
|
||||
The Q6_K_HIFI_RES8 format (upgraded from Q5_K_HIFI_RES8) provides the precision needed at 8B scale. The INT8 residuals with 6-bit base quantization capture weight distributions that Q4_K_M's standard q6_K enhancement misses.
|
||||
|
||||
💡 **Scale Effect:** At 8B scale, models benefit from higher-precision HIFI enhancement. The Q6_K base (vs Q5_K) in the HIFI format provides the extra headroom needed for accurate outlier representation.
|
||||
|
||||
---
|
||||
|
||||
## Precision Loss Summary (vs F16 Baseline: PPL 10.1114)
|
||||
|
||||
| Model | PPL (no imatrix) | vs F16 | PPL (imatrix) | vs F16 |
|
||||
|-----------|------------------|--------|---------------|--------|
|
||||
| Q4_K_HIFI | 10.4133 | **+3.0%** | 10.2225 | **+1.1%** ✅ |
|
||||
| Q4_K_M | 10.4239 | **+3.1%** | 10.2355 | **+1.2%** |
|
||||
| Q4_K_S | 10.6837 | **+5.7%** | 10.2943 | **+1.8%** |
|
||||
|
||||
**Key Observations:**
|
||||
- Without imatrix: Q4_K variants lose only 3.0-5.7% precision vs F16 — excellent retention at 8B scale
|
||||
- With imatrix: Precision loss drops to just 1.1-1.8% — nearly F16 quality!
|
||||
- Q4_K_HIFI (imatrix) achieves the **lowest precision loss** at only +1.1% vs F16
|
||||
- The 8B model shows the best precision retention across all model sizes tested
|
||||
|
||||
---
|
||||
|
||||
## Tensor Distribution
|
||||
|
||||
| Model | q4_K | q5_K | q6_K | Q6_K_HIFI_RES8 | f32 | Total |
|
||||
|---------|------|------|------|----------------|-----|-------|
|
||||
| Q4_K_S | 245 | 8 | 1 | 0 | 145 | 399 |
|
||||
| Q4_K_M | 217 | 0 | 37 | 0 | 145 | 399 |
|
||||
| Q4_K_HIFI | 213 | 0 | 31 | 10 | 145 | 399 |
|
||||
|
||||
**Q4_K_HIFI Enhancement:** 10 critical tensors use Q6_K_HIFI_RES8 format with INT8 residuals + per-block scale (upgraded from Q5_K_HIFI_RES8 for 8B+ models).
|
||||
|
||||
---
|
||||
|
||||
## Addendum: Impact of imatrix on Q4_K_M and Q4_K_S
|
||||
|
||||
When all models are quantized **with an importance matrix (imatrix)**, quality improves significantly across all variants.
|
||||
|
||||
### imatrix Perplexity Improvements
|
||||
|
||||
| Model | Without imatrix | With imatrix | Improvement | PPL vs F16 (imatrix) |
|
||||
|---------|-----------------|--------------|-------------|----------------------|
|
||||
| Q4_K_HIFI | 10.4133 | **10.2225** | **-0.191 (1.8% better)** | **+1.1%** ✅ |
|
||||
| Q4_K_M | 10.4239 | **10.2355** | **-0.188 (1.8% better)** | **+1.2%** |
|
||||
| Q4_K_S | 10.6837 | **10.2943** | **-0.389 (3.6% better)** | **+1.8%** |
|
||||
|
||||
### Revised Comparison (All with imatrix)
|
||||
|
||||
| Model | PPL (imatrix) | vs Q4_K_HIFI | vs F16 | Size |
|
||||
|---------|---------------|------------|--------|------|
|
||||
| Q4_K_HIFI | 10.2225 | baseline | **+1.1%** ✅ | 4.93 GiB |
|
||||
| Q4_K_M | 10.2355 | +0.013 (+0.1%) | +1.2% | 4.68 GiB |
|
||||
| Q4_K_S | 10.2943 | +0.072 (+0.7%) | +1.8% | 4.47 GiB |
|
||||
|
||||
### Key Findings
|
||||
|
||||
**Q4_K_HIFI is the best choice with imatrix:**
|
||||
|
||||
| Comparison | Without imatrix | With imatrix |
|
||||
|------------|-----------------|--------------|
|
||||
| Q4_K_HIFI vs Q4_K_M | **-0.1%** (Q4_K_HIFI better) ✅ | **-0.1%** (Q4_K_HIFI better) ✅ |
|
||||
| Q4_K_HIFI vs Q4_K_S | **-2.5%** (Q4_K_HIFI better) ✅ | **-0.7%** (Q4_K_HIFI better) ✅ |
|
||||
|
||||
### Revised Recommendations (When Using imatrix)
|
||||
|
||||
| Priority | Without imatrix | With imatrix |
|
||||
|-------------------|-----------------|--------------|
|
||||
| **Quality First** | Q4_K_HIFI ✅ | **Q4_K_HIFI** ✅ (best perplexity) |
|
||||
| **Best Balance** | Q4_K_HIFI ✅ | **Q4_K_HIFI** ✅ |
|
||||
| **Size/Speed** | Q4_K_S | Q4_K_S |
|
||||
|
||||
### Conclusion
|
||||
|
||||
**For both imatrix and non-imatrix quantization:**
|
||||
- Q4_K_HIFI consistently outperforms Q4_K_M in perplexity (10.2225 vs 10.2355 with imatrix, 10.4133 vs 10.4239 without)
|
||||
- Q4_K_HIFI is 5.3% larger than Q4_K_M (4.93 GiB vs 4.68 GiB)
|
||||
- **Q4_K_HIFI + imatrix** offers the best quality with acceptable size/speed tradeoffs
|
||||
|
||||
**Q4_K_HIFI advantages:**
|
||||
- Best perplexity in all scenarios (with and without imatrix)
|
||||
- Q6_K_HIFI_RES8 format captures weight outliers that standard quantization misses
|
||||
- Significantly better quality than Q4_K_S (2.5% without imatrix, 0.7% with imatrix)
|
||||
|
||||
---
|
||||
|
||||
## Appendix (Test Environment Details)
|
||||
|
||||
| Component | Specification |
|
||||
|---------------|----------------------------------------|
|
||||
| **OS** | Ubuntu 24.04.3 LTS |
|
||||
| **CPU** | AMD EPYC 9254 24-Core Processor |
|
||||
| **CPU Cores** | 96 cores (2 threads/core) |
|
||||
| **RAM** | 1.0Ti |
|
||||
| **GPU** | NVIDIA L40S × 2 |
|
||||
| **VRAM** | 46068 MiB per GPU |
|
||||
| **CUDA** | 12.9 |
|
||||
| **Test Data** | wikitext-2-raw, 584 chunks |
|
||||
| **Context** | 512 tokens |
|
||||
| **Samples** | 100 per speed benchmark |
|
||||
| **imatrix** | mixed-imatrix-dataset.txt, 4697 chunks |
|
||||
1989
Qwen3-8B-analysis.md
Normal file
1989
Qwen3-8B-analysis.md
Normal file
File diff suppressed because it is too large
Load Diff
206
Qwen3-8B-f16-Q2_K/README.md
Normal file
206
Qwen3-8B-f16-Q2_K/README.md
Normal file
@@ -0,0 +1,206 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
tags:
|
||||
- gguf
|
||||
- qwen
|
||||
- qwen3-8b
|
||||
- qwen3-8b-q2
|
||||
- qwen3-8b-q2_k
|
||||
- qwen3-8b-q2_k-gguf
|
||||
- llama.cpp
|
||||
- quantized
|
||||
- text-generation
|
||||
- reasoning
|
||||
- agent
|
||||
- chat
|
||||
- multilingual
|
||||
- imatrix
|
||||
- q3_hifi
|
||||
base_model: Qwen/Qwen3-8B
|
||||
author: geoffmunn
|
||||
---
|
||||
|
||||
# Qwen3-8B-f16:Q2_K
|
||||
|
||||
Quantized version of [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B) at **Q2_K** level, derived from **f16** base weights.
|
||||
|
||||
## Model Info
|
||||
|
||||
- **Format**: GGUF (for llama.cpp and compatible runtimes)
|
||||
- **Size**: 3.28 GB
|
||||
- **Precision**: Q2_K
|
||||
- **Base Model**: [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B)
|
||||
- **Conversion Tool**: [llama.cpp](https://github.com/ggerganov/llama.cpp)
|
||||
|
||||
## Quality & Performance
|
||||
|
||||
| Metric | Value |
|
||||
|--------------------|-------------------------------------------------------------------------------|
|
||||
| **Speed** | ⚡ Fast |
|
||||
| **RAM Required** | ~3.0 GB |
|
||||
| **Recommendation** | Not recommended. Came first in the bat & ball question, no other appearances. |
|
||||
|
||||
## Prompt Template (ChatML)
|
||||
|
||||
This model uses the **ChatML** format used by Qwen:
|
||||
|
||||
```text
|
||||
<|im_start|>system
|
||||
You are a helpful assistant.<|im_end|>
|
||||
<|im_start|>user
|
||||
{prompt}<|im_end|>
|
||||
<|im_start|>assistant
|
||||
```
|
||||
|
||||
Set this in your app (LM Studio, OpenWebUI, etc.) for best results.
|
||||
|
||||
## Generation Parameters
|
||||
|
||||
### Thinking Mode (Recommended for Logic)
|
||||
Use when solving math, coding, or logical problems.
|
||||
|
||||
| Parameter | Value |
|
||||
|----------------|-------|
|
||||
| Temperature | 0.6 |
|
||||
| Top-P | 0.95 |
|
||||
| Top-K | 20 |
|
||||
| Min-P | 0.0 |
|
||||
| Repeat Penalty | 1.1 |
|
||||
|
||||
> ❗ DO NOT use greedy decoding — it causes infinite loops.
|
||||
|
||||
Enable via:
|
||||
- `enable_thinking=True` in tokenizer
|
||||
- Or add `/think` in user input during conversation
|
||||
|
||||
### Non-Thinking Mode (Fast Dialogue)
|
||||
For casual chat and quick replies.
|
||||
|
||||
| Parameter | Value |
|
||||
|----------------|-------|
|
||||
| Temperature | 0.7 |
|
||||
| Top-P | 0.8 |
|
||||
| Top-K | 20 |
|
||||
| Min-P | 0.0 |
|
||||
| Repeat Penalty | 1.1 |
|
||||
|
||||
Enable via:
|
||||
- `enable_thinking=False`
|
||||
- Or add `/no_think` in prompt
|
||||
|
||||
Stop sequences: `<|im_end|>`, `<|im_start|>`
|
||||
|
||||
## 💡 Usage Tips
|
||||
|
||||
> This model supports two operational modes:
|
||||
>
|
||||
> ### 🔍 Thinking Mode (Recommended for Logic)
|
||||
> Activate with `enable_thinking=True` or append `/think` in prompt.
|
||||
>
|
||||
> - Ideal for: math, coding, planning, analysis
|
||||
> - Use sampling: `temp=0.6`, `top_p=0.95`, `top_k=20`
|
||||
> - Avoid greedy decoding
|
||||
>
|
||||
> ### ⚡ Non-Thinking Mode (Fast Chat)
|
||||
> Use `enable_thinking=False` or `/no_think`.
|
||||
>
|
||||
> - Best for: casual conversation, quick answers
|
||||
> - Sampling: `temp=0.7`, `top_p=0.8`
|
||||
>
|
||||
> ---
|
||||
>
|
||||
> 🔄 **Switch Dynamically**
|
||||
> In multi-turn chats, the last `/think` or `/no_think` directive takes precedence.
|
||||
>
|
||||
> 🔁 **Avoid Repetition**
|
||||
> Set `presence_penalty=1.5` if stuck in loops.
|
||||
>
|
||||
> 📏 **Use Full Context**
|
||||
> Allow up to 32,768 output tokens for complex tasks.
|
||||
>
|
||||
> 🧰 **Agent Ready**
|
||||
> Works with Qwen-Agent, MCP servers, and custom tools.
|
||||
|
||||
## Customisation & Troubleshooting
|
||||
|
||||
Importing directly into Ollama should work, but you might encounter this error: `Error: invalid character '<' looking for beginning of value`.
|
||||
In this case try these steps:
|
||||
|
||||
1. `wget https://huggingface.co/geoffmunn/Qwen3-8B-f16/resolve/main/Qwen3-8B-f16%3AQ2_K.gguf`
|
||||
2. `nano Modelfile` and enter these details:
|
||||
```text
|
||||
FROM ./Qwen3-8B-f16:Q2_K.gguf
|
||||
|
||||
# Chat template using ChatML (used by Qwen)
|
||||
SYSTEM You are a helpful assistant
|
||||
|
||||
TEMPLATE "{{ if .System }}<|im_start|>system
|
||||
{{ .System }}<|im_end|>{{ end }}<|im_start|>user
|
||||
{{ .Prompt }}<|im_end|>
|
||||
<|im_start|>assistant
|
||||
"
|
||||
PARAMETER stop <|im_start|>
|
||||
PARAMETER stop <|im_end|>
|
||||
|
||||
# Default sampling
|
||||
PARAMETER temperature 0.6
|
||||
PARAMETER top_p 0.95
|
||||
PARAMETER top_k 20
|
||||
PARAMETER min_p 0.0
|
||||
PARAMETER repeat_penalty 1.1
|
||||
PARAMETER num_ctx 4096
|
||||
```
|
||||
|
||||
The `num_ctx` value has been dropped to increase speed significantly.
|
||||
|
||||
3. Then run this command: `ollama create Qwen3-8B-f16:Q2_K -f Modelfile`
|
||||
|
||||
You will now see "Qwen3-8B-f16:Q2_K" in your Ollama model list.
|
||||
|
||||
These import steps are also useful if you want to customise the default parameters or system prompt.
|
||||
|
||||
## 🖥️ CLI Example Using Ollama or TGI Server
|
||||
|
||||
Here’s how you can query this model via API using `curl` and `jq`. Replace the endpoint with your local server (e.g., Ollama, Text Generation Inference).
|
||||
|
||||
```bash
|
||||
curl http://localhost:11434/api/generate -s -N -d '{
|
||||
"model": "hf.co/geoffmunn/Qwen3-8B-f16:Q2_K",
|
||||
"prompt": "Repeat the following instruction exactly as given: Summarize what a neural network is in one sentence.",
|
||||
"temperature": 0.5,
|
||||
"top_p": 0.95,
|
||||
"top_k": 20,
|
||||
"min_p": 0.0,
|
||||
"repeat_penalty": 1.1,
|
||||
"stream": false
|
||||
}' | jq -r '.response'
|
||||
```
|
||||
|
||||
🎯 **Why this works well**:
|
||||
- The prompt is meaningful and demonstrates either **reasoning**, **creativity**, or **clarity** depending on quant level.
|
||||
- Temperature is tuned appropriately: lower for factual responses (`0.4`), higher for creative ones (`0.7`).
|
||||
- Uses `jq` to extract clean output.
|
||||
|
||||
> 💬 Tip: For interactive streaming, set `"stream": true` and process line-by-line.
|
||||
|
||||
## Verification
|
||||
|
||||
Check integrity:
|
||||
|
||||
```bash
|
||||
sha256sum -c ../SHA256SUMS.txt
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
Compatible with:
|
||||
- [LM Studio](https://lmstudio.ai) – local AI model runner with GPU acceleration
|
||||
- [OpenWebUI](https://openwebui.com) – self-hosted AI platform with RAG and tools
|
||||
- [GPT4All](https://gpt4all.io) – private, offline AI chatbot
|
||||
- Directly via `llama.cpp`
|
||||
|
||||
Supports dynamic switching between thinking modes via `/think` and `/no_think` in multi-turn conversations.
|
||||
|
||||
## License
|
||||
|
||||
Apache 2.0 – see base model for full terms.
|
||||
204
Qwen3-8B-f16-Q3_K_M/README.md
Normal file
204
Qwen3-8B-f16-Q3_K_M/README.md
Normal file
@@ -0,0 +1,204 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
tags:
|
||||
- gguf
|
||||
- qwen
|
||||
- qwen3-8b
|
||||
- qwen3-8b-q3
|
||||
- qwen3-8b-q3_k_m
|
||||
- qwen3-8b-q3_k_m-gguf
|
||||
- llama.cpp
|
||||
- quantized
|
||||
- text-generation
|
||||
- reasoning
|
||||
- agent
|
||||
- chat
|
||||
- multilingual
|
||||
base_model: Qwen/Qwen3-8B
|
||||
author: geoffmunn
|
||||
---
|
||||
|
||||
# Qwen3-8B-f16:Q3_K_M
|
||||
|
||||
Quantized version of [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B) at **Q3_K_M** level, derived from **f16** base weights.
|
||||
|
||||
## Model Info
|
||||
|
||||
- **Format**: GGUF (for llama.cpp and compatible runtimes)
|
||||
- **Size**: 4.12 GB
|
||||
- **Precision**: Q3_K_M
|
||||
- **Base Model**: [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B)
|
||||
- **Conversion Tool**: [llama.cpp](https://github.com/ggerganov/llama.cpp)
|
||||
|
||||
## Quality & Performance
|
||||
|
||||
| Metric | Value |
|
||||
|--------------------|-------------------------------------------------------------------------------------|
|
||||
| **Speed** | ⚡ Fast |
|
||||
| **RAM Required** | ~3.6 GB |
|
||||
| **Recommendation** | 🥇 **Best overall model.** Was a top 3 finisher for all questions except the haiku. |
|
||||
|
||||
## Prompt Template (ChatML)
|
||||
|
||||
This model uses the **ChatML** format used by Qwen:
|
||||
|
||||
```text
|
||||
<|im_start|>system
|
||||
You are a helpful assistant.<|im_end|>
|
||||
<|im_start|>user
|
||||
{prompt}<|im_end|>
|
||||
<|im_start|>assistant
|
||||
```
|
||||
|
||||
Set this in your app (LM Studio, OpenWebUI, etc.) for best results.
|
||||
|
||||
## Generation Parameters
|
||||
|
||||
### Thinking Mode (Recommended for Logic)
|
||||
Use when solving math, coding, or logical problems.
|
||||
|
||||
| Parameter | Value |
|
||||
|----------------|-------|
|
||||
| Temperature | 0.6 |
|
||||
| Top-P | 0.95 |
|
||||
| Top-K | 20 |
|
||||
| Min-P | 0.0 |
|
||||
| Repeat Penalty | 1.1 |
|
||||
|
||||
> ❗ DO NOT use greedy decoding — it causes infinite loops.
|
||||
|
||||
Enable via:
|
||||
- `enable_thinking=True` in tokenizer
|
||||
- Or add `/think` in user input during conversation
|
||||
|
||||
### Non-Thinking Mode (Fast Dialogue)
|
||||
For casual chat and quick replies.
|
||||
|
||||
| Parameter | Value |
|
||||
|----------------|-------|
|
||||
| Temperature | 0.7 |
|
||||
| Top-P | 0.8 |
|
||||
| Top-K | 20 |
|
||||
| Min-P | 0.0 |
|
||||
| Repeat Penalty | 1.1 |
|
||||
|
||||
Enable via:
|
||||
- `enable_thinking=False`
|
||||
- Or add `/no_think` in prompt
|
||||
|
||||
Stop sequences: `<|im_end|>`, `<|im_start|>`
|
||||
|
||||
## 💡 Usage Tips
|
||||
|
||||
> This model supports two operational modes:
|
||||
>
|
||||
> ### 🔍 Thinking Mode (Recommended for Logic)
|
||||
> Activate with `enable_thinking=True` or append `/think` in prompt.
|
||||
>
|
||||
> - Ideal for: math, coding, planning, analysis
|
||||
> - Use sampling: `temp=0.6`, `top_p=0.95`, `top_k=20`
|
||||
> - Avoid greedy decoding
|
||||
>
|
||||
> ### ⚡ Non-Thinking Mode (Fast Chat)
|
||||
> Use `enable_thinking=False` or `/no_think`.
|
||||
>
|
||||
> - Best for: casual conversation, quick answers
|
||||
> - Sampling: `temp=0.7`, `top_p=0.8`
|
||||
>
|
||||
> ---
|
||||
>
|
||||
> 🔄 **Switch Dynamically**
|
||||
> In multi-turn chats, the last `/think` or `/no_think` directive takes precedence.
|
||||
>
|
||||
> 🔁 **Avoid Repetition**
|
||||
> Set `presence_penalty=1.5` if stuck in loops.
|
||||
>
|
||||
> 📏 **Use Full Context**
|
||||
> Allow up to 32,768 output tokens for complex tasks.
|
||||
>
|
||||
> 🧰 **Agent Ready**
|
||||
> Works with Qwen-Agent, MCP servers, and custom tools.
|
||||
|
||||
## Customisation & Troubleshooting
|
||||
|
||||
Importing directly into Ollama should work, but you might encounter this error: `Error: invalid character '<' looking for beginning of value`.
|
||||
In this case try these steps:
|
||||
|
||||
1. `wget https://huggingface.co/geoffmunn/Qwen3-8B-f16/resolve/main/Qwen3-8B-f16%3AQ3_K_M.gguf`
|
||||
2. `nano Modelfile` and enter these details:
|
||||
```text
|
||||
FROM ./Qwen3-8B-f16:Q3_K_M.gguf
|
||||
|
||||
# Chat template using ChatML (used by Qwen)
|
||||
SYSTEM You are a helpful assistant
|
||||
|
||||
TEMPLATE "{{ if .System }}<|im_start|>system
|
||||
{{ .System }}<|im_end|>{{ end }}<|im_start|>user
|
||||
{{ .Prompt }}<|im_end|>
|
||||
<|im_start|>assistant
|
||||
"
|
||||
PARAMETER stop <|im_start|>
|
||||
PARAMETER stop <|im_end|>
|
||||
|
||||
# Default sampling
|
||||
PARAMETER temperature 0.6
|
||||
PARAMETER top_p 0.95
|
||||
PARAMETER top_k 20
|
||||
PARAMETER min_p 0.0
|
||||
PARAMETER repeat_penalty 1.1
|
||||
PARAMETER num_ctx 4096
|
||||
```
|
||||
|
||||
The `num_ctx` value has been dropped to increase speed significantly.
|
||||
|
||||
3. Then run this command: `ollama create Qwen3-8B-f16:Q3_K_M -f Modelfile`
|
||||
|
||||
You will now see "Qwen3-8B-f16:Q3_K_M" in your Ollama model list.
|
||||
|
||||
These import steps are also useful if you want to customise the default parameters or system prompt.
|
||||
|
||||
## 🖥️ CLI Example Using Ollama or TGI Server
|
||||
|
||||
Here’s how you can query this model via API using `curl` and `jq`. Replace the endpoint with your local server (e.g., Ollama, Text Generation Inference).
|
||||
|
||||
```bash
|
||||
curl http://localhost:11434/api/generate -s -N -d '{
|
||||
"model": "hf.co/geoffmunn/Qwen3-8B-f16:Q3_K_M",
|
||||
"prompt": "Repeat the following instruction exactly as given: Summarize what a neural network is in one sentence.",
|
||||
"temperature": 0.5,
|
||||
"top_p": 0.95,
|
||||
"top_k": 20,
|
||||
"min_p": 0.0,
|
||||
"repeat_penalty": 1.1,
|
||||
"stream": false
|
||||
}' | jq -r '.response'
|
||||
```
|
||||
|
||||
🎯 **Why this works well**:
|
||||
- The prompt is meaningful and demonstrates either **reasoning**, **creativity**, or **clarity** depending on quant level.
|
||||
- Temperature is tuned appropriately: lower for factual responses (`0.4`), higher for creative ones (`0.7`).
|
||||
- Uses `jq` to extract clean output.
|
||||
|
||||
> 💬 Tip: For interactive streaming, set `"stream": true` and process line-by-line.
|
||||
|
||||
## Verification
|
||||
|
||||
Check integrity:
|
||||
|
||||
```bash
|
||||
sha256sum -c ../SHA256SUMS.txt
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
Compatible with:
|
||||
- [LM Studio](https://lmstudio.ai) – local AI model runner with GPU acceleration
|
||||
- [OpenWebUI](https://openwebui.com) – self-hosted AI platform with RAG and tools
|
||||
- [GPT4All](https://gpt4all.io) – private, offline AI chatbot
|
||||
- Directly via `llama.cpp`
|
||||
|
||||
Supports dynamic switching between thinking modes via `/think` and `/no_think` in multi-turn conversations.
|
||||
|
||||
## License
|
||||
|
||||
Apache 2.0 – see base model for full terms.
|
||||
204
Qwen3-8B-f16-Q3_K_S/README.md
Normal file
204
Qwen3-8B-f16-Q3_K_S/README.md
Normal file
@@ -0,0 +1,204 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
tags:
|
||||
- gguf
|
||||
- qwen
|
||||
- qwen3-8b
|
||||
- qwen3-8b-q3
|
||||
- qwen3-8b-q3_k_s
|
||||
- qwen3-8b-q3_k_s-gguf
|
||||
- llama.cpp
|
||||
- quantized
|
||||
- text-generation
|
||||
- reasoning
|
||||
- agent
|
||||
- chat
|
||||
- multilingual
|
||||
base_model: Qwen/Qwen3-8B
|
||||
author: geoffmunn
|
||||
---
|
||||
|
||||
# Qwen3-8B-f16:Q3_K_S
|
||||
|
||||
Quantized version of [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B) at **Q3_K_S** level, derived from **f16** base weights.
|
||||
|
||||
## Model Info
|
||||
|
||||
- **Format**: GGUF (for llama.cpp and compatible runtimes)
|
||||
- **Size**: 3.77 GB
|
||||
- **Precision**: Q3_K_S
|
||||
- **Base Model**: [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B)
|
||||
- **Conversion Tool**: [llama.cpp](https://github.com/ggerganov/llama.cpp)
|
||||
|
||||
## Quality & Performance
|
||||
|
||||
| Metric | Value |
|
||||
|--------------------|---------------------------------------------------------------------------------------|
|
||||
| **Speed** | ⚡ Fast |
|
||||
| **RAM Required** | ~3.4 GB |
|
||||
| **Recommendation** | 🥉 Came first and second in questions covering both ends of the temperature spectrum. |
|
||||
|
||||
## Prompt Template (ChatML)
|
||||
|
||||
This model uses the **ChatML** format used by Qwen:
|
||||
|
||||
```text
|
||||
<|im_start|>system
|
||||
You are a helpful assistant.<|im_end|>
|
||||
<|im_start|>user
|
||||
{prompt}<|im_end|>
|
||||
<|im_start|>assistant
|
||||
```
|
||||
|
||||
Set this in your app (LM Studio, OpenWebUI, etc.) for best results.
|
||||
|
||||
## Generation Parameters
|
||||
|
||||
### Thinking Mode (Recommended for Logic)
|
||||
Use when solving math, coding, or logical problems.
|
||||
|
||||
| Parameter | Value |
|
||||
|----------------|-------|
|
||||
| Temperature | 0.6 |
|
||||
| Top-P | 0.95 |
|
||||
| Top-K | 20 |
|
||||
| Min-P | 0.0 |
|
||||
| Repeat Penalty | 1.1 |
|
||||
|
||||
> ❗ DO NOT use greedy decoding — it causes infinite loops.
|
||||
|
||||
Enable via:
|
||||
- `enable_thinking=True` in tokenizer
|
||||
- Or add `/think` in user input during conversation
|
||||
|
||||
### Non-Thinking Mode (Fast Dialogue)
|
||||
For casual chat and quick replies.
|
||||
|
||||
| Parameter | Value |
|
||||
|----------------|-------|
|
||||
| Temperature | 0.7 |
|
||||
| Top-P | 0.8 |
|
||||
| Top-K | 20 |
|
||||
| Min-P | 0.0 |
|
||||
| Repeat Penalty | 1.1 |
|
||||
|
||||
Enable via:
|
||||
- `enable_thinking=False`
|
||||
- Or add `/no_think` in prompt
|
||||
|
||||
Stop sequences: `<|im_end|>`, `<|im_start|>`
|
||||
|
||||
## 💡 Usage Tips
|
||||
|
||||
> This model supports two operational modes:
|
||||
>
|
||||
> ### 🔍 Thinking Mode (Recommended for Logic)
|
||||
> Activate with `enable_thinking=True` or append `/think` in prompt.
|
||||
>
|
||||
> - Ideal for: math, coding, planning, analysis
|
||||
> - Use sampling: `temp=0.6`, `top_p=0.95`, `top_k=20`
|
||||
> - Avoid greedy decoding
|
||||
>
|
||||
> ### ⚡ Non-Thinking Mode (Fast Chat)
|
||||
> Use `enable_thinking=False` or `/no_think`.
|
||||
>
|
||||
> - Best for: casual conversation, quick answers
|
||||
> - Sampling: `temp=0.7`, `top_p=0.8`
|
||||
>
|
||||
> ---
|
||||
>
|
||||
> 🔄 **Switch Dynamically**
|
||||
> In multi-turn chats, the last `/think` or `/no_think` directive takes precedence.
|
||||
>
|
||||
> 🔁 **Avoid Repetition**
|
||||
> Set `presence_penalty=1.5` if stuck in loops.
|
||||
>
|
||||
> 📏 **Use Full Context**
|
||||
> Allow up to 32,768 output tokens for complex tasks.
|
||||
>
|
||||
> 🧰 **Agent Ready**
|
||||
> Works with Qwen-Agent, MCP servers, and custom tools.
|
||||
|
||||
## Customisation & Troubleshooting
|
||||
|
||||
Importing directly into Ollama should work, but you might encounter this error: `Error: invalid character '<' looking for beginning of value`.
|
||||
In this case try these steps:
|
||||
|
||||
1. `wget https://huggingface.co/geoffmunn/Qwen3-8B-f16/resolve/main/Qwen3-8B-f16%3AQ3_K_S.gguf`
|
||||
2. `nano Modelfile` and enter these details:
|
||||
```text
|
||||
FROM ./Qwen3-8B-f16:Q3_K_S.gguf
|
||||
|
||||
# Chat template using ChatML (used by Qwen)
|
||||
SYSTEM You are a helpful assistant
|
||||
|
||||
TEMPLATE "{{ if .System }}<|im_start|>system
|
||||
{{ .System }}<|im_end|>{{ end }}<|im_start|>user
|
||||
{{ .Prompt }}<|im_end|>
|
||||
<|im_start|>assistant
|
||||
"
|
||||
PARAMETER stop <|im_start|>
|
||||
PARAMETER stop <|im_end|>
|
||||
|
||||
# Default sampling
|
||||
PARAMETER temperature 0.6
|
||||
PARAMETER top_p 0.95
|
||||
PARAMETER top_k 20
|
||||
PARAMETER min_p 0.0
|
||||
PARAMETER repeat_penalty 1.1
|
||||
PARAMETER num_ctx 4096
|
||||
```
|
||||
|
||||
The `num_ctx` value has been dropped to increase speed significantly.
|
||||
|
||||
3. Then run this command: `ollama create Qwen3-8B-f16:Q3_K_S -f Modelfile`
|
||||
|
||||
You will now see "Qwen3-8B-f16:Q3_K_S" in your Ollama model list.
|
||||
|
||||
These import steps are also useful if you want to customise the default parameters or system prompt.
|
||||
|
||||
## 🖥️ CLI Example Using Ollama or TGI Server
|
||||
|
||||
Here’s how you can query this model via API using `curl` and `jq`. Replace the endpoint with your local server (e.g., Ollama, Text Generation Inference).
|
||||
|
||||
```bash
|
||||
curl http://localhost:11434/api/generate -s -N -d '{
|
||||
"model": "hf.co/geoffmunn/Qwen3-8B-f16:Q3_K_S",
|
||||
"prompt": "Repeat the following instruction exactly as given: Summarize what a neural network is in one sentence.",
|
||||
"temperature": 0.5,
|
||||
"top_p": 0.95,
|
||||
"top_k": 20,
|
||||
"min_p": 0.0,
|
||||
"repeat_penalty": 1.1,
|
||||
"stream": false
|
||||
}' | jq -r '.response'
|
||||
```
|
||||
|
||||
🎯 **Why this works well**:
|
||||
- The prompt is meaningful and demonstrates either **reasoning**, **creativity**, or **clarity** depending on quant level.
|
||||
- Temperature is tuned appropriately: lower for factual responses (`0.4`), higher for creative ones (`0.7`).
|
||||
- Uses `jq` to extract clean output.
|
||||
|
||||
> 💬 Tip: For interactive streaming, set `"stream": true` and process line-by-line.
|
||||
|
||||
## Verification
|
||||
|
||||
Check integrity:
|
||||
|
||||
```bash
|
||||
sha256sum -c ../SHA256SUMS.txt
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
Compatible with:
|
||||
- [LM Studio](https://lmstudio.ai) – local AI model runner with GPU acceleration
|
||||
- [OpenWebUI](https://openwebui.com) – self-hosted AI platform with RAG and tools
|
||||
- [GPT4All](https://gpt4all.io) – private, offline AI chatbot
|
||||
- Directly via `llama.cpp`
|
||||
|
||||
Supports dynamic switching between thinking modes via `/think` and `/no_think` in multi-turn conversations.
|
||||
|
||||
## License
|
||||
|
||||
Apache 2.0 – see base model for full terms.
|
||||
206
Qwen3-8B-f16-Q4_K_M/README.md
Normal file
206
Qwen3-8B-f16-Q4_K_M/README.md
Normal file
@@ -0,0 +1,206 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
tags:
|
||||
- gguf
|
||||
- qwen
|
||||
- qwen3-8b
|
||||
- qwen3-8b-q4
|
||||
- qwen3-8b-q4_k_m
|
||||
- qwen3-8b-q4_k_m-gguf
|
||||
- llama.cpp
|
||||
- quantized
|
||||
- text-generation
|
||||
- reasoning
|
||||
- agent
|
||||
- chat
|
||||
- multilingual
|
||||
- imatrix
|
||||
- q3_hifi
|
||||
base_model: Qwen/Qwen3-8B
|
||||
author: geoffmunn
|
||||
---
|
||||
|
||||
# Qwen3-8B-f16:Q4_K_M
|
||||
|
||||
Quantized version of [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B) at **Q4_K_M** level, derived from **f16** base weights.
|
||||
|
||||
## Model Info
|
||||
|
||||
- **Format**: GGUF (for llama.cpp and compatible runtimes)
|
||||
- **Size**: 5.85 GB
|
||||
- **Precision**: Q4_K_M
|
||||
- **Base Model**: [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B)
|
||||
- **Conversion Tool**: [llama.cpp](https://github.com/ggerganov/llama.cpp)
|
||||
|
||||
## Quality & Performance
|
||||
|
||||
| Metric | Value |
|
||||
|--------------------|-------------------------------------------------------------------------|
|
||||
| **Speed** | 🚀 Fast |
|
||||
| **RAM Required** | ~4.3 GB |
|
||||
| **Recommendation** | Came first and second in questions covering high temperature questions. |
|
||||
|
||||
## Prompt Template (ChatML)
|
||||
|
||||
This model uses the **ChatML** format used by Qwen:
|
||||
|
||||
```text
|
||||
<|im_start|>system
|
||||
You are a helpful assistant.<|im_end|>
|
||||
<|im_start|>user
|
||||
{prompt}<|im_end|>
|
||||
<|im_start|>assistant
|
||||
```
|
||||
|
||||
Set this in your app (LM Studio, OpenWebUI, etc.) for best results.
|
||||
|
||||
## Generation Parameters
|
||||
|
||||
### Thinking Mode (Recommended for Logic)
|
||||
Use when solving math, coding, or logical problems.
|
||||
|
||||
| Parameter | Value |
|
||||
|----------------|-------|
|
||||
| Temperature | 0.6 |
|
||||
| Top-P | 0.95 |
|
||||
| Top-K | 20 |
|
||||
| Min-P | 0.0 |
|
||||
| Repeat Penalty | 1.1 |
|
||||
|
||||
> ❗ DO NOT use greedy decoding — it causes infinite loops.
|
||||
|
||||
Enable via:
|
||||
- `enable_thinking=True` in tokenizer
|
||||
- Or add `/think` in user input during conversation
|
||||
|
||||
### Non-Thinking Mode (Fast Dialogue)
|
||||
For casual chat and quick replies.
|
||||
|
||||
| Parameter | Value |
|
||||
|----------------|-------|
|
||||
| Temperature | 0.7 |
|
||||
| Top-P | 0.8 |
|
||||
| Top-K | 20 |
|
||||
| Min-P | 0.0 |
|
||||
| Repeat Penalty | 1.1 |
|
||||
|
||||
Enable via:
|
||||
- `enable_thinking=False`
|
||||
- Or add `/no_think` in prompt
|
||||
|
||||
Stop sequences: `<|im_end|>`, `<|im_start|>`
|
||||
|
||||
## 💡 Usage Tips
|
||||
|
||||
> This model supports two operational modes:
|
||||
>
|
||||
> ### 🔍 Thinking Mode (Recommended for Logic)
|
||||
> Activate with `enable_thinking=True` or append `/think` in prompt.
|
||||
>
|
||||
> - Ideal for: math, coding, planning, analysis
|
||||
> - Use sampling: `temp=0.6`, `top_p=0.95`, `top_k=20`
|
||||
> - Avoid greedy decoding
|
||||
>
|
||||
> ### ⚡ Non-Thinking Mode (Fast Chat)
|
||||
> Use `enable_thinking=False` or `/no_think`.
|
||||
>
|
||||
> - Best for: casual conversation, quick answers
|
||||
> - Sampling: `temp=0.7`, `top_p=0.8`
|
||||
>
|
||||
> ---
|
||||
>
|
||||
> 🔄 **Switch Dynamically**
|
||||
> In multi-turn chats, the last `/think` or `/no_think` directive takes precedence.
|
||||
>
|
||||
> 🔁 **Avoid Repetition**
|
||||
> Set `presence_penalty=1.5` if stuck in loops.
|
||||
>
|
||||
> 📏 **Use Full Context**
|
||||
> Allow up to 32,768 output tokens for complex tasks.
|
||||
>
|
||||
> 🧰 **Agent Ready**
|
||||
> Works with Qwen-Agent, MCP servers, and custom tools.
|
||||
|
||||
## Customisation & Troubleshooting
|
||||
|
||||
Importing directly into Ollama should work, but you might encounter this error: `Error: invalid character '<' looking for beginning of value`.
|
||||
In this case try these steps:
|
||||
|
||||
1. `wget https://huggingface.co/geoffmunn/Qwen3-8B-f16/resolve/main/Qwen3-8B-f16%3AQ4_K_M.gguf`
|
||||
2. `nano Modelfile` and enter these details:
|
||||
```text
|
||||
FROM ./Qwen3-8B-f16:Q4_K_M.gguf
|
||||
|
||||
# Chat template using ChatML (used by Qwen)
|
||||
SYSTEM You are a helpful assistant
|
||||
|
||||
TEMPLATE "{{ if .System }}<|im_start|>system
|
||||
{{ .System }}<|im_end|>{{ end }}<|im_start|>user
|
||||
{{ .Prompt }}<|im_end|>
|
||||
<|im_start|>assistant
|
||||
"
|
||||
PARAMETER stop <|im_start|>
|
||||
PARAMETER stop <|im_end|>
|
||||
|
||||
# Default sampling
|
||||
PARAMETER temperature 0.6
|
||||
PARAMETER top_p 0.95
|
||||
PARAMETER top_k 20
|
||||
PARAMETER min_p 0.0
|
||||
PARAMETER repeat_penalty 1.1
|
||||
PARAMETER num_ctx 4096
|
||||
```
|
||||
|
||||
The `num_ctx` value has been dropped to increase speed significantly.
|
||||
|
||||
3. Then run this command: `ollama create Qwen3-8B-f16:Q4_K_M -f Modelfile`
|
||||
|
||||
You will now see "Qwen3-8B-f16:Q4_K_M" in your Ollama model list.
|
||||
|
||||
These import steps are also useful if you want to customise the default parameters or system prompt.
|
||||
|
||||
## 🖥️ CLI Example Using Ollama or TGI Server
|
||||
|
||||
Here’s how you can query this model via API using `curl` and `jq`. Replace the endpoint with your local server (e.g., Ollama, Text Generation Inference).
|
||||
|
||||
```bash
|
||||
curl http://localhost:11434/api/generate -s -N -d '{
|
||||
"model": "hf.co/geoffmunn/Qwen3-8B-f16:Q4_K_M",
|
||||
"prompt": "Repeat the following instruction exactly as given: Write a short haiku about autumn leaves falling gently in a quiet forest.",
|
||||
"temperature": 0.7,
|
||||
"top_p": 0.95,
|
||||
"top_k": 20,
|
||||
"min_p": 0.0,
|
||||
"repeat_penalty": 1.1,
|
||||
"stream": false
|
||||
}' | jq -r '.response'
|
||||
```
|
||||
|
||||
🎯 **Why this works well**:
|
||||
- The prompt is meaningful and demonstrates either **reasoning**, **creativity**, or **clarity** depending on quant level.
|
||||
- Temperature is tuned appropriately: lower for factual responses (`0.4`), higher for creative ones (`0.7`).
|
||||
- Uses `jq` to extract clean output.
|
||||
|
||||
> 💬 Tip: For interactive streaming, set `"stream": true` and process line-by-line.
|
||||
|
||||
## Verification
|
||||
|
||||
Check integrity:
|
||||
|
||||
```bash
|
||||
sha256sum -c ../SHA256SUMS.txt
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
Compatible with:
|
||||
- [LM Studio](https://lmstudio.ai) – local AI model runner with GPU acceleration
|
||||
- [OpenWebUI](https://openwebui.com) – self-hosted AI platform with RAG and tools
|
||||
- [GPT4All](https://gpt4all.io) – private, offline AI chatbot
|
||||
- Directly via `llama.cpp`
|
||||
|
||||
Supports dynamic switching between thinking modes via `/think` and `/no_think` in multi-turn conversations.
|
||||
|
||||
## License
|
||||
|
||||
Apache 2.0 – see base model for full terms.
|
||||
206
Qwen3-8B-f16-Q4_K_S/README.md
Normal file
206
Qwen3-8B-f16-Q4_K_S/README.md
Normal file
@@ -0,0 +1,206 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
tags:
|
||||
- gguf
|
||||
- qwen
|
||||
- qwen3-8b
|
||||
- qwen3-8b-q4
|
||||
- qwen3-8b-q4_k_s
|
||||
- qwen3-8b-q4_k_s-gguf
|
||||
- llama.cpp
|
||||
- quantized
|
||||
- text-generation
|
||||
- reasoning
|
||||
- agent
|
||||
- chat
|
||||
- multilingual
|
||||
- imatrix
|
||||
- q3_hifi
|
||||
base_model: Qwen/Qwen3-8B
|
||||
author: geoffmunn
|
||||
---
|
||||
|
||||
# Qwen3-8B-f16:Q4_K_S
|
||||
|
||||
Quantized version of [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B) at **Q4_K_S** level, derived from **f16** base weights.
|
||||
|
||||
## Model Info
|
||||
|
||||
- **Format**: GGUF (for llama.cpp and compatible runtimes)
|
||||
- **Size**: 4.8 GB
|
||||
- **Precision**: Q4_K_S
|
||||
- **Base Model**: [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B)
|
||||
- **Conversion Tool**: [llama.cpp](https://github.com/ggerganov/llama.cpp)
|
||||
|
||||
## Quality & Performance
|
||||
|
||||
| Metric | Value |
|
||||
|--------------------|---------------------------------------------------------------------------------------|
|
||||
| **Speed** | 🚀 Fast |
|
||||
| **RAM Required** | ~4.1 GB |
|
||||
| **Recommendation** | 🥉 Came first and second in questions covering both ends of the temperature spectrum. |
|
||||
|
||||
## Prompt Template (ChatML)
|
||||
|
||||
This model uses the **ChatML** format used by Qwen:
|
||||
|
||||
```text
|
||||
<|im_start|>system
|
||||
You are a helpful assistant.<|im_end|>
|
||||
<|im_start|>user
|
||||
{prompt}<|im_end|>
|
||||
<|im_start|>assistant
|
||||
```
|
||||
|
||||
Set this in your app (LM Studio, OpenWebUI, etc.) for best results.
|
||||
|
||||
## Generation Parameters
|
||||
|
||||
### Thinking Mode (Recommended for Logic)
|
||||
Use when solving math, coding, or logical problems.
|
||||
|
||||
| Parameter | Value |
|
||||
|----------------|-------|
|
||||
| Temperature | 0.6 |
|
||||
| Top-P | 0.95 |
|
||||
| Top-K | 20 |
|
||||
| Min-P | 0.0 |
|
||||
| Repeat Penalty | 1.1 |
|
||||
|
||||
> ❗ DO NOT use greedy decoding — it causes infinite loops.
|
||||
|
||||
Enable via:
|
||||
- `enable_thinking=True` in tokenizer
|
||||
- Or add `/think` in user input during conversation
|
||||
|
||||
### Non-Thinking Mode (Fast Dialogue)
|
||||
For casual chat and quick replies.
|
||||
|
||||
| Parameter | Value |
|
||||
|----------------|-------|
|
||||
| Temperature | 0.7 |
|
||||
| Top-P | 0.8 |
|
||||
| Top-K | 20 |
|
||||
| Min-P | 0.0 |
|
||||
| Repeat Penalty | 1.1 |
|
||||
|
||||
Enable via:
|
||||
- `enable_thinking=False`
|
||||
- Or add `/no_think` in prompt
|
||||
|
||||
Stop sequences: `<|im_end|>`, `<|im_start|>`
|
||||
|
||||
## 💡 Usage Tips
|
||||
|
||||
> This model supports two operational modes:
|
||||
>
|
||||
> ### 🔍 Thinking Mode (Recommended for Logic)
|
||||
> Activate with `enable_thinking=True` or append `/think` in prompt.
|
||||
>
|
||||
> - Ideal for: math, coding, planning, analysis
|
||||
> - Use sampling: `temp=0.6`, `top_p=0.95`, `top_k=20`
|
||||
> - Avoid greedy decoding
|
||||
>
|
||||
> ### ⚡ Non-Thinking Mode (Fast Chat)
|
||||
> Use `enable_thinking=False` or `/no_think`.
|
||||
>
|
||||
> - Best for: casual conversation, quick answers
|
||||
> - Sampling: `temp=0.7`, `top_p=0.8`
|
||||
>
|
||||
> ---
|
||||
>
|
||||
> 🔄 **Switch Dynamically**
|
||||
> In multi-turn chats, the last `/think` or `/no_think` directive takes precedence.
|
||||
>
|
||||
> 🔁 **Avoid Repetition**
|
||||
> Set `presence_penalty=1.5` if stuck in loops.
|
||||
>
|
||||
> 📏 **Use Full Context**
|
||||
> Allow up to 32,768 output tokens for complex tasks.
|
||||
>
|
||||
> 🧰 **Agent Ready**
|
||||
> Works with Qwen-Agent, MCP servers, and custom tools.
|
||||
|
||||
## Customisation & Troubleshooting
|
||||
|
||||
Importing directly into Ollama should work, but you might encounter this error: `Error: invalid character '<' looking for beginning of value`.
|
||||
In this case try these steps:
|
||||
|
||||
1. `wget https://huggingface.co/geoffmunn/Qwen3-8B-f16/resolve/main/Qwen3-8B-f16%3AQ4_K_S.gguf`
|
||||
2. `nano Modelfile` and enter these details:
|
||||
```text
|
||||
FROM ./Qwen3-8B-f16:Q4_K_S.gguf
|
||||
|
||||
# Chat template using ChatML (used by Qwen)
|
||||
SYSTEM You are a helpful assistant
|
||||
|
||||
TEMPLATE "{{ if .System }}<|im_start|>system
|
||||
{{ .System }}<|im_end|>{{ end }}<|im_start|>user
|
||||
{{ .Prompt }}<|im_end|>
|
||||
<|im_start|>assistant
|
||||
"
|
||||
PARAMETER stop <|im_start|>
|
||||
PARAMETER stop <|im_end|>
|
||||
|
||||
# Default sampling
|
||||
PARAMETER temperature 0.6
|
||||
PARAMETER top_p 0.95
|
||||
PARAMETER top_k 20
|
||||
PARAMETER min_p 0.0
|
||||
PARAMETER repeat_penalty 1.1
|
||||
PARAMETER num_ctx 4096
|
||||
```
|
||||
|
||||
The `num_ctx` value has been dropped to increase speed significantly.
|
||||
|
||||
3. Then run this command: `ollama create Qwen3-8B-f16:Q4_K_S -f Modelfile`
|
||||
|
||||
You will now see "Qwen3-8B-f16:Q4_K_S" in your Ollama model list.
|
||||
|
||||
These import steps are also useful if you want to customise the default parameters or system prompt.
|
||||
|
||||
## 🖥️ CLI Example Using Ollama or TGI Server
|
||||
|
||||
Here’s how you can query this model via API using `curl` and `jq`. Replace the endpoint with your local server (e.g., Ollama, Text Generation Inference).
|
||||
|
||||
```bash
|
||||
curl http://localhost:11434/api/generate -s -N -d '{
|
||||
"model": "hf.co/geoffmunn/Qwen3-8B-f16:Q4_K_S",
|
||||
"prompt": "Repeat the following instruction exactly as given: Summarize what a neural network is in one sentence.",
|
||||
"temperature": 0.5,
|
||||
"top_p": 0.95,
|
||||
"top_k": 20,
|
||||
"min_p": 0.0,
|
||||
"repeat_penalty": 1.1,
|
||||
"stream": false
|
||||
}' | jq -r '.response'
|
||||
```
|
||||
|
||||
🎯 **Why this works well**:
|
||||
- The prompt is meaningful and demonstrates either **reasoning**, **creativity**, or **clarity** depending on quant level.
|
||||
- Temperature is tuned appropriately: lower for factual responses (`0.4`), higher for creative ones (`0.7`).
|
||||
- Uses `jq` to extract clean output.
|
||||
|
||||
> 💬 Tip: For interactive streaming, set `"stream": true` and process line-by-line.
|
||||
|
||||
## Verification
|
||||
|
||||
Check integrity:
|
||||
|
||||
```bash
|
||||
sha256sum -c ../SHA256SUMS.txt
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
Compatible with:
|
||||
- [LM Studio](https://lmstudio.ai) – local AI model runner with GPU acceleration
|
||||
- [OpenWebUI](https://openwebui.com) – self-hosted AI platform with RAG and tools
|
||||
- [GPT4All](https://gpt4all.io) – private, offline AI chatbot
|
||||
- Directly via `llama.cpp`
|
||||
|
||||
Supports dynamic switching between thinking modes via `/think` and `/no_think` in multi-turn conversations.
|
||||
|
||||
## License
|
||||
|
||||
Apache 2.0 – see base model for full terms.
|
||||
206
Qwen3-8B-f16-Q5_K_M/README.md
Normal file
206
Qwen3-8B-f16-Q5_K_M/README.md
Normal file
@@ -0,0 +1,206 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
tags:
|
||||
- gguf
|
||||
- qwen
|
||||
- qwen3-8b
|
||||
- qwen3-8b-q5
|
||||
- qwen3-8b-q5_k_m
|
||||
- qwen3-8b-q5_k_m-gguf
|
||||
- llama.cpp
|
||||
- quantized
|
||||
- text-generation
|
||||
- reasoning
|
||||
- agent
|
||||
- chat
|
||||
- multilingual
|
||||
- imatrix
|
||||
- q3_hifi
|
||||
base_model: Qwen/Qwen3-8B
|
||||
author: geoffmunn
|
||||
---
|
||||
|
||||
# Qwen3-8B-f16:Q5_K_M
|
||||
|
||||
Quantized version of [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B) at **Q5_K_M** level, derived from **f16** base weights.
|
||||
|
||||
## Model Info
|
||||
|
||||
- **Format**: GGUF (for llama.cpp and compatible runtimes)
|
||||
- **Size**: 5.85 GB
|
||||
- **Precision**: Q5_K_M
|
||||
- **Base Model**: [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B)
|
||||
- **Conversion Tool**: [llama.cpp](https://github.com/ggerganov/llama.cpp)
|
||||
|
||||
## Quality & Performance
|
||||
|
||||
| Metric | Value |
|
||||
|--------------------|-----------------------------------------------------------------|
|
||||
| **Speed** | 🐢 Medium |
|
||||
| **RAM Required** | ~4.9 GB |
|
||||
| **Recommendation** | Not recommended, no appeareances in the top 3 for any question. |
|
||||
|
||||
## Prompt Template (ChatML)
|
||||
|
||||
This model uses the **ChatML** format used by Qwen:
|
||||
|
||||
```text
|
||||
<|im_start|>system
|
||||
You are a helpful assistant.<|im_end|>
|
||||
<|im_start|>user
|
||||
{prompt}<|im_end|>
|
||||
<|im_start|>assistant
|
||||
```
|
||||
|
||||
Set this in your app (LM Studio, OpenWebUI, etc.) for best results.
|
||||
|
||||
## Generation Parameters
|
||||
|
||||
### Thinking Mode (Recommended for Logic)
|
||||
Use when solving math, coding, or logical problems.
|
||||
|
||||
| Parameter | Value |
|
||||
|----------------|-------|
|
||||
| Temperature | 0.6 |
|
||||
| Top-P | 0.95 |
|
||||
| Top-K | 20 |
|
||||
| Min-P | 0.0 |
|
||||
| Repeat Penalty | 1.1 |
|
||||
|
||||
> ❗ DO NOT use greedy decoding — it causes infinite loops.
|
||||
|
||||
Enable via:
|
||||
- `enable_thinking=True` in tokenizer
|
||||
- Or add `/think` in user input during conversation
|
||||
|
||||
### Non-Thinking Mode (Fast Dialogue)
|
||||
For casual chat and quick replies.
|
||||
|
||||
| Parameter | Value |
|
||||
|----------------|-------|
|
||||
| Temperature | 0.7 |
|
||||
| Top-P | 0.8 |
|
||||
| Top-K | 20 |
|
||||
| Min-P | 0.0 |
|
||||
| Repeat Penalty | 1.1 |
|
||||
|
||||
Enable via:
|
||||
- `enable_thinking=False`
|
||||
- Or add `/no_think` in prompt
|
||||
|
||||
Stop sequences: `<|im_end|>`, `<|im_start|>`
|
||||
|
||||
## 💡 Usage Tips
|
||||
|
||||
> This model supports two operational modes:
|
||||
>
|
||||
> ### 🔍 Thinking Mode (Recommended for Logic)
|
||||
> Activate with `enable_thinking=True` or append `/think` in prompt.
|
||||
>
|
||||
> - Ideal for: math, coding, planning, analysis
|
||||
> - Use sampling: `temp=0.6`, `top_p=0.95`, `top_k=20`
|
||||
> - Avoid greedy decoding
|
||||
>
|
||||
> ### ⚡ Non-Thinking Mode (Fast Chat)
|
||||
> Use `enable_thinking=False` or `/no_think`.
|
||||
>
|
||||
> - Best for: casual conversation, quick answers
|
||||
> - Sampling: `temp=0.7`, `top_p=0.8`
|
||||
>
|
||||
> ---
|
||||
>
|
||||
> 🔄 **Switch Dynamically**
|
||||
> In multi-turn chats, the last `/think` or `/no_think` directive takes precedence.
|
||||
>
|
||||
> 🔁 **Avoid Repetition**
|
||||
> Set `presence_penalty=1.5` if stuck in loops.
|
||||
>
|
||||
> 📏 **Use Full Context**
|
||||
> Allow up to 32,768 output tokens for complex tasks.
|
||||
>
|
||||
> 🧰 **Agent Ready**
|
||||
> Works with Qwen-Agent, MCP servers, and custom tools.
|
||||
|
||||
## Customisation & Troubleshooting
|
||||
|
||||
Importing directly into Ollama should work, but you might encounter this error: `Error: invalid character '<' looking for beginning of value`.
|
||||
In this case try these steps:
|
||||
|
||||
1. `wget https://huggingface.co/geoffmunn/Qwen3-8B-f16/resolve/main/Qwen3-8B-f16%3AQ5_K_M.gguf`
|
||||
2. `nano Modelfile` and enter these details:
|
||||
```text
|
||||
FROM ./Qwen3-8B-f16:Q5_K_M.gguf
|
||||
|
||||
# Chat template using ChatML (used by Qwen)
|
||||
SYSTEM You are a helpful assistant
|
||||
|
||||
TEMPLATE "{{ if .System }}<|im_start|>system
|
||||
{{ .System }}<|im_end|>{{ end }}<|im_start|>user
|
||||
{{ .Prompt }}<|im_end|>
|
||||
<|im_start|>assistant
|
||||
"
|
||||
PARAMETER stop <|im_start|>
|
||||
PARAMETER stop <|im_end|>
|
||||
|
||||
# Default sampling
|
||||
PARAMETER temperature 0.6
|
||||
PARAMETER top_p 0.95
|
||||
PARAMETER top_k 20
|
||||
PARAMETER min_p 0.0
|
||||
PARAMETER repeat_penalty 1.1
|
||||
PARAMETER num_ctx 4096
|
||||
```
|
||||
|
||||
The `num_ctx` value has been dropped to increase speed significantly.
|
||||
|
||||
3. Then run this command: `ollama create Qwen3-8B-f16:Q5_K_M -f Modelfile`
|
||||
|
||||
You will now see "Qwen3-8B-f16:Q5_K_M" in your Ollama model list.
|
||||
|
||||
These import steps are also useful if you want to customise the default parameters or system prompt.
|
||||
|
||||
## 🖥️ CLI Example Using Ollama or TGI Server
|
||||
|
||||
Here’s how you can query this model via API using `curl` and `jq`. Replace the endpoint with your local server (e.g., Ollama, Text Generation Inference).
|
||||
|
||||
```bash
|
||||
curl http://localhost:11434/api/generate -s -N -d '{
|
||||
"model": "hf.co/geoffmunn/Qwen3-8B-f16:Q5_K_M",
|
||||
"prompt": "Repeat the following instruction exactly as given: Explain why the sky appears blue during the day but red at sunrise and sunset, using physics principles like Rayleigh scattering.",
|
||||
"temperature": 0.4,
|
||||
"top_p": 0.95,
|
||||
"top_k": 20,
|
||||
"min_p": 0.0,
|
||||
"repeat_penalty": 1.1,
|
||||
"stream": false
|
||||
}' | jq -r '.response'
|
||||
```
|
||||
|
||||
🎯 **Why this works well**:
|
||||
- The prompt is meaningful and demonstrates either **reasoning**, **creativity**, or **clarity** depending on quant level.
|
||||
- Temperature is tuned appropriately: lower for factual responses (`0.4`), higher for creative ones (`0.7`).
|
||||
- Uses `jq` to extract clean output.
|
||||
|
||||
> 💬 Tip: For interactive streaming, set `"stream": true` and process line-by-line.
|
||||
|
||||
## Verification
|
||||
|
||||
Check integrity:
|
||||
|
||||
```bash
|
||||
sha256sum -c ../SHA256SUMS.txt
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
Compatible with:
|
||||
- [LM Studio](https://lmstudio.ai) – local AI model runner with GPU acceleration
|
||||
- [OpenWebUI](https://openwebui.com) – self-hosted AI platform with RAG and tools
|
||||
- [GPT4All](https://gpt4all.io) – private, offline AI chatbot
|
||||
- Directly via `llama.cpp`
|
||||
|
||||
Supports dynamic switching between thinking modes via `/think` and `/no_think` in multi-turn conversations.
|
||||
|
||||
## License
|
||||
|
||||
Apache 2.0 – see base model for full terms.
|
||||
206
Qwen3-8B-f16-Q5_K_S/README.md
Normal file
206
Qwen3-8B-f16-Q5_K_S/README.md
Normal file
@@ -0,0 +1,206 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
tags:
|
||||
- gguf
|
||||
- qwen
|
||||
- qwen3-8b
|
||||
- qwen3-8b-q5
|
||||
- qwen3-8b-q5_k_s
|
||||
- qwen3-8b-q5_k_s-gguf
|
||||
- llama.cpp
|
||||
- quantized
|
||||
- text-generation
|
||||
- reasoning
|
||||
- agent
|
||||
- chat
|
||||
- multilingual
|
||||
- imatrix
|
||||
- q3_hifi
|
||||
base_model: Qwen/Qwen3-8B
|
||||
author: geoffmunn
|
||||
---
|
||||
|
||||
# Qwen3-8B-f16:Q5_K_S
|
||||
|
||||
Quantized version of [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B) at **Q5_K_S** level, derived from **f16** base weights.
|
||||
|
||||
## Model Info
|
||||
|
||||
- **Format**: GGUF (for llama.cpp and compatible runtimes)
|
||||
- **Size**: 5.72 GB
|
||||
- **Precision**: Q5_K_S
|
||||
- **Base Model**: [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B)
|
||||
- **Conversion Tool**: [llama.cpp](https://github.com/ggerganov/llama.cpp)
|
||||
|
||||
## Quality & Performance
|
||||
|
||||
| Metric | Value |
|
||||
|--------------------|---------------------------------------------------|
|
||||
| **Speed** | 🐢 Medium |
|
||||
| **RAM Required** | ~4.8 GB |
|
||||
| **Recommendation** | 🥈 A good second place. Good for all query types. |
|
||||
|
||||
## Prompt Template (ChatML)
|
||||
|
||||
This model uses the **ChatML** format used by Qwen:
|
||||
|
||||
```text
|
||||
<|im_start|>system
|
||||
You are a helpful assistant.<|im_end|>
|
||||
<|im_start|>user
|
||||
{prompt}<|im_end|>
|
||||
<|im_start|>assistant
|
||||
```
|
||||
|
||||
Set this in your app (LM Studio, OpenWebUI, etc.) for best results.
|
||||
|
||||
## Generation Parameters
|
||||
|
||||
### Thinking Mode (Recommended for Logic)
|
||||
Use when solving math, coding, or logical problems.
|
||||
|
||||
| Parameter | Value |
|
||||
|----------------|-------|
|
||||
| Temperature | 0.6 |
|
||||
| Top-P | 0.95 |
|
||||
| Top-K | 20 |
|
||||
| Min-P | 0.0 |
|
||||
| Repeat Penalty | 1.1 |
|
||||
|
||||
> ❗ DO NOT use greedy decoding — it causes infinite loops.
|
||||
|
||||
Enable via:
|
||||
- `enable_thinking=True` in tokenizer
|
||||
- Or add `/think` in user input during conversation
|
||||
|
||||
### Non-Thinking Mode (Fast Dialogue)
|
||||
For casual chat and quick replies.
|
||||
|
||||
| Parameter | Value |
|
||||
|----------------|-------|
|
||||
| Temperature | 0.7 |
|
||||
| Top-P | 0.8 |
|
||||
| Top-K | 20 |
|
||||
| Min-P | 0.0 |
|
||||
| Repeat Penalty | 1.1 |
|
||||
|
||||
Enable via:
|
||||
- `enable_thinking=False`
|
||||
- Or add `/no_think` in prompt
|
||||
|
||||
Stop sequences: `<|im_end|>`, `<|im_start|>`
|
||||
|
||||
## 💡 Usage Tips
|
||||
|
||||
> This model supports two operational modes:
|
||||
>
|
||||
> ### 🔍 Thinking Mode (Recommended for Logic)
|
||||
> Activate with `enable_thinking=True` or append `/think` in prompt.
|
||||
>
|
||||
> - Ideal for: math, coding, planning, analysis
|
||||
> - Use sampling: `temp=0.6`, `top_p=0.95`, `top_k=20`
|
||||
> - Avoid greedy decoding
|
||||
>
|
||||
> ### ⚡ Non-Thinking Mode (Fast Chat)
|
||||
> Use `enable_thinking=False` or `/no_think`.
|
||||
>
|
||||
> - Best for: casual conversation, quick answers
|
||||
> - Sampling: `temp=0.7`, `top_p=0.8`
|
||||
>
|
||||
> ---
|
||||
>
|
||||
> 🔄 **Switch Dynamically**
|
||||
> In multi-turn chats, the last `/think` or `/no_think` directive takes precedence.
|
||||
>
|
||||
> 🔁 **Avoid Repetition**
|
||||
> Set `presence_penalty=1.5` if stuck in loops.
|
||||
>
|
||||
> 📏 **Use Full Context**
|
||||
> Allow up to 32,768 output tokens for complex tasks.
|
||||
>
|
||||
> 🧰 **Agent Ready**
|
||||
> Works with Qwen-Agent, MCP servers, and custom tools.
|
||||
|
||||
## Customisation & Troubleshooting
|
||||
|
||||
Importing directly into Ollama should work, but you might encounter this error: `Error: invalid character '<' looking for beginning of value`.
|
||||
In this case try these steps:
|
||||
|
||||
1. `wget https://huggingface.co/geoffmunn/Qwen3-8B-f16/resolve/main/Qwen3-8B-f16%3AQ5_K_S.gguf`
|
||||
2. `nano Modelfile` and enter these details:
|
||||
```text
|
||||
FROM ./Qwen3-8B-f16:Q5_K_S.gguf
|
||||
|
||||
# Chat template using ChatML (used by Qwen)
|
||||
SYSTEM You are a helpful assistant
|
||||
|
||||
TEMPLATE "{{ if .System }}<|im_start|>system
|
||||
{{ .System }}<|im_end|>{{ end }}<|im_start|>user
|
||||
{{ .Prompt }}<|im_end|>
|
||||
<|im_start|>assistant
|
||||
"
|
||||
PARAMETER stop <|im_start|>
|
||||
PARAMETER stop <|im_end|>
|
||||
|
||||
# Default sampling
|
||||
PARAMETER temperature 0.6
|
||||
PARAMETER top_p 0.95
|
||||
PARAMETER top_k 20
|
||||
PARAMETER min_p 0.0
|
||||
PARAMETER repeat_penalty 1.1
|
||||
PARAMETER num_ctx 4096
|
||||
```
|
||||
|
||||
The `num_ctx` value has been dropped to increase speed significantly.
|
||||
|
||||
3. Then run this command: `ollama create Qwen3-8B-f16:Q5_K_S -f Modelfile`
|
||||
|
||||
You will now see "Qwen3-8B-f16:Q5_K_S" in your Ollama model list.
|
||||
|
||||
These import steps are also useful if you want to customise the default parameters or system prompt.
|
||||
|
||||
## 🖥️ CLI Example Using Ollama or TGI Server
|
||||
|
||||
Here’s how you can query this model via API using `curl` and `jq`. Replace the endpoint with your local server (e.g., Ollama, Text Generation Inference).
|
||||
|
||||
```bash
|
||||
curl http://localhost:11434/api/generate -s -N -d '{
|
||||
"model": "hf.co/geoffmunn/Qwen3-8B-f16:Q5_K_S",
|
||||
"prompt": "Repeat the following instruction exactly as given: Write a short haiku about autumn leaves falling gently in a quiet forest.",
|
||||
"temperature": 0.7,
|
||||
"top_p": 0.95,
|
||||
"top_k": 20,
|
||||
"min_p": 0.0,
|
||||
"repeat_penalty": 1.1,
|
||||
"stream": false
|
||||
}' | jq -r '.response'
|
||||
```
|
||||
|
||||
🎯 **Why this works well**:
|
||||
- The prompt is meaningful and demonstrates either **reasoning**, **creativity**, or **clarity** depending on quant level.
|
||||
- Temperature is tuned appropriately: lower for factual responses (`0.4`), higher for creative ones (`0.7`).
|
||||
- Uses `jq` to extract clean output.
|
||||
|
||||
> 💬 Tip: For interactive streaming, set `"stream": true` and process line-by-line.
|
||||
|
||||
## Verification
|
||||
|
||||
Check integrity:
|
||||
|
||||
```bash
|
||||
sha256sum -c ../SHA256SUMS.txt
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
Compatible with:
|
||||
- [LM Studio](https://lmstudio.ai) – local AI model runner with GPU acceleration
|
||||
- [OpenWebUI](https://openwebui.com) – self-hosted AI platform with RAG and tools
|
||||
- [GPT4All](https://gpt4all.io) – private, offline AI chatbot
|
||||
- Directly via `llama.cpp`
|
||||
|
||||
Supports dynamic switching between thinking modes via `/think` and `/no_think` in multi-turn conversations.
|
||||
|
||||
## License
|
||||
|
||||
Apache 2.0 – see base model for full terms.
|
||||
206
Qwen3-8B-f16-Q6_K/README.md
Normal file
206
Qwen3-8B-f16-Q6_K/README.md
Normal file
@@ -0,0 +1,206 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
tags:
|
||||
- gguf
|
||||
- qwen
|
||||
- qwen3-8b
|
||||
- qwen3-8b-q6
|
||||
- qwen3-8b-q6_k
|
||||
- qwen3-8b-q6_k-gguf
|
||||
- llama.cpp
|
||||
- quantized
|
||||
- text-generation
|
||||
- reasoning
|
||||
- agent
|
||||
- chat
|
||||
- multilingual
|
||||
- imatrix
|
||||
- q3_hifi
|
||||
base_model: Qwen/Qwen3-8B
|
||||
author: geoffmunn
|
||||
---
|
||||
|
||||
# Qwen3-8B-f16:Q6_K
|
||||
|
||||
Quantized version of [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B) at **Q6_K** level, derived from **f16** base weights.
|
||||
|
||||
## Model Info
|
||||
|
||||
- **Format**: GGUF (for llama.cpp and compatible runtimes)
|
||||
- **Size**: 6.73 GB
|
||||
- **Precision**: Q6_K
|
||||
- **Base Model**: [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B)
|
||||
- **Conversion Tool**: [llama.cpp](https://github.com/ggerganov/llama.cpp)
|
||||
|
||||
## Quality & Performance
|
||||
|
||||
| Metric | Value |
|
||||
|--------------------|---------------------------------------------------|
|
||||
| **Speed** | 🐌 Slow |
|
||||
| **RAM Required** | ~5.5 GB |
|
||||
| **Recommendation** | Showed up in a few results, but not recommended. |
|
||||
|
||||
## Prompt Template (ChatML)
|
||||
|
||||
This model uses the **ChatML** format used by Qwen:
|
||||
|
||||
```text
|
||||
<|im_start|>system
|
||||
You are a helpful assistant.<|im_end|>
|
||||
<|im_start|>user
|
||||
{prompt}<|im_end|>
|
||||
<|im_start|>assistant
|
||||
```
|
||||
|
||||
Set this in your app (LM Studio, OpenWebUI, etc.) for best results.
|
||||
|
||||
## Generation Parameters
|
||||
|
||||
### Thinking Mode (Recommended for Logic)
|
||||
Use when solving math, coding, or logical problems.
|
||||
|
||||
| Parameter | Value |
|
||||
|----------------|-------|
|
||||
| Temperature | 0.6 |
|
||||
| Top-P | 0.95 |
|
||||
| Top-K | 20 |
|
||||
| Min-P | 0.0 |
|
||||
| Repeat Penalty | 1.1 |
|
||||
|
||||
> ❗ DO NOT use greedy decoding — it causes infinite loops.
|
||||
|
||||
Enable via:
|
||||
- `enable_thinking=True` in tokenizer
|
||||
- Or add `/think` in user input during conversation
|
||||
|
||||
### Non-Thinking Mode (Fast Dialogue)
|
||||
For casual chat and quick replies.
|
||||
|
||||
| Parameter | Value |
|
||||
|----------------|-------|
|
||||
| Temperature | 0.7 |
|
||||
| Top-P | 0.8 |
|
||||
| Top-K | 20 |
|
||||
| Min-P | 0.0 |
|
||||
| Repeat Penalty | 1.1 |
|
||||
|
||||
Enable via:
|
||||
- `enable_thinking=False`
|
||||
- Or add `/no_think` in prompt
|
||||
|
||||
Stop sequences: `<|im_end|>`, `<|im_start|>`
|
||||
|
||||
## 💡 Usage Tips
|
||||
|
||||
> This model supports two operational modes:
|
||||
>
|
||||
> ### 🔍 Thinking Mode (Recommended for Logic)
|
||||
> Activate with `enable_thinking=True` or append `/think` in prompt.
|
||||
>
|
||||
> - Ideal for: math, coding, planning, analysis
|
||||
> - Use sampling: `temp=0.6`, `top_p=0.95`, `top_k=20`
|
||||
> - Avoid greedy decoding
|
||||
>
|
||||
> ### ⚡ Non-Thinking Mode (Fast Chat)
|
||||
> Use `enable_thinking=False` or `/no_think`.
|
||||
>
|
||||
> - Best for: casual conversation, quick answers
|
||||
> - Sampling: `temp=0.7`, `top_p=0.8`
|
||||
>
|
||||
> ---
|
||||
>
|
||||
> 🔄 **Switch Dynamically**
|
||||
> In multi-turn chats, the last `/think` or `/no_think` directive takes precedence.
|
||||
>
|
||||
> 🔁 **Avoid Repetition**
|
||||
> Set `presence_penalty=1.5` if stuck in loops.
|
||||
>
|
||||
> 📏 **Use Full Context**
|
||||
> Allow up to 32,768 output tokens for complex tasks.
|
||||
>
|
||||
> 🧰 **Agent Ready**
|
||||
> Works with Qwen-Agent, MCP servers, and custom tools.
|
||||
|
||||
## Customisation & Troubleshooting
|
||||
|
||||
Importing directly into Ollama should work, but you might encounter this error: `Error: invalid character '<' looking for beginning of value`.
|
||||
In this case try these steps:
|
||||
|
||||
1. `wget https://huggingface.co/geoffmunn/Qwen3-8B-f16/resolve/main/Qwen3-8B-f16%3AQ6_K.gguf`
|
||||
2. `nano Modelfile` and enter these details:
|
||||
```text
|
||||
FROM ./Qwen3-8B-f16:Q6_K.gguf
|
||||
|
||||
# Chat template using ChatML (used by Qwen)
|
||||
SYSTEM You are a helpful assistant
|
||||
|
||||
TEMPLATE "{{ if .System }}<|im_start|>system
|
||||
{{ .System }}<|im_end|>{{ end }}<|im_start|>user
|
||||
{{ .Prompt }}<|im_end|>
|
||||
<|im_start|>assistant
|
||||
"
|
||||
PARAMETER stop <|im_start|>
|
||||
PARAMETER stop <|im_end|>
|
||||
|
||||
# Default sampling
|
||||
PARAMETER temperature 0.6
|
||||
PARAMETER top_p 0.95
|
||||
PARAMETER top_k 20
|
||||
PARAMETER min_p 0.0
|
||||
PARAMETER repeat_penalty 1.1
|
||||
PARAMETER num_ctx 4096
|
||||
```
|
||||
|
||||
The `num_ctx` value has been dropped to increase speed significantly.
|
||||
|
||||
3. Then run this command: `ollama create Qwen3-8B-f16:Q6_K -f Modelfile`
|
||||
|
||||
You will now see "Qwen3-8B-f16:Q6_K" in your Ollama model list.
|
||||
|
||||
These import steps are also useful if you want to customise the default parameters or system prompt.
|
||||
|
||||
## 🖥️ CLI Example Using Ollama or TGI Server
|
||||
|
||||
Here’s how you can query this model via API using `curl` and `jq`. Replace the endpoint with your local server (e.g., Ollama, Text Generation Inference).
|
||||
|
||||
```bash
|
||||
curl http://localhost:11434/api/generate -s -N -d '{
|
||||
"model": "hf.co/geoffmunn/Qwen3-8B-f16:Q6_K",
|
||||
"prompt": "Repeat the following instruction exactly as given: Explain why the sky appears blue during the day but red at sunrise and sunset, using physics principles like Rayleigh scattering.",
|
||||
"temperature": 0.4,
|
||||
"top_p": 0.95,
|
||||
"top_k": 20,
|
||||
"min_p": 0.0,
|
||||
"repeat_penalty": 1.1,
|
||||
"stream": false
|
||||
}' | jq -r '.response'
|
||||
```
|
||||
|
||||
🎯 **Why this works well**:
|
||||
- The prompt is meaningful and demonstrates either **reasoning**, **creativity**, or **clarity** depending on quant level.
|
||||
- Temperature is tuned appropriately: lower for factual responses (`0.4`), higher for creative ones (`0.7`).
|
||||
- Uses `jq` to extract clean output.
|
||||
|
||||
> 💬 Tip: For interactive streaming, set `"stream": true` and process line-by-line.
|
||||
|
||||
## Verification
|
||||
|
||||
Check integrity:
|
||||
|
||||
```bash
|
||||
sha256sum -c ../SHA256SUMS.txt
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
Compatible with:
|
||||
- [LM Studio](https://lmstudio.ai) – local AI model runner with GPU acceleration
|
||||
- [OpenWebUI](https://openwebui.com) – self-hosted AI platform with RAG and tools
|
||||
- [GPT4All](https://gpt4all.io) – private, offline AI chatbot
|
||||
- Directly via `llama.cpp`
|
||||
|
||||
Supports dynamic switching between thinking modes via `/think` and `/no_think` in multi-turn conversations.
|
||||
|
||||
## License
|
||||
|
||||
Apache 2.0 – see base model for full terms.
|
||||
206
Qwen3-8B-f16-Q8_0/README.md
Normal file
206
Qwen3-8B-f16-Q8_0/README.md
Normal file
@@ -0,0 +1,206 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
tags:
|
||||
- gguf
|
||||
- qwen
|
||||
- qwen3-8b
|
||||
- qwen3-8b-q8
|
||||
- qwen3-8b-q8_0
|
||||
- qwen3-8b-q8_0-gguf
|
||||
- llama.cpp
|
||||
- quantized
|
||||
- text-generation
|
||||
- reasoning
|
||||
- agent
|
||||
- chat
|
||||
- multilingual
|
||||
- imatrix
|
||||
- q3_hifi
|
||||
base_model: Qwen/Qwen3-8B
|
||||
author: geoffmunn
|
||||
---
|
||||
|
||||
# Qwen3-8B-f16:Q8_0
|
||||
|
||||
Quantized version of [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B) at **Q8_0** level, derived from **f16** base weights.
|
||||
|
||||
## Model Info
|
||||
|
||||
- **Format**: GGUF (for llama.cpp and compatible runtimes)
|
||||
- **Size**: 8.71 GB
|
||||
- **Precision**: Q8_0
|
||||
- **Base Model**: [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B)
|
||||
- **Conversion Tool**: [llama.cpp](https://github.com/ggerganov/llama.cpp)
|
||||
|
||||
## Quality & Performance
|
||||
|
||||
| Metric | Value |
|
||||
|--------------------|-----------------------------------------|
|
||||
| **Speed** | 🐌 Slow |
|
||||
| **RAM Required** | ~7.1 GB |
|
||||
| **Recommendation** | Not recommended, Only one top 3 finish. |
|
||||
|
||||
## Prompt Template (ChatML)
|
||||
|
||||
This model uses the **ChatML** format used by Qwen:
|
||||
|
||||
```text
|
||||
<|im_start|>system
|
||||
You are a helpful assistant.<|im_end|>
|
||||
<|im_start|>user
|
||||
{prompt}<|im_end|>
|
||||
<|im_start|>assistant
|
||||
```
|
||||
|
||||
Set this in your app (LM Studio, OpenWebUI, etc.) for best results.
|
||||
|
||||
## Generation Parameters
|
||||
|
||||
### Thinking Mode (Recommended for Logic)
|
||||
Use when solving math, coding, or logical problems.
|
||||
|
||||
| Parameter | Value |
|
||||
|----------------|-------|
|
||||
| Temperature | 0.6 |
|
||||
| Top-P | 0.95 |
|
||||
| Top-K | 20 |
|
||||
| Min-P | 0.0 |
|
||||
| Repeat Penalty | 1.1 |
|
||||
|
||||
> ❗ DO NOT use greedy decoding — it causes infinite loops.
|
||||
|
||||
Enable via:
|
||||
- `enable_thinking=True` in tokenizer
|
||||
- Or add `/think` in user input during conversation
|
||||
|
||||
### Non-Thinking Mode (Fast Dialogue)
|
||||
For casual chat and quick replies.
|
||||
|
||||
| Parameter | Value |
|
||||
|----------------|-------|
|
||||
| Temperature | 0.7 |
|
||||
| Top-P | 0.8 |
|
||||
| Top-K | 20 |
|
||||
| Min-P | 0.0 |
|
||||
| Repeat Penalty | 1.1 |
|
||||
|
||||
Enable via:
|
||||
- `enable_thinking=False`
|
||||
- Or add `/no_think` in prompt
|
||||
|
||||
Stop sequences: `<|im_end|>`, `<|im_start|>`
|
||||
|
||||
## 💡 Usage Tips
|
||||
|
||||
> This model supports two operational modes:
|
||||
>
|
||||
> ### 🔍 Thinking Mode (Recommended for Logic)
|
||||
> Activate with `enable_thinking=True` or append `/think` in prompt.
|
||||
>
|
||||
> - Ideal for: math, coding, planning, analysis
|
||||
> - Use sampling: `temp=0.6`, `top_p=0.95`, `top_k=20`
|
||||
> - Avoid greedy decoding
|
||||
>
|
||||
> ### ⚡ Non-Thinking Mode (Fast Chat)
|
||||
> Use `enable_thinking=False` or `/no_think`.
|
||||
>
|
||||
> - Best for: casual conversation, quick answers
|
||||
> - Sampling: `temp=0.7`, `top_p=0.8`
|
||||
>
|
||||
> ---
|
||||
>
|
||||
> 🔄 **Switch Dynamically**
|
||||
> In multi-turn chats, the last `/think` or `/no_think` directive takes precedence.
|
||||
>
|
||||
> 🔁 **Avoid Repetition**
|
||||
> Set `presence_penalty=1.5` if stuck in loops.
|
||||
>
|
||||
> 📏 **Use Full Context**
|
||||
> Allow up to 32,768 output tokens for complex tasks.
|
||||
>
|
||||
> 🧰 **Agent Ready**
|
||||
> Works with Qwen-Agent, MCP servers, and custom tools.
|
||||
|
||||
## Customisation & Troubleshooting
|
||||
|
||||
Importing directly into Ollama should work, but you might encounter this error: `Error: invalid character '<' looking for beginning of value`.
|
||||
In this case try these steps:
|
||||
|
||||
1. `wget https://huggingface.co/geoffmunn/Qwen3-8B-f16/resolve/main/Qwen3-8B-f16%3AQ8_0.gguf`
|
||||
2. `nano Modelfile` and enter these details:
|
||||
```text
|
||||
FROM ./Qwen3-8B-f16:Q8_0.gguf
|
||||
|
||||
# Chat template using ChatML (used by Qwen)
|
||||
SYSTEM You are a helpful assistant
|
||||
|
||||
TEMPLATE "{{ if .System }}<|im_start|>system
|
||||
{{ .System }}<|im_end|>{{ end }}<|im_start|>user
|
||||
{{ .Prompt }}<|im_end|>
|
||||
<|im_start|>assistant
|
||||
"
|
||||
PARAMETER stop <|im_start|>
|
||||
PARAMETER stop <|im_end|>
|
||||
|
||||
# Default sampling
|
||||
PARAMETER temperature 0.6
|
||||
PARAMETER top_p 0.95
|
||||
PARAMETER top_k 20
|
||||
PARAMETER min_p 0.0
|
||||
PARAMETER repeat_penalty 1.1
|
||||
PARAMETER num_ctx 4096
|
||||
```
|
||||
|
||||
The `num_ctx` value has been dropped to increase speed significantly.
|
||||
|
||||
3. Then run this command: `ollama create Qwen3-8B-f16:Q8_0 -f Modelfile`
|
||||
|
||||
You will now see "Qwen3-8B-f16:Q8_0" in your Ollama model list.
|
||||
|
||||
These import steps are also useful if you want to customise the default parameters or system prompt.
|
||||
|
||||
## 🖥️ CLI Example Using Ollama or TGI Server
|
||||
|
||||
Here’s how you can query this model via API using `curl` and `jq`. Replace the endpoint with your local server (e.g., Ollama, Text Generation Inference).
|
||||
|
||||
```bash
|
||||
curl http://localhost:11434/api/generate -s -N -d '{
|
||||
"model": "hf.co/geoffmunn/Qwen3-8B-f16:Q8_0",
|
||||
"prompt": "Repeat the following instruction exactly as given: Explain why the sky appears blue during the day but red at sunrise and sunset, using physics principles like Rayleigh scattering.",
|
||||
"temperature": 0.4,
|
||||
"top_p": 0.95,
|
||||
"top_k": 20,
|
||||
"min_p": 0.0,
|
||||
"repeat_penalty": 1.1,
|
||||
"stream": false
|
||||
}' | jq -r '.response'
|
||||
```
|
||||
|
||||
🎯 **Why this works well**:
|
||||
- The prompt is meaningful and demonstrates either **reasoning**, **creativity**, or **clarity** depending on quant level.
|
||||
- Temperature is tuned appropriately: lower for factual responses (`0.4`), higher for creative ones (`0.7`).
|
||||
- Uses `jq` to extract clean output.
|
||||
|
||||
> 💬 Tip: For interactive streaming, set `"stream": true` and process line-by-line.
|
||||
|
||||
## Verification
|
||||
|
||||
Check integrity:
|
||||
|
||||
```bash
|
||||
sha256sum -c ../SHA256SUMS.txt
|
||||
```
|
||||
|
||||
## Usage
|
||||
|
||||
Compatible with:
|
||||
- [LM Studio](https://lmstudio.ai) – local AI model runner with GPU acceleration
|
||||
- [OpenWebUI](https://openwebui.com) – self-hosted AI platform with RAG and tools
|
||||
- [GPT4All](https://gpt4all.io) – private, offline AI chatbot
|
||||
- Directly via `llama.cpp`
|
||||
|
||||
Supports dynamic switching between thinking modes via `/think` and `/no_think` in multi-turn conversations.
|
||||
|
||||
## License
|
||||
|
||||
Apache 2.0 – see base model for full terms.
|
||||
3
Qwen3-8B-f16-imatrix-4697-coder.gguf
Normal file
3
Qwen3-8B-f16-imatrix-4697-coder.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:d2ac1c7c3bdf55f85355b01c144f6c5a1a65c9ffd28a9355453d95c444d099df
|
||||
size 5347200
|
||||
3
Qwen3-8B-f16-imatrix-4697-generic.gguf
Normal file
3
Qwen3-8B-f16-imatrix-4697-generic.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:7e59d35c1c4d3114f8ed305e2ab19139d2eaa615d5badab9a978025829d3daa2
|
||||
size 5347200
|
||||
3
Qwen3-8B-f16-imatrix-8843-coder.gguf
Normal file
3
Qwen3-8B-f16-imatrix-8843-coder.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:d2ac1c7c3bdf55f85355b01c144f6c5a1a65c9ffd28a9355453d95c444d099df
|
||||
size 5347200
|
||||
3
Qwen3-8B-f16-imatrix-9343-generic.gguf
Normal file
3
Qwen3-8B-f16-imatrix-9343-generic.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:7e59d35c1c4d3114f8ed305e2ab19139d2eaa615d5badab9a978025829d3daa2
|
||||
size 5347200
|
||||
3
Qwen3-8B-f16-imatrix:Q2_K.gguf
Normal file
3
Qwen3-8B-f16-imatrix:Q2_K.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:85db3fdb95a5155b04fe5297f07ba20b2eb1714fd4b73eebde67e56b3ae130f8
|
||||
size 3281733216
|
||||
3
Qwen3-8B-f16-imatrix:Q2_K_HIFI.gguf
Normal file
3
Qwen3-8B-f16-imatrix:Q2_K_HIFI.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:02cf559368fea369c3442ef58d6c8436b1175d343e805600476f3836a6c7278f
|
||||
size 3442927200
|
||||
3
Qwen3-8B-f16-imatrix:Q3_K_HIFI.gguf
Normal file
3
Qwen3-8B-f16-imatrix:Q3_K_HIFI.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:e85a7c1984ab12122261702e3d5ee18f55c251920ac11c5902d90bf513ec253a
|
||||
size 4828370528
|
||||
3
Qwen3-8B-f16-imatrix:Q3_K_M.gguf
Normal file
3
Qwen3-8B-f16-imatrix:Q3_K_M.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:3b0dfac8d304aede9580a307cf42abc9f1eee0599a7115d75ba9dee40bed0dc0
|
||||
size 4124161632
|
||||
3
Qwen3-8B-f16-imatrix:Q3_K_S.gguf
Normal file
3
Qwen3-8B-f16-imatrix:Q3_K_S.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:d6bcf405ed302a736c96d2ac5193a94111e07577dc187ebf3e7f876bc0011c50
|
||||
size 3769611872
|
||||
3
Qwen3-8B-f16-imatrix:Q4_K_HIFI.gguf
Normal file
3
Qwen3-8B-f16-imatrix:Q4_K_HIFI.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:57b73385284b268a88568ef54d8dd4a10f3868a02da799c2000a2c1963017254
|
||||
size 5703579264
|
||||
3
Qwen3-8B-f16-imatrix:Q4_K_M.gguf
Normal file
3
Qwen3-8B-f16-imatrix:Q4_K_M.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:c09d414416f245c8d20759d7629806aeb6250b5de3b2c159f37f9ac168d4b2b5
|
||||
size 5027784288
|
||||
3
Qwen3-8B-f16-imatrix:Q4_K_S.gguf
Normal file
3
Qwen3-8B-f16-imatrix:Q4_K_S.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:21cd101368068873d9b6ce431aba451e47c4b6b0cbafcd957ad34f673ac30601
|
||||
size 4802012768
|
||||
3
Qwen3-8B-f16-imatrix:Q5_K_HIFI.gguf
Normal file
3
Qwen3-8B-f16-imatrix:Q5_K_HIFI.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:f2ddf0729977105a55cdda882e9d799bff3bac36d1e7fef2a1982f6da00cec0f
|
||||
size 6041810560
|
||||
3
Qwen3-8B-f16-imatrix:Q5_K_M.gguf
Normal file
3
Qwen3-8B-f16-imatrix:Q5_K_M.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:e5625e30f36988e76e3886ac5cfc97e8e746d50dd1f413f08677f70b099e831c
|
||||
size 5851113056
|
||||
3
Qwen3-8B-f16-imatrix:Q5_K_S.gguf
Normal file
3
Qwen3-8B-f16-imatrix:Q5_K_S.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:a085d45efe804436e1d078957c4abbc96fe799622a544232ef0b7069a88f4186
|
||||
size 5720761952
|
||||
3
Qwen3-8B-f16:Q2_K.gguf
Normal file
3
Qwen3-8B-f16:Q2_K.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:93551e45c3c4dd15ca31420a824662c2e5ec0f80310786bcb4364c2c54077d11
|
||||
size 3281732960
|
||||
3
Qwen3-8B-f16:Q2_K_HIFI.gguf
Normal file
3
Qwen3-8B-f16:Q2_K_HIFI.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:8a3013d2518eb264ec3e98fbfcf64b3d8e61786f4845cdd13c702a2900a09091
|
||||
size 3442926944
|
||||
3
Qwen3-8B-f16:Q3_K_HIFI.gguf
Normal file
3
Qwen3-8B-f16:Q3_K_HIFI.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:5caa61a0b2c4f2e008a3a700cc4d2c7958623677ba0415a4e0652f931bcd20f4
|
||||
size 4811986272
|
||||
3
Qwen3-8B-f16:Q3_K_M.gguf
Normal file
3
Qwen3-8B-f16:Q3_K_M.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:a85cd12d9887f0de21c0df6dda0601061e569a580b8121dd94096ea8fbacab2f
|
||||
size 4124161376
|
||||
3
Qwen3-8B-f16:Q3_K_S.gguf
Normal file
3
Qwen3-8B-f16:Q3_K_S.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:34e25be87d35fb3f00779710b22b2b69e6210f67eefcf269f1b3758274a9b2ee
|
||||
size 3769611616
|
||||
3
Qwen3-8B-f16:Q4_K_HIFI.gguf
Normal file
3
Qwen3-8B-f16:Q4_K_HIFI.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:b5cd60d8a390afcee5a5f443ef25235b2592ed72487405d9666d88c6cfb33806
|
||||
size 5703579040
|
||||
3
Qwen3-8B-f16:Q4_K_M.gguf
Normal file
3
Qwen3-8B-f16:Q4_K_M.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:10ded9291f250596ef7149438dd5da80bf3780b0eebcd2a1922c2b4a08c36e54
|
||||
size 5027784032
|
||||
3
Qwen3-8B-f16:Q4_K_S.gguf
Normal file
3
Qwen3-8B-f16:Q4_K_S.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:07567ef46246e0f7bca5dad8fccc7c383e592d5912edd362b3ade8351ad30aab
|
||||
size 4802012512
|
||||
3
Qwen3-8B-f16:Q5_K_HIFI.gguf
Normal file
3
Qwen3-8B-f16:Q5_K_HIFI.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:36d20cea46aaa9bcce80ea48bc0da74bac483218179bba82967829cc734ce1b8
|
||||
size 6041810336
|
||||
3
Qwen3-8B-f16:Q5_K_M.gguf
Normal file
3
Qwen3-8B-f16:Q5_K_M.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:560737e3d4f9182d664858aea9c0a164f751154f78e32abc3276d1b4e606f17d
|
||||
size 5851112800
|
||||
3
Qwen3-8B-f16:Q5_K_S.gguf
Normal file
3
Qwen3-8B-f16:Q5_K_S.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:3c7625faba5fee4a39d627c1e6b2331d21423437a1240ca336a676973a3d8824
|
||||
size 5720761696
|
||||
3
Qwen3-8B-f16:Q6_K.gguf
Normal file
3
Qwen3-8B-f16:Q6_K.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:863da90f35b7c22b1cd184a20d04cc0eaf0df67f8f52ab0a6d4f68d192600898
|
||||
size 6725899552
|
||||
3
Qwen3-8B-f16:Q8_0.gguf
Normal file
3
Qwen3-8B-f16:Q8_0.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:21962d706e584d0058a3f078ba42a40b7c3c82a1aa1a25588d372d41c99a8b6e
|
||||
size 8709518624
|
||||
327
README.md
Normal file
327
README.md
Normal file
@@ -0,0 +1,327 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
tags:
|
||||
- gguf
|
||||
- qwen
|
||||
- qwen3
|
||||
- qwen3-8b
|
||||
- qwen3-8b-gguf
|
||||
- llama.cpp
|
||||
- quantized
|
||||
- text-generation
|
||||
- reasoning
|
||||
- agent
|
||||
- chat
|
||||
- multilingual
|
||||
- matrix
|
||||
- q3_hifi
|
||||
- q4_hifi
|
||||
- q5_hifi
|
||||
base_model: Qwen/Qwen3-8B
|
||||
author: geoffmunn
|
||||
pipeline_tag: text-generation
|
||||
language:
|
||||
- en
|
||||
- zh
|
||||
- es
|
||||
- fr
|
||||
- de
|
||||
- ru
|
||||
- ar
|
||||
- ja
|
||||
- ko
|
||||
- hi
|
||||
---
|
||||
|
||||
# Qwen3-8B-f16-GGUF
|
||||
|
||||
This is a **GGUF-quantized version** of the **[Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B)** language model - an **8-billion-parameter** LLM from Alibaba's Qwen series, designed for **advanced reasoning, agentic behavior, and multilingual tasks**.
|
||||
|
||||
Converted for use with `llama.cpp` and compatible tools like OpenWebUI, LM Studio, GPT4All, and more.
|
||||
|
||||
## Why Use an 8B Model?
|
||||
|
||||
The **Qwen3-8B** model represents a significant leap in capability while remaining remarkably accessible for local and edge deployment. It offers:
|
||||
- **Near-state-of-the-art reasoning, coding, and multilingual performance** among open 8B-class models
|
||||
- **Smooth inference on a single consumer GPU** (e.g., 16–24 GB VRAM) or fast CPU runtime with quantization
|
||||
- **Quantized versions (e.g., GGUF Q4_K_M, AWQ) that fit within ~6–8 GB of memory**, enabling use on mid-range hardware
|
||||
- **Strong performance on complex tasks** like document summarization, structured output generation, and agentic workflows
|
||||
|
||||
It’s ideal for:
|
||||
- Local AI assistants that handle nuanced, multi-turn conversations
|
||||
- Self-hosted RAG pipelines with deep document understanding
|
||||
- Developers building production-grade on-prem AI features without cloud dependencies
|
||||
- Researchers and tinkerers seeking a capable yet manageable open-weight foundation
|
||||
|
||||
Choose Qwen3-8B when you need high-quality output and robust general intelligence - but still value efficiency, privacy, and full control over your deployment environment.
|
||||
|
||||
# Qwen3 8B Quantization Guide: Cross-Bit Summary & Recommendations
|
||||
|
||||
## Executive Summary
|
||||
|
||||
At 8B scale, **quantization achieves exceptional resilience**—all bit widths deliver production-ready quality with imatrix, and even Q2_K becomes viable (+13.4% loss). The model's parameter redundancy provides a "sweet spot" where aggressive compression meets robust architecture. Q5_K_HIFI + imatrix achieves near-lossless fidelity (+0.27% vs F16), while Q4_K_M + imatrix offers the best balance of quality, speed, and compatibility:
|
||||
|
||||
| Bit Width | Best Variant (+ imatrix) | Quality vs F16 | File Size | Speed | Memory | Viability |
|
||||
|-----------|--------------------------|----------------|-----------|-------|--------|-----------|
|
||||
| **Q5_K** | Q5_K_HIFI + imatrix | **+0.27%** ✅✅✅ | 5.62 GiB | 109.7 TPS | 5,754 MiB | Exceptional |
|
||||
| **Q4_K** | Q4_K_M + imatrix | **+1.3%** ✅✅ | 4.68 GiB | 125.5 TPS | 4,792 MiB | Excellent |
|
||||
| **Q3_K** | Q3_K_HIFI + imatrix | **+3.5%** ✅ | 2.15 GiB | 151.3 TPS | 2,202 MiB | Very Good |
|
||||
| **Q2_K** | Q2_K + imatrix | **+13.4%** ⚠️ | 3.05 GiB | 169.9 TPS | 3,134 MiB | Fair (viable) |
|
||||
|
||||
💡 **Critical insight**: 8B represents the **inflection point** where Q2_K becomes genuinely viable with imatrix (+13.4% loss vs +35% at 1.7B). Q5_K_HIFI + imatrix achieves near-lossless quality (+0.27%), while Q4_K_M + imatrix provides the best practical balance. All variants are production-ready with imatrix.
|
||||
|
||||
---
|
||||
|
||||
## Bit-Width Recommendations by Use Case
|
||||
|
||||
### ✅ Quality-Critical Applications
|
||||
**→ Q5_K_HIFI + imatrix**
|
||||
- Best perplexity at **10.1377 PPL (+0.27% vs F16)** — near-lossless fidelity
|
||||
- Only 0.27% precision loss represents the closest approach to F16 quality across all quantization levels
|
||||
- Requires custom llama.cpp build with `Q6_K_HIFI_RES8` support
|
||||
- ⚠️ **Never use Q5_K_S without imatrix** — quality degrades to +1.62% vs F16
|
||||
|
||||
### ⚖️ Best Overall Balance (Recommended Default)
|
||||
**→ Q4_K_M + imatrix**
|
||||
- Excellent +1.3% precision loss vs F16 (PPL 10.2384)
|
||||
- Strong 125.5 TPS speed (+171% vs F16)
|
||||
- Compact 4.68 GiB file size (69.3% smaller than F16)
|
||||
- **Standard llama.cpp compatibility** — no custom build required
|
||||
- Ideal for most development and production scenarios
|
||||
|
||||
### 🚀 Maximum Speed
|
||||
**→ Q2_K + imatrix**
|
||||
- Fastest variant at **169.9 TPS** (+267% vs F16)
|
||||
- Surprisingly viable quality at +13.4% loss with imatrix
|
||||
- ⚠️ **Never use without imatrix** — quality degrades catastrophically to +57.9% loss
|
||||
|
||||
### 💎 Near-Lossless 3-Bit Option
|
||||
**→ Q3_K_HIFI + imatrix**
|
||||
- **Remarkable +3.5% precision loss** — exceptional for 3-bit quantization
|
||||
- 71.2% memory reduction (2,202 MiB vs 7,670 MiB)
|
||||
- Unique value: When you need maximum compression but cannot accept Q3_K_S quality
|
||||
- ⚠️ **27–38% slower than Q3_K_M** — significant speed trade-off
|
||||
|
||||
### 📱 Extreme Memory Constraints (< 2.0 GiB)
|
||||
**→ Q3_K_S + imatrix**
|
||||
- Absolute smallest footprint (1.75 GiB file, 1,792 MiB runtime)
|
||||
- Acceptable +9.0% precision loss with imatrix
|
||||
- Only viable option under 2.0 GiB budget
|
||||
|
||||
---
|
||||
|
||||
## Critical Warnings for 8B Scale
|
||||
|
||||
⚠️ **Q5_K quality ranking reversal with imatrix** — Q5_K_S + imatrix (10.1538 PPL) actually beats Q5_K_M + imatrix (10.1612 PPL) by 0.07 PPL points. This makes Q5_K_S + imatrix viable for speed-constrained deployments where the 3.2% speed advantage matters.
|
||||
|
||||
⚠️ **Q4_K_S without imatrix is unusable** — Suffers +5.7% precision loss (10.6893 PPL) — the highest degradation of any Q4 variant at 8B scale. **Always pair Q4_K_S with imatrix** (reduces loss to +1.9%).
|
||||
|
||||
⚠️ **Q2_K requires imatrix** — Without it, Q2_K suffers +57.9% precision loss (completely unusable). With imatrix, quality improves to +13.4% — viable for non-critical tasks.
|
||||
|
||||
⚠️ **Q2_K_HIFI is strictly worse than Q2_K** — At 8B scale, Q2_K_HIFI loses to Q2_K on every metric (quality, speed, size, memory). Always prefer standard Q2_K over Q2_K_HIFI.
|
||||
|
||||
⚠️ **Q3_K_HIFI requires no special handling** — Unlike at 0.6B/1.7B scales, Q3_K_HIFI at 8B delivers substantial quality gains (+3.5% vs F16 with imatrix) that justify its 13.5% memory premium over Q3_K_M.
|
||||
|
||||
⚠️ **All Q3 variants are production-ready** — Even Q3_K_S with imatrix (+9.0% loss) remains usable for non-critical tasks — a dramatic improvement over smaller scales where Q3 quantization often fails.
|
||||
|
||||
---
|
||||
|
||||
## Memory Budget Guide
|
||||
|
||||
| Available VRAM | Recommended Variant | Expected Quality | Why |
|
||||
|----------------|---------------------|------------------|-----|
|
||||
| **< 2.0 GiB** | Q3_K_S + imatrix | PPL 11.02, +9.0% loss ⚠️ | Only option that fits; quality acceptable for non-critical tasks |
|
||||
| **2.0 – 2.5 GiB** | Q3_K_M + imatrix | PPL 10.62, +5.1% loss ✅ | Best Q3 balance; production-ready quality |
|
||||
| **2.5 – 3.5 GiB** | Q2_K + imatrix | PPL 11.46, +13.4% loss ⚠️ | Maximum speed at 169.9 TPS; quality acceptable for simple tasks |
|
||||
| **3.5 – 5.0 GiB** | Q4_K_M + imatrix | PPL 10.24, +1.3% loss ✅ | Best balance of quality/speed/size; standard compatibility |
|
||||
| **5.0 – 6.5 GiB** | Q5_K_HIFI + imatrix | PPL 10.14, +0.27% loss ✅ | Near-lossless quality; requires custom build |
|
||||
| **> 15.3 GiB** | F16 | Best quality (baseline) | Only if absolute precision required |
|
||||
|
||||
---
|
||||
|
||||
## Cross-Bit Performance Comparison
|
||||
|
||||
| Priority | Q2_K Best | Q3_K Best | Q4_K Best | Q5_K Best | Winner |
|
||||
|----------|-----------|-----------|-----------|-----------|--------|
|
||||
| **Quality (with imat)** | Q2_K (+13.4%) | Q3_K_HIFI (+3.5%) | Q4_K_M (+1.3%) | **Q5_K_HIFI (+0.27%)** ✅ | **Q5_K_HIFI** |
|
||||
| **Speed** | **Q2_K (169.9 TPS)** ✅ | Q3_K_S (223.5 TPS) | Q4_K_S (131.0 TPS) | Q5_K_S (113.3 TPS) | **Q2_K** |
|
||||
| **Smallest Size** | Q2_K (3.05 GiB) | **Q3_K_S (1.75 GiB)** ✅ | Q4_K_S (4.47 GiB) | Q5_K_S (5.32 GiB) | **Q3_K_S** |
|
||||
| **Best Balance** | Q2_K + imat | Q3_K_M + imat | **Q4_K_M + imat** ✅ | Q5_K_HIFI + imat | **Q4_K_M** |
|
||||
|
||||
✅ = Recommended for general use
|
||||
⚠️ = Context-dependent (see warnings above)
|
||||
|
||||
---
|
||||
|
||||
## Scale-Specific Insights: Why 8B Quantizes So Well
|
||||
|
||||
1. **Model redundancy threshold**: 8B represents the point where parameter count provides sufficient redundancy that quantization errors average out rather than accumulating catastrophically (unlike 0.6B/1.7B)
|
||||
|
||||
2. **Q2_K viability inflection**: 8B is the smallest scale where Q2_K becomes genuinely viable with imatrix (+13.4% loss). At 4B, Q2_K + imatrix is +18.7%; at 1.7B, +35.0%. This demonstrates a clear scale-dependent improvement curve.
|
||||
|
||||
3. **imatrix effectiveness plateau**: imatrix recovers 62–76% of precision loss at 8B — less dramatic than at 1.7B (70–78%) but more consistent across bit widths. Q5_K_S benefits most (74.1% recovery), making it competitive with Q5_K_M when imatrix is used.
|
||||
|
||||
4. **Residual quantization sweet spot**: Q5_K_HIFI's `Q6_K_HIFI_RES8` tensors provide maximal benefit at 8B scale — the 5 residual tensors capture precisely the right amount of quantization error without overhead.
|
||||
|
||||
5. **Q4_K_HIFI behavior shift**: Unlike at 14B where imatrix *harms* Q4_K_HIFI, at 8B imatrix *helps* it (-1.1% PPL improvement) — demonstrating non-linear scale effects.
|
||||
|
||||
6. **Q3_K viability threshold**: 8B is the smallest scale where Q3_K_HIFI achieves truly production-ready quality (+3.5% with imatrix) — below this, Q3 quantization requires careful validation.
|
||||
|
||||
---
|
||||
|
||||
## Decision Flowchart
|
||||
|
||||
```mermaid
|
||||
Need best quality?
|
||||
├─ Yes → Q5_K_HIFI + imatrix (+0.27% loss)
|
||||
└─ No → Need max speed?
|
||||
├─ Yes → Q2_K + imatrix (169.9 TPS, +13.4% loss)
|
||||
└─ No → Need smallest size?
|
||||
├─ Yes → Memory < 2.0 GiB?
|
||||
│ ├─ Yes → Q3_K_S + imatrix (1,792 MiB, +9.0% loss)
|
||||
│ └─ No → Q2_K + imatrix (3,134 MiB, +13.4% loss, fastest)
|
||||
└─ No → Q4_K_M + imatrix (best balance, +1.3% loss, standard build)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Practical Deployment Recommendations
|
||||
|
||||
### For Most Users
|
||||
**→ Q4_K_M + imatrix**
|
||||
Delivers excellent quality (+1.3% vs F16), strong speed (125.5 TPS), compact size (4.68 GiB), and universal llama.cpp compatibility. The safe, practical choice for 95% of deployments.
|
||||
|
||||
### For Quality-Critical Work
|
||||
**→ Q5_K_HIFI + imatrix**
|
||||
Achieves near-lossless quantization (+0.27% vs F16) with 64% memory reduction and 2.4× speedup. Requires custom build but worth it for research, content generation, or any task where output fidelity is non-negotiable.
|
||||
|
||||
### For Edge/Mobile Deployment
|
||||
**→ Q3_K_HIFI + imatrix**
|
||||
Best Q3 quality (+3.5% vs F16) with smallest viable footprint (2.15 GiB). Production-ready even without imatrix (+8.6% loss) — valuable for environments where imatrix generation isn't feasible.
|
||||
|
||||
### For High-Throughput Serving
|
||||
**→ Q5_K_S + imatrix**
|
||||
Fastest Q5 variant (113.3 TPS) with surprisingly good quality (+0.42% vs F16) that actually beats Q5_K_M with imatrix. Ideal when every TPS matters and marginal quality differences are acceptable.
|
||||
|
||||
### For Maximum Compression
|
||||
**→ Q2_K + imatrix**
|
||||
Only consider when memory/speed are absolutely critical and quality degradation is acceptable. At 8B scale, Q2_K + imatrix achieves +13.4% loss — viable for simple chatbots or non-critical inference.
|
||||
|
||||
---
|
||||
|
||||
## Bottom Line Recommendations
|
||||
|
||||
| Scenario | Recommended Variant | Rationale |
|
||||
|----------|---------------------|-----------|
|
||||
| **Default / General Purpose** | Q4_K_M + imatrix | Best balance of quality (+1.3%), speed (125.5 TPS), size (4.68 GiB), and compatibility |
|
||||
| **Maximum Quality** | Q5_K_HIFI + imatrix | Near-lossless (+0.27% vs F16) with 64% memory reduction and 2.4× speedup |
|
||||
| **Maximum Speed** | Q2_K + imatrix | Fastest (169.9 TPS, +267% vs F16) with acceptable quality (+13.4% loss) |
|
||||
| **Minimum Size** | Q3_K_S + imatrix | Smallest footprint (1.75 GiB) with acceptable quality (+9.0% loss) |
|
||||
| **No imatrix available** | Q5_K_HIFI (no imat) | Still excellent (+1.11% vs F16); all variants usable but quality reduced |
|
||||
| **Extreme constraints** | Q3_K_S + imatrix | Only if memory < 2.0 GiB; +9.0% loss acceptable for non-critical tasks |
|
||||
|
||||
⚠️ **Golden rules for 8B**:
|
||||
1. **Always use imatrix** — provides 62–76% precision recovery across all bit widths
|
||||
2. **Never use Q2_K without imatrix** — completely unusable (+57.9% loss)
|
||||
3. **Prefer Q2_K over Q2_K_HIFI** — HIFI is strictly worse on all metrics at 8B
|
||||
4. **Q5_K_S + imatrix beats Q5_K_M + imatrix** — unexpected quality ranking reversal
|
||||
5. **All four bit widths are viable** — choose based on constraints, not quality cliffs
|
||||
|
||||
✅ **8B is the quantization sweet spot**: Large enough for robustness across all bit widths (even Q2_K), small enough for dramatic efficiency gains. This scale demonstrates that intelligent quantization can deliver near-F16 quality at 1/3 the memory with 2.4–3.5× speed — a compelling value proposition for nearly all deployments.
|
||||
|
||||
## Non-technical model anaysis and rankings
|
||||
|
||||
**NOTE:** This analysis does not include the HIFI models.
|
||||
|
||||
There are numerous good candidates - lots of different models showed up in the top 3 across all the quesionts. However, **Qwen3-8B-f16:Q3_K_M** was a finalist in all but one question so is the recommended model (or Qwen3-8B-f16:Q3_HIFI). **Qwen3-8B-f16:Q5_K_S** did nearly as well and is worth considering,
|
||||
|
||||
The 'hello' question is the first time that all models got it exactly right. All models in the 8B range did well and it's mainly a question of what one works best on your hardware.
|
||||
|
||||
You can read the results here: [Qwen3-8B-analysis.md](Qwen3-8B-analysis.md)
|
||||
|
||||
If you find this useful, please give the project a ❤️ like.
|
||||
|
||||
## Non-HIFI recommentation table based on output
|
||||
|
||||
| Level | Speed | Size | Recommendation |
|
||||
|-----------|-----------|-------------|---------------------------------------------------------------------------------------|
|
||||
| Q2_K | ⚡ Fastest | 3.28 GB | Not recommended. Came first in the bat & ball question, no other appearances. |
|
||||
| 🥉Q3_K_S | ⚡ Fast | 3.77 GB | 🥉 Came first and second in questions covering both ends of the temperature spectrum. |
|
||||
| 🥇 Q3_K_M | ⚡ Fast | 4.12 GB | 🥇 **Best overall model.** Was a top 3 finisher for all questions except the haiku. |
|
||||
| 🥉Q4_K_S | 🚀 Fast | 4.8 GB | 🥉 Came first and second in questions covering both ends of the temperature spectrum. |
|
||||
| Q4_K_M | 🚀 Fast | 5.85 GB | Came first and second in questions covering high temperature questions. |
|
||||
| 🥈 Q5_K_S | 🐢 Medium | 5.72 GB | 🥈 A good second place. Good for all query types. |
|
||||
| Q5_K_M | 🐢 Medium | 5.85 GB | Not recommended, no appeareances in the top 3 for any question. |
|
||||
| Q6_K | 🐌 Slow | 6.73 GB | Showed up in a few results, but not recommended. |
|
||||
| Q8_0 | 🐌 Slow | 8.71 GB | Not recommended, Only one top 3 finish. |
|
||||
|
||||
## Build notes
|
||||
|
||||
You can read the guide for building llama.cpp here: [HIFI_BUILD_GUIDE.md](https://github.com/geoffmunn/llama.cpp/blob/master/HIFI_BUILD_GUIDE.md).
|
||||
|
||||
The HIFI quantization also used a very large 4697 chunk imatrix file for extra precision. You can re-use it here: [Qwen3-8B-f16-imatrix-4697-generic.gguf](https://huggingface.co/geoffmunn/Qwen3-8B-f16/blob/main/Qwen3-8B-f16-imatrix-4697-generic.gguf)
|
||||
|
||||
The imatrix was created as a generic mix of Wikipedia, mathmatics, and coding examples.
|
||||
|
||||
### Source code
|
||||
|
||||
You can use the HIFI GitHub repository to build it from source if you're interested: [https://github.com/geoffmunn/llama.cpp](https://github.com/geoffmunn/llama.cpp).
|
||||
|
||||
Build notes: [HIFI_BUILD_GUIDE.md](https://github.com/geoffmunn/llama.cpp/blob/master/HIFI_BUILD_GUIDE.md)
|
||||
|
||||
Improvements and feedback are welcome.
|
||||
|
||||
## Usage
|
||||
|
||||
Load this model using:
|
||||
- [OpenWebUI](https://openwebui.com) – self-hosted AI interface with RAG & tools
|
||||
- [LM Studio](https://lmstudio.ai) – desktop app with GPU support and chat templates
|
||||
- [GPT4All](https://gpt4all.io) – private, local AI chatbot (offline-first)
|
||||
- Or directly via `llama.cpp`
|
||||
|
||||
Each quantized model includes its own `README.md` and shares a common `MODELFILE` for optimal configuration.
|
||||
|
||||
Importing directly into Ollama should work, but you might encounter this error: `Error: invalid character '<' looking for beginning of value`.
|
||||
In this case try these steps:
|
||||
|
||||
1. `wget https://huggingface.co/geoffmunn/Qwen3-8B-f16/resolve/main/Qwen3-8B-f16%3AQ3_K_M.gguf` (replace the quantised version with the one you want)
|
||||
2. `nano Modelfile` and enter these details (again, replacing Q3_K_M with the version you want):
|
||||
```text
|
||||
FROM ./Qwen3-8B-f16:Q3_K_M.gguf
|
||||
|
||||
# Chat template using ChatML (used by Qwen)
|
||||
SYSTEM You are a helpful assistant
|
||||
|
||||
TEMPLATE "{{ if .System }}<|im_start|>system
|
||||
{{ .System }}<|im_end|>{{ end }}<|im_start|>user
|
||||
{{ .Prompt }}<|im_end|>
|
||||
<|im_start|>assistant
|
||||
"
|
||||
PARAMETER stop <|im_start|>
|
||||
PARAMETER stop <|im_end|>
|
||||
|
||||
# Default sampling
|
||||
PARAMETER temperature 0.6
|
||||
PARAMETER top_p 0.95
|
||||
PARAMETER top_k 20
|
||||
PARAMETER min_p 0.0
|
||||
PARAMETER repeat_penalty 1.1
|
||||
PARAMETER num_ctx 4096
|
||||
```
|
||||
|
||||
The `num_ctx` value has been dropped to increase speed significantly.
|
||||
|
||||
3. Then run this command: `ollama create Qwen3-8B-f16:Q3_K_M -f Modelfile`
|
||||
|
||||
You will now see "Qwen3-8B-f16:Q3_K_M" in your Ollama model list.
|
||||
|
||||
These import steps are also useful if you want to customise the default parameters or system prompt.
|
||||
|
||||
## Author
|
||||
|
||||
👤 Geoff Munn (@geoffmunn)
|
||||
🔗 [Hugging Face Profile](https://huggingface.co/geoffmunn)
|
||||
|
||||
## Disclaimer
|
||||
|
||||
This is a community conversion for local inference. Not affiliated with Alibaba Cloud or the Qwen team.
|
||||
40000
mixed-imatrix-dataset.txt
Normal file
40000
mixed-imatrix-dataset.txt
Normal file
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user