初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Qwen_Qwen3-14B-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-01 21:53:13 +08:00
commit 5301c41981
30 changed files with 303 additions and 0 deletions

49
.gitattributes vendored Normal file
View File

@@ -0,0 +1,49 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bin.* filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zstandard filter=lfs diff=lfs merge=lfs -text
*.tfevents* filter=lfs diff=lfs merge=lfs -text
*.db* filter=lfs diff=lfs merge=lfs -text
*.ark* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*data* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.meta filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.index filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.gguf* filter=lfs diff=lfs merge=lfs -text
*.ggml filter=lfs diff=lfs merge=lfs -text
*.llamafile* filter=lfs diff=lfs merge=lfs -text
*.pt2 filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Qwen_Qwen3-14B.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2c4d9143e2cd3827a1d4cf56da13e64918bdd664f8908ae4b6d86e55cf8d4bff
size 5322941472

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cfd11747daa50d47d3eeaf2af7503906c2fe966635458d8447dfb8ee40485a8b
size 4963312672

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7277e485b898f8fa7945f701c406401f683e5b49107e3c2f42907e829375f4d1
size 4691589152

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:86b555a31a6d8bd928d5ee4b6dbf8a73ffafe419b8d7247ced73fb2be1f3c9d1
size 6883409952

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:363c51dbaec3e2944f99738568960a92488487e218f1300918af71e9e3802ce5
size 6375301152

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a814fb65a1600c17db566cfc8496cadcb5be0883a58b33d289a9e3eefc157c14
size 5942666272

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fcb00c64571daba61218903cf0af72a4077f1edaa58823db476011c053dde86c
size 8541363232

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:448afbe2d83a40aa38e03019c37fec002d0e42da1f6a0a8d250876915b84a00d
size 8110730272

3
Qwen_Qwen3-14B-Q2_K.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4496a2dc7805120ae48c381e41bc81abee89eb7a82c5f33f72868f5d973e2cb2
size 5753984032

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a14999ffc47a2a98d2eecf3976a2fffe0fea1a67b11d6d88242f3649913ea62d
size 6513664032

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4a0fa4eabfc6a0ae490e35343d214b5d236285acd7543a8325fe99ba778a02df
size 7900651552

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3c49c800881cf97b77d6a6a7d3f653e35b6134e0cb6fd6f7c379b354b48af781
size 7321313312

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:758b8e6a06f55368ea6cf151176dc0d3b931894b8630ec20cf704d2d6be279b4
size 6657105952

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4b34ed0ed1b710121fc93cfb2aea162a0140bae01e503512fff546e2ab6fdaab
size 8581324832

3
Qwen_Qwen3-14B-Q4_0.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:639a6df3ef3231561e0a97e0d683812aa9025f401035c0c80aec7c368b02e18c
size 8543001632

3
Qwen_Qwen3-14B-Q4_1.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:02d8933245c909d63526421eb3c4563c73bd2b0d4b2aa2f2524b3005128094e9
size 9389521952

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0414094a48638a4c4862d7b822fed5ae8a1643e0c7caa799f452c01a00978eb9
size 9579110432

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:915913e22399475dbe6c968ac014d9f1fbe08975e489279aede9d5c7b2c98eb6
size 9001753632

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5c185d2643e1e8c550c1c85991908f160f862864e8d6e7b0bc12e5f2aeef2c6b
size 8573475872

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e8641e156ab6386074f7d3ee37ad89f45639ba3b5e1d6e2ac4f24d3db3ae16e9
size 10994688032

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:bf19bf5c77c530762012f71215b62aa8f0bae08111ee0d537b8496479e17cfa3
size 10514570272

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e13b3a83698cda6ebff869ab366f80b9fddf5d76f02f13bf8236d77d3833d10b
size 10263895072

3
Qwen_Qwen3-14B-Q6_K.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:de571a7d7c72b99de6d67e7f2b780d774c74f998e088c77133069568cad54294
size 12121937952

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ae52c8a419483e366c2db0b4b1bf911b24a78fb14f22564f98c0b7bae0e62248
size 12498739232

3
Qwen_Qwen3-14B-Q8_0.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:62e390154916e1dc6b00f63d997bda39e8f9679c209dcabb69bdff5043fac2e0
size 15698534432

3
Qwen_Qwen3-14B-bf16.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9677a58b0fa8da7771a4d8cc8080208ce02a32ab21305ded10a3798766132d3a
size 29543423776

3
Qwen_Qwen3-14B.imatrix Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:03eb98f405e02dc67d043f15150b362073477c9a7fb1c7d6539fbc3e2aa49787
size 7709778

172
README.md Normal file
View File

@@ -0,0 +1,172 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
base_model: Qwen/Qwen3-14B
base_model_relation: quantized
---
## Llamacpp imatrix Quantizations of Qwen3-14B by Qwen
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b5200">b5200</a> for quantization.
Original model: https://huggingface.co/Qwen/Qwen3-14B
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
Run them directly with [llama.cpp](https://github.com/ggerganov/llama.cpp), or any other llama.cpp based project
## Prompt format
```
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Qwen3-14B-bf16.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-14B-GGUF/blob/main/Qwen_Qwen3-14B-bf16.gguf) | bf16 | 29.54GB | false | Full BF16 weights. |
| [Qwen3-14B-Q8_0.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-14B-GGUF/blob/main/Qwen_Qwen3-14B-Q8_0.gguf) | Q8_0 | 15.70GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Qwen3-14B-Q6_K_L.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-14B-GGUF/blob/main/Qwen_Qwen3-14B-Q6_K_L.gguf) | Q6_K_L | 12.50GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Qwen3-14B-Q6_K.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-14B-GGUF/blob/main/Qwen_Qwen3-14B-Q6_K.gguf) | Q6_K | 12.12GB | false | Very high quality, near perfect, *recommended*. |
| [Qwen3-14B-Q5_K_L.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-14B-GGUF/blob/main/Qwen_Qwen3-14B-Q5_K_L.gguf) | Q5_K_L | 10.99GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Qwen3-14B-Q5_K_M.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-14B-GGUF/blob/main/Qwen_Qwen3-14B-Q5_K_M.gguf) | Q5_K_M | 10.51GB | false | High quality, *recommended*. |
| [Qwen3-14B-Q5_K_S.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-14B-GGUF/blob/main/Qwen_Qwen3-14B-Q5_K_S.gguf) | Q5_K_S | 10.26GB | false | High quality, *recommended*. |
| [Qwen3-14B-Q4_K_L.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-14B-GGUF/blob/main/Qwen_Qwen3-14B-Q4_K_L.gguf) | Q4_K_L | 9.58GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Qwen3-14B-Q4_1.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-14B-GGUF/blob/main/Qwen_Qwen3-14B-Q4_1.gguf) | Q4_1 | 9.39GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| [Qwen3-14B-Q4_K_M.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-14B-GGUF/blob/main/Qwen_Qwen3-14B-Q4_K_M.gguf) | Q4_K_M | 9.00GB | false | Good quality, default size for most use cases, *recommended*. |
| [Qwen3-14B-Q3_K_XL.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-14B-GGUF/blob/main/Qwen_Qwen3-14B-Q3_K_XL.gguf) | Q3_K_XL | 8.58GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Qwen3-14B-Q4_K_S.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-14B-GGUF/blob/main/Qwen_Qwen3-14B-Q4_K_S.gguf) | Q4_K_S | 8.57GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Qwen3-14B-Q4_0.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-14B-GGUF/blob/main/Qwen_Qwen3-14B-Q4_0.gguf) | Q4_0 | 8.54GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| [Qwen3-14B-IQ4_NL.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-14B-GGUF/blob/main/Qwen_Qwen3-14B-IQ4_NL.gguf) | IQ4_NL | 8.54GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [Qwen3-14B-IQ4_XS.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-14B-GGUF/blob/main/Qwen_Qwen3-14B-IQ4_XS.gguf) | IQ4_XS | 8.11GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Qwen3-14B-Q3_K_L.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-14B-GGUF/blob/main/Qwen_Qwen3-14B-Q3_K_L.gguf) | Q3_K_L | 7.90GB | false | Lower quality but usable, good for low RAM availability. |
| [Qwen3-14B-Q3_K_M.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-14B-GGUF/blob/main/Qwen_Qwen3-14B-Q3_K_M.gguf) | Q3_K_M | 7.32GB | false | Low quality. |
| [Qwen3-14B-IQ3_M.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-14B-GGUF/blob/main/Qwen_Qwen3-14B-IQ3_M.gguf) | IQ3_M | 6.88GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Qwen3-14B-Q3_K_S.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-14B-GGUF/blob/main/Qwen_Qwen3-14B-Q3_K_S.gguf) | Q3_K_S | 6.66GB | false | Low quality, not recommended. |
| [Qwen3-14B-Q2_K_L.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-14B-GGUF/blob/main/Qwen_Qwen3-14B-Q2_K_L.gguf) | Q2_K_L | 6.51GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Qwen3-14B-IQ3_XS.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-14B-GGUF/blob/main/Qwen_Qwen3-14B-IQ3_XS.gguf) | IQ3_XS | 6.38GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Qwen3-14B-IQ3_XXS.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-14B-GGUF/blob/main/Qwen_Qwen3-14B-IQ3_XXS.gguf) | IQ3_XXS | 5.94GB | false | Lower quality, new method with decent performance, comparable to Q3 quants. |
| [Qwen3-14B-Q2_K.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-14B-GGUF/blob/main/Qwen_Qwen3-14B-Q2_K.gguf) | Q2_K | 5.75GB | false | Very low quality but surprisingly usable. |
| [Qwen3-14B-IQ2_M.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-14B-GGUF/blob/main/Qwen_Qwen3-14B-IQ2_M.gguf) | IQ2_M | 5.32GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
| [Qwen3-14B-IQ2_S.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-14B-GGUF/blob/main/Qwen_Qwen3-14B-IQ2_S.gguf) | IQ2_S | 4.96GB | false | Low quality, uses SOTA techniques to be usable. |
| [Qwen3-14B-IQ2_XS.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-14B-GGUF/blob/main/Qwen_Qwen3-14B-IQ2_XS.gguf) | IQ2_XS | 4.69GB | false | Low quality, uses SOTA techniques to be usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Qwen_Qwen3-14B-GGUF --include "Qwen_Qwen3-14B-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Qwen_Qwen3-14B-GGUF --include "Qwen_Qwen3-14B-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Qwen_Qwen3-14B-Q8_0) or download them all in place (./)
</details>
## ARM/AVX information
Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.
Now, however, there is something called "online repacking" for weights. details in [this PR](https://github.com/ggerganov/llama.cpp/pull/9921). If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.
As of llama.cpp build [b4282](https://github.com/ggerganov/llama.cpp/releases/tag/b4282) you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.
Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to [this PR](https://github.com/ggerganov/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information (deprecated</summary>
I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Thank you to LM Studio for sponsoring my work.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}