初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Qwen2.5-Coder-14B-Instruct-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-12 11:28:06 +08:00
commit ae1e300a0d
31 changed files with 285 additions and 0 deletions

63
.gitattributes vendored Normal file
View File

@@ -0,0 +1,63 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-14B-Instruct-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-14B-Instruct-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-14B-Instruct-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-14B-Instruct-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-14B-Instruct-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-14B-Instruct-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-14B-Instruct-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-14B-Instruct-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-14B-Instruct-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-14B-Instruct-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-14B-Instruct-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-14B-Instruct-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-14B-Instruct-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-14B-Instruct-IQ4_NL.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-14B-Instruct-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-14B-Instruct-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-14B-Instruct-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-14B-Instruct-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-14B-Instruct-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-14B-Instruct-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-14B-Instruct-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-14B-Instruct-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-14B-Instruct-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-14B-Instruct-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-14B-Instruct-IQ2_S.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-14B-Instruct-IQ2_XS.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-14B-Instruct-f16.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-14B-Instruct.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:aa09df429d40831720204f5e037ddd96317f733f70a0377ad01fa9e86be43fa7
size 5356146912

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d160f97228feca84fe63c4f97640ef106e38f393d0158bee3d389e4e0383e37b
size 5003727072

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cc80bc1705a4fed1a1e3d8bc61e26d8019f74e549976eb6cdf267fcdc9054c51
size 4704575712

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:00d67af6337b7b9c5ddcb196ff2f8f1d4517636060763cba6fb5cc999642da11
size 6916538592

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:57e5a8eae6aa33ccf23516dde9f8e4eb70160af0f389671e6ee97377382152bb
size 6383362272

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:20f0cdeefed18f0b9c465c705f100d65b3cb74427e2872c8b32d751954c00b16
size 8549183712

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b412bc590b7766b53c6966df5613403746d0d14a4510e9e12dac49e9f370c5d7
size 8119840992

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:15325932435e5efc33120a4199963298bca521b961a60ea050eeba7543acab9d
size 5770498272

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a0bc58efadb8f89c2c85a9b17e8fb9ab16ecc986f929a4b5ab9fdc4b8e08a946
size 6530818272

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:02cea12488ffa7b9d00d53c8342799c3f2ba3de252d722613931e84c56689795
size 7924768992

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2eda43d86ba2670e73b2d3b1d819054124d4f266270a8242a7cd12c409903004
size 7339204832

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:34bfe17406ee6b5b2f1b1d4842a4d29795134b91737bcc2c2708ed706763c055
size 6659596512

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2844973456fdef961efa8089b38db174e43a48873f0592fb92eef6455c3e68a2
size 8606015712

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b7e3cdb603ca21b058119802d89c0242ea2105250eb69a28a2b9855c4fa98146
size 8544268512

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:87a186c68da1c17d5827733149181a9e05cefea1cba9d9dde6d258702bab4d45
size 8517726432

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1b936b53a8416cceb19ac7aaae14ad4edf42bdd47f86897fed09e8c37fe36ebe
size 8517726432

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f692aca58d2ac8d9a713d688a1ca4abf26e424331403ea41a6fbd6f00ed70801
size 8517726432

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ea5baf3f6a058f0c39d25e0db8e69b17e024ae64da064535de89b04c88305379
size 9565954272

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2946d28c9e1bb2bcae6d42e8678863a31775df6f740315c7d7e6d6b6411f5937
size 8988111072

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5ce35db75ab313d4a84da6f3f97cb892859bf27e3520991a3b38b28f0b14827f
size 8573432032

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:97b54f2e7d6443ecad08aed97af76b46a95df853fb2e53f75bf260176e294729
size 10989396192

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:68701e9ffaab69e42827524f604bf71b375b7ed8f0c59f6a8c3d646c963b3e2c
size 10508873952

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:467f23a2e05c9ccba102bb9a85254239992f26e4b27adcf389e9308b2560b22c
size 10266554592

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b576e13fe5c36652c9277c2e649a4950bfb8bac04eb63417a512cc902dc032d5
size 12124684512

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6629be0a282d71d8d05d760198f763b2fd7a004d88c4a8b1646e8e39267069f0
size 12501803232

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ceed8d9126648af71bb5b3d1e69c43b25689752ea2c42204a48fea202605adc5
size 15701598432

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:db7dbb7f767612141d005e1434f3444aa9c9634da63c56439a37f9e8ff7cf5f6
size 29547716544

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2eb86d936201a83f79f7fd8bb01844d0e31d75522d5d4f1c93fff74c786895cc
size 8563610

137
README.md Normal file
View File

@@ -0,0 +1,137 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
language:
- en
license_link: https://huggingface.co/Qwen/Qwen2.5-Coder-14B-Instruct/blob/main/LICENSE
tags:
- code
- codeqwen
- chat
- qwen
- qwen-coder
base_model: Qwen/Qwen2.5-Coder-14B-Instruct
license: apache-2.0
---
## Llamacpp imatrix Quantizations of Qwen2.5-Coder-14B-Instruct
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4014">b4014</a> for quantization.
Original model: https://huggingface.co/Qwen/Qwen2.5-Coder-14B-Instruct
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Qwen2.5-Coder-14B-Instruct-f16.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-14B-Instruct-GGUF/blob/main/Qwen2.5-Coder-14B-Instruct-f16.gguf) | f16 | 29.55GB | false | Full F16 weights. |
| [Qwen2.5-Coder-14B-Instruct-Q8_0.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-14B-Instruct-GGUF/blob/main/Qwen2.5-Coder-14B-Instruct-Q8_0.gguf) | Q8_0 | 15.70GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Qwen2.5-Coder-14B-Instruct-Q6_K_L.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-14B-Instruct-GGUF/blob/main/Qwen2.5-Coder-14B-Instruct-Q6_K_L.gguf) | Q6_K_L | 12.50GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Qwen2.5-Coder-14B-Instruct-Q6_K.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-14B-Instruct-GGUF/blob/main/Qwen2.5-Coder-14B-Instruct-Q6_K.gguf) | Q6_K | 12.12GB | false | Very high quality, near perfect, *recommended*. |
| [Qwen2.5-Coder-14B-Instruct-Q5_K_L.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-14B-Instruct-GGUF/blob/main/Qwen2.5-Coder-14B-Instruct-Q5_K_L.gguf) | Q5_K_L | 10.99GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Qwen2.5-Coder-14B-Instruct-Q5_K_M.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-14B-Instruct-GGUF/blob/main/Qwen2.5-Coder-14B-Instruct-Q5_K_M.gguf) | Q5_K_M | 10.51GB | false | High quality, *recommended*. |
| [Qwen2.5-Coder-14B-Instruct-Q5_K_S.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-14B-Instruct-GGUF/blob/main/Qwen2.5-Coder-14B-Instruct-Q5_K_S.gguf) | Q5_K_S | 10.27GB | false | High quality, *recommended*. |
| [Qwen2.5-Coder-14B-Instruct-Q4_K_L.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-14B-Instruct-GGUF/blob/main/Qwen2.5-Coder-14B-Instruct-Q4_K_L.gguf) | Q4_K_L | 9.57GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Qwen2.5-Coder-14B-Instruct-Q4_K_M.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-14B-Instruct-GGUF/blob/main/Qwen2.5-Coder-14B-Instruct-Q4_K_M.gguf) | Q4_K_M | 8.99GB | false | Good quality, default size for most use cases, *recommended*. |
| [Qwen2.5-Coder-14B-Instruct-Q3_K_XL.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-14B-Instruct-GGUF/blob/main/Qwen2.5-Coder-14B-Instruct-Q3_K_XL.gguf) | Q3_K_XL | 8.61GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Qwen2.5-Coder-14B-Instruct-Q4_K_S.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-14B-Instruct-GGUF/blob/main/Qwen2.5-Coder-14B-Instruct-Q4_K_S.gguf) | Q4_K_S | 8.57GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Qwen2.5-Coder-14B-Instruct-IQ4_NL.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-14B-Instruct-GGUF/blob/main/Qwen2.5-Coder-14B-Instruct-IQ4_NL.gguf) | IQ4_NL | 8.55GB | false | Similar to IQ4_XS, but slightly larger. |
| [Qwen2.5-Coder-14B-Instruct-Q4_0.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-14B-Instruct-GGUF/blob/main/Qwen2.5-Coder-14B-Instruct-Q4_0.gguf) | Q4_0 | 8.54GB | false | Legacy format, generally not worth using over similarly sized formats |
| [Qwen2.5-Coder-14B-Instruct-Q4_0_8_8.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-14B-Instruct-GGUF/blob/main/Qwen2.5-Coder-14B-Instruct-Q4_0_8_8.gguf) | Q4_0_8_8 | 8.52GB | false | Optimized for ARM inference. Requires 'sve' support (see link below). *Don't use on Mac or Windows*. |
| [Qwen2.5-Coder-14B-Instruct-Q4_0_4_8.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-14B-Instruct-GGUF/blob/main/Qwen2.5-Coder-14B-Instruct-Q4_0_4_8.gguf) | Q4_0_4_8 | 8.52GB | false | Optimized for ARM inference. Requires 'i8mm' support (see link below). *Don't use on Mac or Windows*. |
| [Qwen2.5-Coder-14B-Instruct-Q4_0_4_4.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-14B-Instruct-GGUF/blob/main/Qwen2.5-Coder-14B-Instruct-Q4_0_4_4.gguf) | Q4_0_4_4 | 8.52GB | false | Optimized for ARM inference. Should work well on all ARM chips, pick this if you're unsure. *Don't use on Mac or Windows*. |
| [Qwen2.5-Coder-14B-Instruct-IQ4_XS.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-14B-Instruct-GGUF/blob/main/Qwen2.5-Coder-14B-Instruct-IQ4_XS.gguf) | IQ4_XS | 8.12GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Qwen2.5-Coder-14B-Instruct-Q3_K_L.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-14B-Instruct-GGUF/blob/main/Qwen2.5-Coder-14B-Instruct-Q3_K_L.gguf) | Q3_K_L | 7.92GB | false | Lower quality but usable, good for low RAM availability. |
| [Qwen2.5-Coder-14B-Instruct-Q3_K_M.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-14B-Instruct-GGUF/blob/main/Qwen2.5-Coder-14B-Instruct-Q3_K_M.gguf) | Q3_K_M | 7.34GB | false | Low quality. |
| [Qwen2.5-Coder-14B-Instruct-IQ3_M.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-14B-Instruct-GGUF/blob/main/Qwen2.5-Coder-14B-Instruct-IQ3_M.gguf) | IQ3_M | 6.92GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Qwen2.5-Coder-14B-Instruct-Q3_K_S.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-14B-Instruct-GGUF/blob/main/Qwen2.5-Coder-14B-Instruct-Q3_K_S.gguf) | Q3_K_S | 6.66GB | false | Low quality, not recommended. |
| [Qwen2.5-Coder-14B-Instruct-Q2_K_L.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-14B-Instruct-GGUF/blob/main/Qwen2.5-Coder-14B-Instruct-Q2_K_L.gguf) | Q2_K_L | 6.53GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Qwen2.5-Coder-14B-Instruct-IQ3_XS.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-14B-Instruct-GGUF/blob/main/Qwen2.5-Coder-14B-Instruct-IQ3_XS.gguf) | IQ3_XS | 6.38GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Qwen2.5-Coder-14B-Instruct-Q2_K.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-14B-Instruct-GGUF/blob/main/Qwen2.5-Coder-14B-Instruct-Q2_K.gguf) | Q2_K | 5.77GB | false | Very low quality but surprisingly usable. |
| [Qwen2.5-Coder-14B-Instruct-IQ2_M.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-14B-Instruct-GGUF/blob/main/Qwen2.5-Coder-14B-Instruct-IQ2_M.gguf) | IQ2_M | 5.36GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
| [Qwen2.5-Coder-14B-Instruct-IQ2_S.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-14B-Instruct-GGUF/blob/main/Qwen2.5-Coder-14B-Instruct-IQ2_S.gguf) | IQ2_S | 5.00GB | false | Low quality, uses SOTA techniques to be usable. |
| [Qwen2.5-Coder-14B-Instruct-IQ2_XS.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-14B-Instruct-GGUF/blob/main/Qwen2.5-Coder-14B-Instruct-IQ2_XS.gguf) | IQ2_XS | 4.70GB | false | Low quality, uses SOTA techniques to be usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Qwen2.5-Coder-14B-Instruct-GGUF --include "Qwen2.5-Coder-14B-Instruct-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Qwen2.5-Coder-14B-Instruct-GGUF --include "Qwen2.5-Coder-14B-Instruct-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Qwen2.5-Coder-14B-Instruct-Q8_0) or download them all in place (./)
## Q4_0_X_X
These are *NOT* for Metal (Apple) offloading, only ARM chips.
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}