初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-25 06:40:06 +08:00
commit 473ec7d905
31 changed files with 286 additions and 0 deletions

63
.gitattributes vendored Normal file
View File

@@ -0,0 +1,63 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-32B-Instruct-abliterated-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-32B-Instruct-abliterated-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-32B-Instruct-abliterated-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-32B-Instruct-abliterated-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-32B-Instruct-abliterated-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-32B-Instruct-abliterated-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-32B-Instruct-abliterated-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-32B-Instruct-abliterated-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-32B-Instruct-abliterated-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-32B-Instruct-abliterated-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-32B-Instruct-abliterated-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-32B-Instruct-abliterated-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-32B-Instruct-abliterated-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-32B-Instruct-abliterated-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-32B-Instruct-abliterated-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-32B-Instruct-abliterated-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-32B-Instruct-abliterated-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-32B-Instruct-abliterated-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-32B-Instruct-abliterated-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-32B-Instruct-abliterated-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-32B-Instruct-abliterated-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-32B-Instruct-abliterated-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-32B-Instruct-abliterated-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-32B-Instruct-abliterated-IQ2_S.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-32B-Instruct-abliterated-IQ2_XS.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-32B-Instruct-abliterated-f16/Qwen2.5-Coder-32B-Instruct-abliterated-f16-00001-of-00002.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-32B-Instruct-abliterated-f16/Qwen2.5-Coder-32B-Instruct-abliterated-f16-00002-of-00002.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-32B-Instruct-abliterated.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e642932f8ecb95cc4143a7054fa46819af59d05f6dfb39599e93b846ab96630d
size 11264441440

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c36129c843faec68154fd18bb5fdee2b70823dbf17085170f01afda985e1da86
size 10387569760

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c2018ec26556f512ebb2f124f12aad002dbf35bf0c6f23d66125fbca209c5840
size 9957551200

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c39e98b481836f066c01922924d2cbba90449efa3f4f160d874e150ce7e139e0
size 14810123360

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0b0967d04516f60ab572b21c32ebca0e7d54a94f5b3ba7a498764e1d52e02fbe
size 13705514080

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0859ac53fe46f32e058eaa28e541c0c29aed9155f3dcaac8c2156e5b91384d8d
size 17693154400

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:97921eb2e9de9c7fac6d683e911706f2d47fea2a8824101de7bce01796e1bdda
size 12313099360

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:302b62c4d3161400202ce13e5499c944208197d7250c9fa059ab738801e4f0b1
size 13073419360

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:92393aaa2916c55f53c81e82622f3f7efca067089f1505ba8a86b5ac442f7dcd
size 17247079520

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fc39bca9fd1019dc25162c2061ea2070f8e605f12adbb7b9427a84f9afdd5450
size 15935048800

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:30ecc831a87899d251d13ab0824dd47aa17c7ed6f44f433484f59324e1fa4f12
size 14392331360

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4e62229ec989295353f6c8f220cd810def8dfbf0cec6cbb87f72dec6ecec5b9b
size 17928326240

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e0ecfa5c9aedc23c2c58ff436ed6d698c623c760a6cdbc672b803ca7a119223d
size 18711010400

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8661260aefa85e4f1c09230733c3f1406d51722ab960ac04751b39ffe8b72feb
size 18640231520

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f38d57f7daf7100545bbebc6c16399f8fb1eecddbf4acc9195220b17d5b6f6ef
size 18640231520

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0cb24726949c16f9c14631037ad30293c12ff924cacc4a4113af5dfec371d794
size 18640231520

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:dee2e1a35f3d48742b29fa26aa4e12d079b10a2a773426ab132c3c8e9b9378fd
size 20429180000

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:54fd911daea5c940acd0f275fce19ce4f5bfd900ea03a7764ab33030f4d4122d
size 19851336800

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b338b8c0a4eefac80182b194e7b7737d047153c4f053e4123237d8918d4e37b1
size 18784410720

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d65f7d871319742715b4abda586bcc379dc4c7644cef9bf2341742a579d120e2
size 23742680160

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3818dcbf7bb6a24561f87e404ab2d33ec904a32565229e976c919d1c9fd73db9
size 23262157920

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7f36d3f01c7fc01cbcdb3b72b22f380bf9051cf3e9bda6b1c87a81274e34b335
size 22638255200

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:45bd59bb45ba522033803ad20d843a88c995131c75d063648f9a1f00f117736c
size 26886155360

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e18303a40c114bfcb68f227e373f0e5144e74fe2555a5c433385e6d78c74eafb
size 27263274080

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:73cf5dc4cb5b47bc98251d614c7b40bf03e0a996cd1bfe7c6c13854fde6966cb
size 34820885280

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f1e616ef6452c0477a0267b7872206fbeafab7533829895f84a6b749f21d7ecc
size 39880800320

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:42f15f35ebf76c9bbc1381372e191d8f4c95cebe2c0d840be790b3b044eca55d
size 25655169952

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:af1822f7ce3437156355ddd98c3ed82e1bc8270da57a92ffe739eae488ed315c
size 14957098

138
README.md Normal file
View File

@@ -0,0 +1,138 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
language:
- en
license_link: https://huggingface.co/huihui-ai/Qwen2.5-Coder-32B-Instruct-abliterate/blob/main/LICENSE
tags:
- code
- codeqwen
- chat
- qwen
- qwen-coder
- abliterated
- uncensored
base_model: huihui-ai/Qwen2.5-Coder-32B-Instruct-abliterated
license: apache-2.0
---
## Llamacpp imatrix Quantizations of Qwen2.5-Coder-32B-Instruct-abliterated
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4058">b4058</a> for quantization.
Original model: https://huggingface.co/huihui-ai/Qwen2.5-Coder-32B-Instruct-abliterated
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Qwen2.5-Coder-32B-Instruct-abliterated-f16.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF/tree/main/Qwen2.5-Coder-32B-Instruct-abliterated-f16) | f16 | 65.54GB | true | Full F16 weights. |
| [Qwen2.5-Coder-32B-Instruct-abliterated-Q8_0.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-32B-Instruct-abliterated-Q8_0.gguf) | Q8_0 | 34.82GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Qwen2.5-Coder-32B-Instruct-abliterated-Q6_K_L.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-32B-Instruct-abliterated-Q6_K_L.gguf) | Q6_K_L | 27.26GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Qwen2.5-Coder-32B-Instruct-abliterated-Q6_K.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-32B-Instruct-abliterated-Q6_K.gguf) | Q6_K | 26.89GB | false | Very high quality, near perfect, *recommended*. |
| [Qwen2.5-Coder-32B-Instruct-abliterated-Q5_K_L.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-32B-Instruct-abliterated-Q5_K_L.gguf) | Q5_K_L | 23.74GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Qwen2.5-Coder-32B-Instruct-abliterated-Q5_K_M.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-32B-Instruct-abliterated-Q5_K_M.gguf) | Q5_K_M | 23.26GB | false | High quality, *recommended*. |
| [Qwen2.5-Coder-32B-Instruct-abliterated-Q5_K_S.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-32B-Instruct-abliterated-Q5_K_S.gguf) | Q5_K_S | 22.64GB | false | High quality, *recommended*. |
| [Qwen2.5-Coder-32B-Instruct-abliterated-Q4_K_L.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-32B-Instruct-abliterated-Q4_K_L.gguf) | Q4_K_L | 20.43GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Qwen2.5-Coder-32B-Instruct-abliterated-Q4_K_M.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-32B-Instruct-abliterated-Q4_K_M.gguf) | Q4_K_M | 19.85GB | false | Good quality, default size for most use cases, *recommended*. |
| [Qwen2.5-Coder-32B-Instruct-abliterated-Q4_K_S.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-32B-Instruct-abliterated-Q4_K_S.gguf) | Q4_K_S | 18.78GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Qwen2.5-Coder-32B-Instruct-abliterated-Q4_0.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-32B-Instruct-abliterated-Q4_0.gguf) | Q4_0 | 18.71GB | false | Legacy format, generally not worth using over similarly sized formats |
| [Qwen2.5-Coder-32B-Instruct-abliterated-Q4_0_8_8.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-32B-Instruct-abliterated-Q4_0_8_8.gguf) | Q4_0_8_8 | 18.64GB | false | Optimized for ARM inference. Requires 'sve' support (see link below). *Don't use on Mac or Windows*. |
| [Qwen2.5-Coder-32B-Instruct-abliterated-Q4_0_4_8.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-32B-Instruct-abliterated-Q4_0_4_8.gguf) | Q4_0_4_8 | 18.64GB | false | Optimized for ARM inference. Requires 'i8mm' support (see link below). *Don't use on Mac or Windows*. |
| [Qwen2.5-Coder-32B-Instruct-abliterated-Q4_0_4_4.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-32B-Instruct-abliterated-Q4_0_4_4.gguf) | Q4_0_4_4 | 18.64GB | false | Optimized for ARM inference. Should work well on all ARM chips, pick this if you're unsure. *Don't use on Mac or Windows*. |
| [Qwen2.5-Coder-32B-Instruct-abliterated-Q3_K_XL.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-32B-Instruct-abliterated-Q3_K_XL.gguf) | Q3_K_XL | 17.93GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Qwen2.5-Coder-32B-Instruct-abliterated-IQ4_XS.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-32B-Instruct-abliterated-IQ4_XS.gguf) | IQ4_XS | 17.69GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Qwen2.5-Coder-32B-Instruct-abliterated-Q3_K_L.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-32B-Instruct-abliterated-Q3_K_L.gguf) | Q3_K_L | 17.25GB | false | Lower quality but usable, good for low RAM availability. |
| [Qwen2.5-Coder-32B-Instruct-abliterated-Q3_K_M.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-32B-Instruct-abliterated-Q3_K_M.gguf) | Q3_K_M | 15.94GB | false | Low quality. |
| [Qwen2.5-Coder-32B-Instruct-abliterated-IQ3_M.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-32B-Instruct-abliterated-IQ3_M.gguf) | IQ3_M | 14.81GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Qwen2.5-Coder-32B-Instruct-abliterated-Q3_K_S.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-32B-Instruct-abliterated-Q3_K_S.gguf) | Q3_K_S | 14.39GB | false | Low quality, not recommended. |
| [Qwen2.5-Coder-32B-Instruct-abliterated-IQ3_XS.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-32B-Instruct-abliterated-IQ3_XS.gguf) | IQ3_XS | 13.71GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Qwen2.5-Coder-32B-Instruct-abliterated-Q2_K_L.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-32B-Instruct-abliterated-Q2_K_L.gguf) | Q2_K_L | 13.07GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Qwen2.5-Coder-32B-Instruct-abliterated-Q2_K.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-32B-Instruct-abliterated-Q2_K.gguf) | Q2_K | 12.31GB | false | Very low quality but surprisingly usable. |
| [Qwen2.5-Coder-32B-Instruct-abliterated-IQ2_M.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-32B-Instruct-abliterated-IQ2_M.gguf) | IQ2_M | 11.26GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
| [Qwen2.5-Coder-32B-Instruct-abliterated-IQ2_S.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-32B-Instruct-abliterated-IQ2_S.gguf) | IQ2_S | 10.39GB | false | Low quality, uses SOTA techniques to be usable. |
| [Qwen2.5-Coder-32B-Instruct-abliterated-IQ2_XS.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-32B-Instruct-abliterated-IQ2_XS.gguf) | IQ2_XS | 9.96GB | false | Low quality, uses SOTA techniques to be usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF --include "Qwen2.5-Coder-32B-Instruct-abliterated-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Qwen2.5-Coder-32B-Instruct-abliterated-GGUF --include "Qwen2.5-Coder-32B-Instruct-abliterated-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Qwen2.5-Coder-32B-Instruct-abliterated-Q8_0) or download them all in place (./)
## Q4_0_X_X
These are *NOT* for Metal (Apple) offloading, only ARM chips.
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}