初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Qwen2.5-Coder-1.5B-Instruct-abliterated-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-05 10:43:12 +08:00
commit 0dbc9bcabf
28 changed files with 268 additions and 0 deletions

60
.gitattributes vendored Normal file
View File

@@ -0,0 +1,60 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-1.5B-Instruct-abliterated-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-1.5B-Instruct-abliterated-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-1.5B-Instruct-abliterated-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-1.5B-Instruct-abliterated-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-1.5B-Instruct-abliterated-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-1.5B-Instruct-abliterated-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-1.5B-Instruct-abliterated-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-1.5B-Instruct-abliterated-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-1.5B-Instruct-abliterated-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-1.5B-Instruct-abliterated-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-1.5B-Instruct-abliterated-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-1.5B-Instruct-abliterated-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-1.5B-Instruct-abliterated-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-1.5B-Instruct-abliterated-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-1.5B-Instruct-abliterated-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-1.5B-Instruct-abliterated-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-1.5B-Instruct-abliterated-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-1.5B-Instruct-abliterated-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-1.5B-Instruct-abliterated-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-1.5B-Instruct-abliterated-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-1.5B-Instruct-abliterated-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-1.5B-Instruct-abliterated-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-1.5B-Instruct-abliterated-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-1.5B-Instruct-abliterated-f16.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-Coder-1.5B-Instruct-abliterated.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ff3eeb97756a6c173aff59478017a4e23da6c757424aef3f40b6fe8be7d46738
size 701332960

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0b47fd670654ce755f3af4cbad04f52acc4bf842b2c321baddec5652c9cbb629
size 876942304

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7cb0c9512962657bcb1c6265393ea925699b52e49ee8296abfceb7ead0844be3
size 831977440

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:18112952269dd83b2b04a20895de66e475c2cac4831ddbe75341b3182fffee39
size 1019711968

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9a7b865ad68e2cce58b5bed9c6fc65096c0a88b1eceb5a12f55005ffc75c2d1f
size 752881120

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fdf40566806b55fb2c662da43f1735c0f53d2365f9e131d1509e3b7ab5457947
size 980785120

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:64b7a71c96b80526e5d88ffcc3ab9f0f94f8a24f4e138f09bfcfcb449a7c5f80
size 980441056

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1e7d20fc77b8e22e43f11d86246a4858ce97d65fe54bbd80b53814adf6497cac
size 924456928

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2388c887e5b183b1740d5d5a12a8d46e96e16cc456945b2c22e4c7e773f3990b
size 861222880

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2a257d732e978495f6de517717d4dde7a8e55b1c53b572fc8b9212e9fe3a3c46
size 1184643040

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2ed14a569301b5eb82bacaf710c84539f94916d30dec9cd73d1f8294a2c10982
size 1068808672

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:25523476a159e73c61e5eab9f492c149b2a4db54f89fdb80d0719a11bfe5b6f3
size 1066228192

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:22d55e6d7dcc860c228f830db0c57d20a2d4e2e92e45c285186f4b04ff3d283b
size 1066228192

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ce2aea3d4712b81d69e96db3d34ceffdde4a7410b405aa2a234dbf1ace0fc427
size 1066228192

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:108902ab59d8988d6efa27b06eb0675dcbf52b8209c7b49332172e2f5c003535
size 1290528736

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:389b8172028c874b79a11358d258222dbce272b56ab3512ccc743893a25e06b4
size 1117321696

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5593d73b1f6664bd54bac17a7340c571c27adbecb51200eedaf9c9f269e53e2d
size 1071585760

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6044e7fdbde4e016f691ccc9f5108bc2b1223ec841f0eac8edbe0d6637c9c429
size 1429530592

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:12a8ae418ab50c59cd1b5bdbc1d8ab7a30c6239de8e1f622e97d932bc3fc7f04
size 1285495264

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3c69d6f744b26a13b212822a752999229f5514adaf8c77879329399dc332bfbb
size 1259174368

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:36062301cb5ff47837113c616cf7f9983c84434031b44c284d5d125760bda153
size 1464179680

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9f584bcf5e0ccdef12063bd4138d7f809d40c20afd6231ab41ae064b63ceb414
size 1577220064

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:15492089c47f1c36f3a548c067ba20264e377d54dd559292bfaed39eb33bf0d4
size 1894533088

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:52c1f3b872adee66fe4ae97e0fb85167912287135d6548791c311937ebd57456
size 3560416928

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a79302b507a93fa7b118c6a87995cdee21ff3aad3774619c07a27dc020845acf
size 2042214

132
README.md Normal file
View File

@@ -0,0 +1,132 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
language:
- en
license_link: https://huggingface.co/huihui-ai/Qwen2.5-Coder-1.5B-Instruct-abliterated/blob/main/LICENSE
tags:
- chat
- abliterated
- uncensored
base_model: huihui-ai/Qwen2.5-Coder-1.5B-Instruct-abliterated
license: apache-2.0
---
## Llamacpp imatrix Quantizations of Qwen2.5-Coder-1.5B-Instruct-abliterated
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4058">b4058</a> for quantization.
Original model: https://huggingface.co/huihui-ai/Qwen2.5-Coder-1.5B-Instruct-abliterated
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Qwen2.5-Coder-1.5B-Instruct-abliterated-f16.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-1.5B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-1.5B-Instruct-abliterated-f16.gguf) | f16 | 3.56GB | false | Full F16 weights. |
| [Qwen2.5-Coder-1.5B-Instruct-abliterated-Q8_0.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-1.5B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-1.5B-Instruct-abliterated-Q8_0.gguf) | Q8_0 | 1.89GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Qwen2.5-Coder-1.5B-Instruct-abliterated-Q6_K_L.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-1.5B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-1.5B-Instruct-abliterated-Q6_K_L.gguf) | Q6_K_L | 1.58GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Qwen2.5-Coder-1.5B-Instruct-abliterated-Q6_K.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-1.5B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-1.5B-Instruct-abliterated-Q6_K.gguf) | Q6_K | 1.46GB | false | Very high quality, near perfect, *recommended*. |
| [Qwen2.5-Coder-1.5B-Instruct-abliterated-Q5_K_L.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-1.5B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-1.5B-Instruct-abliterated-Q5_K_L.gguf) | Q5_K_L | 1.43GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Qwen2.5-Coder-1.5B-Instruct-abliterated-Q5_K_M.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-1.5B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-1.5B-Instruct-abliterated-Q5_K_M.gguf) | Q5_K_M | 1.29GB | false | High quality, *recommended*. |
| [Qwen2.5-Coder-1.5B-Instruct-abliterated-Q4_K_L.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-1.5B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-1.5B-Instruct-abliterated-Q4_K_L.gguf) | Q4_K_L | 1.29GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Qwen2.5-Coder-1.5B-Instruct-abliterated-Q5_K_S.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-1.5B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-1.5B-Instruct-abliterated-Q5_K_S.gguf) | Q5_K_S | 1.26GB | false | High quality, *recommended*. |
| [Qwen2.5-Coder-1.5B-Instruct-abliterated-Q3_K_XL.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-1.5B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-1.5B-Instruct-abliterated-Q3_K_XL.gguf) | Q3_K_XL | 1.18GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Qwen2.5-Coder-1.5B-Instruct-abliterated-Q4_K_M.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-1.5B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-1.5B-Instruct-abliterated-Q4_K_M.gguf) | Q4_K_M | 1.12GB | false | Good quality, default size for most use cases, *recommended*. |
| [Qwen2.5-Coder-1.5B-Instruct-abliterated-Q4_K_S.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-1.5B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-1.5B-Instruct-abliterated-Q4_K_S.gguf) | Q4_K_S | 1.07GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Qwen2.5-Coder-1.5B-Instruct-abliterated-Q4_0_8_8.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-1.5B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-1.5B-Instruct-abliterated-Q4_0_8_8.gguf) | Q4_0_8_8 | 1.07GB | false | Optimized for ARM inference. Requires 'sve' support (see link below). *Don't use on Mac or Windows*. |
| [Qwen2.5-Coder-1.5B-Instruct-abliterated-Q4_0_4_8.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-1.5B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-1.5B-Instruct-abliterated-Q4_0_4_8.gguf) | Q4_0_4_8 | 1.07GB | false | Optimized for ARM inference. Requires 'i8mm' support (see link below). *Don't use on Mac or Windows*. |
| [Qwen2.5-Coder-1.5B-Instruct-abliterated-Q4_0_4_4.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-1.5B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-1.5B-Instruct-abliterated-Q4_0_4_4.gguf) | Q4_0_4_4 | 1.07GB | false | Optimized for ARM inference. Should work well on all ARM chips, pick this if you're unsure. *Don't use on Mac or Windows*. |
| [Qwen2.5-Coder-1.5B-Instruct-abliterated-Q4_0.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-1.5B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-1.5B-Instruct-abliterated-Q4_0.gguf) | Q4_0 | 1.07GB | false | Legacy format, generally not worth using over similarly sized formats |
| [Qwen2.5-Coder-1.5B-Instruct-abliterated-IQ4_XS.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-1.5B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-1.5B-Instruct-abliterated-IQ4_XS.gguf) | IQ4_XS | 1.02GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Qwen2.5-Coder-1.5B-Instruct-abliterated-Q3_K_L.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-1.5B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-1.5B-Instruct-abliterated-Q3_K_L.gguf) | Q3_K_L | 0.98GB | false | Lower quality but usable, good for low RAM availability. |
| [Qwen2.5-Coder-1.5B-Instruct-abliterated-Q2_K_L.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-1.5B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-1.5B-Instruct-abliterated-Q2_K_L.gguf) | Q2_K_L | 0.98GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Qwen2.5-Coder-1.5B-Instruct-abliterated-Q3_K_M.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-1.5B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-1.5B-Instruct-abliterated-Q3_K_M.gguf) | Q3_K_M | 0.92GB | false | Low quality. |
| [Qwen2.5-Coder-1.5B-Instruct-abliterated-IQ3_M.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-1.5B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-1.5B-Instruct-abliterated-IQ3_M.gguf) | IQ3_M | 0.88GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Qwen2.5-Coder-1.5B-Instruct-abliterated-Q3_K_S.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-1.5B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-1.5B-Instruct-abliterated-Q3_K_S.gguf) | Q3_K_S | 0.86GB | false | Low quality, not recommended. |
| [Qwen2.5-Coder-1.5B-Instruct-abliterated-IQ3_XS.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-1.5B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-1.5B-Instruct-abliterated-IQ3_XS.gguf) | IQ3_XS | 0.83GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Qwen2.5-Coder-1.5B-Instruct-abliterated-Q2_K.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-1.5B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-1.5B-Instruct-abliterated-Q2_K.gguf) | Q2_K | 0.75GB | false | Very low quality but surprisingly usable. |
| [Qwen2.5-Coder-1.5B-Instruct-abliterated-IQ2_M.gguf](https://huggingface.co/bartowski/Qwen2.5-Coder-1.5B-Instruct-abliterated-GGUF/blob/main/Qwen2.5-Coder-1.5B-Instruct-abliterated-IQ2_M.gguf) | IQ2_M | 0.70GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Qwen2.5-Coder-1.5B-Instruct-abliterated-GGUF --include "Qwen2.5-Coder-1.5B-Instruct-abliterated-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Qwen2.5-Coder-1.5B-Instruct-abliterated-GGUF --include "Qwen2.5-Coder-1.5B-Instruct-abliterated-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Qwen2.5-Coder-1.5B-Instruct-abliterated-Q8_0) or download them all in place (./)
## Q4_0_X_X
These are *NOT* for Metal (Apple) offloading, only ARM chips.
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}