初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Qwen2.5.1-Coder-7B-Instruct-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-02 11:58:12 +08:00
commit e856c13426
28 changed files with 274 additions and 0 deletions

60
.gitattributes vendored Normal file
View File

@@ -0,0 +1,60 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Qwen2.5.1-Coder-7B-Instruct-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5.1-Coder-7B-Instruct-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5.1-Coder-7B-Instruct-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5.1-Coder-7B-Instruct-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5.1-Coder-7B-Instruct-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5.1-Coder-7B-Instruct-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5.1-Coder-7B-Instruct-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5.1-Coder-7B-Instruct-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5.1-Coder-7B-Instruct-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5.1-Coder-7B-Instruct-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5.1-Coder-7B-Instruct-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5.1-Coder-7B-Instruct-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5.1-Coder-7B-Instruct-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5.1-Coder-7B-Instruct-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5.1-Coder-7B-Instruct-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5.1-Coder-7B-Instruct-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5.1-Coder-7B-Instruct-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5.1-Coder-7B-Instruct-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5.1-Coder-7B-Instruct-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5.1-Coder-7B-Instruct-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5.1-Coder-7B-Instruct-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5.1-Coder-7B-Instruct-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5.1-Coder-7B-Instruct-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5.1-Coder-7B-Instruct-f16.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5.1-Coder-7B-Instruct.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9a2b5dae5b674e4f6e1a47a568627d24c150a528bf77321ca4f41de6ae24bc5d
size 2780343072

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6e555af3403eb082d4b28bae2948b6731251a7d32665707e632b7cd54d2a10ad
size 3574012704

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e84a9ee698b8f45f960fb858e2960de43d3a24a905dbaf035fede5b8555f050d
size 3346256672

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:bdb211ed825934cc305e3eb7c3465be5998901ad916f7ded5972c2c732ec87dc
size 4218473248

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a9ae3bcac3e9e77cb74dbe749e75cf1018811336dc645e30946b3e97e0ebce30
size 3015940896

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fae1aebb694aa03ce7162bfbd53b5fbcd8020546575aacd351948c088a697a70
size 3548164896

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3c70128b60ccc3e88810b630c61e286433bd6caec91383b5495f3ca48c3b578c
size 4088460064

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9870dc84365657939344de90f563339352f9b2a8e56964d711b827e06a7ff041
size 3808391968

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a57b0a67006a7d26d7df66495c4e240ce3f28e81414cb694cd324d0ac282e40a
size 3492369184

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b30130e46857a6f0e825f197cb7144c1ed24fc6b60084c3eab02724319159815
size 4565332768

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c3682cfbfaf1125f5dab3b54313201eb55b4c2924ccfc92687f2d9aa4702c3b0
size 4444121888

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c3aede667e645dfb4b734c62c5d54732291108f7c915f047d43667d4f7d2c722
size 4431391520

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d6a22b45c702e64a2eb607e8dbaa31257f105ba0a2eae6a2537f80537d19f563
size 4431391520

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:735a842a07b82d8ce38d28ba85e1475046d23c3fc7f8a65ec20fd7dcd21f4914
size 4431391520

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1ecec29948cb85be85b774acebf4c8093a1c34ead9bae7f8f2c53894f8956243
size 5087564576

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5e646e684063d4f3d6512fa9c895b05eaee22ad73b8d19e50518de656ff018e1
size 4683074336

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ca16306b21222078da51236efcc061ad7deebd03c59ce11fb868961f0ec7a74d
size 4457769760

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4fabc47e1248ea3903911ab5da5a27a9283dbeb60789abf44cd9358e857c05a3
size 5781197600

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:94353749827fcf2e807efbe1da38d6e9592e6d571d72f139fda16eb3388e80ac
size 5444832032

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fd716c798c24e7f4f0f69f86571eb71e876f320a708f33a2658c61410a315c11
size 5315177248

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fc063aaf9196bf54acedd9389e4e12410d8dcf801ef125d804249f1e531602b7
size 6254199584

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:110358cc2af007c6278975c90586ca198407391747b4152f67ec1bcee94a6230
size 6518182688

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:61834b88c1a1ce5c277028a98c4a0c94a564210290992a7ba301bbef96ef8eba
size 8098525984

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:04ab04d1fe44110d2438bb15ecaab45594498f6ecb6e704245f86a73de0696cf
size 15237853696

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fa7f9917eaa44fa0249df56f6e5e1dcd2a67430dad1aba3335701d729d76cda5
size 4536678

138
README.md Normal file
View File

@@ -0,0 +1,138 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
license: apache-2.0
base_model: Qwen/Qwen2.5-Coder-7B-Instruct
language:
- en
tags:
- code
- codeqwen
- chat
- qwen
- qwen-coder
license_link: https://huggingface.co/Qwen/Qwen2.5-Coder-7B-Instruct/blob/main/LICENSE
---
## Llamacpp imatrix Quantizations of Qwen2.5.1-Coder-7B-Instruct
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4014">b4014</a> for quantization.
Original model: https://huggingface.co/Qwen/Qwen2.5-Coder-7B-Instruct
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
## What's new:
New weights uploaded in place
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Qwen2.5.1-Coder-7B-Instruct-f16.gguf](https://huggingface.co/bartowski/Qwen2.5.1-Coder-7B-Instruct-GGUF/blob/main/Qwen2.5.1-Coder-7B-Instruct-f16.gguf) | f16 | 15.24GB | false | Full F16 weights. |
| [Qwen2.5.1-Coder-7B-Instruct-Q8_0.gguf](https://huggingface.co/bartowski/Qwen2.5.1-Coder-7B-Instruct-GGUF/blob/main/Qwen2.5.1-Coder-7B-Instruct-Q8_0.gguf) | Q8_0 | 8.10GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Qwen2.5.1-Coder-7B-Instruct-Q6_K_L.gguf](https://huggingface.co/bartowski/Qwen2.5.1-Coder-7B-Instruct-GGUF/blob/main/Qwen2.5.1-Coder-7B-Instruct-Q6_K_L.gguf) | Q6_K_L | 6.52GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Qwen2.5.1-Coder-7B-Instruct-Q6_K.gguf](https://huggingface.co/bartowski/Qwen2.5.1-Coder-7B-Instruct-GGUF/blob/main/Qwen2.5.1-Coder-7B-Instruct-Q6_K.gguf) | Q6_K | 6.25GB | false | Very high quality, near perfect, *recommended*. |
| [Qwen2.5.1-Coder-7B-Instruct-Q5_K_L.gguf](https://huggingface.co/bartowski/Qwen2.5.1-Coder-7B-Instruct-GGUF/blob/main/Qwen2.5.1-Coder-7B-Instruct-Q5_K_L.gguf) | Q5_K_L | 5.78GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Qwen2.5.1-Coder-7B-Instruct-Q5_K_M.gguf](https://huggingface.co/bartowski/Qwen2.5.1-Coder-7B-Instruct-GGUF/blob/main/Qwen2.5.1-Coder-7B-Instruct-Q5_K_M.gguf) | Q5_K_M | 5.44GB | false | High quality, *recommended*. |
| [Qwen2.5.1-Coder-7B-Instruct-Q5_K_S.gguf](https://huggingface.co/bartowski/Qwen2.5.1-Coder-7B-Instruct-GGUF/blob/main/Qwen2.5.1-Coder-7B-Instruct-Q5_K_S.gguf) | Q5_K_S | 5.32GB | false | High quality, *recommended*. |
| [Qwen2.5.1-Coder-7B-Instruct-Q4_K_L.gguf](https://huggingface.co/bartowski/Qwen2.5.1-Coder-7B-Instruct-GGUF/blob/main/Qwen2.5.1-Coder-7B-Instruct-Q4_K_L.gguf) | Q4_K_L | 5.09GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Qwen2.5.1-Coder-7B-Instruct-Q4_K_M.gguf](https://huggingface.co/bartowski/Qwen2.5.1-Coder-7B-Instruct-GGUF/blob/main/Qwen2.5.1-Coder-7B-Instruct-Q4_K_M.gguf) | Q4_K_M | 4.68GB | false | Good quality, default size for must use cases, *recommended*. |
| [Qwen2.5.1-Coder-7B-Instruct-Q3_K_XL.gguf](https://huggingface.co/bartowski/Qwen2.5.1-Coder-7B-Instruct-GGUF/blob/main/Qwen2.5.1-Coder-7B-Instruct-Q3_K_XL.gguf) | Q3_K_XL | 4.57GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Qwen2.5.1-Coder-7B-Instruct-Q4_K_S.gguf](https://huggingface.co/bartowski/Qwen2.5.1-Coder-7B-Instruct-GGUF/blob/main/Qwen2.5.1-Coder-7B-Instruct-Q4_K_S.gguf) | Q4_K_S | 4.46GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Qwen2.5.1-Coder-7B-Instruct-Q4_0.gguf](https://huggingface.co/bartowski/Qwen2.5.1-Coder-7B-Instruct-GGUF/blob/main/Qwen2.5.1-Coder-7B-Instruct-Q4_0.gguf) | Q4_0 | 4.44GB | false | Legacy format, generally not worth using over similarly sized formats |
| [Qwen2.5.1-Coder-7B-Instruct-Q4_0_8_8.gguf](https://huggingface.co/bartowski/Qwen2.5.1-Coder-7B-Instruct-GGUF/blob/main/Qwen2.5.1-Coder-7B-Instruct-Q4_0_8_8.gguf) | Q4_0_8_8 | 4.43GB | false | Optimized for ARM inference. Requires 'sve' support (see link below). *Don't use on Mac or Windows*. |
| [Qwen2.5.1-Coder-7B-Instruct-Q4_0_4_8.gguf](https://huggingface.co/bartowski/Qwen2.5.1-Coder-7B-Instruct-GGUF/blob/main/Qwen2.5.1-Coder-7B-Instruct-Q4_0_4_8.gguf) | Q4_0_4_8 | 4.43GB | false | Optimized for ARM inference. Requires 'i8mm' support (see link below). *Don't use on Mac or Windows*. |
| [Qwen2.5.1-Coder-7B-Instruct-Q4_0_4_4.gguf](https://huggingface.co/bartowski/Qwen2.5.1-Coder-7B-Instruct-GGUF/blob/main/Qwen2.5.1-Coder-7B-Instruct-Q4_0_4_4.gguf) | Q4_0_4_4 | 4.43GB | false | Optimized for ARM inference. Should work well on all ARM chips, pick this if you're unsure. *Don't use on Mac or Windows*. |
| [Qwen2.5.1-Coder-7B-Instruct-IQ4_XS.gguf](https://huggingface.co/bartowski/Qwen2.5.1-Coder-7B-Instruct-GGUF/blob/main/Qwen2.5.1-Coder-7B-Instruct-IQ4_XS.gguf) | IQ4_XS | 4.22GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Qwen2.5.1-Coder-7B-Instruct-Q3_K_L.gguf](https://huggingface.co/bartowski/Qwen2.5.1-Coder-7B-Instruct-GGUF/blob/main/Qwen2.5.1-Coder-7B-Instruct-Q3_K_L.gguf) | Q3_K_L | 4.09GB | false | Lower quality but usable, good for low RAM availability. |
| [Qwen2.5.1-Coder-7B-Instruct-Q3_K_M.gguf](https://huggingface.co/bartowski/Qwen2.5.1-Coder-7B-Instruct-GGUF/blob/main/Qwen2.5.1-Coder-7B-Instruct-Q3_K_M.gguf) | Q3_K_M | 3.81GB | false | Low quality. |
| [Qwen2.5.1-Coder-7B-Instruct-IQ3_M.gguf](https://huggingface.co/bartowski/Qwen2.5.1-Coder-7B-Instruct-GGUF/blob/main/Qwen2.5.1-Coder-7B-Instruct-IQ3_M.gguf) | IQ3_M | 3.57GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Qwen2.5.1-Coder-7B-Instruct-Q2_K_L.gguf](https://huggingface.co/bartowski/Qwen2.5.1-Coder-7B-Instruct-GGUF/blob/main/Qwen2.5.1-Coder-7B-Instruct-Q2_K_L.gguf) | Q2_K_L | 3.55GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Qwen2.5.1-Coder-7B-Instruct-Q3_K_S.gguf](https://huggingface.co/bartowski/Qwen2.5.1-Coder-7B-Instruct-GGUF/blob/main/Qwen2.5.1-Coder-7B-Instruct-Q3_K_S.gguf) | Q3_K_S | 3.49GB | false | Low quality, not recommended. |
| [Qwen2.5.1-Coder-7B-Instruct-IQ3_XS.gguf](https://huggingface.co/bartowski/Qwen2.5.1-Coder-7B-Instruct-GGUF/blob/main/Qwen2.5.1-Coder-7B-Instruct-IQ3_XS.gguf) | IQ3_XS | 3.35GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Qwen2.5.1-Coder-7B-Instruct-Q2_K.gguf](https://huggingface.co/bartowski/Qwen2.5.1-Coder-7B-Instruct-GGUF/blob/main/Qwen2.5.1-Coder-7B-Instruct-Q2_K.gguf) | Q2_K | 3.02GB | false | Very low quality but surprisingly usable. |
| [Qwen2.5.1-Coder-7B-Instruct-IQ2_M.gguf](https://huggingface.co/bartowski/Qwen2.5.1-Coder-7B-Instruct-GGUF/blob/main/Qwen2.5.1-Coder-7B-Instruct-IQ2_M.gguf) | IQ2_M | 2.78GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Qwen2.5.1-Coder-7B-Instruct-GGUF --include "Qwen2.5.1-Coder-7B-Instruct-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Qwen2.5.1-Coder-7B-Instruct-GGUF --include "Qwen2.5.1-Coder-7B-Instruct-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Qwen2.5.1-Coder-7B-Instruct-Q8_0) or download them all in place (./)
## Q4_0_X_X
These are *NOT* for Metal (Apple) offloading, only ARM chips.
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset
Thank you ZeroWw for the inspiration to experiment with embed/output
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}