初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Mistral-Nemo-Gutenberg-Doppel-12B-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-04 04:35:13 +08:00
commit d14723cedf
28 changed files with 261 additions and 0 deletions

60
.gitattributes vendored Normal file
View File

@@ -0,0 +1,60 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-Gutenberg-Doppel-12B-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-Gutenberg-Doppel-12B-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-Gutenberg-Doppel-12B-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-Gutenberg-Doppel-12B-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-Gutenberg-Doppel-12B-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-Gutenberg-Doppel-12B-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-Gutenberg-Doppel-12B-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-Gutenberg-Doppel-12B-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-Gutenberg-Doppel-12B-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-Gutenberg-Doppel-12B-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-Gutenberg-Doppel-12B-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-Gutenberg-Doppel-12B-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-Gutenberg-Doppel-12B-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-Gutenberg-Doppel-12B-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-Gutenberg-Doppel-12B-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-Gutenberg-Doppel-12B-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-Gutenberg-Doppel-12B-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-Gutenberg-Doppel-12B-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-Gutenberg-Doppel-12B-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-Gutenberg-Doppel-12B-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-Gutenberg-Doppel-12B-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-Gutenberg-Doppel-12B-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-Gutenberg-Doppel-12B-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-Gutenberg-Doppel-12B-f16.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-Gutenberg-Doppel-12B.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4384fc45e0eeef08824b6e85048cc09dc466c1428f51e45943b0591458ca13e9
size 4435027104

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ca6f90f2808f7dccc584c06dd83ce62f56797a5125ae9ebe34a1612874c371b2
size 5722236064

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1d98aff4663bf6a77b89d34b94d3e81a98d1b37acf726d5c7ac995dc0c6e6f4d
size 5306492064

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8b5c4dd9f18af02d462d326379434ea6a6f0f31d5d4aaefb7f0f62faf039993f
size 6742713504

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c49b820d7faad4d054732080d36acd0592f8cf5da4a2f02c18a565501530cf2c
size 4791051424

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:199e1d83dc07cecaf26bc95cfd6b73d8f2322f50c4e57bbf8098b7093d69cfa3
size 5446411424

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c7926aafdbbcde8c50fc43847e4ca136edfc4b03f4bb39e0124dfa9c769f3f18
size 6561506464

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c501f33afd1b4541224e9ef42cc66d39b645f8aa15587af1d5e90469649c06b3
size 6083093664

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f1787dce739ee4e7e9f66192099923b52059055dcdff01e4c85177c1d1ec7f11
size 5534229664

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b75fcd7897e0edd2f2417426e7eb625ace6435dbe84210a2e71ba462d6cefa5d
size 7148709024

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:231ccdb96d33d0337c0279d3d370ecdf1a3e63c27291c664ac5f5aa6c1c60b9a
size 7094641824

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2c0e93f76a2632f38bdcad0ebe2373c3b3a45471393f8f12126c47ab9d71d665
size 7071704224

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c0f696e107488da6593bee5cb5723c15a3532695ed95f4ee93418685609a1ea6
size 7071704224

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:61e3034cedabb451defb95944fe250052ba6a979ae4c05b5a975cd3dfd7a1288
size 7071704224

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:bf1f019cbd2752cda38ed172f711404af32f79cb02cf4c2af591a31939970ff8
size 7975281824

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:41af224e422be1e78ddf0c7e44022dbe9b4b5a012ddaa931dcf1eb2624ec638f
size 7477208224

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:31d7b74d34aaf73f7c7201b516e99eaacd7352cb3edc81ce4d4d87cebbae0730
size 7120200864

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f57c9a4496d9a2a132f43a97de59020e1724cbbab7b11219dfa412435341866d
size 9141822624

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d55d1fd48288b24320596cdf1c35efd1f3786f2a15123e04807ccf92a3cf5f72
size 8727635104

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8c1765fe51cad9fa84614c0345dc010837d5d3c90bdc98ce65d1d335b19341ca
size 8518739104

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:498f79e6ac651816ef181cfa41c842b8683cf8302cc6084bfe97fd0fa4af89d3
size 10056213664

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5c5ceb60a0eb11457f9bcfff6d4052b3a1f58777eb502f38642f4f5411b4e21f
size 10381272224

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:018b7ec11b35f8454ce0466b701fe7183adff172fac618d38f796f10675015db
size 13022373024

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d0ed1de54d5ab702c5a1f237cd0d100b480c5415d8f5e4ab8f879b2d7ab4f265
size 24504279936

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a21f2514039a7616ebb8cc46e9379f667b111011b8ee1ad686e6e51fcf97e5e2
size 7054418

125
README.md Normal file
View File

@@ -0,0 +1,125 @@
---
base_model: nbeerbower/Mistral-Nemo-Gutenberg-Doppel-12B
datasets:
- jondurbin/gutenberg-dpo-v0.1
- nbeerbower/gutenberg2-dpo
library_name: transformers
license: apache-2.0
pipeline_tag: text-generation
quantized_by: bartowski
---
## Llamacpp imatrix Quantizations of Mistral-Nemo-Gutenberg-Doppel-12B
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b3825">b3825</a> for quantization.
Original model: https://huggingface.co/nbeerbower/Mistral-Nemo-Gutenberg-Doppel-12B
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<s>[INST]{prompt}[/INST]
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Mistral-Nemo-Gutenberg-Doppel-12B-f16.gguf](https://huggingface.co/bartowski/Mistral-Nemo-Gutenberg-Doppel-12B-GGUF/blob/main/Mistral-Nemo-Gutenberg-Doppel-12B-f16.gguf) | f16 | 24.50GB | false | Full F16 weights. |
| [Mistral-Nemo-Gutenberg-Doppel-12B-Q8_0.gguf](https://huggingface.co/bartowski/Mistral-Nemo-Gutenberg-Doppel-12B-GGUF/blob/main/Mistral-Nemo-Gutenberg-Doppel-12B-Q8_0.gguf) | Q8_0 | 13.02GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Mistral-Nemo-Gutenberg-Doppel-12B-Q6_K_L.gguf](https://huggingface.co/bartowski/Mistral-Nemo-Gutenberg-Doppel-12B-GGUF/blob/main/Mistral-Nemo-Gutenberg-Doppel-12B-Q6_K_L.gguf) | Q6_K_L | 10.38GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Mistral-Nemo-Gutenberg-Doppel-12B-Q6_K.gguf](https://huggingface.co/bartowski/Mistral-Nemo-Gutenberg-Doppel-12B-GGUF/blob/main/Mistral-Nemo-Gutenberg-Doppel-12B-Q6_K.gguf) | Q6_K | 10.06GB | false | Very high quality, near perfect, *recommended*. |
| [Mistral-Nemo-Gutenberg-Doppel-12B-Q5_K_L.gguf](https://huggingface.co/bartowski/Mistral-Nemo-Gutenberg-Doppel-12B-GGUF/blob/main/Mistral-Nemo-Gutenberg-Doppel-12B-Q5_K_L.gguf) | Q5_K_L | 9.14GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Mistral-Nemo-Gutenberg-Doppel-12B-Q5_K_M.gguf](https://huggingface.co/bartowski/Mistral-Nemo-Gutenberg-Doppel-12B-GGUF/blob/main/Mistral-Nemo-Gutenberg-Doppel-12B-Q5_K_M.gguf) | Q5_K_M | 8.73GB | false | High quality, *recommended*. |
| [Mistral-Nemo-Gutenberg-Doppel-12B-Q5_K_S.gguf](https://huggingface.co/bartowski/Mistral-Nemo-Gutenberg-Doppel-12B-GGUF/blob/main/Mistral-Nemo-Gutenberg-Doppel-12B-Q5_K_S.gguf) | Q5_K_S | 8.52GB | false | High quality, *recommended*. |
| [Mistral-Nemo-Gutenberg-Doppel-12B-Q4_K_L.gguf](https://huggingface.co/bartowski/Mistral-Nemo-Gutenberg-Doppel-12B-GGUF/blob/main/Mistral-Nemo-Gutenberg-Doppel-12B-Q4_K_L.gguf) | Q4_K_L | 7.98GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Mistral-Nemo-Gutenberg-Doppel-12B-Q4_K_M.gguf](https://huggingface.co/bartowski/Mistral-Nemo-Gutenberg-Doppel-12B-GGUF/blob/main/Mistral-Nemo-Gutenberg-Doppel-12B-Q4_K_M.gguf) | Q4_K_M | 7.48GB | false | Good quality, default size for must use cases, *recommended*. |
| [Mistral-Nemo-Gutenberg-Doppel-12B-Q3_K_XL.gguf](https://huggingface.co/bartowski/Mistral-Nemo-Gutenberg-Doppel-12B-GGUF/blob/main/Mistral-Nemo-Gutenberg-Doppel-12B-Q3_K_XL.gguf) | Q3_K_XL | 7.15GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Mistral-Nemo-Gutenberg-Doppel-12B-Q4_K_S.gguf](https://huggingface.co/bartowski/Mistral-Nemo-Gutenberg-Doppel-12B-GGUF/blob/main/Mistral-Nemo-Gutenberg-Doppel-12B-Q4_K_S.gguf) | Q4_K_S | 7.12GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Mistral-Nemo-Gutenberg-Doppel-12B-Q4_0.gguf](https://huggingface.co/bartowski/Mistral-Nemo-Gutenberg-Doppel-12B-GGUF/blob/main/Mistral-Nemo-Gutenberg-Doppel-12B-Q4_0.gguf) | Q4_0 | 7.09GB | false | Legacy format, generally not worth using over similarly sized formats |
| [Mistral-Nemo-Gutenberg-Doppel-12B-Q4_0_8_8.gguf](https://huggingface.co/bartowski/Mistral-Nemo-Gutenberg-Doppel-12B-GGUF/blob/main/Mistral-Nemo-Gutenberg-Doppel-12B-Q4_0_8_8.gguf) | Q4_0_8_8 | 7.07GB | false | Optimized for ARM inference. Requires 'sve' support (see link below). |
| [Mistral-Nemo-Gutenberg-Doppel-12B-Q4_0_4_8.gguf](https://huggingface.co/bartowski/Mistral-Nemo-Gutenberg-Doppel-12B-GGUF/blob/main/Mistral-Nemo-Gutenberg-Doppel-12B-Q4_0_4_8.gguf) | Q4_0_4_8 | 7.07GB | false | Optimized for ARM inference. Requires 'i8mm' support (see link below). |
| [Mistral-Nemo-Gutenberg-Doppel-12B-Q4_0_4_4.gguf](https://huggingface.co/bartowski/Mistral-Nemo-Gutenberg-Doppel-12B-GGUF/blob/main/Mistral-Nemo-Gutenberg-Doppel-12B-Q4_0_4_4.gguf) | Q4_0_4_4 | 7.07GB | false | Optimized for ARM inference. Should work well on all ARM chips, pick this if you're unsure. |
| [Mistral-Nemo-Gutenberg-Doppel-12B-IQ4_XS.gguf](https://huggingface.co/bartowski/Mistral-Nemo-Gutenberg-Doppel-12B-GGUF/blob/main/Mistral-Nemo-Gutenberg-Doppel-12B-IQ4_XS.gguf) | IQ4_XS | 6.74GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Mistral-Nemo-Gutenberg-Doppel-12B-Q3_K_L.gguf](https://huggingface.co/bartowski/Mistral-Nemo-Gutenberg-Doppel-12B-GGUF/blob/main/Mistral-Nemo-Gutenberg-Doppel-12B-Q3_K_L.gguf) | Q3_K_L | 6.56GB | false | Lower quality but usable, good for low RAM availability. |
| [Mistral-Nemo-Gutenberg-Doppel-12B-Q3_K_M.gguf](https://huggingface.co/bartowski/Mistral-Nemo-Gutenberg-Doppel-12B-GGUF/blob/main/Mistral-Nemo-Gutenberg-Doppel-12B-Q3_K_M.gguf) | Q3_K_M | 6.08GB | false | Low quality. |
| [Mistral-Nemo-Gutenberg-Doppel-12B-IQ3_M.gguf](https://huggingface.co/bartowski/Mistral-Nemo-Gutenberg-Doppel-12B-GGUF/blob/main/Mistral-Nemo-Gutenberg-Doppel-12B-IQ3_M.gguf) | IQ3_M | 5.72GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Mistral-Nemo-Gutenberg-Doppel-12B-Q3_K_S.gguf](https://huggingface.co/bartowski/Mistral-Nemo-Gutenberg-Doppel-12B-GGUF/blob/main/Mistral-Nemo-Gutenberg-Doppel-12B-Q3_K_S.gguf) | Q3_K_S | 5.53GB | false | Low quality, not recommended. |
| [Mistral-Nemo-Gutenberg-Doppel-12B-Q2_K_L.gguf](https://huggingface.co/bartowski/Mistral-Nemo-Gutenberg-Doppel-12B-GGUF/blob/main/Mistral-Nemo-Gutenberg-Doppel-12B-Q2_K_L.gguf) | Q2_K_L | 5.45GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Mistral-Nemo-Gutenberg-Doppel-12B-IQ3_XS.gguf](https://huggingface.co/bartowski/Mistral-Nemo-Gutenberg-Doppel-12B-GGUF/blob/main/Mistral-Nemo-Gutenberg-Doppel-12B-IQ3_XS.gguf) | IQ3_XS | 5.31GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Mistral-Nemo-Gutenberg-Doppel-12B-Q2_K.gguf](https://huggingface.co/bartowski/Mistral-Nemo-Gutenberg-Doppel-12B-GGUF/blob/main/Mistral-Nemo-Gutenberg-Doppel-12B-Q2_K.gguf) | Q2_K | 4.79GB | false | Very low quality but surprisingly usable. |
| [Mistral-Nemo-Gutenberg-Doppel-12B-IQ2_M.gguf](https://huggingface.co/bartowski/Mistral-Nemo-Gutenberg-Doppel-12B-GGUF/blob/main/Mistral-Nemo-Gutenberg-Doppel-12B-IQ2_M.gguf) | IQ2_M | 4.44GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Mistral-Nemo-Gutenberg-Doppel-12B-GGUF --include "Mistral-Nemo-Gutenberg-Doppel-12B-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Mistral-Nemo-Gutenberg-Doppel-12B-GGUF --include "Mistral-Nemo-Gutenberg-Doppel-12B-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Mistral-Nemo-Gutenberg-Doppel-12B-Q8_0) or download them all in place (./)
## Q4_0_X_X
These are *NOT* for Metal (Apple) offloading, only ARM chips.
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset
Thank you ZeroWw for the inspiration to experiment with embed/output
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}