初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Mahou-1.5-mistral-nemo-12B-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-31 00:46:13 +08:00
commit bad75d09f7
29 changed files with 265 additions and 0 deletions

61
.gitattributes vendored Normal file
View File

@@ -0,0 +1,61 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Mahou-1.5-mistral-nemo-12B-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Mahou-1.5-mistral-nemo-12B-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Mahou-1.5-mistral-nemo-12B-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Mahou-1.5-mistral-nemo-12B-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Mahou-1.5-mistral-nemo-12B-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Mahou-1.5-mistral-nemo-12B-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Mahou-1.5-mistral-nemo-12B-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Mahou-1.5-mistral-nemo-12B-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Mahou-1.5-mistral-nemo-12B-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Mahou-1.5-mistral-nemo-12B-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
Mahou-1.5-mistral-nemo-12B-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
Mahou-1.5-mistral-nemo-12B-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
Mahou-1.5-mistral-nemo-12B-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Mahou-1.5-mistral-nemo-12B-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Mahou-1.5-mistral-nemo-12B-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Mahou-1.5-mistral-nemo-12B-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Mahou-1.5-mistral-nemo-12B-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Mahou-1.5-mistral-nemo-12B-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Mahou-1.5-mistral-nemo-12B-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Mahou-1.5-mistral-nemo-12B-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Mahou-1.5-mistral-nemo-12B-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Mahou-1.5-mistral-nemo-12B-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Mahou-1.5-mistral-nemo-12B-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Mahou-1.5-mistral-nemo-12B-IQ2_S.gguf filter=lfs diff=lfs merge=lfs -text
Mahou-1.5-mistral-nemo-12B-f16.gguf filter=lfs diff=lfs merge=lfs -text
Mahou-1.5-mistral-nemo-12B.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4a0c108a1ea82132955dd3734657e747671bbecd1879a051d3a0f8448d62bd89
size 4435023328

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:90f875400ecc3395991ff67fc05c1f1afbe5d415037188f4d0af51eb9def5ac1
size 4138472928

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1f8d4f24623c73909bc4bf9b12bd637cf39731a59e41a37189d6203c9cd6c436
size 5722232288

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:407b50de58e37ba1963f4914c5bd7a79e0a7094ca8413bb19a064b53d074e9bf
size 5306488288

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:757222b2d2388d50114607a884f014d96ed345b4cc12fb955d6bbca2a1d8a9b8
size 6742709728

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:55588ddc5bb3dc1a45fde73726eebe742c4ec846058a84845574b34b91cedc2d
size 4791047648

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2a60ae8a27f8cda0d6dae8bc375e1f7dfd2f4a8787f8403f0b42b21b8d655ac4
size 5446407648

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fe16f0ea29191e64d51ba8e8279ba080059088a9054a60cf4a4ed0cdbbee921a
size 6561502688

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9449b73e2f9479d003f68ddafa8a53be5ee38005721e9e0933a185077af516ae
size 6083089888

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:87ab162f0e6cabe7067c46a3337d5bb3bf4a878b210b10d6d455830c17d7f5c8
size 5534225888

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:65d48058eac91411eb7800aca5426324b96e6f46397709a72d5b8ed3d604ce8b
size 7148705248

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a9f5c7fdeef8177545ea3243f70eebea263b56a94889c7d9e26addf601f93b14
size 7094638048

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0085848e948dfb594e44979777f5136699c6e3f7de99aec99802fa7c34e66437
size 7071700448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d8fa3c0399632f43e6df2ca41a4c2843ee4f9f9b34d94425c5990134e378757a
size 7071700448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6180906ea6ecba60ae921f5ed3648a34238d28135faf1b036e8418ef62eb038e
size 7071700448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:bf9f9b8ea3159a7bc5b16a3878f00abe3cfec616766dfa8adf9e4027237b6c3a
size 7975278048

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:37bdf3d1a761fe45edd20501fdf87a8eafd3fd8f02b94c2490910b1be6a0fbf8
size 7477204448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:27dab5b9055a18d8d2a7dd763bee2e4d970edc22558b9931038812ff2ef645ca
size 7120197088

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:314c6dff517cf9ee0102f40205456164ac9f76890585a01841ca8e03c446b314
size 9141818848

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:82dbaed2f23e8965171acf129c4d427f6a8635b59f8018a49a613cf0231c5677
size 8727631328

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:49d890d0dbeb16310a7ee72219d3d8ff5e2699d79ff818f0b4e44bcb64665a24
size 8518735328

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6b6d37a8603d90519301ab944db6bc2b49cd814c8eff7f12f049709e1168aaaf
size 10056209888

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:163937858a5d6b4263a88da787ffb0010af51e8a54883aec9cb9de0ca2f3543c
size 10381268448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:db98b27dc9809e4aba7b011d791c9448ed7f1c450cab6b7dfb13bfde9dff1bb8
size 13022369248

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:64bab104f6ed170ad55fbb37593828b4e2d734bea4dfc92f0b2756f75542bfe1
size 24504276160

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a3e55cce011020782979731d2c9cf016919368461c9ce1189657cbabd206b87a
size 7054418

125
README.md Normal file
View File

@@ -0,0 +1,125 @@
---
base_model: flammenai/Mahou-1.5-mistral-nemo-12B
pipeline_tag: text-generation
quantized_by: bartowski
---
## Llamacpp imatrix Quantizations of Mahou-1.5-mistral-nemo-12B
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b3878">b3878</a> for quantization.
Original model: https://huggingface.co/flammenai/Mahou-1.5-mistral-nemo-12B
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Mahou-1.5-mistral-nemo-12B-f16.gguf](https://huggingface.co/bartowski/Mahou-1.5-mistral-nemo-12B-GGUF/blob/main/Mahou-1.5-mistral-nemo-12B-f16.gguf) | f16 | 24.50GB | false | Full F16 weights. |
| [Mahou-1.5-mistral-nemo-12B-Q8_0.gguf](https://huggingface.co/bartowski/Mahou-1.5-mistral-nemo-12B-GGUF/blob/main/Mahou-1.5-mistral-nemo-12B-Q8_0.gguf) | Q8_0 | 13.02GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Mahou-1.5-mistral-nemo-12B-Q6_K_L.gguf](https://huggingface.co/bartowski/Mahou-1.5-mistral-nemo-12B-GGUF/blob/main/Mahou-1.5-mistral-nemo-12B-Q6_K_L.gguf) | Q6_K_L | 10.38GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Mahou-1.5-mistral-nemo-12B-Q6_K.gguf](https://huggingface.co/bartowski/Mahou-1.5-mistral-nemo-12B-GGUF/blob/main/Mahou-1.5-mistral-nemo-12B-Q6_K.gguf) | Q6_K | 10.06GB | false | Very high quality, near perfect, *recommended*. |
| [Mahou-1.5-mistral-nemo-12B-Q5_K_L.gguf](https://huggingface.co/bartowski/Mahou-1.5-mistral-nemo-12B-GGUF/blob/main/Mahou-1.5-mistral-nemo-12B-Q5_K_L.gguf) | Q5_K_L | 9.14GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Mahou-1.5-mistral-nemo-12B-Q5_K_M.gguf](https://huggingface.co/bartowski/Mahou-1.5-mistral-nemo-12B-GGUF/blob/main/Mahou-1.5-mistral-nemo-12B-Q5_K_M.gguf) | Q5_K_M | 8.73GB | false | High quality, *recommended*. |
| [Mahou-1.5-mistral-nemo-12B-Q5_K_S.gguf](https://huggingface.co/bartowski/Mahou-1.5-mistral-nemo-12B-GGUF/blob/main/Mahou-1.5-mistral-nemo-12B-Q5_K_S.gguf) | Q5_K_S | 8.52GB | false | High quality, *recommended*. |
| [Mahou-1.5-mistral-nemo-12B-Q4_K_L.gguf](https://huggingface.co/bartowski/Mahou-1.5-mistral-nemo-12B-GGUF/blob/main/Mahou-1.5-mistral-nemo-12B-Q4_K_L.gguf) | Q4_K_L | 7.98GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Mahou-1.5-mistral-nemo-12B-Q4_K_M.gguf](https://huggingface.co/bartowski/Mahou-1.5-mistral-nemo-12B-GGUF/blob/main/Mahou-1.5-mistral-nemo-12B-Q4_K_M.gguf) | Q4_K_M | 7.48GB | false | Good quality, default size for must use cases, *recommended*. |
| [Mahou-1.5-mistral-nemo-12B-Q3_K_XL.gguf](https://huggingface.co/bartowski/Mahou-1.5-mistral-nemo-12B-GGUF/blob/main/Mahou-1.5-mistral-nemo-12B-Q3_K_XL.gguf) | Q3_K_XL | 7.15GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Mahou-1.5-mistral-nemo-12B-Q4_K_S.gguf](https://huggingface.co/bartowski/Mahou-1.5-mistral-nemo-12B-GGUF/blob/main/Mahou-1.5-mistral-nemo-12B-Q4_K_S.gguf) | Q4_K_S | 7.12GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Mahou-1.5-mistral-nemo-12B-Q4_0.gguf](https://huggingface.co/bartowski/Mahou-1.5-mistral-nemo-12B-GGUF/blob/main/Mahou-1.5-mistral-nemo-12B-Q4_0.gguf) | Q4_0 | 7.09GB | false | Legacy format, generally not worth using over similarly sized formats |
| [Mahou-1.5-mistral-nemo-12B-Q4_0_8_8.gguf](https://huggingface.co/bartowski/Mahou-1.5-mistral-nemo-12B-GGUF/blob/main/Mahou-1.5-mistral-nemo-12B-Q4_0_8_8.gguf) | Q4_0_8_8 | 7.07GB | false | Optimized for ARM inference. Requires 'sve' support (see link below). *Don't use on Mac or Windows*. |
| [Mahou-1.5-mistral-nemo-12B-Q4_0_4_8.gguf](https://huggingface.co/bartowski/Mahou-1.5-mistral-nemo-12B-GGUF/blob/main/Mahou-1.5-mistral-nemo-12B-Q4_0_4_8.gguf) | Q4_0_4_8 | 7.07GB | false | Optimized for ARM inference. Requires 'i8mm' support (see link below). *Don't use on Mac or Windows*. |
| [Mahou-1.5-mistral-nemo-12B-Q4_0_4_4.gguf](https://huggingface.co/bartowski/Mahou-1.5-mistral-nemo-12B-GGUF/blob/main/Mahou-1.5-mistral-nemo-12B-Q4_0_4_4.gguf) | Q4_0_4_4 | 7.07GB | false | Optimized for ARM inference. Should work well on all ARM chips, pick this if you're unsure. *Don't use on Mac or Windows*. |
| [Mahou-1.5-mistral-nemo-12B-IQ4_XS.gguf](https://huggingface.co/bartowski/Mahou-1.5-mistral-nemo-12B-GGUF/blob/main/Mahou-1.5-mistral-nemo-12B-IQ4_XS.gguf) | IQ4_XS | 6.74GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Mahou-1.5-mistral-nemo-12B-Q3_K_L.gguf](https://huggingface.co/bartowski/Mahou-1.5-mistral-nemo-12B-GGUF/blob/main/Mahou-1.5-mistral-nemo-12B-Q3_K_L.gguf) | Q3_K_L | 6.56GB | false | Lower quality but usable, good for low RAM availability. |
| [Mahou-1.5-mistral-nemo-12B-Q3_K_M.gguf](https://huggingface.co/bartowski/Mahou-1.5-mistral-nemo-12B-GGUF/blob/main/Mahou-1.5-mistral-nemo-12B-Q3_K_M.gguf) | Q3_K_M | 6.08GB | false | Low quality. |
| [Mahou-1.5-mistral-nemo-12B-IQ3_M.gguf](https://huggingface.co/bartowski/Mahou-1.5-mistral-nemo-12B-GGUF/blob/main/Mahou-1.5-mistral-nemo-12B-IQ3_M.gguf) | IQ3_M | 5.72GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Mahou-1.5-mistral-nemo-12B-Q3_K_S.gguf](https://huggingface.co/bartowski/Mahou-1.5-mistral-nemo-12B-GGUF/blob/main/Mahou-1.5-mistral-nemo-12B-Q3_K_S.gguf) | Q3_K_S | 5.53GB | false | Low quality, not recommended. |
| [Mahou-1.5-mistral-nemo-12B-Q2_K_L.gguf](https://huggingface.co/bartowski/Mahou-1.5-mistral-nemo-12B-GGUF/blob/main/Mahou-1.5-mistral-nemo-12B-Q2_K_L.gguf) | Q2_K_L | 5.45GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Mahou-1.5-mistral-nemo-12B-IQ3_XS.gguf](https://huggingface.co/bartowski/Mahou-1.5-mistral-nemo-12B-GGUF/blob/main/Mahou-1.5-mistral-nemo-12B-IQ3_XS.gguf) | IQ3_XS | 5.31GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Mahou-1.5-mistral-nemo-12B-Q2_K.gguf](https://huggingface.co/bartowski/Mahou-1.5-mistral-nemo-12B-GGUF/blob/main/Mahou-1.5-mistral-nemo-12B-Q2_K.gguf) | Q2_K | 4.79GB | false | Very low quality but surprisingly usable. |
| [Mahou-1.5-mistral-nemo-12B-IQ2_M.gguf](https://huggingface.co/bartowski/Mahou-1.5-mistral-nemo-12B-GGUF/blob/main/Mahou-1.5-mistral-nemo-12B-IQ2_M.gguf) | IQ2_M | 4.44GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
| [Mahou-1.5-mistral-nemo-12B-IQ2_S.gguf](https://huggingface.co/bartowski/Mahou-1.5-mistral-nemo-12B-GGUF/blob/main/Mahou-1.5-mistral-nemo-12B-IQ2_S.gguf) | IQ2_S | 4.14GB | false | Low quality, uses SOTA techniques to be usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Mahou-1.5-mistral-nemo-12B-GGUF --include "Mahou-1.5-mistral-nemo-12B-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Mahou-1.5-mistral-nemo-12B-GGUF --include "Mahou-1.5-mistral-nemo-12B-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Mahou-1.5-mistral-nemo-12B-Q8_0) or download them all in place (./)
## Q4_0_X_X
These are *NOT* for Metal (Apple) offloading, only ARM chips.
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset
Thank you ZeroWw for the inspiration to experiment with embed/output
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}