初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Mistral-Small-Instruct-2409-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-07 03:06:12 +08:00
commit cbc26c48b7
30 changed files with 286 additions and 0 deletions

62
.gitattributes vendored Normal file
View File

@@ -0,0 +1,62 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Mistral-Small-Instruct-2409-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-Instruct-2409-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-Instruct-2409-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-Instruct-2409-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-Instruct-2409-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-Instruct-2409-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-Instruct-2409-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-Instruct-2409-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-Instruct-2409-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-Instruct-2409-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-Instruct-2409-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-Instruct-2409-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-Instruct-2409-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-Instruct-2409-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-Instruct-2409-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-Instruct-2409-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-Instruct-2409-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-Instruct-2409-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-Instruct-2409-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-Instruct-2409-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-Instruct-2409-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-Instruct-2409-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-Instruct-2409-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-Instruct-2409-IQ2_XS.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-Instruct-2409-IQ2_XXS.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-Instruct-2409-f16.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-Instruct-2409.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b02ae15681135255bdc5a340f6e8f8ee92538d17fbefcab524c2f5093fb74c5a
size 7618966528

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:37e627eee6cf3c1d0b7e051223d8a1df9055699df4b7788fcf1f47f743fa332c
size 6646150144

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:902f463812e27120dd68d340d5b57b38fb1e34a83be7c13a8b2d2073c91d024d
size 5996557312

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:43b61454a7ffad448a2a6b32169cfb118885ef1a6be5186db998bff8dde6c7ab
size 10062410752

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1b73da007e2fa090f38d0daaed0388014fafd56ac3b102d3ae9b2ff9f16c68c5
size 9176101888

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e4b494eb11f275819bcf4c680790403f8614defdaacc4b134170fe9a5bb8338b
size 11935298560

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2328a15242bb1dedeba159e0bd080e67942153d071608fc26e404dea9a460255
size 8272098304

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d07f801a5591a9c97c09a45d19ff0d9de0ae05050f1f3a5f78e21995ea48e5db
size 8468706304

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3ef8e2b904f3fd93cacd88c3b00b5ff0baf5ff852fd79b25465c86fce3ad6af2
size 11730433024

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a85ed7a805d4fcac101652e748f796185c0a118f2f79fd94dd9f114799acac32
size 10756830208

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:15fb4b9c6ab125a0226e444433c7a3a20226a0c7c8f07baec6d0e23879320da5
size 9641276416

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f3c8182b54c5c7359be303e4ce57d1f00ac98c061b415eb57cf4a45195e5b0ae
size 11906593792

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:af8bb8c71c20da453de1c85137e496312c803d9d4adb96e4852e86028e48c1a7
size 12613202944

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b9813b6e3f5eb63aa223b2a32906d00c99108f465c6b75cf5e3594a6b798e6ef
size 12569162752

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:75ee0611d37964f297da79da1c3d2d462d6d0e4df208d8f95497f0d14769705d
size 12569162752

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3bf1031b5557c77faf407f1e6365745d073e134ade30d3fea2c58def22b573c2
size 12569162752

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7d944a8b94153082c59244c99f97b6eebc94bdf59b1da18013dd227086bc16c7
size 13490664448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8091d756d67d518b901e4b3b3986b0786f91356505e3065cba9fe793de5fd0c7
size 13341242368

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4bb1c6688f6455e2f94deee0d5800057776e93b7d486c6063ef3b499848eaa5f
size 12660388864

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7d3237dd43d8acdba60d7aa02637444401fe056febbb9d2b2bb7ba8426927bf4
size 15846814720

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b84ae96c2b36975beb785f09579b411ea68a83f2eabd37bd17f6b64052e8cb75
size 15722558464

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d7f43811b92eac571cf91a910dc6ffc66a134c1d6002ef8624b64fdcd1fc83c6
size 15324820480

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b919ae1cda19d5e15880d4a821de2440a6b2ee6ce82ef1d71eeef220bc6f86eb
size 18252706816

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:385b2944f379522a6ca73a81f7f36fc75fb5b9c94c02f1afd26956ed98cc41d0
size 18350224384

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ba62280bb31abc582a790cd0e63a1946f4864db6053bf885dba63bf5d0f651ac
size 23640552448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:06f23a367d869e9e3a982a0ed2d8da8ea7d513a73565ef3faa5f531cf0e1fd3a
size 44496728768

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a3f7ef62e9db96f9b919a23eee9bf047d2af391fee78f2d7f69187b422fa2a3b
size 11940578

142
README.md Normal file
View File

@@ -0,0 +1,142 @@
---
base_model: mistralai/Mistral-Small-Instruct-2409
language:
- en
- fr
- de
- es
- it
- pt
- zh
- ja
- ru
- ko
license: other
license_name: mrl
license_link: https://mistral.ai/licenses/MRL-0.1.md
pipeline_tag: text-generation
quantized_by: bartowski
extra_gated_description: If you want to learn more about how we process your personal
data, please read our <a href="https://mistral.ai/terms/">Privacy Policy</a>.
---
## Llamacpp imatrix Quantizations of Mistral-Small-Instruct-2409
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b3772">b3772</a> for quantization.
Original model: https://huggingface.co/mistralai/Mistral-Small-Instruct-2409
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<s>[INST] {prompt}[/INST]
```
## What's new:
Update prompt template
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Mistral-Small-Instruct-2409-f16.gguf](https://huggingface.co/bartowski/Mistral-Small-Instruct-2409-GGUF/blob/main/Mistral-Small-Instruct-2409-f16.gguf) | f16 | 44.50GB | false | Full F16 weights. |
| [Mistral-Small-Instruct-2409-Q8_0.gguf](https://huggingface.co/bartowski/Mistral-Small-Instruct-2409-GGUF/blob/main/Mistral-Small-Instruct-2409-Q8_0.gguf) | Q8_0 | 23.64GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Mistral-Small-Instruct-2409-Q6_K_L.gguf](https://huggingface.co/bartowski/Mistral-Small-Instruct-2409-GGUF/blob/main/Mistral-Small-Instruct-2409-Q6_K_L.gguf) | Q6_K_L | 18.35GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Mistral-Small-Instruct-2409-Q6_K.gguf](https://huggingface.co/bartowski/Mistral-Small-Instruct-2409-GGUF/blob/main/Mistral-Small-Instruct-2409-Q6_K.gguf) | Q6_K | 18.25GB | false | Very high quality, near perfect, *recommended*. |
| [Mistral-Small-Instruct-2409-Q5_K_L.gguf](https://huggingface.co/bartowski/Mistral-Small-Instruct-2409-GGUF/blob/main/Mistral-Small-Instruct-2409-Q5_K_L.gguf) | Q5_K_L | 15.85GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Mistral-Small-Instruct-2409-Q5_K_M.gguf](https://huggingface.co/bartowski/Mistral-Small-Instruct-2409-GGUF/blob/main/Mistral-Small-Instruct-2409-Q5_K_M.gguf) | Q5_K_M | 15.72GB | false | High quality, *recommended*. |
| [Mistral-Small-Instruct-2409-Q5_K_S.gguf](https://huggingface.co/bartowski/Mistral-Small-Instruct-2409-GGUF/blob/main/Mistral-Small-Instruct-2409-Q5_K_S.gguf) | Q5_K_S | 15.32GB | false | High quality, *recommended*. |
| [Mistral-Small-Instruct-2409-Q4_K_L.gguf](https://huggingface.co/bartowski/Mistral-Small-Instruct-2409-GGUF/blob/main/Mistral-Small-Instruct-2409-Q4_K_L.gguf) | Q4_K_L | 13.49GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Mistral-Small-Instruct-2409-Q4_K_M.gguf](https://huggingface.co/bartowski/Mistral-Small-Instruct-2409-GGUF/blob/main/Mistral-Small-Instruct-2409-Q4_K_M.gguf) | Q4_K_M | 13.34GB | false | Good quality, default size for must use cases, *recommended*. |
| [Mistral-Small-Instruct-2409-Q4_K_S.gguf](https://huggingface.co/bartowski/Mistral-Small-Instruct-2409-GGUF/blob/main/Mistral-Small-Instruct-2409-Q4_K_S.gguf) | Q4_K_S | 12.66GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Mistral-Small-Instruct-2409-Q4_0.gguf](https://huggingface.co/bartowski/Mistral-Small-Instruct-2409-GGUF/blob/main/Mistral-Small-Instruct-2409-Q4_0.gguf) | Q4_0 | 12.61GB | false | Legacy format, generally not worth using over similarly sized formats |
| [Mistral-Small-Instruct-2409-Q4_0_8_8.gguf](https://huggingface.co/bartowski/Mistral-Small-Instruct-2409-GGUF/blob/main/Mistral-Small-Instruct-2409-Q4_0_8_8.gguf) | Q4_0_8_8 | 12.57GB | false | Optimized for ARM inference. Requires 'sve' support (see link below). |
| [Mistral-Small-Instruct-2409-Q4_0_4_8.gguf](https://huggingface.co/bartowski/Mistral-Small-Instruct-2409-GGUF/blob/main/Mistral-Small-Instruct-2409-Q4_0_4_8.gguf) | Q4_0_4_8 | 12.57GB | false | Optimized for ARM inference. Requires 'i8mm' support (see link below). |
| [Mistral-Small-Instruct-2409-Q4_0_4_4.gguf](https://huggingface.co/bartowski/Mistral-Small-Instruct-2409-GGUF/blob/main/Mistral-Small-Instruct-2409-Q4_0_4_4.gguf) | Q4_0_4_4 | 12.57GB | false | Optimized for ARM inference. Should work well on all ARM chips, pick this if you're unsure. |
| [Mistral-Small-Instruct-2409-IQ4_XS.gguf](https://huggingface.co/bartowski/Mistral-Small-Instruct-2409-GGUF/blob/main/Mistral-Small-Instruct-2409-IQ4_XS.gguf) | IQ4_XS | 11.94GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Mistral-Small-Instruct-2409-Q3_K_XL.gguf](https://huggingface.co/bartowski/Mistral-Small-Instruct-2409-GGUF/blob/main/Mistral-Small-Instruct-2409-Q3_K_XL.gguf) | Q3_K_XL | 11.91GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Mistral-Small-Instruct-2409-Q3_K_L.gguf](https://huggingface.co/bartowski/Mistral-Small-Instruct-2409-GGUF/blob/main/Mistral-Small-Instruct-2409-Q3_K_L.gguf) | Q3_K_L | 11.73GB | false | Lower quality but usable, good for low RAM availability. |
| [Mistral-Small-Instruct-2409-Q3_K_M.gguf](https://huggingface.co/bartowski/Mistral-Small-Instruct-2409-GGUF/blob/main/Mistral-Small-Instruct-2409-Q3_K_M.gguf) | Q3_K_M | 10.76GB | false | Low quality. |
| [Mistral-Small-Instruct-2409-IQ3_M.gguf](https://huggingface.co/bartowski/Mistral-Small-Instruct-2409-GGUF/blob/main/Mistral-Small-Instruct-2409-IQ3_M.gguf) | IQ3_M | 10.06GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Mistral-Small-Instruct-2409-Q3_K_S.gguf](https://huggingface.co/bartowski/Mistral-Small-Instruct-2409-GGUF/blob/main/Mistral-Small-Instruct-2409-Q3_K_S.gguf) | Q3_K_S | 9.64GB | false | Low quality, not recommended. |
| [Mistral-Small-Instruct-2409-IQ3_XS.gguf](https://huggingface.co/bartowski/Mistral-Small-Instruct-2409-GGUF/blob/main/Mistral-Small-Instruct-2409-IQ3_XS.gguf) | IQ3_XS | 9.18GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Mistral-Small-Instruct-2409-Q2_K_L.gguf](https://huggingface.co/bartowski/Mistral-Small-Instruct-2409-GGUF/blob/main/Mistral-Small-Instruct-2409-Q2_K_L.gguf) | Q2_K_L | 8.47GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Mistral-Small-Instruct-2409-Q2_K.gguf](https://huggingface.co/bartowski/Mistral-Small-Instruct-2409-GGUF/blob/main/Mistral-Small-Instruct-2409-Q2_K.gguf) | Q2_K | 8.27GB | false | Very low quality but surprisingly usable. |
| [Mistral-Small-Instruct-2409-IQ2_M.gguf](https://huggingface.co/bartowski/Mistral-Small-Instruct-2409-GGUF/blob/main/Mistral-Small-Instruct-2409-IQ2_M.gguf) | IQ2_M | 7.62GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
| [Mistral-Small-Instruct-2409-IQ2_XS.gguf](https://huggingface.co/bartowski/Mistral-Small-Instruct-2409-GGUF/blob/main/Mistral-Small-Instruct-2409-IQ2_XS.gguf) | IQ2_XS | 6.65GB | false | Low quality, uses SOTA techniques to be usable. |
| [Mistral-Small-Instruct-2409-IQ2_XXS.gguf](https://huggingface.co/bartowski/Mistral-Small-Instruct-2409-GGUF/blob/main/Mistral-Small-Instruct-2409-IQ2_XXS.gguf) | IQ2_XXS | 6.00GB | false | Very low quality, uses SOTA techniques to be usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Mistral-Small-Instruct-2409-GGUF --include "Mistral-Small-Instruct-2409-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Mistral-Small-Instruct-2409-GGUF --include "Mistral-Small-Instruct-2409-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Mistral-Small-Instruct-2409-Q8_0) or download them all in place (./)
## Q4_0_X_X
These are *NOT* for Metal (Apple) offloading, only ARM chips.
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset
Thank you ZeroWw for the inspiration to experiment with embed/output
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}