初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Mistral-Small-24B-Instruct-2501-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-09-07 14:55:13 +08:00
commit 2d88fcf9c2
32 changed files with 336 additions and 0 deletions

64
.gitattributes vendored Normal file
View File

@@ -0,0 +1,64 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501-f16.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501-Q4_1.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501-IQ4_NL.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501-IQ2_S.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501-IQ2_XS.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501-f32/Mistral-Small-24B-Instruct-2501-f32-00001-of-00003.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501-f32/Mistral-Small-24B-Instruct-2501-f32-00002-of-00003.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501-f32/Mistral-Small-24B-Instruct-2501-f32-00003-of-00003.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Small-24B-Instruct-2501.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fe018ef3c08096f2ff34c1ea5adc57aa60ea748e8e379045de1518176db1e148
size 8114050752

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fadc0fc19c902b942e73fe3f599c03d5fc6456d4664f7499c9e9104dfa4fc716
size 7478351552

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:baa57e2dac6f4e89495cabb098bcd4bf5f5906b40259b20440e92ef87a45aeae
size 7207032512

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9f12808f5ce5eff4a7b2afef0b5de8364823b379e47b2d040d418a12a4a8e711
size 10650949312

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:00de244eaf7895c9847a031cd6a80b8d0b4243e47fdd05472bf6da33c10e4165
size 9907115712

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0a386ad1ce3b21b6a6144d4233f9dce8dcff95cf38043b95b297819704902abe
size 13468014272

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4ebbddd57a7c7817fdf7a1a9b156bd96e8a2c5f4a65a11c633068ee2be7d52b9
size 12758914752

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7c6ec472d1ebda53fb95e6cadf8d58859ae814ac320d9a8a2275dfa8f4632632
size 8890324672

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:802de1a45a452eee74449747cdd88cfc6d64a9433176ce018beedd8115c8f952
size 9545684672

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ec9100f078695fd81cd7ce3f940c14bbd15768294b1283c76c2c3b7287376265
size 12400760512

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:15f1e4a554567837320af890e11b7eaa9eb9f929c5a2b5417277d1510fafe9bc
size 11474081472

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:591f636fc59d1291fae674714481e6bb2e65e57699efd2a918769823cf83ee88
size 10400274112

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7fa244ad7c3ef0c5cf07c1fe44851f340b43c88fdebcedb70e8e3ecd010b9f70
size 12987963072

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b3ecb85d6c4f1ecce19db5fed2d549778454a930b05381d24b715596e13c40ec
size 13494228672

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9541863ab5baa22a6d2b77d998b322bac8195566882dc8cb8195a4b791f68faf
size 14873106112

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ab7706916df1832700ef5abc53aac039d15f0f173acaa94bb22a2f13ca21fc5e
size 14831982272

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d1a6d049f09730c3f8ba26cf6b0b60c89790b5fdafa9a59c819acdfe93fffd1b
size 14333908672

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:919a8b31c63c6d349feef4f3bd0b74a3681cb2d15e2587ddedb928ab3023c2e4
size 13549278912

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2ae69c617a65170d2a2bb945304760239dfa101b9a2ee19800c6a81793e5b080
size 17178171072

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a9470aae6112f3dee88d7d2a4469c315696f09ce6d95b6de503cdc546b9713b4
size 16763983552

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9c388f675f79be147675584ba1a9151c0b0226fd2a0d041cfbf741d2291d5f3c
size 16304412352

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7c187a4f701ae209e5eabe656c4af071f25150f6fc147403cfe40bd7cbbb7c0a
size 19345938112

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ca2aeb8913df96621d06ae8957257eebd3fa78882bde08fdafc7399da4b487ff
size 19670996672

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b4a40c793acbea52ac00236e445a4a17c9443b6ca0910b6bc43ffb1fde6fa758
size 25054779072

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c5d082bbeae78ee960f4637c170e553bae85895390bc3326b45de6d62523b62c
size 47153518272

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:21d153fa065180a812676392211ce950ca4ac8895e2fa9f3c11add01a5864d79
size 39812490432

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8ccd2f0a03d83c27a3f49e81f89f27469c128abd3e1c44b4ce43e8494f89cd10
size 39804691808

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d58eba87a240114851230519aafaf253f9f4bfa4fa13a439ce9de4f5786b4995
size 14680313024

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7a3f3b3d232f0cb6e6c31c1d527e9cb47b247f257db59d60da07ef8000758368
size 10003538

184
README.md Normal file
View File

@@ -0,0 +1,184 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
extra_gated_description: If you want to learn more about how we process your personal
data, please read our <a href="https://mistral.ai/terms/">Privacy Policy</a>.
base_model: mistralai/Mistral-Small-24B-Instruct-2501
inference: false
language:
- en
- fr
- de
- es
- it
- pt
- zh
- ja
- ru
- ko
license: apache-2.0
---
## Llamacpp imatrix Quantizations of Mistral-Small-24B-Instruct-2501
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4585">b4585</a> for quantization.
Original model: https://huggingface.co/mistralai/Mistral-Small-24B-Instruct-2501
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
Run them directly with [llama.cpp](https://github.com/ggerganov/llama.cpp), or any other llama.cpp based project
## Prompt format
```
<s>[SYSTEM_PROMPT]{system_prompt}[/SYSTEM_PROMPT][INST]{prompt}[/INST]
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Mistral-Small-24B-Instruct-2501-f32.gguf](https://huggingface.co/bartowski/Mistral-Small-24B-Instruct-2501-GGUF/tree/main/Mistral-Small-24B-Instruct-2501-f32) | f32 | 94.30GB | true | Full F32 weights. |
| [Mistral-Small-24B-Instruct-2501-f16.gguf](https://huggingface.co/bartowski/Mistral-Small-24B-Instruct-2501-GGUF/blob/main/Mistral-Small-24B-Instruct-2501-f16.gguf) | f16 | 47.15GB | false | Full F16 weights. |
| [Mistral-Small-24B-Instruct-2501-Q8_0.gguf](https://huggingface.co/bartowski/Mistral-Small-24B-Instruct-2501-GGUF/blob/main/Mistral-Small-24B-Instruct-2501-Q8_0.gguf) | Q8_0 | 25.05GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Mistral-Small-24B-Instruct-2501-Q6_K_L.gguf](https://huggingface.co/bartowski/Mistral-Small-24B-Instruct-2501-GGUF/blob/main/Mistral-Small-24B-Instruct-2501-Q6_K_L.gguf) | Q6_K_L | 19.67GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Mistral-Small-24B-Instruct-2501-Q6_K.gguf](https://huggingface.co/bartowski/Mistral-Small-24B-Instruct-2501-GGUF/blob/main/Mistral-Small-24B-Instruct-2501-Q6_K.gguf) | Q6_K | 19.35GB | false | Very high quality, near perfect, *recommended*. |
| [Mistral-Small-24B-Instruct-2501-Q5_K_L.gguf](https://huggingface.co/bartowski/Mistral-Small-24B-Instruct-2501-GGUF/blob/main/Mistral-Small-24B-Instruct-2501-Q5_K_L.gguf) | Q5_K_L | 17.18GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Mistral-Small-24B-Instruct-2501-Q5_K_M.gguf](https://huggingface.co/bartowski/Mistral-Small-24B-Instruct-2501-GGUF/blob/main/Mistral-Small-24B-Instruct-2501-Q5_K_M.gguf) | Q5_K_M | 16.76GB | false | High quality, *recommended*. |
| [Mistral-Small-24B-Instruct-2501-Q5_K_S.gguf](https://huggingface.co/bartowski/Mistral-Small-24B-Instruct-2501-GGUF/blob/main/Mistral-Small-24B-Instruct-2501-Q5_K_S.gguf) | Q5_K_S | 16.30GB | false | High quality, *recommended*. |
| [Mistral-Small-24B-Instruct-2501-Q4_1.gguf](https://huggingface.co/bartowski/Mistral-Small-24B-Instruct-2501-GGUF/blob/main/Mistral-Small-24B-Instruct-2501-Q4_1.gguf) | Q4_1 | 14.87GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| [Mistral-Small-24B-Instruct-2501-Q4_K_L.gguf](https://huggingface.co/bartowski/Mistral-Small-24B-Instruct-2501-GGUF/blob/main/Mistral-Small-24B-Instruct-2501-Q4_K_L.gguf) | Q4_K_L | 14.83GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Mistral-Small-24B-Instruct-2501-Q4_K_M.gguf](https://huggingface.co/bartowski/Mistral-Small-24B-Instruct-2501-GGUF/blob/main/Mistral-Small-24B-Instruct-2501-Q4_K_M.gguf) | Q4_K_M | 14.33GB | false | Good quality, default size for most use cases, *recommended*. |
| [Mistral-Small-24B-Instruct-2501-Q4_K_S.gguf](https://huggingface.co/bartowski/Mistral-Small-24B-Instruct-2501-GGUF/blob/main/Mistral-Small-24B-Instruct-2501-Q4_K_S.gguf) | Q4_K_S | 13.55GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Mistral-Small-24B-Instruct-2501-Q4_0.gguf](https://huggingface.co/bartowski/Mistral-Small-24B-Instruct-2501-GGUF/blob/main/Mistral-Small-24B-Instruct-2501-Q4_0.gguf) | Q4_0 | 13.49GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| [Mistral-Small-24B-Instruct-2501-IQ4_NL.gguf](https://huggingface.co/bartowski/Mistral-Small-24B-Instruct-2501-GGUF/blob/main/Mistral-Small-24B-Instruct-2501-IQ4_NL.gguf) | IQ4_NL | 13.47GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [Mistral-Small-24B-Instruct-2501-Q3_K_XL.gguf](https://huggingface.co/bartowski/Mistral-Small-24B-Instruct-2501-GGUF/blob/main/Mistral-Small-24B-Instruct-2501-Q3_K_XL.gguf) | Q3_K_XL | 12.99GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Mistral-Small-24B-Instruct-2501-IQ4_XS.gguf](https://huggingface.co/bartowski/Mistral-Small-24B-Instruct-2501-GGUF/blob/main/Mistral-Small-24B-Instruct-2501-IQ4_XS.gguf) | IQ4_XS | 12.76GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Mistral-Small-24B-Instruct-2501-Q3_K_L.gguf](https://huggingface.co/bartowski/Mistral-Small-24B-Instruct-2501-GGUF/blob/main/Mistral-Small-24B-Instruct-2501-Q3_K_L.gguf) | Q3_K_L | 12.40GB | false | Lower quality but usable, good for low RAM availability. |
| [Mistral-Small-24B-Instruct-2501-Q3_K_M.gguf](https://huggingface.co/bartowski/Mistral-Small-24B-Instruct-2501-GGUF/blob/main/Mistral-Small-24B-Instruct-2501-Q3_K_M.gguf) | Q3_K_M | 11.47GB | false | Low quality. |
| [Mistral-Small-24B-Instruct-2501-IQ3_M.gguf](https://huggingface.co/bartowski/Mistral-Small-24B-Instruct-2501-GGUF/blob/main/Mistral-Small-24B-Instruct-2501-IQ3_M.gguf) | IQ3_M | 10.65GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Mistral-Small-24B-Instruct-2501-Q3_K_S.gguf](https://huggingface.co/bartowski/Mistral-Small-24B-Instruct-2501-GGUF/blob/main/Mistral-Small-24B-Instruct-2501-Q3_K_S.gguf) | Q3_K_S | 10.40GB | false | Low quality, not recommended. |
| [Mistral-Small-24B-Instruct-2501-IQ3_XS.gguf](https://huggingface.co/bartowski/Mistral-Small-24B-Instruct-2501-GGUF/blob/main/Mistral-Small-24B-Instruct-2501-IQ3_XS.gguf) | IQ3_XS | 9.91GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Mistral-Small-24B-Instruct-2501-Q2_K_L.gguf](https://huggingface.co/bartowski/Mistral-Small-24B-Instruct-2501-GGUF/blob/main/Mistral-Small-24B-Instruct-2501-Q2_K_L.gguf) | Q2_K_L | 9.55GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Mistral-Small-24B-Instruct-2501-Q2_K.gguf](https://huggingface.co/bartowski/Mistral-Small-24B-Instruct-2501-GGUF/blob/main/Mistral-Small-24B-Instruct-2501-Q2_K.gguf) | Q2_K | 8.89GB | false | Very low quality but surprisingly usable. |
| [Mistral-Small-24B-Instruct-2501-IQ2_M.gguf](https://huggingface.co/bartowski/Mistral-Small-24B-Instruct-2501-GGUF/blob/main/Mistral-Small-24B-Instruct-2501-IQ2_M.gguf) | IQ2_M | 8.11GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
| [Mistral-Small-24B-Instruct-2501-IQ2_S.gguf](https://huggingface.co/bartowski/Mistral-Small-24B-Instruct-2501-GGUF/blob/main/Mistral-Small-24B-Instruct-2501-IQ2_S.gguf) | IQ2_S | 7.48GB | false | Low quality, uses SOTA techniques to be usable. |
| [Mistral-Small-24B-Instruct-2501-IQ2_XS.gguf](https://huggingface.co/bartowski/Mistral-Small-24B-Instruct-2501-GGUF/blob/main/Mistral-Small-24B-Instruct-2501-IQ2_XS.gguf) | IQ2_XS | 7.21GB | false | Low quality, uses SOTA techniques to be usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Mistral-Small-24B-Instruct-2501-GGUF --include "Mistral-Small-24B-Instruct-2501-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Mistral-Small-24B-Instruct-2501-GGUF --include "Mistral-Small-24B-Instruct-2501-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Mistral-Small-24B-Instruct-2501-Q8_0) or download them all in place (./)
</details>
## ARM/AVX information
Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.
Now, however, there is something called "online repacking" for weights. details in [this PR](https://github.com/ggerganov/llama.cpp/pull/9921). If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.
As of llama.cpp build [b4282](https://github.com/ggerganov/llama.cpp/releases/tag/b4282) you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.
Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to [this PR](https://github.com/ggerganov/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information (deprecated</summary>
I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Thank you to LM Studio for sponsoring my work.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}