初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-19 13:10:05 +08:00
commit cec9838b84
28 changed files with 257 additions and 0 deletions

60
.gitattributes vendored Normal file
View File

@@ -0,0 +1,60 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-12B-ArliAI-RPMax-v1.1-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-12B-ArliAI-RPMax-v1.1-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-12B-ArliAI-RPMax-v1.1-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-12B-ArliAI-RPMax-v1.1-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-12B-ArliAI-RPMax-v1.1-f16.gguf filter=lfs diff=lfs merge=lfs -text
Mistral-Nemo-12B-ArliAI-RPMax-v1.1.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:de70269977da091e5dcd2b0d3836fd1abd2c9a00297bed5ca72f2885a89b3662
size 4435026688

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7c0e94799c1c5b5ef6d4c8adbda16364d8a5182b9649b01e5e8c4b0bfb91745a
size 5722235648

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cbe0c58158888bcb1d4f38d38e059cc307e91b621d5daa5daa889a5a8dbef400
size 5306491648

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d5259c6fe43a07d8ce691d792adddc786836105741211ea874f4f46822db5bfd
size 6742713088

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9c616bddf5fdfa2956b08803e9a715358813085e24d8f9bcfb7bbf774c3dc24d
size 4791051008

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d904ae3a14dfa5adfedb1d502cb032c7a0abfab78c35e66da5a092e079e8c50d
size 5446411008

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:699bfb56186695f75ab227b8c9f1a1e491b0c3b75f3552951dcf3c69817f2a92
size 6561506048

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:56f994545b528b1a73c965b54743c020adf30ff00ef40eff9fd4739018b0c6e6
size 6083093248

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9d6c7069aaea8826c05248e5ec912d64933684d2221a4c93f4d055addf5729a2
size 5534229248

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:18cd86279b000e3dc798a48e3fc0ee87d1cd49a6604c20d4be39ba38cb33dd3c
size 7148708608

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1d152f4ad393650c9b854dd8ace4b68b14902e3971b38bcc96ba68017befcf13
size 7094641408

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5d1aa0f1e42d97d927acfa42aa264a758086c5486fbe01d93c38d06ed56ae6f6
size 7071703808

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:93c952888dce57dc0b3884cd6fda0454df6553ce033aad6fbe5e696a223c3dc6
size 7071703808

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ca4bc60204e70244005a46d317312f3a3bd8bc5d4276b66147fd88f46493cfea
size 7071703808

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7d2f2c1f151d62abf74b98d506314b0ce5c2ace9ca3b3be8f45c746f9076e3e9
size 7975281408

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ae911014c6ca813c3fc661ce967e18895092ffa5fd4b4528c26edb3d9765efc0
size 7477207808

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:33d9dfe921db20bbe808d6b88c80c3f4e75e639a731f268dd20a5cc1f78ed636
size 7120200448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9988b2c71f461c913add5c5ba74ace7bad3588720bfe48a68849c6c6ce75caee
size 9141822208

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f57aaf1b39a044ef941aae7e08e2eacfd10898c85fbdbd592462c347c10a0c40
size 8727634688

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4e75815bec67fd6077c107058b57fa292851d735402fb7153004fe4fedb1d546
size 8518738688

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7c0e557d51658ed392ff8a522d81db6fbfa520ba3eb3335dc0f17d59170b3041
size 10056213248

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a764df33b45a38eeedf861aa615c8f614a1fc6fe6a9c893e6b151d444365730d
size 10381271808

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5806560406d8b210e2663977c4f32095274918f45af41635bca3f4936508f405
size 13022372608

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:eef3009178869fbddfe383b1e73ffcf3a17424983a8afeca909bd4e8da3a1fc5
size 24504279488

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b771821867aebd84e0b43173426454d2c4579413c58e8a8294c7c82598bacdd7
size 7054418

121
README.md Normal file
View File

@@ -0,0 +1,121 @@
---
base_model: ArliAI/Mistral-Nemo-12B-ArliAI-RPMax-v1.1
license: apache-2.0
pipeline_tag: text-generation
quantized_by: bartowski
---
## Llamacpp imatrix Quantizations of Mistral-Nemo-12B-ArliAI-RPMax-v1.1
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b3715">b3715</a> for quantization.
Original model: https://huggingface.co/ArliAI/Mistral-Nemo-12B-ArliAI-RPMax-v1.1
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<s>[INST]{prompt}[/INST]
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Mistral-Nemo-12B-ArliAI-RPMax-v1.1-f16.gguf](https://huggingface.co/bartowski/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-GGUF/blob/main/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-f16.gguf) | f16 | 24.50GB | false | Full F16 weights. |
| [Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q8_0.gguf](https://huggingface.co/bartowski/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-GGUF/blob/main/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q8_0.gguf) | Q8_0 | 13.02GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q6_K_L.gguf](https://huggingface.co/bartowski/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-GGUF/blob/main/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q6_K_L.gguf) | Q6_K_L | 10.38GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q6_K.gguf](https://huggingface.co/bartowski/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-GGUF/blob/main/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q6_K.gguf) | Q6_K | 10.06GB | false | Very high quality, near perfect, *recommended*. |
| [Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q5_K_L.gguf](https://huggingface.co/bartowski/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-GGUF/blob/main/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q5_K_L.gguf) | Q5_K_L | 9.14GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q5_K_M.gguf](https://huggingface.co/bartowski/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-GGUF/blob/main/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q5_K_M.gguf) | Q5_K_M | 8.73GB | false | High quality, *recommended*. |
| [Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q5_K_S.gguf](https://huggingface.co/bartowski/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-GGUF/blob/main/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q5_K_S.gguf) | Q5_K_S | 8.52GB | false | High quality, *recommended*. |
| [Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q4_K_L.gguf](https://huggingface.co/bartowski/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-GGUF/blob/main/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q4_K_L.gguf) | Q4_K_L | 7.98GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q4_K_M.gguf](https://huggingface.co/bartowski/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-GGUF/blob/main/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q4_K_M.gguf) | Q4_K_M | 7.48GB | false | Good quality, default size for must use cases, *recommended*. |
| [Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q3_K_XL.gguf](https://huggingface.co/bartowski/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-GGUF/blob/main/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q3_K_XL.gguf) | Q3_K_XL | 7.15GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q4_K_S.gguf](https://huggingface.co/bartowski/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-GGUF/blob/main/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q4_K_S.gguf) | Q4_K_S | 7.12GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q4_0.gguf](https://huggingface.co/bartowski/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-GGUF/blob/main/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q4_0.gguf) | Q4_0 | 7.09GB | false | Legacy format, generally not worth using over similarly sized formats |
| [Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q4_0_8_8.gguf](https://huggingface.co/bartowski/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-GGUF/blob/main/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q4_0_8_8.gguf) | Q4_0_8_8 | 7.07GB | false | Optimized for ARM inference. Requires 'sve' support (see link below). |
| [Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q4_0_4_8.gguf](https://huggingface.co/bartowski/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-GGUF/blob/main/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q4_0_4_8.gguf) | Q4_0_4_8 | 7.07GB | false | Optimized for ARM inference. Requires 'i8mm' support (see link below). |
| [Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q4_0_4_4.gguf](https://huggingface.co/bartowski/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-GGUF/blob/main/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q4_0_4_4.gguf) | Q4_0_4_4 | 7.07GB | false | Optimized for ARM inference. Should work well on all ARM chips, pick this if you're unsure. |
| [Mistral-Nemo-12B-ArliAI-RPMax-v1.1-IQ4_XS.gguf](https://huggingface.co/bartowski/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-GGUF/blob/main/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-IQ4_XS.gguf) | IQ4_XS | 6.74GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q3_K_L.gguf](https://huggingface.co/bartowski/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-GGUF/blob/main/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q3_K_L.gguf) | Q3_K_L | 6.56GB | false | Lower quality but usable, good for low RAM availability. |
| [Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q3_K_M.gguf](https://huggingface.co/bartowski/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-GGUF/blob/main/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q3_K_M.gguf) | Q3_K_M | 6.08GB | false | Low quality. |
| [Mistral-Nemo-12B-ArliAI-RPMax-v1.1-IQ3_M.gguf](https://huggingface.co/bartowski/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-GGUF/blob/main/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-IQ3_M.gguf) | IQ3_M | 5.72GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q3_K_S.gguf](https://huggingface.co/bartowski/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-GGUF/blob/main/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q3_K_S.gguf) | Q3_K_S | 5.53GB | false | Low quality, not recommended. |
| [Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q2_K_L.gguf](https://huggingface.co/bartowski/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-GGUF/blob/main/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q2_K_L.gguf) | Q2_K_L | 5.45GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Mistral-Nemo-12B-ArliAI-RPMax-v1.1-IQ3_XS.gguf](https://huggingface.co/bartowski/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-GGUF/blob/main/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-IQ3_XS.gguf) | IQ3_XS | 5.31GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q2_K.gguf](https://huggingface.co/bartowski/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-GGUF/blob/main/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q2_K.gguf) | Q2_K | 4.79GB | false | Very low quality but surprisingly usable. |
| [Mistral-Nemo-12B-ArliAI-RPMax-v1.1-IQ2_M.gguf](https://huggingface.co/bartowski/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-GGUF/blob/main/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-IQ2_M.gguf) | IQ2_M | 4.44GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-GGUF --include "Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Mistral-Nemo-12B-ArliAI-RPMax-v1.1-GGUF --include "Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Mistral-Nemo-12B-ArliAI-RPMax-v1.1-Q8_0) or download them all in place (./)
## Q4_0_X_X
These are *NOT* for Metal (Apple) offloading, only ARM chips.
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset
Thank you ZeroWw for the inspiration to experiment with embed/output
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}