初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Pantheon-RP-1.6.1-12b-Nemo-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-19 13:25:06 +08:00
commit 44939233bc
28 changed files with 269 additions and 0 deletions

60
.gitattributes vendored Normal file
View File

@@ -0,0 +1,60 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6.1-12b-Nemo-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6.1-12b-Nemo-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6.1-12b-Nemo-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6.1-12b-Nemo-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6.1-12b-Nemo-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6.1-12b-Nemo-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6.1-12b-Nemo-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6.1-12b-Nemo-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6.1-12b-Nemo-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6.1-12b-Nemo-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6.1-12b-Nemo-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6.1-12b-Nemo-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6.1-12b-Nemo-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6.1-12b-Nemo-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6.1-12b-Nemo-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6.1-12b-Nemo-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6.1-12b-Nemo-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6.1-12b-Nemo-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6.1-12b-Nemo-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6.1-12b-Nemo-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6.1-12b-Nemo-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6.1-12b-Nemo-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6.1-12b-Nemo-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6.1-12b-Nemo-f16.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6.1-12b-Nemo.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:417fdb937feef66815a28adcdc9a4f51da42e93bfbe1578fc753597fa4811315
size 4435023456

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:838bea83a42ac7de910ff806068a2713f6fb2dad30d57f9d0767ddc89647fc0a
size 5722232416

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ea2af8ee16d3349f826d0af91321e1f0dd51e0239086af6e96d707189d8d4959
size 5306488416

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a5795354938bf51940ab3eb68bb9025709fda9c9329b261d0c25d5ce13858774
size 6742709856

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:27de6bcd7e34585476a4f00e5fe427aebc07714c8215dd4aaadcf9e74cb4100a
size 4791047776

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d88c195b0dbb458b2d9876f4def1aa6e4b33e2fac3820b302caa8496a5c25495
size 5446407776

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:582c2b214f73cd350ae6a406bdf240ae5a630e71778c971cac314cbb3a460832
size 6561502816

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6cd8a32ca3894d17b6f001596fdcca4b02c2762db145643467f6c996d380e54c
size 6083090016

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7947198544469651ee13e60557d62d5bdd2752607603e48f64e8baf617ce1611
size 5534226016

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7fb95c91b7a78dd0a88b19debe8f79d7d3e3d326560f2380f4ba0f7825e65ba6
size 7148705376

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3c078cd024c2a895e331705fe3c134bdbf9fb84b85f55ad9d6c179bbb0038e21
size 7094638176

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7130fe21d9af41d7f4268e7431b79b1ac59c1f111b2edc58678e917a3f702299
size 7071700576

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2bc21f23df6777443f95fcec3126f20c43353c58e9f056c49b01ff9b34474bfd
size 7071700576

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f221815c8bdc628936da58e07ea90037a68c4d018af0be3a6e581ec22be534cc
size 7071700576

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7869b7d57e931616c8c5500021d92e64bab073a7f682c9fb591c03ef84b0dbdb
size 7975278176

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b61b1410c319ccdb0be04a08322ccf4619625bff57ac77328ea33b09e61611bc
size 7477204576

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:feb337725c4e9a0769fd43eb4e472187878d09db7d8051ff16efd8c085b64f2e
size 7120197216

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3d5c1659fd7f8ea288b97b4814beeac5c01b9f8d70f646bdaf782b1469b85f29
size 9141818976

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9f91b05b92a86c05633357ef4a9057ffb93ce29a3d51b87b22c839de2d46151f
size 8727631456

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b6eb8aa5dadb6e05975f91277ea5522c2961b0ebdce2d2e539eca88c42b3bba0
size 8518735456

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c7faef2826e252bdd767fdd995670c36234158309d25c292b18698e394328f09
size 10056210016

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7e4f47e6f2f00a86bb2f3346e7e5099e31474b138317da28c342ffed92551152
size 10381268576

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:28497e5c91f0de2fa822cb2b33a7cfc877ce43810f5e095873b56460f9ba2e48
size 13022369376

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:aae2d498ad51632e3d154bd2997ab1018e2984412750cc84ef26e9b8ea58861b
size 24504276256

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cd0c3afa89309f0a2c197239f4503c08828290215bfa3a9335108a806a144361
size 7054418

133
README.md Normal file
View File

@@ -0,0 +1,133 @@
---
base_model: Gryphe/Pantheon-RP-1.6.1-12b-Nemo
language:
- en
license: apache-2.0
pipeline_tag: text-generation
tags:
- instruct
- finetune
- chatml
- axolotl
- roleplay
quantized_by: bartowski
---
## Llamacpp imatrix Quantizations of Pantheon-RP-1.6.1-12b-Nemo
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b3658">b3658</a> for quantization.
Original model: https://huggingface.co/Gryphe/Pantheon-RP-1.6.1-12b-Nemo
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Pantheon-RP-1.6.1-12b-Nemo-f16.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6.1-12b-Nemo-GGUF/blob/main/Pantheon-RP-1.6.1-12b-Nemo-f16.gguf) | f16 | 24.50GB | false | Full F16 weights. |
| [Pantheon-RP-1.6.1-12b-Nemo-Q8_0.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6.1-12b-Nemo-GGUF/blob/main/Pantheon-RP-1.6.1-12b-Nemo-Q8_0.gguf) | Q8_0 | 13.02GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Pantheon-RP-1.6.1-12b-Nemo-Q6_K_L.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6.1-12b-Nemo-GGUF/blob/main/Pantheon-RP-1.6.1-12b-Nemo-Q6_K_L.gguf) | Q6_K_L | 10.38GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Pantheon-RP-1.6.1-12b-Nemo-Q6_K.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6.1-12b-Nemo-GGUF/blob/main/Pantheon-RP-1.6.1-12b-Nemo-Q6_K.gguf) | Q6_K | 10.06GB | false | Very high quality, near perfect, *recommended*. |
| [Pantheon-RP-1.6.1-12b-Nemo-Q5_K_L.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6.1-12b-Nemo-GGUF/blob/main/Pantheon-RP-1.6.1-12b-Nemo-Q5_K_L.gguf) | Q5_K_L | 9.14GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Pantheon-RP-1.6.1-12b-Nemo-Q5_K_M.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6.1-12b-Nemo-GGUF/blob/main/Pantheon-RP-1.6.1-12b-Nemo-Q5_K_M.gguf) | Q5_K_M | 8.73GB | false | High quality, *recommended*. |
| [Pantheon-RP-1.6.1-12b-Nemo-Q5_K_S.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6.1-12b-Nemo-GGUF/blob/main/Pantheon-RP-1.6.1-12b-Nemo-Q5_K_S.gguf) | Q5_K_S | 8.52GB | false | High quality, *recommended*. |
| [Pantheon-RP-1.6.1-12b-Nemo-Q4_K_L.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6.1-12b-Nemo-GGUF/blob/main/Pantheon-RP-1.6.1-12b-Nemo-Q4_K_L.gguf) | Q4_K_L | 7.98GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Pantheon-RP-1.6.1-12b-Nemo-Q4_K_M.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6.1-12b-Nemo-GGUF/blob/main/Pantheon-RP-1.6.1-12b-Nemo-Q4_K_M.gguf) | Q4_K_M | 7.48GB | false | Good quality, default size for must use cases, *recommended*. |
| [Pantheon-RP-1.6.1-12b-Nemo-Q3_K_XL.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6.1-12b-Nemo-GGUF/blob/main/Pantheon-RP-1.6.1-12b-Nemo-Q3_K_XL.gguf) | Q3_K_XL | 7.15GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Pantheon-RP-1.6.1-12b-Nemo-Q4_K_S.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6.1-12b-Nemo-GGUF/blob/main/Pantheon-RP-1.6.1-12b-Nemo-Q4_K_S.gguf) | Q4_K_S | 7.12GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Pantheon-RP-1.6.1-12b-Nemo-Q4_0.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6.1-12b-Nemo-GGUF/blob/main/Pantheon-RP-1.6.1-12b-Nemo-Q4_0.gguf) | Q4_0 | 7.09GB | false | Legacy format, generally not worth using over similarly sized formats |
| [Pantheon-RP-1.6.1-12b-Nemo-Q4_0_8_8.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6.1-12b-Nemo-GGUF/blob/main/Pantheon-RP-1.6.1-12b-Nemo-Q4_0_8_8.gguf) | Q4_0_8_8 | 7.07GB | false | Optimized for ARM inference. Requires 'sve' support (see link below). |
| [Pantheon-RP-1.6.1-12b-Nemo-Q4_0_4_8.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6.1-12b-Nemo-GGUF/blob/main/Pantheon-RP-1.6.1-12b-Nemo-Q4_0_4_8.gguf) | Q4_0_4_8 | 7.07GB | false | Optimized for ARM inference. Requires 'i8mm' support (see link below). |
| [Pantheon-RP-1.6.1-12b-Nemo-Q4_0_4_4.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6.1-12b-Nemo-GGUF/blob/main/Pantheon-RP-1.6.1-12b-Nemo-Q4_0_4_4.gguf) | Q4_0_4_4 | 7.07GB | false | Optimized for ARM inference. Should work well on all ARM chips, pick this if you're unsure. |
| [Pantheon-RP-1.6.1-12b-Nemo-IQ4_XS.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6.1-12b-Nemo-GGUF/blob/main/Pantheon-RP-1.6.1-12b-Nemo-IQ4_XS.gguf) | IQ4_XS | 6.74GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Pantheon-RP-1.6.1-12b-Nemo-Q3_K_L.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6.1-12b-Nemo-GGUF/blob/main/Pantheon-RP-1.6.1-12b-Nemo-Q3_K_L.gguf) | Q3_K_L | 6.56GB | false | Lower quality but usable, good for low RAM availability. |
| [Pantheon-RP-1.6.1-12b-Nemo-Q3_K_M.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6.1-12b-Nemo-GGUF/blob/main/Pantheon-RP-1.6.1-12b-Nemo-Q3_K_M.gguf) | Q3_K_M | 6.08GB | false | Low quality. |
| [Pantheon-RP-1.6.1-12b-Nemo-IQ3_M.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6.1-12b-Nemo-GGUF/blob/main/Pantheon-RP-1.6.1-12b-Nemo-IQ3_M.gguf) | IQ3_M | 5.72GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Pantheon-RP-1.6.1-12b-Nemo-Q3_K_S.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6.1-12b-Nemo-GGUF/blob/main/Pantheon-RP-1.6.1-12b-Nemo-Q3_K_S.gguf) | Q3_K_S | 5.53GB | false | Low quality, not recommended. |
| [Pantheon-RP-1.6.1-12b-Nemo-Q2_K_L.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6.1-12b-Nemo-GGUF/blob/main/Pantheon-RP-1.6.1-12b-Nemo-Q2_K_L.gguf) | Q2_K_L | 5.45GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Pantheon-RP-1.6.1-12b-Nemo-IQ3_XS.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6.1-12b-Nemo-GGUF/blob/main/Pantheon-RP-1.6.1-12b-Nemo-IQ3_XS.gguf) | IQ3_XS | 5.31GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Pantheon-RP-1.6.1-12b-Nemo-Q2_K.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6.1-12b-Nemo-GGUF/blob/main/Pantheon-RP-1.6.1-12b-Nemo-Q2_K.gguf) | Q2_K | 4.79GB | false | Very low quality but surprisingly usable. |
| [Pantheon-RP-1.6.1-12b-Nemo-IQ2_M.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6.1-12b-Nemo-GGUF/blob/main/Pantheon-RP-1.6.1-12b-Nemo-IQ2_M.gguf) | IQ2_M | 4.44GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Pantheon-RP-1.6.1-12b-Nemo-GGUF --include "Pantheon-RP-1.6.1-12b-Nemo-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Pantheon-RP-1.6.1-12b-Nemo-GGUF --include "Pantheon-RP-1.6.1-12b-Nemo-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Pantheon-RP-1.6.1-12b-Nemo-Q8_0) or download them all in place (./)
## Q4_0_X_X
These are *NOT* for Metal (Apple) offloading, only ARM chips.
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset
Thank you ZeroWw for the inspiration to experiment with embed/output
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}