初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Pantheon-RP-1.6-12b-Nemo-KTO-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-09-12 16:40:12 +08:00
commit 9390f4ff9d
28 changed files with 263 additions and 0 deletions

60
.gitattributes vendored Normal file
View File

@@ -0,0 +1,60 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6-12b-Nemo-KTO-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6-12b-Nemo-KTO-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6-12b-Nemo-KTO-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6-12b-Nemo-KTO-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6-12b-Nemo-KTO-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6-12b-Nemo-KTO-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6-12b-Nemo-KTO-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6-12b-Nemo-KTO-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6-12b-Nemo-KTO-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6-12b-Nemo-KTO-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6-12b-Nemo-KTO-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6-12b-Nemo-KTO-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6-12b-Nemo-KTO-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6-12b-Nemo-KTO-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6-12b-Nemo-KTO-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6-12b-Nemo-KTO-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6-12b-Nemo-KTO-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6-12b-Nemo-KTO-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6-12b-Nemo-KTO-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6-12b-Nemo-KTO-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6-12b-Nemo-KTO-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6-12b-Nemo-KTO-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6-12b-Nemo-KTO-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6-12b-Nemo-KTO-f16.gguf filter=lfs diff=lfs merge=lfs -text
Pantheon-RP-1.6-12b-Nemo-KTO.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c1fadf0ba69c55ccc1916f6acbcb06cfb9b8cfbc39d9ea575d5324f31e5f706c
size 4435023712

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:28b1900eee5eb47ccd881c15bcdc22240645e91de391d76c91608548a897e050
size 5722232672

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5928ee6b9355f6d6a07cbc0abee25eba1327497295c0b37c3b8096b335a9a42d
size 5306488672

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1383d1b7fe874cc1dbd79ffb2d2389417fef29051e3a65725ada69fe7e832904
size 6742710112

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0a8ecf4daf7b2a410d0f7338376e8dc992e02d0e6c4da28be78b1a993e54119a
size 4791048032

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:57764f2bdff43b298af35389d3543a2dcef54bd0c324ee31a836fd764bcafc98
size 5446408032

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fb375ed852add9b1d23817ee3e335573cd6beba04cf0c594fa9b41b42ef2254a
size 6561503072

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2d1a410d84a6c1996967f7fe3e15c64aa00b27b6e589d4c20f386f6f121af139
size 6083090272

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:068a7021c763d4486814566b5d7496205e6c612a9c6426d36b155eca95eddc15
size 5534226272

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ceae997ab8859f93f3422d76e12206213c9d010167b56883f4ff1686f46ee0a1
size 7148705632

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6711919128a6b519b2e8e496f3b0e9fc4f4d169f02088ece7ac348b7c75055d8
size 7094638432

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d1882084bac2c20a17148ec35b4b99f6aa5eb55f6f15788d5ec7c911dd8da524
size 7071700832

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4028624321d473bc7da47df589b2b061fded21938e2cde984ffe600b4b1f85a1
size 7071700832

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:53acce43ceaf112086e64c9f264a765d8b541d8b8ffe10b921919814064e4c84
size 7071700832

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8dcc93a2a1b6ac207afabaffd804331cc9cf325c1a7d2d2d9a05c708d628c9fd
size 7975278432

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:baf570abc6b5fa5c26f25bde0ed5dfc08ac78d311a0f77dba12f0632ddae0f68
size 7477204832

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:174e6bf14588160552b56ec20a2f159fa206a16b9c1aaefcce89f0d0ec98f57c
size 7120197472

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a93333e694f62d600b91d8db5560c39b584374657a32e871b3b119fe4781ebf2
size 9141819232

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d21e1ada65e6616f4e876a7687de6437160f0a7b3329522dc2f5134c4041faa4
size 8727631712

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2dea599082b41fa10405352073b0705dc49879cee4a9d317d418868c45f7932a
size 8518735712

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e2daf694dc389b77bea4c914b05466d898cb09d3162a3cf704f170180d0e398f
size 10056210272

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d3a1f84075360feb3571bd0fb36e0cdae3d9a24070f8744d59497e7917a8adb6
size 10381268832

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:45d34463f0b8b49a7ad9071a3148ee5fbf9503a20152cffd01ebc8b7246b021b
size 13022369632

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f044a530efdaceaad955e769a6fb2c7e18372090ff0850484da80354c89a0446
size 24504276544

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:142a814424dea3d05f5712536929936e83097fe41f0d1aaa11eed3e881471c78
size 7054418

127
README.md Normal file
View File

@@ -0,0 +1,127 @@
---
base_model: Gryphe/Pantheon-RP-1.6-12b-Nemo-KTO
language:
- en
license: apache-2.0
pipeline_tag: text-generation
tags:
- instruct
- finetune
- chatml
- axolotl
- roleplay
quantized_by: bartowski
---
## Llamacpp imatrix Quantizations of Pantheon-RP-1.6-12b-Nemo-KTO
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b3634">b3634</a> for quantization.
Original model: https://huggingface.co/Gryphe/Pantheon-RP-1.6-12b-Nemo-KTO
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Pantheon-RP-1.6-12b-Nemo-KTO-f16.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6-12b-Nemo-KTO-GGUF/blob/main/Pantheon-RP-1.6-12b-Nemo-KTO-f16.gguf) | f16 | 24.50GB | false | Full F16 weights. |
| [Pantheon-RP-1.6-12b-Nemo-KTO-Q8_0.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6-12b-Nemo-KTO-GGUF/blob/main/Pantheon-RP-1.6-12b-Nemo-KTO-Q8_0.gguf) | Q8_0 | 13.02GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Pantheon-RP-1.6-12b-Nemo-KTO-Q6_K_L.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6-12b-Nemo-KTO-GGUF/blob/main/Pantheon-RP-1.6-12b-Nemo-KTO-Q6_K_L.gguf) | Q6_K_L | 10.38GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Pantheon-RP-1.6-12b-Nemo-KTO-Q6_K.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6-12b-Nemo-KTO-GGUF/blob/main/Pantheon-RP-1.6-12b-Nemo-KTO-Q6_K.gguf) | Q6_K | 10.06GB | false | Very high quality, near perfect, *recommended*. |
| [Pantheon-RP-1.6-12b-Nemo-KTO-Q5_K_L.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6-12b-Nemo-KTO-GGUF/blob/main/Pantheon-RP-1.6-12b-Nemo-KTO-Q5_K_L.gguf) | Q5_K_L | 9.14GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Pantheon-RP-1.6-12b-Nemo-KTO-Q5_K_M.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6-12b-Nemo-KTO-GGUF/blob/main/Pantheon-RP-1.6-12b-Nemo-KTO-Q5_K_M.gguf) | Q5_K_M | 8.73GB | false | High quality, *recommended*. |
| [Pantheon-RP-1.6-12b-Nemo-KTO-Q5_K_S.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6-12b-Nemo-KTO-GGUF/blob/main/Pantheon-RP-1.6-12b-Nemo-KTO-Q5_K_S.gguf) | Q5_K_S | 8.52GB | false | High quality, *recommended*. |
| [Pantheon-RP-1.6-12b-Nemo-KTO-Q4_K_L.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6-12b-Nemo-KTO-GGUF/blob/main/Pantheon-RP-1.6-12b-Nemo-KTO-Q4_K_L.gguf) | Q4_K_L | 7.98GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Pantheon-RP-1.6-12b-Nemo-KTO-Q4_K_M.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6-12b-Nemo-KTO-GGUF/blob/main/Pantheon-RP-1.6-12b-Nemo-KTO-Q4_K_M.gguf) | Q4_K_M | 7.48GB | false | Good quality, default size for must use cases, *recommended*. |
| [Pantheon-RP-1.6-12b-Nemo-KTO-Q3_K_XL.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6-12b-Nemo-KTO-GGUF/blob/main/Pantheon-RP-1.6-12b-Nemo-KTO-Q3_K_XL.gguf) | Q3_K_XL | 7.15GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Pantheon-RP-1.6-12b-Nemo-KTO-Q4_K_S.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6-12b-Nemo-KTO-GGUF/blob/main/Pantheon-RP-1.6-12b-Nemo-KTO-Q4_K_S.gguf) | Q4_K_S | 7.12GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Pantheon-RP-1.6-12b-Nemo-KTO-Q4_0.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6-12b-Nemo-KTO-GGUF/blob/main/Pantheon-RP-1.6-12b-Nemo-KTO-Q4_0.gguf) | Q4_0 | 7.09GB | false | Legacy format, generally not worth using over similarly sized formats |
| [Pantheon-RP-1.6-12b-Nemo-KTO-Q4_0_8_8.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6-12b-Nemo-KTO-GGUF/blob/main/Pantheon-RP-1.6-12b-Nemo-KTO-Q4_0_8_8.gguf) | Q4_0_8_8 | 7.07GB | false | Optimized for ARM and CPU inference, much faster than Q4_0 at similar quality. |
| [Pantheon-RP-1.6-12b-Nemo-KTO-Q4_0_4_8.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6-12b-Nemo-KTO-GGUF/blob/main/Pantheon-RP-1.6-12b-Nemo-KTO-Q4_0_4_8.gguf) | Q4_0_4_8 | 7.07GB | false | Optimized for ARM and CPU inference, much faster than Q4_0 at similar quality. |
| [Pantheon-RP-1.6-12b-Nemo-KTO-Q4_0_4_4.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6-12b-Nemo-KTO-GGUF/blob/main/Pantheon-RP-1.6-12b-Nemo-KTO-Q4_0_4_4.gguf) | Q4_0_4_4 | 7.07GB | false | Optimized for ARM and CPU inference, much faster than Q4_0 at similar quality. |
| [Pantheon-RP-1.6-12b-Nemo-KTO-IQ4_XS.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6-12b-Nemo-KTO-GGUF/blob/main/Pantheon-RP-1.6-12b-Nemo-KTO-IQ4_XS.gguf) | IQ4_XS | 6.74GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Pantheon-RP-1.6-12b-Nemo-KTO-Q3_K_L.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6-12b-Nemo-KTO-GGUF/blob/main/Pantheon-RP-1.6-12b-Nemo-KTO-Q3_K_L.gguf) | Q3_K_L | 6.56GB | false | Lower quality but usable, good for low RAM availability. |
| [Pantheon-RP-1.6-12b-Nemo-KTO-Q3_K_M.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6-12b-Nemo-KTO-GGUF/blob/main/Pantheon-RP-1.6-12b-Nemo-KTO-Q3_K_M.gguf) | Q3_K_M | 6.08GB | false | Low quality. |
| [Pantheon-RP-1.6-12b-Nemo-KTO-IQ3_M.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6-12b-Nemo-KTO-GGUF/blob/main/Pantheon-RP-1.6-12b-Nemo-KTO-IQ3_M.gguf) | IQ3_M | 5.72GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Pantheon-RP-1.6-12b-Nemo-KTO-Q3_K_S.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6-12b-Nemo-KTO-GGUF/blob/main/Pantheon-RP-1.6-12b-Nemo-KTO-Q3_K_S.gguf) | Q3_K_S | 5.53GB | false | Low quality, not recommended. |
| [Pantheon-RP-1.6-12b-Nemo-KTO-Q2_K_L.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6-12b-Nemo-KTO-GGUF/blob/main/Pantheon-RP-1.6-12b-Nemo-KTO-Q2_K_L.gguf) | Q2_K_L | 5.45GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Pantheon-RP-1.6-12b-Nemo-KTO-IQ3_XS.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6-12b-Nemo-KTO-GGUF/blob/main/Pantheon-RP-1.6-12b-Nemo-KTO-IQ3_XS.gguf) | IQ3_XS | 5.31GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Pantheon-RP-1.6-12b-Nemo-KTO-Q2_K.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6-12b-Nemo-KTO-GGUF/blob/main/Pantheon-RP-1.6-12b-Nemo-KTO-Q2_K.gguf) | Q2_K | 4.79GB | false | Very low quality but surprisingly usable. |
| [Pantheon-RP-1.6-12b-Nemo-KTO-IQ2_M.gguf](https://huggingface.co/bartowski/Pantheon-RP-1.6-12b-Nemo-KTO-GGUF/blob/main/Pantheon-RP-1.6-12b-Nemo-KTO-IQ2_M.gguf) | IQ2_M | 4.44GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset
Thank you ZeroWw for the inspiration to experiment with embed/output
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Pantheon-RP-1.6-12b-Nemo-KTO-GGUF --include "Pantheon-RP-1.6-12b-Nemo-KTO-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Pantheon-RP-1.6-12b-Nemo-KTO-GGUF --include "Pantheon-RP-1.6-12b-Nemo-KTO-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Pantheon-RP-1.6-12b-Nemo-KTO-Q8_0) or download them all in place (./)
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}