初始化项目,由ModelHub XC社区提供模型

Model: bartowski/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-19 09:51:16 +08:00
commit d7942b2901
28 changed files with 322 additions and 0 deletions

61
.gitattributes vendored Normal file
View File

@@ -0,0 +1,61 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q4_1.gguf filter=lfs diff=lfs merge=lfs -text
ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-IQ4_NL.gguf filter=lfs diff=lfs merge=lfs -text
ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-IQ3_XXS.gguf filter=lfs diff=lfs merge=lfs -text
ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-IQ2_S.gguf filter=lfs diff=lfs merge=lfs -text
ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-bf16.gguf filter=lfs diff=lfs merge=lfs -text
ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0.imatrix filter=lfs diff=lfs merge=lfs -text

183
README.md Normal file
View File

@@ -0,0 +1,183 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
language:
- en
base_model_relation: quantized
tags:
- nsfw
- explicit
- roleplay
- unaligned
- dangerous
- ERP
license: gemma
base_model: ReadyArt/The-Omega-Abomination-Gemma3-12B-v1.0
---
## Llamacpp imatrix Quantizations of The-Omega-Abomination-Gemma3-12B-v1.0 by ReadyArt
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b5074">b5074</a> for quantization.
Original model: https://huggingface.co/ReadyArt/The-Omega-Abomination-Gemma3-12B-v1.0
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
Run them directly with [llama.cpp](https://github.com/ggerganov/llama.cpp), or any other llama.cpp based project
## Prompt format
```
<bos><start_of_turn>user
{system_prompt}
{prompt}<end_of_turn>
<start_of_turn>model
<end_of_turn>
<start_of_turn>model
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [The-Omega-Abomination-Gemma3-12B-v1.0-bf16.gguf](https://huggingface.co/bartowski/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-GGUF/blob/main/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-bf16.gguf) | bf16 | 23.54GB | false | Full BF16 weights. |
| [The-Omega-Abomination-Gemma3-12B-v1.0-Q8_0.gguf](https://huggingface.co/bartowski/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-GGUF/blob/main/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q8_0.gguf) | Q8_0 | 12.51GB | false | Extremely high quality, generally unneeded but max available quant. |
| [The-Omega-Abomination-Gemma3-12B-v1.0-Q6_K_L.gguf](https://huggingface.co/bartowski/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-GGUF/blob/main/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q6_K_L.gguf) | Q6_K_L | 9.90GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [The-Omega-Abomination-Gemma3-12B-v1.0-Q6_K.gguf](https://huggingface.co/bartowski/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-GGUF/blob/main/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q6_K.gguf) | Q6_K | 9.66GB | false | Very high quality, near perfect, *recommended*. |
| [The-Omega-Abomination-Gemma3-12B-v1.0-Q5_K_L.gguf](https://huggingface.co/bartowski/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-GGUF/blob/main/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q5_K_L.gguf) | Q5_K_L | 8.69GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [The-Omega-Abomination-Gemma3-12B-v1.0-Q5_K_M.gguf](https://huggingface.co/bartowski/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-GGUF/blob/main/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q5_K_M.gguf) | Q5_K_M | 8.44GB | false | High quality, *recommended*. |
| [The-Omega-Abomination-Gemma3-12B-v1.0-Q5_K_S.gguf](https://huggingface.co/bartowski/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-GGUF/blob/main/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q5_K_S.gguf) | Q5_K_S | 8.23GB | false | High quality, *recommended*. |
| [The-Omega-Abomination-Gemma3-12B-v1.0-Q4_1.gguf](https://huggingface.co/bartowski/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-GGUF/blob/main/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q4_1.gguf) | Q4_1 | 7.56GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| [The-Omega-Abomination-Gemma3-12B-v1.0-Q4_K_L.gguf](https://huggingface.co/bartowski/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-GGUF/blob/main/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q4_K_L.gguf) | Q4_K_L | 7.54GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [The-Omega-Abomination-Gemma3-12B-v1.0-Q4_K_M.gguf](https://huggingface.co/bartowski/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-GGUF/blob/main/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q4_K_M.gguf) | Q4_K_M | 7.30GB | false | Good quality, default size for most use cases, *recommended*. |
| [The-Omega-Abomination-Gemma3-12B-v1.0-Q4_K_S.gguf](https://huggingface.co/bartowski/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-GGUF/blob/main/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q4_K_S.gguf) | Q4_K_S | 6.94GB | false | Slightly lower quality with more space savings, *recommended*. |
| [The-Omega-Abomination-Gemma3-12B-v1.0-Q4_0.gguf](https://huggingface.co/bartowski/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-GGUF/blob/main/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q4_0.gguf) | Q4_0 | 6.91GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| [The-Omega-Abomination-Gemma3-12B-v1.0-IQ4_NL.gguf](https://huggingface.co/bartowski/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-GGUF/blob/main/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-IQ4_NL.gguf) | IQ4_NL | 6.89GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [The-Omega-Abomination-Gemma3-12B-v1.0-Q3_K_XL.gguf](https://huggingface.co/bartowski/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-GGUF/blob/main/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q3_K_XL.gguf) | Q3_K_XL | 6.72GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [The-Omega-Abomination-Gemma3-12B-v1.0-IQ4_XS.gguf](https://huggingface.co/bartowski/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-GGUF/blob/main/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-IQ4_XS.gguf) | IQ4_XS | 6.55GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [The-Omega-Abomination-Gemma3-12B-v1.0-Q3_K_L.gguf](https://huggingface.co/bartowski/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-GGUF/blob/main/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q3_K_L.gguf) | Q3_K_L | 6.48GB | false | Lower quality but usable, good for low RAM availability. |
| [The-Omega-Abomination-Gemma3-12B-v1.0-Q3_K_M.gguf](https://huggingface.co/bartowski/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-GGUF/blob/main/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q3_K_M.gguf) | Q3_K_M | 6.01GB | false | Low quality. |
| [The-Omega-Abomination-Gemma3-12B-v1.0-IQ3_M.gguf](https://huggingface.co/bartowski/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-GGUF/blob/main/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-IQ3_M.gguf) | IQ3_M | 5.66GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [The-Omega-Abomination-Gemma3-12B-v1.0-Q3_K_S.gguf](https://huggingface.co/bartowski/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-GGUF/blob/main/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q3_K_S.gguf) | Q3_K_S | 5.46GB | false | Low quality, not recommended. |
| [The-Omega-Abomination-Gemma3-12B-v1.0-IQ3_XS.gguf](https://huggingface.co/bartowski/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-GGUF/blob/main/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-IQ3_XS.gguf) | IQ3_XS | 5.21GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [The-Omega-Abomination-Gemma3-12B-v1.0-Q2_K_L.gguf](https://huggingface.co/bartowski/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-GGUF/blob/main/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q2_K_L.gguf) | Q2_K_L | 5.01GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [The-Omega-Abomination-Gemma3-12B-v1.0-IQ3_XXS.gguf](https://huggingface.co/bartowski/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-GGUF/blob/main/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-IQ3_XXS.gguf) | IQ3_XXS | 4.78GB | false | Lower quality, new method with decent performance, comparable to Q3 quants. |
| [The-Omega-Abomination-Gemma3-12B-v1.0-Q2_K.gguf](https://huggingface.co/bartowski/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-GGUF/blob/main/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q2_K.gguf) | Q2_K | 4.77GB | false | Very low quality but surprisingly usable. |
| [The-Omega-Abomination-Gemma3-12B-v1.0-IQ2_M.gguf](https://huggingface.co/bartowski/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-GGUF/blob/main/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-IQ2_M.gguf) | IQ2_M | 4.31GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
| [The-Omega-Abomination-Gemma3-12B-v1.0-IQ2_S.gguf](https://huggingface.co/bartowski/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-GGUF/blob/main/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-IQ2_S.gguf) | IQ2_S | 4.02GB | false | Low quality, uses SOTA techniques to be usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-GGUF --include "ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-GGUF --include "ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (ReadyArt_The-Omega-Abomination-Gemma3-12B-v1.0-Q8_0) or download them all in place (./)
</details>
## ARM/AVX information
Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.
Now, however, there is something called "online repacking" for weights. details in [this PR](https://github.com/ggerganov/llama.cpp/pull/9921). If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.
As of llama.cpp build [b4282](https://github.com/ggerganov/llama.cpp/releases/tag/b4282) you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.
Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to [this PR](https://github.com/ggerganov/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information (deprecated</summary>
I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Thank you to LM Studio for sponsoring my work.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a3ac36cf664677c2cf8c12272bb44fc8577fc10f5146106ecb5e8ed9fa636f26
size 4310293440

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e557f196b7f228fe3439a74a84e1831720c61bfe3028b1dcb0c05a3c4ad6c630
size 4020542400

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1349e7eb1a720de9e51c21612621565868b917029f99768a712b980cf676a1ab
size 5655522752

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8d0ba70e792ae5271df0e2f43e21a655317126b0e9bf676512df3c67f3412ee9
size 5205966272

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3a0e4377217bc0f52420b19b2acd17a628cb341b834f3c50509c708e5ffe5a9f
size 4784733120

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9cc711cb8bbdab2e2ddd9b50e0f4072b27012e800bc11e0e033ece70b9adaa45
size 6886964672

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:46ba00f580206cc39cf955930353a2778bf8cc57d98cbf179aca8c2d4be42b3e
size 6550764992

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f7d110766acc2f9f81e9b673dd5c6746755238b53cbdd0f078e457241530fb07
size 4768021952

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e212c2e16ec35cfb8e3b0137d704e25958a04717474f205e637e7e92e4884679
size 5011816800

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:89417358c11d9d50371985c312b0bfcead11874f592ea7ec8c0f7e57b3a004f3
size 6479986112

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:32309947647b195b4e45ebfebca381fb598f1c927903eacc24098cae3e0c54bf
size 6008618432

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2e913cbf53ce9b4e7a3d5e0b4880fcd8d824915480d051d7bb9dcaec4b2d47b5
size 5458116032

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2e44c8606c03f6c0664ccab822c813212df457bc658c23e7d32bd1d505cce1b0
size 6723780960

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f1d718c6c752d6086fc8448239a7bba69423d92e4eac3d5625b41c379ee39e0c
size 6909083072

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2d8deb707a2fe3ccbe24cbe3a35d2db5f3cb0c83acfad983f7a8db660abff0b6
size 7559364032

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:765437972a78eec8166c20e764ac4114ac6b41c66bd9e5f99af00697f4c1fece
size 7544373600

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:66bc9fe47325351a2bbc910198bcca98c29679d5d15f2846ba5c2cbcafd8b67d
size 7300578752

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f5e4d2c1393f3b08726c5e4dbb1bfe3c29250ac8abd15f92043d9756f7427fbd
size 6935133632

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b9dc5cf97606ec90b67a1e91cc288aaea2f8d52691e80d04f60f1a9d782fde93
size 8688632160

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:95bfb9ce543a99040d6fbd81625ed7e8f4ecd74350495c66605b464299de9790
size 8444837312

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b15dfb35a0d206f1d5b1443d74c34bda44bda073c34cc01bd470793d151d3b8e
size 8231763392

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5850b0d6bcfa5e1d7f1c0753dd116aa37ed63d0515380a1536ad3ba36d97c855
size 9660612032

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:853b5e0c62a538173c58e9c28d6ee1066d07589f06965b0cbfda2eae6bc6a28e
size 9904406880

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fdb7c0ad04cc8afa621807195c0172b3c99461d5b424cb21402e9f9f698c761a
size 12509954400

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2da745588aacc4512c1711385bdf799a40b929d837ddf784dd5ef19ceca40bf5
size 23539666464

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:678fb1d791b0a8bea451bf6caf2f04fcf1baf52badd3dc03ec8c38a87c22cb4d
size 7433114