初始化项目,由ModelHub XC社区提供模型

Model: bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-13 10:46:09 +08:00
commit f528452d25
31 changed files with 319 additions and 0 deletions

64
.gitattributes vendored Normal file
View File

@@ -0,0 +1,64 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q4_1.gguf filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-MXFP4_MOE.gguf filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-IQ4_NL.gguf filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-IQ3_XXS.gguf filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-IQ2_S.gguf filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-IQ2_XS.gguf filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-IQ2_XXS.gguf filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-bf16.gguf filter=lfs diff=lfs merge=lfs -text
huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-imatrix.gguf filter=lfs diff=lfs merge=lfs -text

168
README.md Normal file
View File

@@ -0,0 +1,168 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
base_model_relation: quantized
base_model: huihui-ai/Huihui-gpt-oss-20b-BF16-abliterated
---
## Llamacpp imatrix Quantizations of Huihui-gpt-oss-20b-BF16-abliterated by huihui-ai
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b6115">b6115</a> for quantization.
Original model: https://huggingface.co/huihui-ai/Huihui-gpt-oss-20b-BF16-abliterated
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
Run them directly with [llama.cpp](https://github.com/ggerganov/llama.cpp), or any other llama.cpp based project
## Prompt format
No prompt format found, check original model page
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Huihui-gpt-oss-20b-BF16-abliterated-bf16.gguf](https://huggingface.co/bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF/blob/main/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-bf16.gguf) | bf16 | 41.86GB | false | Full BF16 weights. |
| [Huihui-gpt-oss-20b-BF16-abliterated-Q8_0.gguf](https://huggingface.co/bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF/blob/main/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q8_0.gguf) | Q8_0 | 22.26GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Huihui-gpt-oss-20b-BF16-abliterated-MXFP4_MOE.gguf](https://huggingface.co/bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF/blob/main/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-MXFP4_MOE.gguf) | MXFP4_MOE | 12.11GB | false | Special format for OpenAI's gpt-oss models, see: https://github.com/ggml-org/llama.cpp/pull/15091 |
| [Huihui-gpt-oss-20b-BF16-abliterated-Q6_K_L.gguf](https://huggingface.co/bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF/blob/main/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q6_K_L.gguf) | Q6_K_L | 12.04GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Huihui-gpt-oss-20b-BF16-abliterated-Q6_K.gguf](https://huggingface.co/bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF/blob/main/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q6_K.gguf) | Q6_K | 12.04GB | false | Very high quality, near perfect, *recommended*. |
| [Huihui-gpt-oss-20b-BF16-abliterated-Q5_K_L.gguf](https://huggingface.co/bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF/blob/main/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q5_K_L.gguf) | Q5_K_L | 11.91GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Huihui-gpt-oss-20b-BF16-abliterated-Q4_K_L.gguf](https://huggingface.co/bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF/blob/main/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q4_K_L.gguf) | Q4_K_L | 11.89GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Huihui-gpt-oss-20b-BF16-abliterated-Q2_K_L.gguf](https://huggingface.co/bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF/blob/main/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q2_K_L.gguf) | Q2_K_L | 11.85GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Huihui-gpt-oss-20b-BF16-abliterated-Q3_K_XL.gguf](https://huggingface.co/bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF/blob/main/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q3_K_XL.gguf) | Q3_K_XL | 11.78GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Huihui-gpt-oss-20b-BF16-abliterated-Q5_K_M.gguf](https://huggingface.co/bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF/blob/main/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q5_K_M.gguf) | Q5_K_M | 11.73GB | false | High quality, *recommended*. |
| [Huihui-gpt-oss-20b-BF16-abliterated-Q5_K_S.gguf](https://huggingface.co/bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF/blob/main/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q5_K_S.gguf) | Q5_K_S | 11.72GB | false | High quality, *recommended*. |
| [Huihui-gpt-oss-20b-BF16-abliterated-Q4_K_M.gguf](https://huggingface.co/bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF/blob/main/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q4_K_M.gguf) | Q4_K_M | 11.67GB | false | Good quality, default size for most use cases, *recommended*. |
| [Huihui-gpt-oss-20b-BF16-abliterated-Q4_K_S.gguf](https://huggingface.co/bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF/blob/main/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q4_K_S.gguf) | Q4_K_S | 11.67GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Huihui-gpt-oss-20b-BF16-abliterated-Q4_1.gguf](https://huggingface.co/bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF/blob/main/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q4_1.gguf) | Q4_1 | 11.59GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| [Huihui-gpt-oss-20b-BF16-abliterated-IQ4_NL.gguf](https://huggingface.co/bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF/blob/main/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-IQ4_NL.gguf) | IQ4_NL | 11.56GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [Huihui-gpt-oss-20b-BF16-abliterated-IQ4_XS.gguf](https://huggingface.co/bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF/blob/main/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-IQ4_XS.gguf) | IQ4_XS | 11.56GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Huihui-gpt-oss-20b-BF16-abliterated-Q3_K_M.gguf](https://huggingface.co/bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF/blob/main/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q3_K_M.gguf) | Q3_K_M | 11.56GB | false | Low quality. |
| [Huihui-gpt-oss-20b-BF16-abliterated-IQ3_M.gguf](https://huggingface.co/bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF/blob/main/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-IQ3_M.gguf) | IQ3_M | 11.56GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Huihui-gpt-oss-20b-BF16-abliterated-IQ3_XS.gguf](https://huggingface.co/bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF/blob/main/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-IQ3_XS.gguf) | IQ3_XS | 11.56GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Huihui-gpt-oss-20b-BF16-abliterated-IQ3_XXS.gguf](https://huggingface.co/bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF/blob/main/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-IQ3_XXS.gguf) | IQ3_XXS | 11.56GB | false | Lower quality, new method with decent performance, comparable to Q3 quants. |
| [Huihui-gpt-oss-20b-BF16-abliterated-Q2_K.gguf](https://huggingface.co/bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF/blob/main/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q2_K.gguf) | Q2_K | 11.56GB | false | Very low quality but surprisingly usable. |
| [Huihui-gpt-oss-20b-BF16-abliterated-Q3_K_S.gguf](https://huggingface.co/bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF/blob/main/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q3_K_S.gguf) | Q3_K_S | 11.55GB | false | Low quality, not recommended. |
| [Huihui-gpt-oss-20b-BF16-abliterated-IQ2_M.gguf](https://huggingface.co/bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF/blob/main/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-IQ2_M.gguf) | IQ2_M | 11.55GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
| [Huihui-gpt-oss-20b-BF16-abliterated-IQ2_S.gguf](https://huggingface.co/bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF/blob/main/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-IQ2_S.gguf) | IQ2_S | 11.55GB | false | Low quality, uses SOTA techniques to be usable. |
| [Huihui-gpt-oss-20b-BF16-abliterated-Q4_0.gguf](https://huggingface.co/bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF/blob/main/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q4_0.gguf) | Q4_0 | 11.52GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| [Huihui-gpt-oss-20b-BF16-abliterated-IQ2_XS.gguf](https://huggingface.co/bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF/blob/main/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-IQ2_XS.gguf) | IQ2_XS | 11.51GB | false | Low quality, uses SOTA techniques to be usable. |
| [Huihui-gpt-oss-20b-BF16-abliterated-IQ2_XXS.gguf](https://huggingface.co/bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF/blob/main/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-IQ2_XXS.gguf) | IQ2_XXS | 11.51GB | false | Very low quality, uses SOTA techniques to be usable. |
| [Huihui-gpt-oss-20b-BF16-abliterated-Q3_K_L.gguf](https://huggingface.co/bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF/blob/main/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q3_K_L.gguf) | Q3_K_L | 11.49GB | false | Lower quality but usable, good for low RAM availability. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF --include "huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-GGUF --include "huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (huihui-ai_Huihui-gpt-oss-20b-BF16-abliterated-Q8_0) or download them all in place (./)
</details>
## ARM/AVX information
Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.
Now, however, there is something called "online repacking" for weights. details in [this PR](https://github.com/ggerganov/llama.cpp/pull/9921). If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.
As of llama.cpp build [b4282](https://github.com/ggerganov/llama.cpp/releases/tag/b4282) you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.
Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to [this PR](https://github.com/ggerganov/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information (deprecated</summary>
I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Thank you to LM Studio for sponsoring my work.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6295073d986cb73b72a5132ec55e5bde5465818320701eebdee63ab2d14910b8
size 11545730720

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1fd5c9eb8a0da8d3b861a97dba4ee45765e12765454577c32b7d00afe2206630
size 11545730720

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b97134849de5fae1838da20a2e51c91b3862ffa4ed2569a949876f8df223bfde
size 11510341280

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:76535bfb3b58862b0317993806198c0b372eda6cf80ebc2c42a7980c0487475e
size 11510341280

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6ed7990c1900d82d6dafef252e69319b37db6d28a710f6adfdf72e5ad307aae3
size 11559001760

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:58f2f72a9e89bee6c45e0c722a1d2556e62916bedad1e46625fc759e822209a4
size 11559001760

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:936bf2c953d9182d78c0e1307a37af56d7ce549447980ef0f3bf86daa07ad0af
size 11559001760

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4e9aca89ae5e34ca32fddf4d99bf50e0f23c9906a47cc281477ed71d50e223c5
size 11561213600

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:99d4f575c9346a2a4efcbaec7f2761f87ebcb5fc41d92df35dcd2b6a37d57e15
size 11561213600

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:abca50d1bd95c49d71db36aad0f38090ea5465ce148634c496a48bc87030bdd9
size 12109565600

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c1e1e8708d40a30452f3a9317f1c7a9d763415ae58e6a91fff447f7a752291b9
size 11559001760

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e25d983af6cf3f5b041eeb1fb8aaf1e043797efa5d6334b425c7fd342e421ec0
size 11848568480

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:962536f2d29f1af1581b69c5685d202c8e60bbf0ae1a4d5d9d8be43e9d30ad78
size 11488222880

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2f633b5110a1d0b1484e5b0ebc25e6ad0784e61fa6b21239456e65ad6bfdacdb
size 11559186080

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b1b36ba6c219de6913acc130ae590ffa0e336414d9163a546f2d0d17cbfc6de7
size 11554578080

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:92b70cc35e19f9b039c2d4c26f3973c47e74aa9ae351386f97fe29f59ecac6e5
size 11777789600

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6b1e8408a83d5619082703797282399dc40a1100336f72704dbd0b208742b307
size 11519188640

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:910fdd48be64c11fd04a16a8912e51bb8efc4f0e5e539530a2f369eb86fe1086
size 11592985760

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:201577fe1c7a8bb17c311455cd56215e81fb1cb1ad8c5318de290e6526c1301c
size 11890593440

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e3dfaacc1b7131c9790f5240ecbfca70a449ef8da6f68cf04577de89c55eea7b
size 11673418400

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:550d00994093580749a63a6038f41351e290d40beff0bd0f3fbb541a46f0b111
size 11667151520

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3297956cc17472dfe000c16f7fd9ac1c79ce187ab059d624ce9dac99882bd757
size 11909394080

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b900354893d7ac55c5c889ebd0bc6726be352293b5fb07bb4bc7e2b2d6480819
size 11728414880

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:578829605005ef68c8a423809df9267b437b02ca37551e4bd37c61eca9f785e3
size 11722885280

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:24f679be856db154e4d106de5d25627fe68c96aff2a7072f6bcc9e7b66d0faf8
size 12040998560

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:24f679be856db154e4d106de5d25627fe68c96aff2a7072f6bcc9e7b66d0faf8
size 12040998560

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:341f05face523f6fbc9b5563234d1c57911be49726a46d07111ab9c30d72e6d6
size 22261911200

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1cd29875637894d48006545ef3ca20110633a136f06bd236de0bf3878af74f6f
size 41860886880

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:277129922edfabadefb2428a13829198d867f83578657a5cd45b238d16da75f7
size 28079776