初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Quest-AI_quest-corruption-7b-s375-v3-GRPO-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-07 12:33:13 +08:00
commit 2a8deb95dc
28 changed files with 290 additions and 0 deletions

47
.gitattributes vendored Normal file
View File

@@ -0,0 +1,47 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bin.* filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zstandard filter=lfs diff=lfs merge=lfs -text
*.tfevents* filter=lfs diff=lfs merge=lfs -text
*.db* filter=lfs diff=lfs merge=lfs -text
*.ark* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*data* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.meta filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.index filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.gguf* filter=lfs diff=lfs merge=lfs -text
*.ggml filter=lfs diff=lfs merge=lfs -text
*.llamafile* filter=lfs diff=lfs merge=lfs -text
*.pt2 filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:78d515c1c6cc36534c0c1082f4275155ab21f157b94e1ffcf2a9692b69b4b81c
size 2780340704

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7ad3ee1a3e28d8058e23259511ff0c8732c9330a45a5fc516f16261eee177dad
size 3574010336

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:357721c8985f2e90b8d13062732545ca55332237e6b13f207e578e2300566e5c
size 3346254304

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4ce64754c5f63453f418b264876dc39f9eec85b1b2d4fc9c019bed7fc0084cb9
size 3114512864

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1949eb2a566a8fd1533a62472dc738cbb6f18ff0f836001f01ae6bd2c1eb3e6a
size 4437811680

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0b1b69cf17fcc2065f92bbeaf1de6eb4a31710568aaef94cc800379216039d85
size 4218470880

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1e3dfbdf6235e26c11dd7502305184bb30872417912374dd52b4303eab2ae52f
size 3015938528

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:30068c4d62a2987780db8dd96fe99237aa5391c127b4bd2c66d45993504a687b
size 3548162528

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:51ca3f9c16e0286403fbb6c78d77e9332e1680f171534283da1a0f6af935223d
size 4088457696

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c891cd1a8445c6d3ed12870588bba706ea91a024da1a8d70c397fd4b1a200839
size 3808389600

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f26ab3bad23aaee4b40b4b7f6cbd52421243fc400dee217acea20aeff163d92a
size 3492366816

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0c7c80e76b1e1e2e5db5898c0178aa73ffa94c4aa4495fa7be88af15e74b66ed
size 4565330400

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:583dd0f68f4d2f045a59db61e0f74694c8c7dc6a12889a3fe228e0a2b7d779fd
size 4444119520

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b228bea62ec960bca14809ccc07996997fa2b18976a9d7a5ca62d6b9fcf42756
size 4873282016

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:18bb7c4bb42225ffe56015447d25825b3acaf250196a15b9eebffbfdf7d4ed99
size 5087562208

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:998ec40b72260cb842b88900175708a768712a1d377928f3729bbfbb07fa40b0
size 4683071968

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:18a15fd0a48c6019f94698779063aff006af7e2688e9f4540d0e9be41125e183
size 4457767392

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:621a09ac6d63eac246d33bfde8ed743effcd9abd29f429d406e182d03c84a4e1
size 5781195232

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:dd1c5511a50786d093d51911ce76a06a7b161441c3c4f391dcf62a4305db509d
size 5444829664

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f69eb25cb26596cbeabd51c9622e8b9af5e58928b5c05b24321068380ea344eb
size 5315174880

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b346da908a32316ac8411c004bc70455aa38ea0cb042a4cb1236ab53704a6cbf
size 6254197216

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5b6d0beb80d6a3c49b8f76ea910d4e54963e0db890cd7fe8deebdf57cc166ede
size 6518180320

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2b17c04fc295aa48a9ed846ccd6f3060a56999a3d13074412e4ffa60b288be8d
size 8098523616

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ac9dc152354eca971820d17d5d5f952bc5547c730e4078d5d54cfc2b73aa00ca
size 15237851616

Binary file not shown.

170
README.md Normal file
View File

@@ -0,0 +1,170 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
base_model: Quest-AI/quest-corruption-7b-s375-v3-GRPO
license: apache-2.0
datasets:
- Quest-AI/quest-corruption-truncated4grpo-6k-dataset-v1
---
## Llamacpp imatrix Quantizations of quest-corruption-7b-s375-v3-GRPO by Quest-AI
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4738">b4738</a> for quantization.
Original model: https://huggingface.co/Quest-AI/quest-corruption-7b-s375-v3-GRPO
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
Run them directly with [llama.cpp](https://github.com/ggerganov/llama.cpp), or any other llama.cpp based project
## Prompt format
```
{system_prompt}
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [quest-corruption-7b-s375-v3-GRPO-f16.gguf](https://huggingface.co/bartowski/Quest-AI_quest-corruption-7b-s375-v3-GRPO-GGUF/blob/main/Quest-AI_quest-corruption-7b-s375-v3-GRPO-f16.gguf) | f16 | 15.24GB | false | Full F16 weights. |
| [quest-corruption-7b-s375-v3-GRPO-Q8_0.gguf](https://huggingface.co/bartowski/Quest-AI_quest-corruption-7b-s375-v3-GRPO-GGUF/blob/main/Quest-AI_quest-corruption-7b-s375-v3-GRPO-Q8_0.gguf) | Q8_0 | 8.10GB | false | Extremely high quality, generally unneeded but max available quant. |
| [quest-corruption-7b-s375-v3-GRPO-Q6_K_L.gguf](https://huggingface.co/bartowski/Quest-AI_quest-corruption-7b-s375-v3-GRPO-GGUF/blob/main/Quest-AI_quest-corruption-7b-s375-v3-GRPO-Q6_K_L.gguf) | Q6_K_L | 6.52GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [quest-corruption-7b-s375-v3-GRPO-Q6_K.gguf](https://huggingface.co/bartowski/Quest-AI_quest-corruption-7b-s375-v3-GRPO-GGUF/blob/main/Quest-AI_quest-corruption-7b-s375-v3-GRPO-Q6_K.gguf) | Q6_K | 6.25GB | false | Very high quality, near perfect, *recommended*. |
| [quest-corruption-7b-s375-v3-GRPO-Q5_K_L.gguf](https://huggingface.co/bartowski/Quest-AI_quest-corruption-7b-s375-v3-GRPO-GGUF/blob/main/Quest-AI_quest-corruption-7b-s375-v3-GRPO-Q5_K_L.gguf) | Q5_K_L | 5.78GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [quest-corruption-7b-s375-v3-GRPO-Q5_K_M.gguf](https://huggingface.co/bartowski/Quest-AI_quest-corruption-7b-s375-v3-GRPO-GGUF/blob/main/Quest-AI_quest-corruption-7b-s375-v3-GRPO-Q5_K_M.gguf) | Q5_K_M | 5.44GB | false | High quality, *recommended*. |
| [quest-corruption-7b-s375-v3-GRPO-Q5_K_S.gguf](https://huggingface.co/bartowski/Quest-AI_quest-corruption-7b-s375-v3-GRPO-GGUF/blob/main/Quest-AI_quest-corruption-7b-s375-v3-GRPO-Q5_K_S.gguf) | Q5_K_S | 5.32GB | false | High quality, *recommended*. |
| [quest-corruption-7b-s375-v3-GRPO-Q4_K_L.gguf](https://huggingface.co/bartowski/Quest-AI_quest-corruption-7b-s375-v3-GRPO-GGUF/blob/main/Quest-AI_quest-corruption-7b-s375-v3-GRPO-Q4_K_L.gguf) | Q4_K_L | 5.09GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [quest-corruption-7b-s375-v3-GRPO-Q4_1.gguf](https://huggingface.co/bartowski/Quest-AI_quest-corruption-7b-s375-v3-GRPO-GGUF/blob/main/Quest-AI_quest-corruption-7b-s375-v3-GRPO-Q4_1.gguf) | Q4_1 | 4.87GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| [quest-corruption-7b-s375-v3-GRPO-Q4_K_M.gguf](https://huggingface.co/bartowski/Quest-AI_quest-corruption-7b-s375-v3-GRPO-GGUF/blob/main/Quest-AI_quest-corruption-7b-s375-v3-GRPO-Q4_K_M.gguf) | Q4_K_M | 4.68GB | false | Good quality, default size for most use cases, *recommended*. |
| [quest-corruption-7b-s375-v3-GRPO-Q3_K_XL.gguf](https://huggingface.co/bartowski/Quest-AI_quest-corruption-7b-s375-v3-GRPO-GGUF/blob/main/Quest-AI_quest-corruption-7b-s375-v3-GRPO-Q3_K_XL.gguf) | Q3_K_XL | 4.57GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [quest-corruption-7b-s375-v3-GRPO-Q4_K_S.gguf](https://huggingface.co/bartowski/Quest-AI_quest-corruption-7b-s375-v3-GRPO-GGUF/blob/main/Quest-AI_quest-corruption-7b-s375-v3-GRPO-Q4_K_S.gguf) | Q4_K_S | 4.46GB | false | Slightly lower quality with more space savings, *recommended*. |
| [quest-corruption-7b-s375-v3-GRPO-Q4_0.gguf](https://huggingface.co/bartowski/Quest-AI_quest-corruption-7b-s375-v3-GRPO-GGUF/blob/main/Quest-AI_quest-corruption-7b-s375-v3-GRPO-Q4_0.gguf) | Q4_0 | 4.44GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| [quest-corruption-7b-s375-v3-GRPO-IQ4_NL.gguf](https://huggingface.co/bartowski/Quest-AI_quest-corruption-7b-s375-v3-GRPO-GGUF/blob/main/Quest-AI_quest-corruption-7b-s375-v3-GRPO-IQ4_NL.gguf) | IQ4_NL | 4.44GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [quest-corruption-7b-s375-v3-GRPO-IQ4_XS.gguf](https://huggingface.co/bartowski/Quest-AI_quest-corruption-7b-s375-v3-GRPO-GGUF/blob/main/Quest-AI_quest-corruption-7b-s375-v3-GRPO-IQ4_XS.gguf) | IQ4_XS | 4.22GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [quest-corruption-7b-s375-v3-GRPO-Q3_K_L.gguf](https://huggingface.co/bartowski/Quest-AI_quest-corruption-7b-s375-v3-GRPO-GGUF/blob/main/Quest-AI_quest-corruption-7b-s375-v3-GRPO-Q3_K_L.gguf) | Q3_K_L | 4.09GB | false | Lower quality but usable, good for low RAM availability. |
| [quest-corruption-7b-s375-v3-GRPO-Q3_K_M.gguf](https://huggingface.co/bartowski/Quest-AI_quest-corruption-7b-s375-v3-GRPO-GGUF/blob/main/Quest-AI_quest-corruption-7b-s375-v3-GRPO-Q3_K_M.gguf) | Q3_K_M | 3.81GB | false | Low quality. |
| [quest-corruption-7b-s375-v3-GRPO-IQ3_M.gguf](https://huggingface.co/bartowski/Quest-AI_quest-corruption-7b-s375-v3-GRPO-GGUF/blob/main/Quest-AI_quest-corruption-7b-s375-v3-GRPO-IQ3_M.gguf) | IQ3_M | 3.57GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [quest-corruption-7b-s375-v3-GRPO-Q2_K_L.gguf](https://huggingface.co/bartowski/Quest-AI_quest-corruption-7b-s375-v3-GRPO-GGUF/blob/main/Quest-AI_quest-corruption-7b-s375-v3-GRPO-Q2_K_L.gguf) | Q2_K_L | 3.55GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [quest-corruption-7b-s375-v3-GRPO-Q3_K_S.gguf](https://huggingface.co/bartowski/Quest-AI_quest-corruption-7b-s375-v3-GRPO-GGUF/blob/main/Quest-AI_quest-corruption-7b-s375-v3-GRPO-Q3_K_S.gguf) | Q3_K_S | 3.49GB | false | Low quality, not recommended. |
| [quest-corruption-7b-s375-v3-GRPO-IQ3_XS.gguf](https://huggingface.co/bartowski/Quest-AI_quest-corruption-7b-s375-v3-GRPO-GGUF/blob/main/Quest-AI_quest-corruption-7b-s375-v3-GRPO-IQ3_XS.gguf) | IQ3_XS | 3.35GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [quest-corruption-7b-s375-v3-GRPO-IQ3_XXS.gguf](https://huggingface.co/bartowski/Quest-AI_quest-corruption-7b-s375-v3-GRPO-GGUF/blob/main/Quest-AI_quest-corruption-7b-s375-v3-GRPO-IQ3_XXS.gguf) | IQ3_XXS | 3.11GB | false | Lower quality, new method with decent performance, comparable to Q3 quants. |
| [quest-corruption-7b-s375-v3-GRPO-Q2_K.gguf](https://huggingface.co/bartowski/Quest-AI_quest-corruption-7b-s375-v3-GRPO-GGUF/blob/main/Quest-AI_quest-corruption-7b-s375-v3-GRPO-Q2_K.gguf) | Q2_K | 3.02GB | false | Very low quality but surprisingly usable. |
| [quest-corruption-7b-s375-v3-GRPO-IQ2_M.gguf](https://huggingface.co/bartowski/Quest-AI_quest-corruption-7b-s375-v3-GRPO-GGUF/blob/main/Quest-AI_quest-corruption-7b-s375-v3-GRPO-IQ2_M.gguf) | IQ2_M | 2.78GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Quest-AI_quest-corruption-7b-s375-v3-GRPO-GGUF --include "Quest-AI_quest-corruption-7b-s375-v3-GRPO-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Quest-AI_quest-corruption-7b-s375-v3-GRPO-GGUF --include "Quest-AI_quest-corruption-7b-s375-v3-GRPO-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Quest-AI_quest-corruption-7b-s375-v3-GRPO-Q8_0) or download them all in place (./)
</details>
## ARM/AVX information
Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.
Now, however, there is something called "online repacking" for weights. details in [this PR](https://github.com/ggerganov/llama.cpp/pull/9921). If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.
As of llama.cpp build [b4282](https://github.com/ggerganov/llama.cpp/releases/tag/b4282) you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.
Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to [this PR](https://github.com/ggerganov/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information (deprecated</summary>
I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Thank you to LM Studio for sponsoring my work.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}