初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Qwen_Qwen3-4B-Thinking-2507-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-18 15:24:07 +08:00
commit 6cdb4a4870
28 changed files with 294 additions and 0 deletions

47
.gitattributes vendored Normal file
View File

@@ -0,0 +1,47 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bin.* filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zstandard filter=lfs diff=lfs merge=lfs -text
*.tfevents* filter=lfs diff=lfs merge=lfs -text
*.db* filter=lfs diff=lfs merge=lfs -text
*.ark* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*data* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.meta filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.index filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.gguf* filter=lfs diff=lfs merge=lfs -text
*.ggml filter=lfs diff=lfs merge=lfs -text
*.llamafile* filter=lfs diff=lfs merge=lfs -text
*.pt2 filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e0ac7e5d9c2afc8cf9536574b0a8099927a7391cbbed26ae91969623d3bb3e1c
size 1512983776

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7fc2876d7bb5781b2959170f7f9f0dc2562ddc20cc83074ae7a790004f495734
size 1962896096

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0a9218f9a7d3c79507b8cf5739c73765651140bff3dc96ccbc6314044199c2a9
size 1814375136

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1fd1d19e024ad730bb07123ee392a0798de0fecae40ea5339d2b1206b325d3ba
size 1670188256

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cd67c01235ccc7229a9598f8886727c0cb5e98156193e0d33ac4fa06a911cc7a
size 2381343456

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0bfef7797a9a8ed02cd62a0953f3151d995c5968fde13a6ee703c874e1af2a10
size 2270751456

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0e1d2abd41836cffab54d7effb12cb6e4db7e38172c07a27595fcee2b25b7e05
size 1669499616

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8471e10484570ac077c502ad8a4f9f2dcd595ce01c2c40da2da560cd982ca6c6
size 1763699936

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:303dc33bd29abd181cba4aaa2237dc7d65241c11e536473abdae0d6959c62060
size 2239785696

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1721b213111c7c12411c32e27fdb7da8692f39dcabaf2263bf80a34d6767995a
size 2075618016

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8278d9f6f04e260d6844cec2fb2c1a9d2a3effaf21ddf7b9ce24d444ff02b9d0
size 1886997216

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3e1ef2e2394b8e1b1c1bf38477884aab59d12d0278a9b7492b09b2825760ee02
size 2333986016

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:82d03153a3a58235f23e9c9e3b53a0b2f4abdebbaf31b1aefd90827cfa6194af
size 2375772896

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8e064f0207d44cf8eac44e8be8fcfcad42879ecabd153fb4b1c5ec7769ea8c18
size 2596629216

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ab08bff07d15fd95b9e718eae14e31c7ca003b5b6301d7cdfaddb544855d4559
size 2591481056

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:64ce5838a76b4d63071289c17ed04525265a47abf51546cd747a30e41b7687fa
size 2497280736

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0856c1f57b21187c647941b6f64838d1858e478ca9281ab2d8237d16d14c9b26
size 2383309536

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f31a185f665b5ded5eb09405b4093f91466d9c486a5508f6eb2dda0fa0d394f7
size 2983714016

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5e2e831affbbae2c383e5a4de90549144980315026def43e4b101ade46e463bc
size 2889513696

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:76c994d07e401260f69918c57f67ba40da4ec18db29556a1ecdbde107d10957a
size 2823711456

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f3f0b80140e7e41d965339fdefd9c98fb1453095cf4077fa587ab9266b627488
size 3306261216

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c124636068658c87864fe5e02eed24f1cd0f2e11e67315865bfb2ac01fc218dd
size 3400461536

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2c08db093bc57c2c77222d27ffe8d41cb0b5648e66ba84e5fb9ceab429f6735c
size 4280405216

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e8cb8e45c766c2f341c1a345767e48c5b83802e35faa9302922378d345a986c8
size 8051284928

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d2b99e1370d392119850d9061c27c659fe4fa4077609de06c98cb7bfa64631e5
size 3872640

171
README.md Normal file
View File

@@ -0,0 +1,171 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
base_model_relation: quantized
base_model: Qwen/Qwen3-4B-Thinking-2507
---
## Llamacpp imatrix Quantizations of Qwen3-4B-Thinking-2507 by Qwen
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b6096">b6096</a> for quantization.
Original model: https://huggingface.co/Qwen/Qwen3-4B-Thinking-2507
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
Run them directly with [llama.cpp](https://github.com/ggerganov/llama.cpp), or any other llama.cpp based project
## Prompt format
```
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
<think>
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Qwen3-4B-Thinking-2507-bf16.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-4B-Thinking-2507-GGUF/blob/main/Qwen_Qwen3-4B-Thinking-2507-bf16.gguf) | bf16 | 8.05GB | false | Full BF16 weights. |
| [Qwen3-4B-Thinking-2507-Q8_0.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-4B-Thinking-2507-GGUF/blob/main/Qwen_Qwen3-4B-Thinking-2507-Q8_0.gguf) | Q8_0 | 4.28GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Qwen3-4B-Thinking-2507-Q6_K_L.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-4B-Thinking-2507-GGUF/blob/main/Qwen_Qwen3-4B-Thinking-2507-Q6_K_L.gguf) | Q6_K_L | 3.40GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Qwen3-4B-Thinking-2507-Q6_K.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-4B-Thinking-2507-GGUF/blob/main/Qwen_Qwen3-4B-Thinking-2507-Q6_K.gguf) | Q6_K | 3.31GB | false | Very high quality, near perfect, *recommended*. |
| [Qwen3-4B-Thinking-2507-Q5_K_L.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-4B-Thinking-2507-GGUF/blob/main/Qwen_Qwen3-4B-Thinking-2507-Q5_K_L.gguf) | Q5_K_L | 2.98GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Qwen3-4B-Thinking-2507-Q5_K_M.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-4B-Thinking-2507-GGUF/blob/main/Qwen_Qwen3-4B-Thinking-2507-Q5_K_M.gguf) | Q5_K_M | 2.89GB | false | High quality, *recommended*. |
| [Qwen3-4B-Thinking-2507-Q5_K_S.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-4B-Thinking-2507-GGUF/blob/main/Qwen_Qwen3-4B-Thinking-2507-Q5_K_S.gguf) | Q5_K_S | 2.82GB | false | High quality, *recommended*. |
| [Qwen3-4B-Thinking-2507-Q4_1.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-4B-Thinking-2507-GGUF/blob/main/Qwen_Qwen3-4B-Thinking-2507-Q4_1.gguf) | Q4_1 | 2.60GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| [Qwen3-4B-Thinking-2507-Q4_K_L.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-4B-Thinking-2507-GGUF/blob/main/Qwen_Qwen3-4B-Thinking-2507-Q4_K_L.gguf) | Q4_K_L | 2.59GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Qwen3-4B-Thinking-2507-Q4_K_M.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-4B-Thinking-2507-GGUF/blob/main/Qwen_Qwen3-4B-Thinking-2507-Q4_K_M.gguf) | Q4_K_M | 2.50GB | false | Good quality, default size for most use cases, *recommended*. |
| [Qwen3-4B-Thinking-2507-Q4_K_S.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-4B-Thinking-2507-GGUF/blob/main/Qwen_Qwen3-4B-Thinking-2507-Q4_K_S.gguf) | Q4_K_S | 2.38GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Qwen3-4B-Thinking-2507-Q4_0.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-4B-Thinking-2507-GGUF/blob/main/Qwen_Qwen3-4B-Thinking-2507-Q4_0.gguf) | Q4_0 | 2.38GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| [Qwen3-4B-Thinking-2507-IQ4_NL.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-4B-Thinking-2507-GGUF/blob/main/Qwen_Qwen3-4B-Thinking-2507-IQ4_NL.gguf) | IQ4_NL | 2.38GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [Qwen3-4B-Thinking-2507-Q3_K_XL.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-4B-Thinking-2507-GGUF/blob/main/Qwen_Qwen3-4B-Thinking-2507-Q3_K_XL.gguf) | Q3_K_XL | 2.33GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Qwen3-4B-Thinking-2507-IQ4_XS.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-4B-Thinking-2507-GGUF/blob/main/Qwen_Qwen3-4B-Thinking-2507-IQ4_XS.gguf) | IQ4_XS | 2.27GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Qwen3-4B-Thinking-2507-Q3_K_L.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-4B-Thinking-2507-GGUF/blob/main/Qwen_Qwen3-4B-Thinking-2507-Q3_K_L.gguf) | Q3_K_L | 2.24GB | false | Lower quality but usable, good for low RAM availability. |
| [Qwen3-4B-Thinking-2507-Q3_K_M.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-4B-Thinking-2507-GGUF/blob/main/Qwen_Qwen3-4B-Thinking-2507-Q3_K_M.gguf) | Q3_K_M | 2.08GB | false | Low quality. |
| [Qwen3-4B-Thinking-2507-IQ3_M.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-4B-Thinking-2507-GGUF/blob/main/Qwen_Qwen3-4B-Thinking-2507-IQ3_M.gguf) | IQ3_M | 1.96GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Qwen3-4B-Thinking-2507-Q3_K_S.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-4B-Thinking-2507-GGUF/blob/main/Qwen_Qwen3-4B-Thinking-2507-Q3_K_S.gguf) | Q3_K_S | 1.89GB | false | Low quality, not recommended. |
| [Qwen3-4B-Thinking-2507-IQ3_XS.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-4B-Thinking-2507-GGUF/blob/main/Qwen_Qwen3-4B-Thinking-2507-IQ3_XS.gguf) | IQ3_XS | 1.81GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Qwen3-4B-Thinking-2507-Q2_K_L.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-4B-Thinking-2507-GGUF/blob/main/Qwen_Qwen3-4B-Thinking-2507-Q2_K_L.gguf) | Q2_K_L | 1.76GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Qwen3-4B-Thinking-2507-IQ3_XXS.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-4B-Thinking-2507-GGUF/blob/main/Qwen_Qwen3-4B-Thinking-2507-IQ3_XXS.gguf) | IQ3_XXS | 1.67GB | false | Lower quality, new method with decent performance, comparable to Q3 quants. |
| [Qwen3-4B-Thinking-2507-Q2_K.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-4B-Thinking-2507-GGUF/blob/main/Qwen_Qwen3-4B-Thinking-2507-Q2_K.gguf) | Q2_K | 1.67GB | false | Very low quality but surprisingly usable. |
| [Qwen3-4B-Thinking-2507-IQ2_M.gguf](https://huggingface.co/bartowski/Qwen_Qwen3-4B-Thinking-2507-GGUF/blob/main/Qwen_Qwen3-4B-Thinking-2507-IQ2_M.gguf) | IQ2_M | 1.51GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Qwen_Qwen3-4B-Thinking-2507-GGUF --include "Qwen_Qwen3-4B-Thinking-2507-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Qwen_Qwen3-4B-Thinking-2507-GGUF --include "Qwen_Qwen3-4B-Thinking-2507-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Qwen_Qwen3-4B-Thinking-2507-Q8_0) or download them all in place (./)
</details>
## ARM/AVX information
Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.
Now, however, there is something called "online repacking" for weights. details in [this PR](https://github.com/ggerganov/llama.cpp/pull/9921). If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.
As of llama.cpp build [b4282](https://github.com/ggerganov/llama.cpp/releases/tag/b4282) you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.
Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to [this PR](https://github.com/ggerganov/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information (deprecated</summary>
I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Thank you to LM Studio for sponsoring my work.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}