初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Akhil-Theerthala_Kuvera-8B-v0.1.0-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-29 03:40:13 +08:00
commit 891abbcd3f
28 changed files with 305 additions and 0 deletions

49
.gitattributes vendored Normal file
View File

@@ -0,0 +1,49 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bin.* filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zstandard filter=lfs diff=lfs merge=lfs -text
*.tfevents* filter=lfs diff=lfs merge=lfs -text
*.db* filter=lfs diff=lfs merge=lfs -text
*.ark* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*data* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.meta filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.index filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.gguf* filter=lfs diff=lfs merge=lfs -text
*.ggml filter=lfs diff=lfs merge=lfs -text
*.llamafile* filter=lfs diff=lfs merge=lfs -text
*.pt2 filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Akhil-Theerthala_Kuvera-8B-v0.1.0.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:93dfa1dd420cb2b27fff57091e2b2d62e9f5eb45f8bce1f253fd04f320947be9
size 3051915136

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6a64c4c261f94e72ba72a68e442a5199839bdce6dec9b5ea82d8d0bf0ebad3e7
size 3896620928

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b93bb091f68879a02cb26510b45f7e0c8fab2121094c9e750fd0119d26905d60
size 3626874752

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b79d79d60cd78e16486a4e7f04c134a0da6cc2e483d84f72e98df6dd219c1c2c
size 3369633664

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c28367cdc10f2aba8db14cfc5d96c7ed94e8027b50034d626bbcca8a16464d8f
size 4793624448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:79b4ea225683347220834fb64b674e88ffd79691f7bf3250cc7dc402860bd32d
size 4561840000

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2b82e9d0bfb98c2507a8498672a28b22b0fdb3e7681343ef6d29cfdc7138683e
size 3281733504

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d25fc8571614892055af052d80d133e3c2f6b5d7020222808ae0cf41b4533713
size 3889477504

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7a634ebca750ef8800c6926747894b25e70e839e834f2970895a3c48478501f4
size 4431394688

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6a4697f0228a42516395c73f218829da09bf7618a2af2e4dd720d5bddd206e01
size 4124161920

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:591cbd6c70524a4ce40a338694d979294bdfa04b1b5d09f754d82f36da7e56a9
size 3769612160

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:18df458767af90223225ff156846303bc1374e21f94269144fd19ba84c9d5924
size 4975933312

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4e962f2966bf9b49ec8abb9e38ee80c9688da2d9b6f3548c213f634608bd78a9
size 4787332992

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1db678b1c0d990bd4c67bc8ff77a0cc40392dff6df1f38984f2bb690d6c17524
size 5247756160

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a5923ccd0283e92d7610964e97e9ad749a72f0983fd2057bcc0abd90ab73ef63
size 5489670016

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a4e5f379ad58b4225620b664f2c67470f40b43d49a6cf05c83d10ab34ddceb85
size 5027784576

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:50ca3c477104f455f18db1f3e6c7f223b0786cb14b883dc851d2eecdd87f30d8
size 4802013056

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c9ade5728aaa4bfdc7541ccbfa799dcf92cf00f7ad0999eb551b097f23c75cd6
size 6235207552

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e36eaf27e8b72f54eccb9c9bf4c1662a37bbc6666e28b4eed32da9a4eee92c67
size 5851113344

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8980422a25b13f6f701cdc6f1abfdc7fa78ca1f73a290b29d3cbf8aef3e0a88e
size 5720762240

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8a8586026ddb8eb03a08343435ab6b9271dc66e01b54b5900699ff801b0a0e1e
size 6725900160

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:490de2d4aaf5fe472b25d9c6f4ca3df6b2d06bde111623db73d7648fe5833945
size 7027341184

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e71a91f52c535855664a5d4d2e69f7b669372605a64797db98359ed5a553c257
size 8709519232

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:063a4a56a1854b24b758404ecc155df6ddfd56a52298c9532f58b2e9f2962dd4
size 16388044384

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4d21595e85007d65cd1bfa028cf180fb521cc44b7d833d5ea91ed1346c2ba682
size 5316782

180
README.md Normal file
View File

@@ -0,0 +1,180 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
base_model: Akhil-Theerthala/Kuvera-8B-v0.1.0
base_model_relation: quantized
tags:
- finance
- personal_finance
- Qwen3
- text-generation-inference
datasets:
- Akhil-Theerthala/Personal-Finance-Queries
license: mit
language:
- en
---
## Llamacpp imatrix Quantizations of Kuvera-8B-v0.1.0 by Akhil-Theerthala
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b5596">b5596</a> for quantization.
Original model: https://huggingface.co/Akhil-Theerthala/Kuvera-8B-v0.1.0
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
Run them directly with [llama.cpp](https://github.com/ggerganov/llama.cpp), or any other llama.cpp based project
## Prompt format
```
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Kuvera-8B-v0.1.0-bf16.gguf](https://huggingface.co/bartowski/Akhil-Theerthala_Kuvera-8B-v0.1.0-GGUF/blob/main/Akhil-Theerthala_Kuvera-8B-v0.1.0-bf16.gguf) | bf16 | 16.39GB | false | Full BF16 weights. |
| [Kuvera-8B-v0.1.0-Q8_0.gguf](https://huggingface.co/bartowski/Akhil-Theerthala_Kuvera-8B-v0.1.0-GGUF/blob/main/Akhil-Theerthala_Kuvera-8B-v0.1.0-Q8_0.gguf) | Q8_0 | 8.71GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Kuvera-8B-v0.1.0-Q6_K_L.gguf](https://huggingface.co/bartowski/Akhil-Theerthala_Kuvera-8B-v0.1.0-GGUF/blob/main/Akhil-Theerthala_Kuvera-8B-v0.1.0-Q6_K_L.gguf) | Q6_K_L | 7.03GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Kuvera-8B-v0.1.0-Q6_K.gguf](https://huggingface.co/bartowski/Akhil-Theerthala_Kuvera-8B-v0.1.0-GGUF/blob/main/Akhil-Theerthala_Kuvera-8B-v0.1.0-Q6_K.gguf) | Q6_K | 6.73GB | false | Very high quality, near perfect, *recommended*. |
| [Kuvera-8B-v0.1.0-Q5_K_L.gguf](https://huggingface.co/bartowski/Akhil-Theerthala_Kuvera-8B-v0.1.0-GGUF/blob/main/Akhil-Theerthala_Kuvera-8B-v0.1.0-Q5_K_L.gguf) | Q5_K_L | 6.24GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Kuvera-8B-v0.1.0-Q5_K_M.gguf](https://huggingface.co/bartowski/Akhil-Theerthala_Kuvera-8B-v0.1.0-GGUF/blob/main/Akhil-Theerthala_Kuvera-8B-v0.1.0-Q5_K_M.gguf) | Q5_K_M | 5.85GB | false | High quality, *recommended*. |
| [Kuvera-8B-v0.1.0-Q5_K_S.gguf](https://huggingface.co/bartowski/Akhil-Theerthala_Kuvera-8B-v0.1.0-GGUF/blob/main/Akhil-Theerthala_Kuvera-8B-v0.1.0-Q5_K_S.gguf) | Q5_K_S | 5.72GB | false | High quality, *recommended*. |
| [Kuvera-8B-v0.1.0-Q4_K_L.gguf](https://huggingface.co/bartowski/Akhil-Theerthala_Kuvera-8B-v0.1.0-GGUF/blob/main/Akhil-Theerthala_Kuvera-8B-v0.1.0-Q4_K_L.gguf) | Q4_K_L | 5.49GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Kuvera-8B-v0.1.0-Q4_1.gguf](https://huggingface.co/bartowski/Akhil-Theerthala_Kuvera-8B-v0.1.0-GGUF/blob/main/Akhil-Theerthala_Kuvera-8B-v0.1.0-Q4_1.gguf) | Q4_1 | 5.25GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| [Kuvera-8B-v0.1.0-Q4_K_M.gguf](https://huggingface.co/bartowski/Akhil-Theerthala_Kuvera-8B-v0.1.0-GGUF/blob/main/Akhil-Theerthala_Kuvera-8B-v0.1.0-Q4_K_M.gguf) | Q4_K_M | 5.03GB | false | Good quality, default size for most use cases, *recommended*. |
| [Kuvera-8B-v0.1.0-Q3_K_XL.gguf](https://huggingface.co/bartowski/Akhil-Theerthala_Kuvera-8B-v0.1.0-GGUF/blob/main/Akhil-Theerthala_Kuvera-8B-v0.1.0-Q3_K_XL.gguf) | Q3_K_XL | 4.98GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Kuvera-8B-v0.1.0-Q4_K_S.gguf](https://huggingface.co/bartowski/Akhil-Theerthala_Kuvera-8B-v0.1.0-GGUF/blob/main/Akhil-Theerthala_Kuvera-8B-v0.1.0-Q4_K_S.gguf) | Q4_K_S | 4.80GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Kuvera-8B-v0.1.0-Q4_0.gguf](https://huggingface.co/bartowski/Akhil-Theerthala_Kuvera-8B-v0.1.0-GGUF/blob/main/Akhil-Theerthala_Kuvera-8B-v0.1.0-Q4_0.gguf) | Q4_0 | 4.79GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| [Kuvera-8B-v0.1.0-IQ4_NL.gguf](https://huggingface.co/bartowski/Akhil-Theerthala_Kuvera-8B-v0.1.0-GGUF/blob/main/Akhil-Theerthala_Kuvera-8B-v0.1.0-IQ4_NL.gguf) | IQ4_NL | 4.79GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [Kuvera-8B-v0.1.0-IQ4_XS.gguf](https://huggingface.co/bartowski/Akhil-Theerthala_Kuvera-8B-v0.1.0-GGUF/blob/main/Akhil-Theerthala_Kuvera-8B-v0.1.0-IQ4_XS.gguf) | IQ4_XS | 4.56GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Kuvera-8B-v0.1.0-Q3_K_L.gguf](https://huggingface.co/bartowski/Akhil-Theerthala_Kuvera-8B-v0.1.0-GGUF/blob/main/Akhil-Theerthala_Kuvera-8B-v0.1.0-Q3_K_L.gguf) | Q3_K_L | 4.43GB | false | Lower quality but usable, good for low RAM availability. |
| [Kuvera-8B-v0.1.0-Q3_K_M.gguf](https://huggingface.co/bartowski/Akhil-Theerthala_Kuvera-8B-v0.1.0-GGUF/blob/main/Akhil-Theerthala_Kuvera-8B-v0.1.0-Q3_K_M.gguf) | Q3_K_M | 4.12GB | false | Low quality. |
| [Kuvera-8B-v0.1.0-IQ3_M.gguf](https://huggingface.co/bartowski/Akhil-Theerthala_Kuvera-8B-v0.1.0-GGUF/blob/main/Akhil-Theerthala_Kuvera-8B-v0.1.0-IQ3_M.gguf) | IQ3_M | 3.90GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Kuvera-8B-v0.1.0-Q2_K_L.gguf](https://huggingface.co/bartowski/Akhil-Theerthala_Kuvera-8B-v0.1.0-GGUF/blob/main/Akhil-Theerthala_Kuvera-8B-v0.1.0-Q2_K_L.gguf) | Q2_K_L | 3.89GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Kuvera-8B-v0.1.0-Q3_K_S.gguf](https://huggingface.co/bartowski/Akhil-Theerthala_Kuvera-8B-v0.1.0-GGUF/blob/main/Akhil-Theerthala_Kuvera-8B-v0.1.0-Q3_K_S.gguf) | Q3_K_S | 3.77GB | false | Low quality, not recommended. |
| [Kuvera-8B-v0.1.0-IQ3_XS.gguf](https://huggingface.co/bartowski/Akhil-Theerthala_Kuvera-8B-v0.1.0-GGUF/blob/main/Akhil-Theerthala_Kuvera-8B-v0.1.0-IQ3_XS.gguf) | IQ3_XS | 3.63GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Kuvera-8B-v0.1.0-IQ3_XXS.gguf](https://huggingface.co/bartowski/Akhil-Theerthala_Kuvera-8B-v0.1.0-GGUF/blob/main/Akhil-Theerthala_Kuvera-8B-v0.1.0-IQ3_XXS.gguf) | IQ3_XXS | 3.37GB | false | Lower quality, new method with decent performance, comparable to Q3 quants. |
| [Kuvera-8B-v0.1.0-Q2_K.gguf](https://huggingface.co/bartowski/Akhil-Theerthala_Kuvera-8B-v0.1.0-GGUF/blob/main/Akhil-Theerthala_Kuvera-8B-v0.1.0-Q2_K.gguf) | Q2_K | 3.28GB | false | Very low quality but surprisingly usable. |
| [Kuvera-8B-v0.1.0-IQ2_M.gguf](https://huggingface.co/bartowski/Akhil-Theerthala_Kuvera-8B-v0.1.0-GGUF/blob/main/Akhil-Theerthala_Kuvera-8B-v0.1.0-IQ2_M.gguf) | IQ2_M | 3.05GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Akhil-Theerthala_Kuvera-8B-v0.1.0-GGUF --include "Akhil-Theerthala_Kuvera-8B-v0.1.0-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Akhil-Theerthala_Kuvera-8B-v0.1.0-GGUF --include "Akhil-Theerthala_Kuvera-8B-v0.1.0-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Akhil-Theerthala_Kuvera-8B-v0.1.0-Q8_0) or download them all in place (./)
</details>
## ARM/AVX information
Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.
Now, however, there is something called "online repacking" for weights. details in [this PR](https://github.com/ggerganov/llama.cpp/pull/9921). If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.
As of llama.cpp build [b4282](https://github.com/ggerganov/llama.cpp/releases/tag/b4282) you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.
Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to [this PR](https://github.com/ggerganov/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information (deprecated</summary>
I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Thank you to LM Studio for sponsoring my work.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}