初始化项目,由ModelHub XC社区提供模型

Model: bartowski/HelpingAI_Dhanishtha-2.0-preview-0825-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-09-14 08:24:12 +08:00
commit de5bea47c3
29 changed files with 297 additions and 0 deletions

47
.gitattributes vendored Normal file
View File

@@ -0,0 +1,47 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bin.* filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zstandard filter=lfs diff=lfs merge=lfs -text
*.tfevents* filter=lfs diff=lfs merge=lfs -text
*.db* filter=lfs diff=lfs merge=lfs -text
*.ark* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*data* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.meta filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.index filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.gguf* filter=lfs diff=lfs merge=lfs -text
*.ggml filter=lfs diff=lfs merge=lfs -text
*.llamafile* filter=lfs diff=lfs merge=lfs -text
*.pt2 filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:52ea63b523acd5f85754eebdf64bb4b9b31aecc0feccbb6959cc2bfc38dd08e0
size 5322941600

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:014aaea69afd412f3503e07bc9d291ba2a2604c11f0d750f7f65785301b337ae
size 4963312800

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8560248c06c2e6a680ec5b8ee72316763b52f2e73e6f561f0ba52f73bb7ecc88
size 6883410080

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:63ee96bf96e04ddd2356ce103e200726e6eee313bfa21e8f2ba8d24ee55ed611
size 6375301280

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9498fc3fa2bb0d477bb296381b1a753744cea714b1f66f554e271c5da86198c0
size 5942666400

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:569028b4f410de19408a0fbaffcfb3ca610d7a0cf86f06cb84ad6b526637f283
size 8541363360

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6cb2c3c6dfdf28b5277e45ce1678d316c29645dcd63893cf4707fca75c815e53
size 8110730400

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:19660e02d0f5e714f3f311f4df14bbd70eb63e1a7b95ff63fe22144b5d4d4683
size 5753984160

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:466ce2531a06786aa48e4f36c6b7575b86196d1c79685c8d3075465b9bcfee65
size 6513664160

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cca785f04f3565a7e7a0f168c50ec9b60aade1d1fec2f3ab1fea777bfc9061ce
size 7900651680

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8e0f42d20809ef2096cb79d53c126661d31ea1300a7f3cc2c6a65db8f31cbab7
size 7321313440

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0bd611332f4af6d46c290b104a1373535821db79071f171e20ca0389b8fe56bb
size 6657106080

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:26c28474e84009f6db6ea34132fc455afa8f02e18a0cfddd2fc569414ead2574
size 8581324960

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:37605d39dc55d2b50d599fe6953b7b5816903c4d1983d225f27d99796b6be05f
size 8543001760

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cc0b40920c028fd0c5c09604d394fd07389ec0e321cc49ad6031ffb50ce13bf8
size 9389522080

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:69cfd77c8d8e442a44e1a1c2eca8ab3b62181c6f1e811ea225e478bc7522c42d
size 9579110560

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5b33d0a60847b03e75486c7954ce952ac2a6e677666e93e07d68a8d3e3024d50
size 9001753760

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0ab799931f7c792c9c87b93bc1c9f7a9698985390c91d7d33d991b9f362bf6b3
size 8573476000

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:71663cdc849b06e096992cda95f44f2070affb6d9d051c0ec6160b5c2c4526f3
size 10994688160

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fa6901003b815a13c3892a767ea0cd545e89ab0b3974869b0e7871c4b4b4f68f
size 10514570400

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1da623ad228d9adc9f4ad2c391e3bc41c3e285c51f288af1b075588a4ea5843b
size 10263895200

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a71b61bc0ba5e03aba307de77eb90e555b04da7dc38134cc3929722c1acfc5f0
size 12121938080

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1fe60cc0caa2e3b956971736a97d7bc890f73e87c83ba4b2df81efb6f6357b24
size 12498739360

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cec20d3d9823066f30211aab004152f1c9e357cec90c281240db56b5e70f266d
size 15698534560

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b466ced382c537b0c1ffa594115bbc62bb284250da7f974d948e42325cf78e46
size 29543423840

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:af11c3e7e3d491e8c4e91451f7faf4ea54513bc3d20d64691a48884573082845
size 7743584

171
README.md Normal file
View File

@@ -0,0 +1,171 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
base_model_relation: quantized
base_model: HelpingAI/Dhanishtha-2.0-preview-0825
---
## Llamacpp imatrix Quantizations of Dhanishtha-2.0-preview-0825 by HelpingAI
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b6014">b6014</a> for quantization.
Original model: https://huggingface.co/HelpingAI/Dhanishtha-2.0-preview-0825
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
Run them directly with [llama.cpp](https://github.com/ggerganov/llama.cpp), or any other llama.cpp based project
## Prompt format
```
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Dhanishtha-2.0-preview-0825-bf16.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-0825-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-0825-bf16.gguf) | bf16 | 29.54GB | false | Full BF16 weights. |
| [Dhanishtha-2.0-preview-0825-Q8_0.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-0825-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-0825-Q8_0.gguf) | Q8_0 | 15.70GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Dhanishtha-2.0-preview-0825-Q6_K_L.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-0825-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-0825-Q6_K_L.gguf) | Q6_K_L | 12.50GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Dhanishtha-2.0-preview-0825-Q6_K.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-0825-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-0825-Q6_K.gguf) | Q6_K | 12.12GB | false | Very high quality, near perfect, *recommended*. |
| [Dhanishtha-2.0-preview-0825-Q5_K_L.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-0825-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-0825-Q5_K_L.gguf) | Q5_K_L | 10.99GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Dhanishtha-2.0-preview-0825-Q5_K_M.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-0825-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-0825-Q5_K_M.gguf) | Q5_K_M | 10.51GB | false | High quality, *recommended*. |
| [Dhanishtha-2.0-preview-0825-Q5_K_S.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-0825-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-0825-Q5_K_S.gguf) | Q5_K_S | 10.26GB | false | High quality, *recommended*. |
| [Dhanishtha-2.0-preview-0825-Q4_K_L.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-0825-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-0825-Q4_K_L.gguf) | Q4_K_L | 9.58GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Dhanishtha-2.0-preview-0825-Q4_1.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-0825-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-0825-Q4_1.gguf) | Q4_1 | 9.39GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| [Dhanishtha-2.0-preview-0825-Q4_K_M.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-0825-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-0825-Q4_K_M.gguf) | Q4_K_M | 9.00GB | false | Good quality, default size for most use cases, *recommended*. |
| [Dhanishtha-2.0-preview-0825-Q3_K_XL.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-0825-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-0825-Q3_K_XL.gguf) | Q3_K_XL | 8.58GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Dhanishtha-2.0-preview-0825-Q4_K_S.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-0825-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-0825-Q4_K_S.gguf) | Q4_K_S | 8.57GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Dhanishtha-2.0-preview-0825-Q4_0.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-0825-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-0825-Q4_0.gguf) | Q4_0 | 8.54GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| [Dhanishtha-2.0-preview-0825-IQ4_NL.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-0825-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-0825-IQ4_NL.gguf) | IQ4_NL | 8.54GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [Dhanishtha-2.0-preview-0825-IQ4_XS.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-0825-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-0825-IQ4_XS.gguf) | IQ4_XS | 8.11GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Dhanishtha-2.0-preview-0825-Q3_K_L.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-0825-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-0825-Q3_K_L.gguf) | Q3_K_L | 7.90GB | false | Lower quality but usable, good for low RAM availability. |
| [Dhanishtha-2.0-preview-0825-Q3_K_M.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-0825-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-0825-Q3_K_M.gguf) | Q3_K_M | 7.32GB | false | Low quality. |
| [Dhanishtha-2.0-preview-0825-IQ3_M.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-0825-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-0825-IQ3_M.gguf) | IQ3_M | 6.88GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Dhanishtha-2.0-preview-0825-Q3_K_S.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-0825-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-0825-Q3_K_S.gguf) | Q3_K_S | 6.66GB | false | Low quality, not recommended. |
| [Dhanishtha-2.0-preview-0825-Q2_K_L.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-0825-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-0825-Q2_K_L.gguf) | Q2_K_L | 6.51GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Dhanishtha-2.0-preview-0825-IQ3_XS.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-0825-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-0825-IQ3_XS.gguf) | IQ3_XS | 6.38GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Dhanishtha-2.0-preview-0825-IQ3_XXS.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-0825-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-0825-IQ3_XXS.gguf) | IQ3_XXS | 5.94GB | false | Lower quality, new method with decent performance, comparable to Q3 quants. |
| [Dhanishtha-2.0-preview-0825-Q2_K.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-0825-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-0825-Q2_K.gguf) | Q2_K | 5.75GB | false | Very low quality but surprisingly usable. |
| [Dhanishtha-2.0-preview-0825-IQ2_M.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-0825-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-0825-IQ2_M.gguf) | IQ2_M | 5.32GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
| [Dhanishtha-2.0-preview-0825-IQ2_S.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-0825-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-0825-IQ2_S.gguf) | IQ2_S | 4.96GB | false | Low quality, uses SOTA techniques to be usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/HelpingAI_Dhanishtha-2.0-preview-0825-GGUF --include "HelpingAI_Dhanishtha-2.0-preview-0825-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/HelpingAI_Dhanishtha-2.0-preview-0825-GGUF --include "HelpingAI_Dhanishtha-2.0-preview-0825-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (HelpingAI_Dhanishtha-2.0-preview-0825-Q8_0) or download them all in place (./)
</details>
## ARM/AVX information
Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.
Now, however, there is something called "online repacking" for weights. details in [this PR](https://github.com/ggerganov/llama.cpp/pull/9921). If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.
As of llama.cpp build [b4282](https://github.com/ggerganov/llama.cpp/releases/tag/b4282) you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.
Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to [this PR](https://github.com/ggerganov/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information (deprecated</summary>
I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Thank you to LM Studio for sponsoring my work.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}