初始化项目,由ModelHub XC社区提供模型

Model: bartowski/HelpingAI_Dhanishtha-2.0-preview-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-30 14:59:14 +08:00
commit 2ab05f71ce
29 changed files with 364 additions and 0 deletions

49
.gitattributes vendored Normal file
View File

@@ -0,0 +1,49 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bin.* filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zstandard filter=lfs diff=lfs merge=lfs -text
*.tfevents* filter=lfs diff=lfs merge=lfs -text
*.db* filter=lfs diff=lfs merge=lfs -text
*.ark* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*data* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.meta filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.index filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.gguf* filter=lfs diff=lfs merge=lfs -text
*.ggml filter=lfs diff=lfs merge=lfs -text
*.llamafile* filter=lfs diff=lfs merge=lfs -text
*.pt2 filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
HelpingAI_Dhanishtha-2.0-preview.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7065c6960737bb0d92a7c12af850f1a099e5654c36ce3ee4aa228e5a676e04e3
size 5322941984

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:24bf58da1cfa3b6e2663f4d8f81d5e0705695df895c553bbf3787a7852d0fb3d
size 4963313184

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8ef6d1e40fbe4d41e58950b654a6e1fe53b8eadd10b66e97846589d99caae38c
size 6883410464

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f9449cfa57b681ad445b0065fec8e9b2db7d54dd68ea4f16404606e88ed33ace
size 6375301664

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:734be93538aeb1f8e88f2c990b477facaa9362f1188f521abed77e0bf70a3a2e
size 5942666784

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a138be1e5e10712bd2fcaa4b1378f6f715c4a1116496e4d2bd646cc9fc836955
size 8541363744

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1d5139e01b0810277c6c9c8d583135d154986268faaa2c985320690a7f92da1f
size 8110730784

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:218eabeada922fdaeedb3fe4c473092893cdabd745f51ef8d6e7ab6dd2322707
size 5753984544

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2928cd1d64e622bd909104780ed7861dc69a5d7ca2dd2a53f1ec5c8dddd8e6ac
size 6513664544

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:577e1f84515879deed32a70649c3f01b0fd64b38299319b0218a8cf19422edb1
size 7900652064

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8626ab1d1ca9caf3340312eb1cc7f3d671f32511e20109b527a28bd443e43904
size 7321313824

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6e25b26838041e258135d0f7ae4d595f8cf174f2e41fda064fd81b4d237d54a0
size 6657106464

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6abea2f40361980fe6062acbb4c39fe5187278cbcb46c665acb1760a4288bd78
size 8581325344

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c6758df6cc4fab9f0328c8ce528d99da08aa7a0c92ccacf0181d4badc64eb1e1
size 8543002144

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5b2bbae462f234c133a7f8293bdfdd3071e7dd26f0920b54f3a9280b2923db79
size 9389522464

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4dd5cd3d0e7dc7eaf4cc4b50be11630f5a71771e5b6e7cf80ee49b3fc4860e58
size 9579110944

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:026a1f80187c9ecdd0227816a35661f3b6b7abe85971121b4c1c25b6cdd7ab86
size 9001754144

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6696a9b88cb0ed8bfcee91469fd8ab7875dda922d1eafa7931b6c13987916a33
size 8573476384

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a64dd24d209e57a95607da50077d8f8256e9269ee60814955c4d1ed52916a455
size 10994688544

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:479f4417ff8dbe74b0a7e8e7c0578fa7627dbd9d80cf0b14678647122156b68c
size 10514570784

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4e3d6d6ff4eb353f44eedb9ccf16b5e6af718ed42211e612c9a87c95d76db7c1
size 10263895584

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e27476fc6466cd77c03ddd03379607bd719b840a5743250ea3928f6ed95f767c
size 12121938464

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a33ad1d0c80fb6b3869d089b9a43e788cbaf94e96297971207fa0899b9bcf67f
size 12498739744

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:28ebd59368525a8edd04a0101c8c485f9f6b0910e3b70e34fdaec3ff83b7770f
size 15698534944

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d4e40663a0f966ce8e16655818e0716ed1eb56f7f68555bda5d91194c249e4f5
size 29543424256

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:41da09a561ae3d3a68132803241d54a5f85af90e6f5c93b7a3101eb94eb15e87
size 7709778

236
README.md Normal file
View File

@@ -0,0 +1,236 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
base_model: HelpingAI/Dhanishtha-2.0-preview
base_model_relation: quantized
license: apache-2.0
language:
- en
- hi
- zh
- es
- fr
- de
- ja
- ko
- ar
- pt
- ru
- it
- nl
- tr
- pl
- sv
- da
- 'no'
- fi
- he
- th
- vi
- id
- ms
- tl
- sw
- yo
- zu
- am
- bn
- gu
- kn
- ml
- mr
- ne
- or
- pa
- ta
- te
- ur
- multilingual
tags:
- reasoning
- intermediate-thinking
- transformers
- conversational
- bilingual
widget:
- text: 'Solve this riddle step by step: I am taken from a mine, and shut up in a
wooden case, from which I am never released, and yet I am used by almost everybody.
What am I?'
example_title: Complex Riddle Solving
- text: Explain the philosophical implications of artificial consciousness and think
through different perspectives.
example_title: Philosophical Reasoning
- text: Help me understand quantum mechanics, but take your time to think through
the explanation.
example_title: Educational Explanation
datasets:
- Abhaykoul/Dhanishtha-R1
- open-thoughts/OpenThoughts-114k
- Abhaykoul/Dhanishtha-2.0-SUPERTHINKER
- Abhaykoul/Dhanishtha-2.0
---
## Llamacpp imatrix Quantizations of Dhanishtha-2.0-preview by HelpingAI
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b5760">b5760</a> for quantization.
Original model: https://huggingface.co/HelpingAI/Dhanishtha-2.0-preview
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
Run them directly with [llama.cpp](https://github.com/ggerganov/llama.cpp), or any other llama.cpp based project
## Prompt format
```
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Dhanishtha-2.0-preview-bf16.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-bf16.gguf) | bf16 | 29.54GB | false | Full BF16 weights. |
| [Dhanishtha-2.0-preview-Q8_0.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-Q8_0.gguf) | Q8_0 | 15.70GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Dhanishtha-2.0-preview-Q6_K_L.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-Q6_K_L.gguf) | Q6_K_L | 12.50GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Dhanishtha-2.0-preview-Q6_K.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-Q6_K.gguf) | Q6_K | 12.12GB | false | Very high quality, near perfect, *recommended*. |
| [Dhanishtha-2.0-preview-Q5_K_L.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-Q5_K_L.gguf) | Q5_K_L | 10.99GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Dhanishtha-2.0-preview-Q5_K_M.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-Q5_K_M.gguf) | Q5_K_M | 10.51GB | false | High quality, *recommended*. |
| [Dhanishtha-2.0-preview-Q5_K_S.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-Q5_K_S.gguf) | Q5_K_S | 10.26GB | false | High quality, *recommended*. |
| [Dhanishtha-2.0-preview-Q4_K_L.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-Q4_K_L.gguf) | Q4_K_L | 9.58GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Dhanishtha-2.0-preview-Q4_1.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-Q4_1.gguf) | Q4_1 | 9.39GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| [Dhanishtha-2.0-preview-Q4_K_M.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-Q4_K_M.gguf) | Q4_K_M | 9.00GB | false | Good quality, default size for most use cases, *recommended*. |
| [Dhanishtha-2.0-preview-Q3_K_XL.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-Q3_K_XL.gguf) | Q3_K_XL | 8.58GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Dhanishtha-2.0-preview-Q4_K_S.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-Q4_K_S.gguf) | Q4_K_S | 8.57GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Dhanishtha-2.0-preview-Q4_0.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-Q4_0.gguf) | Q4_0 | 8.54GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| [Dhanishtha-2.0-preview-IQ4_NL.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-IQ4_NL.gguf) | IQ4_NL | 8.54GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [Dhanishtha-2.0-preview-IQ4_XS.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-IQ4_XS.gguf) | IQ4_XS | 8.11GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Dhanishtha-2.0-preview-Q3_K_L.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-Q3_K_L.gguf) | Q3_K_L | 7.90GB | false | Lower quality but usable, good for low RAM availability. |
| [Dhanishtha-2.0-preview-Q3_K_M.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-Q3_K_M.gguf) | Q3_K_M | 7.32GB | false | Low quality. |
| [Dhanishtha-2.0-preview-IQ3_M.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-IQ3_M.gguf) | IQ3_M | 6.88GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Dhanishtha-2.0-preview-Q3_K_S.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-Q3_K_S.gguf) | Q3_K_S | 6.66GB | false | Low quality, not recommended. |
| [Dhanishtha-2.0-preview-Q2_K_L.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-Q2_K_L.gguf) | Q2_K_L | 6.51GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Dhanishtha-2.0-preview-IQ3_XS.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-IQ3_XS.gguf) | IQ3_XS | 6.38GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Dhanishtha-2.0-preview-IQ3_XXS.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-IQ3_XXS.gguf) | IQ3_XXS | 5.94GB | false | Lower quality, new method with decent performance, comparable to Q3 quants. |
| [Dhanishtha-2.0-preview-Q2_K.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-Q2_K.gguf) | Q2_K | 5.75GB | false | Very low quality but surprisingly usable. |
| [Dhanishtha-2.0-preview-IQ2_M.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-IQ2_M.gguf) | IQ2_M | 5.32GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
| [Dhanishtha-2.0-preview-IQ2_S.gguf](https://huggingface.co/bartowski/HelpingAI_Dhanishtha-2.0-preview-GGUF/blob/main/HelpingAI_Dhanishtha-2.0-preview-IQ2_S.gguf) | IQ2_S | 4.96GB | false | Low quality, uses SOTA techniques to be usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/HelpingAI_Dhanishtha-2.0-preview-GGUF --include "HelpingAI_Dhanishtha-2.0-preview-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/HelpingAI_Dhanishtha-2.0-preview-GGUF --include "HelpingAI_Dhanishtha-2.0-preview-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (HelpingAI_Dhanishtha-2.0-preview-Q8_0) or download them all in place (./)
</details>
## ARM/AVX information
Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.
Now, however, there is something called "online repacking" for weights. details in [this PR](https://github.com/ggerganov/llama.cpp/pull/9921). If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.
As of llama.cpp build [b4282](https://github.com/ggerganov/llama.cpp/releases/tag/b4282) you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.
Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to [this PR](https://github.com/ggerganov/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information (deprecated</summary>
I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Thank you to LM Studio for sponsoring my work.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}