初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Dolphin3.0-Llama3.2-3B-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-11 18:42:08 +08:00
commit 24ba3ccbaf
28 changed files with 320 additions and 0 deletions

60
.gitattributes vendored Normal file
View File

@@ -0,0 +1,60 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-3B-f16.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-3B-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-3B-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-3B-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-3B-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-3B-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-3B-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-3B-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-3B-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-3B-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-3B-Q4_1.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-3B-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-3B-IQ4_NL.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-3B-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-3B-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-3B-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-3B-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-3B-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-3B-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-3B-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-3B-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-3B-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-3B-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-3B-f32.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-3B.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:dc07e2f0dfd4a0882aa5b3f1e6fba1fde504e14209623e3b38d296384bcd29e5
size 1229035840

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2e8a89732af9cfcb2573f7ff26e83b97f6ee2bec8e86c649ca7aeed96c688cc1
size 1599673472

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1964c35f7df1db779484cb7e71477da2b9ce8bfcb0ca822d636bbcec44776f47
size 1476793472

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a3152aefb150122405d3c41e86d51038b4d970e9f78654dec88e38990f1e6f6f
size 1917195392

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:12b6c0d4f2279b136a4593ada5bf3e6a613241120dfa00aa20b35bebd510ec53
size 1829115008

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c43919526d31d95a9832e2df327e0ebfc2d016413b5d231eda709ac833974860
size 1363940480

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9f5011bc0abaf974c282f7be8d29b69a9317b23ab54c524552f04cd681917016
size 1459364416

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3daf22106d38a44ccb60a6effe5fc39a201d7829b50a02fcae98df46e0807771
size 1815352448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fd7a13e2eb7ec276c24062880b249ac3990fdc38c61abcae9d15dc9b849b35b3
size 1687164032

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:88346f6c1d9bc57ca3d41b3a6c8ab83c65271e8b8927f0f95957f3ce2961ceca
size 1542853760

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:29e34f4107cf2a67062a3db48768a07419acfca97aac5f687bb0a29a951149fa
size 1910776384

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:551618c7c03866cf7966809e83ef0f5d81c4d2eb726a3b05f8399e849e327f0a
size 1921913984

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3719877091aad85b7f346856b7c4777ffca9f7b768d458d3a4fdb7513dc33e18
size 2093356160

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1c550caa35be15af8a7d0c961221b369092ce0f6333e5b93474913bcd0d65017
size 2114806336

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5d6d02eeefa1ab5dbf23f97afdf5c2c95ad3d946dc3b6e9ab72e6c1637d54177
size 2019382400

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8ba95a98e27b5559d3b7a692248ab4fd5399793d80b7e7b531ca726fce819c0b
size 1928205440

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d91e3637a352051f99aaa9eea6550fce82c2914f35dc98b77827dd7989b9f7d7
size 2417582656

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6d80400d2354e95966e324fe0840ce3d0d4f6f72b16f4bea940baa4aa1be4a86
size 2322158720

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:08ff6ec021c1952ba937a20ef942175ace69559ddabadb1ae036320b72896aa2
size 2269516928

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a3197bbc65a383767daef570060802cb6bbafad4470a09386509dcbcd03fec5b
size 2643858560

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:419ca3d68001820496f155643d79ecbce2474033ca4463917b8afbb6995efe07
size 2739282496

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d7a51f65ebd35e7e7da5632274183c9b54f38330e015a6c73adca0bf3523fdae
size 3421905472

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:bde88f7223ea57676aa3a6c138e2463012bcd5b0d501ad43f3837d4e5c065b00
size 6433700032

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5f39913705a149684e470abe5e129d7910460064568628f39bf42e62412a9411
size 12858861472

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:25e7a0c2f9856951dfb227230c42d6dacb6393308ba633abf50fbcb251fc3bc6
size 2988390

184
README.md Normal file
View File

@@ -0,0 +1,184 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
language:
- en
base_model: cognitivecomputations/Dolphin3.0-Llama3.2-3B
license: llama3.2
datasets:
- OpenCoder-LLM/opc-sft-stage1
- OpenCoder-LLM/opc-sft-stage2
- microsoft/orca-agentinstruct-1M-v1
- microsoft/orca-math-word-problems-200k
- NousResearch/hermes-function-calling-v1
- AI-MO/NuminaMath-CoT
- AI-MO/NuminaMath-TIR
- allenai/tulu-3-sft-mixture
- cognitivecomputations/dolphin-coder
- HuggingFaceTB/smoltalk
- cognitivecomputations/samantha-data
- m-a-p/CodeFeedback-Filtered-Instruction
- m-a-p/Code-Feedback
---
## Llamacpp imatrix Quantizations of Dolphin3.0-Llama3.2-3B
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4418">b4418</a> for quantization.
Original model: https://huggingface.co/cognitivecomputations/Dolphin3.0-Llama3.2-3B
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Dolphin3.0-Llama3.2-3B-f32.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-3B-GGUF/blob/main/Dolphin3.0-Llama3.2-3B-f32.gguf) | f32 | 12.86GB | false | Full F32 weights. |
| [Dolphin3.0-Llama3.2-3B-f16.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-3B-GGUF/blob/main/Dolphin3.0-Llama3.2-3B-f16.gguf) | f16 | 6.43GB | false | Full F16 weights. |
| [Dolphin3.0-Llama3.2-3B-Q8_0.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-3B-GGUF/blob/main/Dolphin3.0-Llama3.2-3B-Q8_0.gguf) | Q8_0 | 3.42GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Dolphin3.0-Llama3.2-3B-Q6_K_L.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-3B-GGUF/blob/main/Dolphin3.0-Llama3.2-3B-Q6_K_L.gguf) | Q6_K_L | 2.74GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Dolphin3.0-Llama3.2-3B-Q6_K.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-3B-GGUF/blob/main/Dolphin3.0-Llama3.2-3B-Q6_K.gguf) | Q6_K | 2.64GB | false | Very high quality, near perfect, *recommended*. |
| [Dolphin3.0-Llama3.2-3B-Q5_K_L.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-3B-GGUF/blob/main/Dolphin3.0-Llama3.2-3B-Q5_K_L.gguf) | Q5_K_L | 2.42GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Dolphin3.0-Llama3.2-3B-Q5_K_M.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-3B-GGUF/blob/main/Dolphin3.0-Llama3.2-3B-Q5_K_M.gguf) | Q5_K_M | 2.32GB | false | High quality, *recommended*. |
| [Dolphin3.0-Llama3.2-3B-Q5_K_S.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-3B-GGUF/blob/main/Dolphin3.0-Llama3.2-3B-Q5_K_S.gguf) | Q5_K_S | 2.27GB | false | High quality, *recommended*. |
| [Dolphin3.0-Llama3.2-3B-Q4_K_L.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-3B-GGUF/blob/main/Dolphin3.0-Llama3.2-3B-Q4_K_L.gguf) | Q4_K_L | 2.11GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Dolphin3.0-Llama3.2-3B-Q4_1.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-3B-GGUF/blob/main/Dolphin3.0-Llama3.2-3B-Q4_1.gguf) | Q4_1 | 2.09GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| [Dolphin3.0-Llama3.2-3B-Q4_K_M.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-3B-GGUF/blob/main/Dolphin3.0-Llama3.2-3B-Q4_K_M.gguf) | Q4_K_M | 2.02GB | false | Good quality, default size for most use cases, *recommended*. |
| [Dolphin3.0-Llama3.2-3B-Q4_K_S.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-3B-GGUF/blob/main/Dolphin3.0-Llama3.2-3B-Q4_K_S.gguf) | Q4_K_S | 1.93GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Dolphin3.0-Llama3.2-3B-Q4_0.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-3B-GGUF/blob/main/Dolphin3.0-Llama3.2-3B-Q4_0.gguf) | Q4_0 | 1.92GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| [Dolphin3.0-Llama3.2-3B-IQ4_NL.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-3B-GGUF/blob/main/Dolphin3.0-Llama3.2-3B-IQ4_NL.gguf) | IQ4_NL | 1.92GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [Dolphin3.0-Llama3.2-3B-Q3_K_XL.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-3B-GGUF/blob/main/Dolphin3.0-Llama3.2-3B-Q3_K_XL.gguf) | Q3_K_XL | 1.91GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Dolphin3.0-Llama3.2-3B-IQ4_XS.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-3B-GGUF/blob/main/Dolphin3.0-Llama3.2-3B-IQ4_XS.gguf) | IQ4_XS | 1.83GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Dolphin3.0-Llama3.2-3B-Q3_K_L.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-3B-GGUF/blob/main/Dolphin3.0-Llama3.2-3B-Q3_K_L.gguf) | Q3_K_L | 1.82GB | false | Lower quality but usable, good for low RAM availability. |
| [Dolphin3.0-Llama3.2-3B-Q3_K_M.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-3B-GGUF/blob/main/Dolphin3.0-Llama3.2-3B-Q3_K_M.gguf) | Q3_K_M | 1.69GB | false | Low quality. |
| [Dolphin3.0-Llama3.2-3B-IQ3_M.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-3B-GGUF/blob/main/Dolphin3.0-Llama3.2-3B-IQ3_M.gguf) | IQ3_M | 1.60GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Dolphin3.0-Llama3.2-3B-Q3_K_S.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-3B-GGUF/blob/main/Dolphin3.0-Llama3.2-3B-Q3_K_S.gguf) | Q3_K_S | 1.54GB | false | Low quality, not recommended. |
| [Dolphin3.0-Llama3.2-3B-IQ3_XS.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-3B-GGUF/blob/main/Dolphin3.0-Llama3.2-3B-IQ3_XS.gguf) | IQ3_XS | 1.48GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Dolphin3.0-Llama3.2-3B-Q2_K_L.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-3B-GGUF/blob/main/Dolphin3.0-Llama3.2-3B-Q2_K_L.gguf) | Q2_K_L | 1.46GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Dolphin3.0-Llama3.2-3B-Q2_K.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-3B-GGUF/blob/main/Dolphin3.0-Llama3.2-3B-Q2_K.gguf) | Q2_K | 1.36GB | false | Very low quality but surprisingly usable. |
| [Dolphin3.0-Llama3.2-3B-IQ2_M.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-3B-GGUF/blob/main/Dolphin3.0-Llama3.2-3B-IQ2_M.gguf) | IQ2_M | 1.23GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Dolphin3.0-Llama3.2-3B-GGUF --include "Dolphin3.0-Llama3.2-3B-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Dolphin3.0-Llama3.2-3B-GGUF --include "Dolphin3.0-Llama3.2-3B-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Dolphin3.0-Llama3.2-3B-Q8_0) or download them all in place (./)
</details>
## ARM/AVX information
Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.
Now, however, there is something called "online repacking" for weights. details in [this PR](https://github.com/ggerganov/llama.cpp/pull/9921). If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.
As of llama.cpp build [b4282](https://github.com/ggerganov/llama.cpp/releases/tag/b4282) you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.
Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to [this PR](https://github.com/ggerganov/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information (deprecated</summary>
I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}