初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Dolphin3.0-Llama3.2-1B-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-22 23:08:07 +08:00
commit 92d0a27286
28 changed files with 320 additions and 0 deletions

60
.gitattributes vendored Normal file
View File

@@ -0,0 +1,60 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-1B-f16.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-1B-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-1B-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-1B-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-1B-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-1B-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-1B-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-1B-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-1B-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-1B-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-1B-Q4_1.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-1B-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-1B-IQ4_NL.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-1B-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-1B-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-1B-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-1B-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-1B-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-1B-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-1B-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-1B-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-1B-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-1B-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-1B-f32.gguf filter=lfs diff=lfs merge=lfs -text
Dolphin3.0-Llama3.2-1B.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0a2cca8e9820dc20a57a95fae02ff3c2c0d4f6c64bcd7bcf6ea50db25174a664
size 515451392

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:60d838c8505b3d731669a94d94c7fc7e47482ee6baff2dd5cc53c708f4a79926
size 657292320

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:641698250e52e0f38f0ca5c4e41b2cb112a40001a7fc062d4fc4078123321dde
size 621116448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:06ef4c27ccd93c82f883182ffce49c827068d736a75be74cee3558be7d6fbc3d
size 773028896

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:65108c8359f79a977470ed23bcf36fdd64d93521a133fbaeb54b0c6479b745e7
size 743144480

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1e0681b236e13f1b967f0e9799a6c6c38ddde94fc50adbee9da167a6286b91d4
size 580877344

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:558d7edcd76c2fed05d03fabd741cba30d45daa283481a846a1754ec4dad0376
size 644493312

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4e6bb2f75094302378c3ae12eb8c07549d45ac794096b8f1c52dd435a5ca5b01
size 732527648

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c12a6eddc3393802743a00dccae6df65d387b773ba0924f0b27d017345476e76
size 690846752

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7264d16b4eecd0b2ad2471a61b0697fb199a6a943868d01a75294efca5221ba7
size 641694752

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9dfbe76ceeb36dd4792b5e89191ca60acf7e91800a1d2426583ab73a38de8352
size 796143616

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a4548befa7ea666fedabc7fa5f6ff66ddd6071f26c7fb1e086a5a35aeab9bc5e
size 773028896

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:61e1cbda1fee03b633807f79d31f71f9a0c2a89236230a1e73a57565addb0b90
size 831749152

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d5ee252af1d56aa02c2a4f77aa1430e0a08b585418cb27437a69b63c579311f3
size 871313408

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7ed39ee0638e18d3e47bf12e60e917c792ca5332606a72bd1882ab1f62a13a7a
size 807697440

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8690b8fd253d778b466f9127476a9e95559d4fbeb061577bf278ff7ab03cb9c7
size 775650336

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:758dba0cd4d740d73705265e5a434558597c1a99ba10ade669d30b51b43c070f
size 975122432

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3828d2f762f9442e6bbbf2fa2410c582a045718be2777ad200112554cae70494
size 911506464

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d3a01f5c145162169c352809060112608bcaa65db8f3ab470fc77c12b35f826a
size 892566560

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:88050ebc70d2a62aeb4527457af791c3960bc6a0453b812ada83774e2e744cbe
size 1021803552

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:95fb39326c0d7e995c25ff59211e6f7d4d384b05dd38a112b0a0e2079fcfeb08
size 1085419520

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:86c649e38e27f9bb492daab07b49093997d369d3c91cc1648414688917ceac4d
size 1321086976

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:697c5bd314b57a7e7f8c6f570d452c06cb35eea4bae24e267f07b91cd4ec5e0b
size 2479603456

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:747a98c79e502aaba6b8d3ec194b3f8d230cc1b36a18934f065f63fc04da68cf
size 4951104992

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:79e74c2ab59f0c2b7b4c87226e53fcc0ad8b5fb1bfc5f735ebc7c304811ff891
size 1314426

184
README.md Normal file
View File

@@ -0,0 +1,184 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
language:
- en
base_model: cognitivecomputations/Dolphin3.0-Llama3.2-1B
license: llama3.2
datasets:
- OpenCoder-LLM/opc-sft-stage1
- OpenCoder-LLM/opc-sft-stage2
- microsoft/orca-agentinstruct-1M-v1
- microsoft/orca-math-word-problems-200k
- NousResearch/hermes-function-calling-v1
- AI-MO/NuminaMath-CoT
- AI-MO/NuminaMath-TIR
- allenai/tulu-3-sft-mixture
- cognitivecomputations/dolphin-coder
- HuggingFaceTB/smoltalk
- cognitivecomputations/samantha-data
- m-a-p/CodeFeedback-Filtered-Instruction
- m-a-p/Code-Feedback
---
## Llamacpp imatrix Quantizations of Dolphin3.0-Llama3.2-1B
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4418">b4418</a> for quantization.
Original model: https://huggingface.co/cognitivecomputations/Dolphin3.0-Llama3.2-1B
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Dolphin3.0-Llama3.2-1B-f32.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-1B-GGUF/blob/main/Dolphin3.0-Llama3.2-1B-f32.gguf) | f32 | 4.95GB | false | Full F32 weights. |
| [Dolphin3.0-Llama3.2-1B-f16.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-1B-GGUF/blob/main/Dolphin3.0-Llama3.2-1B-f16.gguf) | f16 | 2.48GB | false | Full F16 weights. |
| [Dolphin3.0-Llama3.2-1B-Q8_0.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-1B-GGUF/blob/main/Dolphin3.0-Llama3.2-1B-Q8_0.gguf) | Q8_0 | 1.32GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Dolphin3.0-Llama3.2-1B-Q6_K_L.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-1B-GGUF/blob/main/Dolphin3.0-Llama3.2-1B-Q6_K_L.gguf) | Q6_K_L | 1.09GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Dolphin3.0-Llama3.2-1B-Q6_K.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-1B-GGUF/blob/main/Dolphin3.0-Llama3.2-1B-Q6_K.gguf) | Q6_K | 1.02GB | false | Very high quality, near perfect, *recommended*. |
| [Dolphin3.0-Llama3.2-1B-Q5_K_L.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-1B-GGUF/blob/main/Dolphin3.0-Llama3.2-1B-Q5_K_L.gguf) | Q5_K_L | 0.98GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Dolphin3.0-Llama3.2-1B-Q5_K_M.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-1B-GGUF/blob/main/Dolphin3.0-Llama3.2-1B-Q5_K_M.gguf) | Q5_K_M | 0.91GB | false | High quality, *recommended*. |
| [Dolphin3.0-Llama3.2-1B-Q5_K_S.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-1B-GGUF/blob/main/Dolphin3.0-Llama3.2-1B-Q5_K_S.gguf) | Q5_K_S | 0.89GB | false | High quality, *recommended*. |
| [Dolphin3.0-Llama3.2-1B-Q4_K_L.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-1B-GGUF/blob/main/Dolphin3.0-Llama3.2-1B-Q4_K_L.gguf) | Q4_K_L | 0.87GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Dolphin3.0-Llama3.2-1B-Q4_1.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-1B-GGUF/blob/main/Dolphin3.0-Llama3.2-1B-Q4_1.gguf) | Q4_1 | 0.83GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| [Dolphin3.0-Llama3.2-1B-Q4_K_M.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-1B-GGUF/blob/main/Dolphin3.0-Llama3.2-1B-Q4_K_M.gguf) | Q4_K_M | 0.81GB | false | Good quality, default size for most use cases, *recommended*. |
| [Dolphin3.0-Llama3.2-1B-Q3_K_XL.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-1B-GGUF/blob/main/Dolphin3.0-Llama3.2-1B-Q3_K_XL.gguf) | Q3_K_XL | 0.80GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Dolphin3.0-Llama3.2-1B-Q4_K_S.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-1B-GGUF/blob/main/Dolphin3.0-Llama3.2-1B-Q4_K_S.gguf) | Q4_K_S | 0.78GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Dolphin3.0-Llama3.2-1B-Q4_0.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-1B-GGUF/blob/main/Dolphin3.0-Llama3.2-1B-Q4_0.gguf) | Q4_0 | 0.77GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| [Dolphin3.0-Llama3.2-1B-IQ4_NL.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-1B-GGUF/blob/main/Dolphin3.0-Llama3.2-1B-IQ4_NL.gguf) | IQ4_NL | 0.77GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [Dolphin3.0-Llama3.2-1B-IQ4_XS.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-1B-GGUF/blob/main/Dolphin3.0-Llama3.2-1B-IQ4_XS.gguf) | IQ4_XS | 0.74GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Dolphin3.0-Llama3.2-1B-Q3_K_L.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-1B-GGUF/blob/main/Dolphin3.0-Llama3.2-1B-Q3_K_L.gguf) | Q3_K_L | 0.73GB | false | Lower quality but usable, good for low RAM availability. |
| [Dolphin3.0-Llama3.2-1B-Q3_K_M.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-1B-GGUF/blob/main/Dolphin3.0-Llama3.2-1B-Q3_K_M.gguf) | Q3_K_M | 0.69GB | false | Low quality. |
| [Dolphin3.0-Llama3.2-1B-IQ3_M.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-1B-GGUF/blob/main/Dolphin3.0-Llama3.2-1B-IQ3_M.gguf) | IQ3_M | 0.66GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Dolphin3.0-Llama3.2-1B-Q3_K_S.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-1B-GGUF/blob/main/Dolphin3.0-Llama3.2-1B-Q3_K_S.gguf) | Q3_K_S | 0.64GB | false | Low quality, not recommended. |
| [Dolphin3.0-Llama3.2-1B-Q2_K_L.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-1B-GGUF/blob/main/Dolphin3.0-Llama3.2-1B-Q2_K_L.gguf) | Q2_K_L | 0.64GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Dolphin3.0-Llama3.2-1B-IQ3_XS.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-1B-GGUF/blob/main/Dolphin3.0-Llama3.2-1B-IQ3_XS.gguf) | IQ3_XS | 0.62GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Dolphin3.0-Llama3.2-1B-Q2_K.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-1B-GGUF/blob/main/Dolphin3.0-Llama3.2-1B-Q2_K.gguf) | Q2_K | 0.58GB | false | Very low quality but surprisingly usable. |
| [Dolphin3.0-Llama3.2-1B-IQ2_M.gguf](https://huggingface.co/bartowski/Dolphin3.0-Llama3.2-1B-GGUF/blob/main/Dolphin3.0-Llama3.2-1B-IQ2_M.gguf) | IQ2_M | 0.52GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Dolphin3.0-Llama3.2-1B-GGUF --include "Dolphin3.0-Llama3.2-1B-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Dolphin3.0-Llama3.2-1B-GGUF --include "Dolphin3.0-Llama3.2-1B-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Dolphin3.0-Llama3.2-1B-Q8_0) or download them all in place (./)
</details>
## ARM/AVX information
Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.
Now, however, there is something called "online repacking" for weights. details in [this PR](https://github.com/ggerganov/llama.cpp/pull/9921). If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.
As of llama.cpp build [b4282](https://github.com/ggerganov/llama.cpp/releases/tag/b4282) you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.
Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to [this PR](https://github.com/ggerganov/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information (deprecated</summary>
I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}