初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Llama-3.1-Tulu-3-8B-DPO-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-21 15:22:07 +08:00
commit c9ce3c2b44
28 changed files with 306 additions and 0 deletions

60
.gitattributes vendored Normal file
View File

@@ -0,0 +1,60 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-DPO-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-DPO-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-DPO-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-DPO-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-DPO-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-DPO-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-DPO-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-DPO-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-DPO-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-DPO-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-DPO-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-DPO-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-DPO-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-DPO-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-DPO-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-DPO-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-DPO-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-DPO-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-DPO-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-DPO-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-DPO-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-DPO-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-DPO-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-DPO-f16.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-DPO.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:dd22e93a379545ab91687d3803b5dc43f588c6bdf0720df2cc2df571f3d19ad4
size 2948318944

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f04deed83cfff0edec5a9cb54c7caaf659f699a09816eb2f7f0605dd968143fa
size 3784865760

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6f0bb2379438894b49d460e5de02d78b52688c05a6fd129df2ade4319b3126ec
size 3518789600

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ed14d0cfef0ae16620f31a860a1b657c6c0560fca1d11f857ccb0cf0f06ef319
size 4447708384

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e43318d4273bc0011af2ba20daad78f3017177371fdac3565666c3f070582761
size 3179170528

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4740f431f5fb5195542dc58fcccbc1691b108c7346f28aaa5d7de5bfc06eeed9
size 3692226528

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4f1f147fa6b6a7d2d20bfcbc85d43d9b1dd5ba983118537588d6c475ec7a1be2
size 4321998816

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d8dc80bec1f632c9bf53c797bebd1854066545db009bdd0f01b620497ad13bcd
size 4018960352

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:728ba0511ee8c48c679c1e1d12b380f2d4c87c210448c04018a60a45cbb6f40a
size 3664541664

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7ca0c1891f4856585a4a382f0ae26e66d1aea3c73cc02375abc10fafe34c2834
size 4781696992

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ef8f385099f8d8adc43ffd3340cd9bd0db55585f4db86ac747fafe2989d22943
size 4675938528

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a939da901fe720ade1b170178539c8a8b6e3a7b999a7212fc296792e00362ba8
size 4661258464

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e0ffe75095cc1ed94aea109a465c87e0001ec0c15d4ab46557b361b0e46f8dfb
size 4661258464

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:15f6c0c37065cf894db2db5203598cbcacf8b3d87f818d93b200dca114ce2e93
size 4661258464

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:efbb46e5a0761fbc5d4702b19ebc902facfbe642a9137a412e8ef466179de655
size 5310703584

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6ca51a88ad7db7fa4dad0f26a11deb674572bfd73765268920d2545b5c716fa5
size 4920781024

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1ff883b69eb746341ced2a3ed75cd927b93c2d88c24dfffda84abf2f8ec1d6b7
size 4692715744

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9fb1c0f8e8072e652f73e4f0938ca611d1ecbc2f3a90e2afcb730b8490a7f224
size 6057289696

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4af4b217082997efd69c3722926581453c43366ca293cb1d44a2d009e729f727
size 5733038304

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4946b245c8be686c9ed185acfe9231f12f9737667fbb562ea5e3fd6d73aabb75
size 5599344864

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:081d6ac7d7b265c3d89794455bc85cfc487084ce1e42db76e9c99fdb2218a714
size 6596061664

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:138b9c206e81a45c5cc056995183057611040cf8c4e77b4999f7eda07f65e9d9
size 6850537440

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6b96226973b84d29d6ccebc21f30ca554bac7778669150dbc8951e29cbec6b16
size 8540841952

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3a0bf67805d1785f8a6a62eeabc9729cc583aa632cd5f177eec427d105fa544d
size 16069023424

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:48120f4b77719f6d4b5772fae733dc096956954fca870cd57363b0c09ea0fbca
size 4988170

170
README.md Normal file
View File

@@ -0,0 +1,170 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
datasets:
- allenai/llama-3.1-tulu-3-8b-preference-mixture
base_model: allenai/Llama-3.1-Tulu-3-8B-DPO
license: llama3.1
language:
- en
---
## Llamacpp imatrix Quantizations of Llama-3.1-Tulu-3-8B-DPO
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4132">b4132</a> for quantization.
Original model: https://huggingface.co/allenai/Llama-3.1-Tulu-3-8B-DPO
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<|system|>
{system_prompt}
<|user|>
{prompt}
<|assistant|>
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Llama-3.1-Tulu-3-8B-DPO-f16.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-DPO-GGUF/blob/main/Llama-3.1-Tulu-3-8B-DPO-f16.gguf) | f16 | 16.07GB | false | Full F16 weights. |
| [Llama-3.1-Tulu-3-8B-DPO-Q8_0.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-DPO-GGUF/blob/main/Llama-3.1-Tulu-3-8B-DPO-Q8_0.gguf) | Q8_0 | 8.54GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Llama-3.1-Tulu-3-8B-DPO-Q6_K_L.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-DPO-GGUF/blob/main/Llama-3.1-Tulu-3-8B-DPO-Q6_K_L.gguf) | Q6_K_L | 6.85GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Llama-3.1-Tulu-3-8B-DPO-Q6_K.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-DPO-GGUF/blob/main/Llama-3.1-Tulu-3-8B-DPO-Q6_K.gguf) | Q6_K | 6.60GB | false | Very high quality, near perfect, *recommended*. |
| [Llama-3.1-Tulu-3-8B-DPO-Q5_K_L.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-DPO-GGUF/blob/main/Llama-3.1-Tulu-3-8B-DPO-Q5_K_L.gguf) | Q5_K_L | 6.06GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Llama-3.1-Tulu-3-8B-DPO-Q5_K_M.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-DPO-GGUF/blob/main/Llama-3.1-Tulu-3-8B-DPO-Q5_K_M.gguf) | Q5_K_M | 5.73GB | false | High quality, *recommended*. |
| [Llama-3.1-Tulu-3-8B-DPO-Q5_K_S.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-DPO-GGUF/blob/main/Llama-3.1-Tulu-3-8B-DPO-Q5_K_S.gguf) | Q5_K_S | 5.60GB | false | High quality, *recommended*. |
| [Llama-3.1-Tulu-3-8B-DPO-Q4_K_L.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-DPO-GGUF/blob/main/Llama-3.1-Tulu-3-8B-DPO-Q4_K_L.gguf) | Q4_K_L | 5.31GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Llama-3.1-Tulu-3-8B-DPO-Q4_K_M.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-DPO-GGUF/blob/main/Llama-3.1-Tulu-3-8B-DPO-Q4_K_M.gguf) | Q4_K_M | 4.92GB | false | Good quality, default size for most use cases, *recommended*. |
| [Llama-3.1-Tulu-3-8B-DPO-Q3_K_XL.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-DPO-GGUF/blob/main/Llama-3.1-Tulu-3-8B-DPO-Q3_K_XL.gguf) | Q3_K_XL | 4.78GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Llama-3.1-Tulu-3-8B-DPO-Q4_K_S.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-DPO-GGUF/blob/main/Llama-3.1-Tulu-3-8B-DPO-Q4_K_S.gguf) | Q4_K_S | 4.69GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Llama-3.1-Tulu-3-8B-DPO-Q4_0.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-DPO-GGUF/blob/main/Llama-3.1-Tulu-3-8B-DPO-Q4_0.gguf) | Q4_0 | 4.68GB | false | Legacy format, generally not worth using over similarly sized formats |
| [Llama-3.1-Tulu-3-8B-DPO-Q4_0_8_8.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-DPO-GGUF/blob/main/Llama-3.1-Tulu-3-8B-DPO-Q4_0_8_8.gguf) | Q4_0_8_8 | 4.66GB | false | Optimized for ARM and AVX inference. Requires 'sve' support for ARM (see details below). *Don't use on Mac*. |
| [Llama-3.1-Tulu-3-8B-DPO-Q4_0_4_8.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-DPO-GGUF/blob/main/Llama-3.1-Tulu-3-8B-DPO-Q4_0_4_8.gguf) | Q4_0_4_8 | 4.66GB | false | Optimized for ARM inference. Requires 'i8mm' support (see details below). *Don't use on Mac*. |
| [Llama-3.1-Tulu-3-8B-DPO-Q4_0_4_4.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-DPO-GGUF/blob/main/Llama-3.1-Tulu-3-8B-DPO-Q4_0_4_4.gguf) | Q4_0_4_4 | 4.66GB | false | Optimized for ARM inference. Should work well on all ARM chips, not for use with GPUs. *Don't use on Mac*. |
| [Llama-3.1-Tulu-3-8B-DPO-IQ4_XS.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-DPO-GGUF/blob/main/Llama-3.1-Tulu-3-8B-DPO-IQ4_XS.gguf) | IQ4_XS | 4.45GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Llama-3.1-Tulu-3-8B-DPO-Q3_K_L.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-DPO-GGUF/blob/main/Llama-3.1-Tulu-3-8B-DPO-Q3_K_L.gguf) | Q3_K_L | 4.32GB | false | Lower quality but usable, good for low RAM availability. |
| [Llama-3.1-Tulu-3-8B-DPO-Q3_K_M.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-DPO-GGUF/blob/main/Llama-3.1-Tulu-3-8B-DPO-Q3_K_M.gguf) | Q3_K_M | 4.02GB | false | Low quality. |
| [Llama-3.1-Tulu-3-8B-DPO-IQ3_M.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-DPO-GGUF/blob/main/Llama-3.1-Tulu-3-8B-DPO-IQ3_M.gguf) | IQ3_M | 3.78GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Llama-3.1-Tulu-3-8B-DPO-Q2_K_L.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-DPO-GGUF/blob/main/Llama-3.1-Tulu-3-8B-DPO-Q2_K_L.gguf) | Q2_K_L | 3.69GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Llama-3.1-Tulu-3-8B-DPO-Q3_K_S.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-DPO-GGUF/blob/main/Llama-3.1-Tulu-3-8B-DPO-Q3_K_S.gguf) | Q3_K_S | 3.66GB | false | Low quality, not recommended. |
| [Llama-3.1-Tulu-3-8B-DPO-IQ3_XS.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-DPO-GGUF/blob/main/Llama-3.1-Tulu-3-8B-DPO-IQ3_XS.gguf) | IQ3_XS | 3.52GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Llama-3.1-Tulu-3-8B-DPO-Q2_K.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-DPO-GGUF/blob/main/Llama-3.1-Tulu-3-8B-DPO-Q2_K.gguf) | Q2_K | 3.18GB | false | Very low quality but surprisingly usable. |
| [Llama-3.1-Tulu-3-8B-DPO-IQ2_M.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-DPO-GGUF/blob/main/Llama-3.1-Tulu-3-8B-DPO-IQ2_M.gguf) | IQ2_M | 2.95GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Llama-3.1-Tulu-3-8B-DPO-GGUF --include "Llama-3.1-Tulu-3-8B-DPO-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Llama-3.1-Tulu-3-8B-DPO-GGUF --include "Llama-3.1-Tulu-3-8B-DPO-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Llama-3.1-Tulu-3-8B-DPO-Q8_0) or download them all in place (./)
</details>
## Q4_0_X_X information
<details>
<summary>Click to view Q4_0_X_X information</summary>
These are *NOT* for Metal (Apple) or GPU (nvidia/AMD/intel) offloading, only ARM chips (and certain AVX2/AVX512 CPUs).
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
If you're using a CPU that supports AVX2 or AVX512 (typically server CPUs and AMD's latest Zen5 CPUs) and are not offloading to a GPU, the Q4_0_8_8 may offer a nice speed as well:
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}