初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Llama-3.1-Tulu-3-8B-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-19 01:45:08 +08:00
commit 2cf3de28e6
28 changed files with 306 additions and 0 deletions

60
.gitattributes vendored Normal file
View File

@@ -0,0 +1,60 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B-f16.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3.1-Tulu-3-8B.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4ebaa3b5ec92760e397d48886f268dac40eff5187c14dc893aae8ca7c617cde8
size 2948318912

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5b67e50048effeb78319e61c25ca58cda9a0d2c6bd60d0aea24ca077770b8445
size 3784865728

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e07083e8f590e268aef0d7a27e363d233a8ffc6767646cff6dc7003cceff385a
size 3518789568

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2b6903b784cd57e5d3e5f668b34a20d10eb34c6c29e9b949905645f0c7266165
size 4447708352

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2ff002c6016a59907104a9052c1b8fac75c13520e0d121a8a078f6de4790d0d9
size 3179170496

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0dca2a97cfffcf2fe6a49d4fbba24efb0286473c37e5f625ef246059d79359ce
size 3692226496

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:25776f1ab791e5134e353031efb3ddbbfe5a198b8b8b8010b1f45bf3c37456d3
size 4321998784

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d6da6d2e3042dbc53d0c507a11c5c89e851d57d6256ce35f460779ddcff5f8c5
size 4018960320

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:56c25db69d3d031b88812c4412bb5e25c509430b421449f36ad5030504560723
size 3664541632

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b9ad1793300e52f7e35b82a29cbb0c756afde06df866577ecd001500ce17869d
size 4781696960

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:936243f4ab34ab3cb73dd4656a9186e61c924b8b25a4fc5b33f7221cbd1d53f6
size 4675938496

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:07d03f095b5c20b86cd49df9530c9b07ae87a369fafb581027e4a47ada154997
size 4661258432

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e4b5e316b819a9169bbf7801f0b5e3415064256a919cc3b0933aef5519c239cb
size 4661258432

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9aa5fbbfd8c673dc5e1efcc197883d9c6e661442d0bb3379a91c4306965ba1b8
size 4661258432

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:72d3d5c6c3d80d97a71be8e6a9b9165bc875294998b8ace3bdf3145e93b747a7
size 5310703552

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:12c3f945020278557768e6ecbdf2b410fe82db8ec179ea09f65d5f73e825383c
size 4920780992

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:429d5839b804af9eb166ea2020222195fbb3666304114493a2a61d0f9d13d2b2
size 4692715712

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f3a5eda38c2d6f01cec60b59612bd7eca85f47e16b45808741749af4f46dbe89
size 6057289664

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:12edecb3798cea1761657df5e42a47eb5fc5cbadec6078fd6225a1afbee0168d
size 5733038272

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2a2de50762a6889f0901013b3fe181b624e30f41859ce8a547a2371130432ec3
size 5599344832

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:42489eb87d86cd44aca63aa89ec26fec16a63940011404db22a014c9173d1da0
size 6596061632

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b4718c3e8bf5a1666a7ebe7f4131ec6ffbb3f9e74989cee66fb3effe7b912333
size 6850537408

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:388db24ac65abaf1cd9dc9a0fc8d5aaebde5df908048b89c8cf3c2cec92562ef
size 8540841632

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d658b36213d21daeffd9c6d3dacebc85537161d277325747ae0b566d8c0c0c4d
size 16069023392

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:687882ada167d588f9652cbdb2c1dfd18e3e6f71d8e2a86fecf5d4a1f16da637
size 4988170

170
README.md Normal file
View File

@@ -0,0 +1,170 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
datasets:
- allenai/RLVR-GSM-MATH-IF-Mixed-Constraints
base_model: allenai/Llama-3.1-Tulu-3-8B
license: llama3.1
language:
- en
---
## Llamacpp imatrix Quantizations of Llama-3.1-Tulu-3-8B
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4132">b4132</a> for quantization.
Original model: https://huggingface.co/allenai/Llama-3.1-Tulu-3-8B
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<|system|>
{system_prompt}
<|user|>
{prompt}
<|assistant|>
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Llama-3.1-Tulu-3-8B-f16.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-GGUF/blob/main/Llama-3.1-Tulu-3-8B-f16.gguf) | f16 | 16.07GB | false | Full F16 weights. |
| [Llama-3.1-Tulu-3-8B-Q8_0.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-GGUF/blob/main/Llama-3.1-Tulu-3-8B-Q8_0.gguf) | Q8_0 | 8.54GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Llama-3.1-Tulu-3-8B-Q6_K_L.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-GGUF/blob/main/Llama-3.1-Tulu-3-8B-Q6_K_L.gguf) | Q6_K_L | 6.85GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Llama-3.1-Tulu-3-8B-Q6_K.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-GGUF/blob/main/Llama-3.1-Tulu-3-8B-Q6_K.gguf) | Q6_K | 6.60GB | false | Very high quality, near perfect, *recommended*. |
| [Llama-3.1-Tulu-3-8B-Q5_K_L.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-GGUF/blob/main/Llama-3.1-Tulu-3-8B-Q5_K_L.gguf) | Q5_K_L | 6.06GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Llama-3.1-Tulu-3-8B-Q5_K_M.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-GGUF/blob/main/Llama-3.1-Tulu-3-8B-Q5_K_M.gguf) | Q5_K_M | 5.73GB | false | High quality, *recommended*. |
| [Llama-3.1-Tulu-3-8B-Q5_K_S.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-GGUF/blob/main/Llama-3.1-Tulu-3-8B-Q5_K_S.gguf) | Q5_K_S | 5.60GB | false | High quality, *recommended*. |
| [Llama-3.1-Tulu-3-8B-Q4_K_L.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-GGUF/blob/main/Llama-3.1-Tulu-3-8B-Q4_K_L.gguf) | Q4_K_L | 5.31GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Llama-3.1-Tulu-3-8B-Q4_K_M.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-GGUF/blob/main/Llama-3.1-Tulu-3-8B-Q4_K_M.gguf) | Q4_K_M | 4.92GB | false | Good quality, default size for most use cases, *recommended*. |
| [Llama-3.1-Tulu-3-8B-Q3_K_XL.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-GGUF/blob/main/Llama-3.1-Tulu-3-8B-Q3_K_XL.gguf) | Q3_K_XL | 4.78GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Llama-3.1-Tulu-3-8B-Q4_K_S.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-GGUF/blob/main/Llama-3.1-Tulu-3-8B-Q4_K_S.gguf) | Q4_K_S | 4.69GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Llama-3.1-Tulu-3-8B-Q4_0.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-GGUF/blob/main/Llama-3.1-Tulu-3-8B-Q4_0.gguf) | Q4_0 | 4.68GB | false | Legacy format, generally not worth using over similarly sized formats |
| [Llama-3.1-Tulu-3-8B-Q4_0_8_8.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-GGUF/blob/main/Llama-3.1-Tulu-3-8B-Q4_0_8_8.gguf) | Q4_0_8_8 | 4.66GB | false | Optimized for ARM and AVX inference. Requires 'sve' support for ARM (see details below). *Don't use on Mac*. |
| [Llama-3.1-Tulu-3-8B-Q4_0_4_8.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-GGUF/blob/main/Llama-3.1-Tulu-3-8B-Q4_0_4_8.gguf) | Q4_0_4_8 | 4.66GB | false | Optimized for ARM inference. Requires 'i8mm' support (see details below). *Don't use on Mac*. |
| [Llama-3.1-Tulu-3-8B-Q4_0_4_4.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-GGUF/blob/main/Llama-3.1-Tulu-3-8B-Q4_0_4_4.gguf) | Q4_0_4_4 | 4.66GB | false | Optimized for ARM inference. Should work well on all ARM chips, not for use with GPUs. *Don't use on Mac*. |
| [Llama-3.1-Tulu-3-8B-IQ4_XS.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-GGUF/blob/main/Llama-3.1-Tulu-3-8B-IQ4_XS.gguf) | IQ4_XS | 4.45GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Llama-3.1-Tulu-3-8B-Q3_K_L.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-GGUF/blob/main/Llama-3.1-Tulu-3-8B-Q3_K_L.gguf) | Q3_K_L | 4.32GB | false | Lower quality but usable, good for low RAM availability. |
| [Llama-3.1-Tulu-3-8B-Q3_K_M.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-GGUF/blob/main/Llama-3.1-Tulu-3-8B-Q3_K_M.gguf) | Q3_K_M | 4.02GB | false | Low quality. |
| [Llama-3.1-Tulu-3-8B-IQ3_M.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-GGUF/blob/main/Llama-3.1-Tulu-3-8B-IQ3_M.gguf) | IQ3_M | 3.78GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Llama-3.1-Tulu-3-8B-Q2_K_L.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-GGUF/blob/main/Llama-3.1-Tulu-3-8B-Q2_K_L.gguf) | Q2_K_L | 3.69GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Llama-3.1-Tulu-3-8B-Q3_K_S.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-GGUF/blob/main/Llama-3.1-Tulu-3-8B-Q3_K_S.gguf) | Q3_K_S | 3.66GB | false | Low quality, not recommended. |
| [Llama-3.1-Tulu-3-8B-IQ3_XS.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-GGUF/blob/main/Llama-3.1-Tulu-3-8B-IQ3_XS.gguf) | IQ3_XS | 3.52GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Llama-3.1-Tulu-3-8B-Q2_K.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-GGUF/blob/main/Llama-3.1-Tulu-3-8B-Q2_K.gguf) | Q2_K | 3.18GB | false | Very low quality but surprisingly usable. |
| [Llama-3.1-Tulu-3-8B-IQ2_M.gguf](https://huggingface.co/bartowski/Llama-3.1-Tulu-3-8B-GGUF/blob/main/Llama-3.1-Tulu-3-8B-IQ2_M.gguf) | IQ2_M | 2.95GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Llama-3.1-Tulu-3-8B-GGUF --include "Llama-3.1-Tulu-3-8B-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Llama-3.1-Tulu-3-8B-GGUF --include "Llama-3.1-Tulu-3-8B-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Llama-3.1-Tulu-3-8B-Q8_0) or download them all in place (./)
</details>
## Q4_0_X_X information
<details>
<summary>Click to view Q4_0_X_X information</summary>
These are *NOT* for Metal (Apple) or GPU (nvidia/AMD/intel) offloading, only ARM chips (and certain AVX2/AVX512 CPUs).
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
If you're using a CPU that supports AVX2 or AVX512 (typically server CPUs and AMD's latest Zen5 CPUs) and are not offloading to a GPU, the Q4_0_8_8 may offer a nice speed as well:
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}