初始化项目,由ModelHub XC社区提供模型

Model: bartowski/EXAONE-3.5-7.8B-Instruct-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-18 18:35:06 +08:00
commit bf949682f0
31 changed files with 327 additions and 0 deletions

63
.gitattributes vendored Normal file
View File

@@ -0,0 +1,63 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
EXAONE-3.5-7.8B-Instruct-f16.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-3.5-7.8B-Instruct-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-3.5-7.8B-Instruct-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-3.5-7.8B-Instruct-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-3.5-7.8B-Instruct-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-3.5-7.8B-Instruct-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-3.5-7.8B-Instruct-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-3.5-7.8B-Instruct-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-3.5-7.8B-Instruct-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-3.5-7.8B-Instruct-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-3.5-7.8B-Instruct-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-3.5-7.8B-Instruct-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-3.5-7.8B-Instruct-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-3.5-7.8B-Instruct-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-3.5-7.8B-Instruct-IQ4_NL.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-3.5-7.8B-Instruct-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-3.5-7.8B-Instruct-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-3.5-7.8B-Instruct-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-3.5-7.8B-Instruct-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-3.5-7.8B-Instruct-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-3.5-7.8B-Instruct-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-3.5-7.8B-Instruct-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-3.5-7.8B-Instruct-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-3.5-7.8B-Instruct-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-3.5-7.8B-Instruct-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-3.5-7.8B-Instruct-IQ2_S.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-3.5-7.8B-Instruct-f32.gguf filter=lfs diff=lfs merge=lfs -text
EXAONE-3.5-7.8B-Instruct.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:70b8982b8f4fdbf896c3b8a80811c7952587879349d25d829927a86cf4c433b3
size 2826328256

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:714a0d8c355c58b1ad79182d826a7910925c47594bfeb130e5cb8bdeef53f8b2
size 2636536000

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4b721e6bc703872465317483f9b8c43272849b9943cc2bee07012bfe0d37d461
size 3648805056

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7dd39676a5179be70dbc948339c1dad5803e0d9857ad041eab1c1422745e7d6c
size 3382728896

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a5768d4556509ec8632c108e46697d005e4d82776e509d162ec46e1066f77b1d
size 4527904960

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8a82c17bdbc6cfdb3492b9beeeed028b8fa60d10d09844b202af385e3fa5ba4a
size 4300888256

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fe1b7ffad5a8cf9e88d64e8bee0681711c38ce30745f27f5c321586e2455bd32
size 3053869248

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9da1109e08bc769a2a98b0e04fa3f3e569efee163d482f375d146a7e94e6baee
size 3463469248

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:eb1c84837884786b0c398763456a8a2a4f3001b55fb106c2718e5c6b11ceecd4
size 4185938112

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8f20d2ce1a83499e815d7516692ba7714cb84cf7d3f617d6129528eeec04dbbe
size 3882899648

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cade3a5cb7b21b4b9b482573d1f81978d075654c92c1a2c526ddc56d13390c7c
size 3528480960

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:934893c8e9d90c2b9157f7aef1fe302c4659e2881fd693106769f731fb8b1047
size 4552939712

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:da51e4cdfbb684a961fbde7ae1188dd618f513ba838971e1bc41ba32cb353fc1
size 4525807808

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0e5aeabd00d583d679cbd50f31a55c09ac61888daa5cbd10e3b2f80b54c09579
size 4511127744

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6fc96b6071c8115b8a9cee5b41bf08d24b39384a589bbfbbfe21b5a296384cc9
size 4511127744

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d59413ffea2b7680c80a621e83f2ec933478103cdc97945d795b04d2ef86f81b
size 4511127744

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:29eada66a7d6c07f939e4bd18052fee5f9b972c801b9429e3bc4863ed964f001
size 5081946304

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3224531828df45db7834521a7310cd4bbd9c998168ba64fbee81d9be0c2cdee0
size 4770650304

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ae582b736fe696b0bfa3e9ee5a96b954a903893fe0c6451f809273cc9691a6ec
size 4542585024

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fa7ffdb14dab8d1dabd94bca4fd97da1f63a977f08eeae16e3a56a4917e6fd25
size 5828532416

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b8cdd7205dbd5eacd68f2fd47b4ad34f986104668b3e8a3b9df410febdb8a8e0
size 5569665216

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:62fa8dd4f4b0ddb2785b1c505de5ee8eb4a4455028774755cd4ee4e0ab508b98
size 5435971776

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:77cf909a5015e710303fc2a56ee9a205dc1015ca1ebfc0a8905d4226ea3745dd
size 6418618560

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:875c195c2edcf6d91d9f58c563feb17171e9694576350f849d80a14c181b313c
size 6621780160

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b50ab9a91dfbe2879553c0b6bd73e703b43e7b79245912414c5b60bb90367f39
size 8312084672

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:dbeda9bf7fe7d0c3ca4bb562438a24d7dd05dd7be70229e4b10c4aa0ba6a7c03
size 15641630912

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0e3f91a75cac14e35055b423ea3c394e3bdc85494f031b139e522c0bce079c8d
size 31277995936

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:22b14d43097e740f3c0c58ccd62b469d59ab22d40b629b37521d8ceefa53a604
size 4988170

179
README.md Normal file
View File

@@ -0,0 +1,179 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
license_link: LICENSE
base_model: LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct
language:
- en
- ko
license: other
license_name: exaone
tags:
- lg-ai
- exaone
- exaone-3.5
---
## Llamacpp imatrix Quantizations of EXAONE-3.5-7.8B-Instruct
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4273">b4273</a> for quantization.
Original model: https://huggingface.co/LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
[|system|]{system_prompt}[|endofturn|]
[|user|]{prompt}
[|assistant|]
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [EXAONE-3.5-7.8B-Instruct-f32.gguf](https://huggingface.co/bartowski/EXAONE-3.5-7.8B-Instruct-GGUF/blob/main/EXAONE-3.5-7.8B-Instruct-f32.gguf) | f32 | 31.28GB | false | Full F32 weights. |
| [EXAONE-3.5-7.8B-Instruct-f16.gguf](https://huggingface.co/bartowski/EXAONE-3.5-7.8B-Instruct-GGUF/blob/main/EXAONE-3.5-7.8B-Instruct-f16.gguf) | f16 | 15.64GB | false | Full F16 weights. |
| [EXAONE-3.5-7.8B-Instruct-Q8_0.gguf](https://huggingface.co/bartowski/EXAONE-3.5-7.8B-Instruct-GGUF/blob/main/EXAONE-3.5-7.8B-Instruct-Q8_0.gguf) | Q8_0 | 8.31GB | false | Extremely high quality, generally unneeded but max available quant. |
| [EXAONE-3.5-7.8B-Instruct-Q6_K_L.gguf](https://huggingface.co/bartowski/EXAONE-3.5-7.8B-Instruct-GGUF/blob/main/EXAONE-3.5-7.8B-Instruct-Q6_K_L.gguf) | Q6_K_L | 6.62GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [EXAONE-3.5-7.8B-Instruct-Q6_K.gguf](https://huggingface.co/bartowski/EXAONE-3.5-7.8B-Instruct-GGUF/blob/main/EXAONE-3.5-7.8B-Instruct-Q6_K.gguf) | Q6_K | 6.42GB | false | Very high quality, near perfect, *recommended*. |
| [EXAONE-3.5-7.8B-Instruct-Q5_K_L.gguf](https://huggingface.co/bartowski/EXAONE-3.5-7.8B-Instruct-GGUF/blob/main/EXAONE-3.5-7.8B-Instruct-Q5_K_L.gguf) | Q5_K_L | 5.83GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [EXAONE-3.5-7.8B-Instruct-Q5_K_M.gguf](https://huggingface.co/bartowski/EXAONE-3.5-7.8B-Instruct-GGUF/blob/main/EXAONE-3.5-7.8B-Instruct-Q5_K_M.gguf) | Q5_K_M | 5.57GB | false | High quality, *recommended*. |
| [EXAONE-3.5-7.8B-Instruct-Q5_K_S.gguf](https://huggingface.co/bartowski/EXAONE-3.5-7.8B-Instruct-GGUF/blob/main/EXAONE-3.5-7.8B-Instruct-Q5_K_S.gguf) | Q5_K_S | 5.44GB | false | High quality, *recommended*. |
| [EXAONE-3.5-7.8B-Instruct-Q4_K_L.gguf](https://huggingface.co/bartowski/EXAONE-3.5-7.8B-Instruct-GGUF/blob/main/EXAONE-3.5-7.8B-Instruct-Q4_K_L.gguf) | Q4_K_L | 5.08GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [EXAONE-3.5-7.8B-Instruct-Q4_K_M.gguf](https://huggingface.co/bartowski/EXAONE-3.5-7.8B-Instruct-GGUF/blob/main/EXAONE-3.5-7.8B-Instruct-Q4_K_M.gguf) | Q4_K_M | 4.77GB | false | Good quality, default size for most use cases, *recommended*. |
| [EXAONE-3.5-7.8B-Instruct-Q3_K_XL.gguf](https://huggingface.co/bartowski/EXAONE-3.5-7.8B-Instruct-GGUF/blob/main/EXAONE-3.5-7.8B-Instruct-Q3_K_XL.gguf) | Q3_K_XL | 4.55GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [EXAONE-3.5-7.8B-Instruct-Q4_K_S.gguf](https://huggingface.co/bartowski/EXAONE-3.5-7.8B-Instruct-GGUF/blob/main/EXAONE-3.5-7.8B-Instruct-Q4_K_S.gguf) | Q4_K_S | 4.54GB | false | Slightly lower quality with more space savings, *recommended*. |
| [EXAONE-3.5-7.8B-Instruct-Q4_0.gguf](https://huggingface.co/bartowski/EXAONE-3.5-7.8B-Instruct-GGUF/blob/main/EXAONE-3.5-7.8B-Instruct-Q4_0.gguf) | Q4_0 | 4.53GB | false | Legacy format, offers online repacking for ARM CPU inference. |
| [EXAONE-3.5-7.8B-Instruct-IQ4_NL.gguf](https://huggingface.co/bartowski/EXAONE-3.5-7.8B-Instruct-GGUF/blob/main/EXAONE-3.5-7.8B-Instruct-IQ4_NL.gguf) | IQ4_NL | 4.53GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [EXAONE-3.5-7.8B-Instruct-Q4_0_8_8.gguf](https://huggingface.co/bartowski/EXAONE-3.5-7.8B-Instruct-GGUF/blob/main/EXAONE-3.5-7.8B-Instruct-Q4_0_8_8.gguf) | Q4_0_8_8 | 4.51GB | false | Optimized for ARM and AVX inference. Requires 'sve' support for ARM (see details below). *Don't use on Mac*. |
| [EXAONE-3.5-7.8B-Instruct-Q4_0_4_8.gguf](https://huggingface.co/bartowski/EXAONE-3.5-7.8B-Instruct-GGUF/blob/main/EXAONE-3.5-7.8B-Instruct-Q4_0_4_8.gguf) | Q4_0_4_8 | 4.51GB | false | Optimized for ARM inference. Requires 'i8mm' support (see details below). *Don't use on Mac*. |
| [EXAONE-3.5-7.8B-Instruct-Q4_0_4_4.gguf](https://huggingface.co/bartowski/EXAONE-3.5-7.8B-Instruct-GGUF/blob/main/EXAONE-3.5-7.8B-Instruct-Q4_0_4_4.gguf) | Q4_0_4_4 | 4.51GB | false | Optimized for ARM inference. Should work well on all ARM chips, not for use with GPUs. *Don't use on Mac*. |
| [EXAONE-3.5-7.8B-Instruct-IQ4_XS.gguf](https://huggingface.co/bartowski/EXAONE-3.5-7.8B-Instruct-GGUF/blob/main/EXAONE-3.5-7.8B-Instruct-IQ4_XS.gguf) | IQ4_XS | 4.30GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [EXAONE-3.5-7.8B-Instruct-Q3_K_L.gguf](https://huggingface.co/bartowski/EXAONE-3.5-7.8B-Instruct-GGUF/blob/main/EXAONE-3.5-7.8B-Instruct-Q3_K_L.gguf) | Q3_K_L | 4.19GB | false | Lower quality but usable, good for low RAM availability. |
| [EXAONE-3.5-7.8B-Instruct-Q3_K_M.gguf](https://huggingface.co/bartowski/EXAONE-3.5-7.8B-Instruct-GGUF/blob/main/EXAONE-3.5-7.8B-Instruct-Q3_K_M.gguf) | Q3_K_M | 3.88GB | false | Low quality. |
| [EXAONE-3.5-7.8B-Instruct-IQ3_M.gguf](https://huggingface.co/bartowski/EXAONE-3.5-7.8B-Instruct-GGUF/blob/main/EXAONE-3.5-7.8B-Instruct-IQ3_M.gguf) | IQ3_M | 3.65GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [EXAONE-3.5-7.8B-Instruct-Q3_K_S.gguf](https://huggingface.co/bartowski/EXAONE-3.5-7.8B-Instruct-GGUF/blob/main/EXAONE-3.5-7.8B-Instruct-Q3_K_S.gguf) | Q3_K_S | 3.53GB | false | Low quality, not recommended. |
| [EXAONE-3.5-7.8B-Instruct-Q2_K_L.gguf](https://huggingface.co/bartowski/EXAONE-3.5-7.8B-Instruct-GGUF/blob/main/EXAONE-3.5-7.8B-Instruct-Q2_K_L.gguf) | Q2_K_L | 3.46GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [EXAONE-3.5-7.8B-Instruct-IQ3_XS.gguf](https://huggingface.co/bartowski/EXAONE-3.5-7.8B-Instruct-GGUF/blob/main/EXAONE-3.5-7.8B-Instruct-IQ3_XS.gguf) | IQ3_XS | 3.38GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [EXAONE-3.5-7.8B-Instruct-Q2_K.gguf](https://huggingface.co/bartowski/EXAONE-3.5-7.8B-Instruct-GGUF/blob/main/EXAONE-3.5-7.8B-Instruct-Q2_K.gguf) | Q2_K | 3.05GB | false | Very low quality but surprisingly usable. |
| [EXAONE-3.5-7.8B-Instruct-IQ2_M.gguf](https://huggingface.co/bartowski/EXAONE-3.5-7.8B-Instruct-GGUF/blob/main/EXAONE-3.5-7.8B-Instruct-IQ2_M.gguf) | IQ2_M | 2.83GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
| [EXAONE-3.5-7.8B-Instruct-IQ2_S.gguf](https://huggingface.co/bartowski/EXAONE-3.5-7.8B-Instruct-GGUF/blob/main/EXAONE-3.5-7.8B-Instruct-IQ2_S.gguf) | IQ2_S | 2.64GB | false | Low quality, uses SOTA techniques to be usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/EXAONE-3.5-7.8B-Instruct-GGUF --include "EXAONE-3.5-7.8B-Instruct-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/EXAONE-3.5-7.8B-Instruct-GGUF --include "EXAONE-3.5-7.8B-Instruct-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (EXAONE-3.5-7.8B-Instruct-Q8_0) or download them all in place (./)
</details>
## Q4_0_X_X information
New: Thanks to efforts made to have online repacking of weights in [this PR](https://github.com/ggerganov/llama.cpp/pull/9921), you can now just use Q4_0 if your llama.cpp has been compiled for your ARM device.
Similarly, if you want to get slightly better performance, you can use IQ4_NL thanks to [this PR](https://github.com/ggerganov/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information</summary>
These are *NOT* for Metal (Apple) or GPU (nvidia/AMD/intel) offloading, only ARM chips (and certain AVX2/AVX512 CPUs).
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
If you're using a CPU that supports AVX2 or AVX512 (typically server CPUs and AMD's latest Zen5 CPUs) and are not offloading to a GPU, the Q4_0_8_8 may offer a nice speed as well:
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}