初始化项目,由ModelHub XC社区提供模型

Model: bartowski/LGAI-EXAONE_EXAONE-Deep-2.4B-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-09-23 22:53:13 +08:00
commit ad130151c8
28 changed files with 300 additions and 0 deletions

47
.gitattributes vendored Normal file
View File

@@ -0,0 +1,47 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bin.* filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zstandard filter=lfs diff=lfs merge=lfs -text
*.tfevents* filter=lfs diff=lfs merge=lfs -text
*.db* filter=lfs diff=lfs merge=lfs -text
*.ark* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*data* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.meta filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.index filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.gguf* filter=lfs diff=lfs merge=lfs -text
*.ggml filter=lfs diff=lfs merge=lfs -text
*.llamafile* filter=lfs diff=lfs merge=lfs -text
*.pt2 filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:07596179b4485e972a12c83e34690947eadb24404d039835b87688878394ca5f
size 910915968

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8aef0cd0d00aea16d90ef23871acc15162f480018ff1d3bb9d63fdd8fe52423e
size 1180647808

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8e2e517043e603071a9b829d527059fcce00768a6bc39145cf9c0179453bbed1
size 1096137088

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c81ad93157d81edaac3c0cae1e6b20d2fa45aab379284938ce7b855e8142c0a4
size 1008114048

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a0cc54bc5d579cd2e38559c77589764a3a4131852140bcdc8b1b7f38ac4eb22b
size 1431461248

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e596adf1188de1ab186aeb7394b88a4f9bb4edfe806c799ffab9eb5f0246d974
size 1366027648

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3e07acf19823dfb34641cfef7742b9dc1a8cd5116e7b41f2ce44df7eb994fcbe
size 1010443648

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e93b62725ffc969d4fddd76a806141b603739b0638da2696cc125d0a004bf0b5
size 1073931648

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:627fee9a3920fb1d70b2c8d2f0db97c08dcfac4bc28675e1b2ab101b4222defb
size 1345982848

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3c099b0f0809a2c53e049b055441dd2cbd8a3b0dd0000a280548b9d6d830b5cf
size 1249153408

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2d7b3ef08ce5c37a43b51bfdff806de2068c06c8ded5f5e2363ebe596e6df411
size 1140696448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8d192f80a84244b20842860cfcfd44ed920181b412ad8b6c5a067fbfcd8d9285
size 1409470848

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4e9d08d5cbcab88baa71722ccbabac6ac9eabf3b1056164d5c0e900d36cb4d22
size 1428757888

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6ff27d464aa0bb0aecbd34414894bf1f241725ea30c738cd89186a5d7aad23af
size 1559256448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3c42333f3d17c9922838b1cfb0c7b6c8a1432fe56c1ccc7d440ba43fe2782867
size 1560951168

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d9fa63cb6fb547d0762e1f6b52393e696b525edd881daaa8e63de64e8e7e00bf
size 1497463168

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c8cf2d9bab2263d523047afd587ad44a99e0356ea76fbcbbff6ed3e91e0a455b
size 1433017728

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:249e81bd3b062da864319f0a9b31edf42136335f9dad977fcdf1cbe84d432d19
size 1793849728

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:544538b287d1c044f7a71160572d20d51b1c89ea9dee59876ea629776b62c7cf
size 1730361728

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:525767f65ce66068f184ab979fd057caa07f6300f610911efb83589c24ef9fd4
size 1693195648

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0c6477806370470173cb7ebce81fb8b2f8943fcdeb01666371cf48973436f2ce
size 1977816448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:68dc63c7f021375a466898e22b05149b159c39de7962f8441834ea43a28aba9d
size 2041304448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:33e7e5fa707bf4a40caec6e3d34718a54181d38b0e3479f32336fb34273a9027
size 2560318848

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5208bedf9991cd03afa72c791f669c35b7803b0220eff90840f1aa8b8f057f0c
size 4815166592

Binary file not shown.

180
README.md Normal file
View File

@@ -0,0 +1,180 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
license_link: LICENSE
language:
- en
- ko
license_name: exaone
tags:
- lg-ai
- exaone
- exaone-deep
base_model: LGAI-EXAONE/EXAONE-Deep-2.4B
base_model_relation: quantized
license: other
---
## Llamacpp imatrix Quantizations of EXAONE-Deep-2.4B by LGAI-EXAONE
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4915">b4915</a> for quantization.
Original model: https://huggingface.co/LGAI-EXAONE/EXAONE-Deep-2.4B
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
Run them directly with [llama.cpp](https://github.com/ggerganov/llama.cpp), or any other llama.cpp based project
## Prompt format
```
[|system|]{system_prompt}[|endofturn|]
[|user|]{prompt}
[|assistant|]<thought>
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [EXAONE-Deep-2.4B-bf16.gguf](https://huggingface.co/bartowski/LGAI-EXAONE_EXAONE-Deep-2.4B-GGUF/blob/main/LGAI-EXAONE_EXAONE-Deep-2.4B-bf16.gguf) | bf16 | 4.82GB | false | Full BF16 weights. |
| [EXAONE-Deep-2.4B-Q8_0.gguf](https://huggingface.co/bartowski/LGAI-EXAONE_EXAONE-Deep-2.4B-GGUF/blob/main/LGAI-EXAONE_EXAONE-Deep-2.4B-Q8_0.gguf) | Q8_0 | 2.56GB | false | Extremely high quality, generally unneeded but max available quant. |
| [EXAONE-Deep-2.4B-Q6_K_L.gguf](https://huggingface.co/bartowski/LGAI-EXAONE_EXAONE-Deep-2.4B-GGUF/blob/main/LGAI-EXAONE_EXAONE-Deep-2.4B-Q6_K_L.gguf) | Q6_K_L | 2.04GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [EXAONE-Deep-2.4B-Q6_K.gguf](https://huggingface.co/bartowski/LGAI-EXAONE_EXAONE-Deep-2.4B-GGUF/blob/main/LGAI-EXAONE_EXAONE-Deep-2.4B-Q6_K.gguf) | Q6_K | 1.98GB | false | Very high quality, near perfect, *recommended*. |
| [EXAONE-Deep-2.4B-Q5_K_L.gguf](https://huggingface.co/bartowski/LGAI-EXAONE_EXAONE-Deep-2.4B-GGUF/blob/main/LGAI-EXAONE_EXAONE-Deep-2.4B-Q5_K_L.gguf) | Q5_K_L | 1.79GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [EXAONE-Deep-2.4B-Q5_K_M.gguf](https://huggingface.co/bartowski/LGAI-EXAONE_EXAONE-Deep-2.4B-GGUF/blob/main/LGAI-EXAONE_EXAONE-Deep-2.4B-Q5_K_M.gguf) | Q5_K_M | 1.73GB | false | High quality, *recommended*. |
| [EXAONE-Deep-2.4B-Q5_K_S.gguf](https://huggingface.co/bartowski/LGAI-EXAONE_EXAONE-Deep-2.4B-GGUF/blob/main/LGAI-EXAONE_EXAONE-Deep-2.4B-Q5_K_S.gguf) | Q5_K_S | 1.69GB | false | High quality, *recommended*. |
| [EXAONE-Deep-2.4B-Q4_K_L.gguf](https://huggingface.co/bartowski/LGAI-EXAONE_EXAONE-Deep-2.4B-GGUF/blob/main/LGAI-EXAONE_EXAONE-Deep-2.4B-Q4_K_L.gguf) | Q4_K_L | 1.56GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [EXAONE-Deep-2.4B-Q4_1.gguf](https://huggingface.co/bartowski/LGAI-EXAONE_EXAONE-Deep-2.4B-GGUF/blob/main/LGAI-EXAONE_EXAONE-Deep-2.4B-Q4_1.gguf) | Q4_1 | 1.56GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| [EXAONE-Deep-2.4B-Q4_K_M.gguf](https://huggingface.co/bartowski/LGAI-EXAONE_EXAONE-Deep-2.4B-GGUF/blob/main/LGAI-EXAONE_EXAONE-Deep-2.4B-Q4_K_M.gguf) | Q4_K_M | 1.50GB | false | Good quality, default size for most use cases, *recommended*. |
| [EXAONE-Deep-2.4B-Q4_K_S.gguf](https://huggingface.co/bartowski/LGAI-EXAONE_EXAONE-Deep-2.4B-GGUF/blob/main/LGAI-EXAONE_EXAONE-Deep-2.4B-Q4_K_S.gguf) | Q4_K_S | 1.43GB | false | Slightly lower quality with more space savings, *recommended*. |
| [EXAONE-Deep-2.4B-Q4_0.gguf](https://huggingface.co/bartowski/LGAI-EXAONE_EXAONE-Deep-2.4B-GGUF/blob/main/LGAI-EXAONE_EXAONE-Deep-2.4B-Q4_0.gguf) | Q4_0 | 1.43GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| [EXAONE-Deep-2.4B-IQ4_NL.gguf](https://huggingface.co/bartowski/LGAI-EXAONE_EXAONE-Deep-2.4B-GGUF/blob/main/LGAI-EXAONE_EXAONE-Deep-2.4B-IQ4_NL.gguf) | IQ4_NL | 1.43GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [EXAONE-Deep-2.4B-Q3_K_XL.gguf](https://huggingface.co/bartowski/LGAI-EXAONE_EXAONE-Deep-2.4B-GGUF/blob/main/LGAI-EXAONE_EXAONE-Deep-2.4B-Q3_K_XL.gguf) | Q3_K_XL | 1.41GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [EXAONE-Deep-2.4B-IQ4_XS.gguf](https://huggingface.co/bartowski/LGAI-EXAONE_EXAONE-Deep-2.4B-GGUF/blob/main/LGAI-EXAONE_EXAONE-Deep-2.4B-IQ4_XS.gguf) | IQ4_XS | 1.37GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [EXAONE-Deep-2.4B-Q3_K_L.gguf](https://huggingface.co/bartowski/LGAI-EXAONE_EXAONE-Deep-2.4B-GGUF/blob/main/LGAI-EXAONE_EXAONE-Deep-2.4B-Q3_K_L.gguf) | Q3_K_L | 1.35GB | false | Lower quality but usable, good for low RAM availability. |
| [EXAONE-Deep-2.4B-Q3_K_M.gguf](https://huggingface.co/bartowski/LGAI-EXAONE_EXAONE-Deep-2.4B-GGUF/blob/main/LGAI-EXAONE_EXAONE-Deep-2.4B-Q3_K_M.gguf) | Q3_K_M | 1.25GB | false | Low quality. |
| [EXAONE-Deep-2.4B-IQ3_M.gguf](https://huggingface.co/bartowski/LGAI-EXAONE_EXAONE-Deep-2.4B-GGUF/blob/main/LGAI-EXAONE_EXAONE-Deep-2.4B-IQ3_M.gguf) | IQ3_M | 1.18GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [EXAONE-Deep-2.4B-Q3_K_S.gguf](https://huggingface.co/bartowski/LGAI-EXAONE_EXAONE-Deep-2.4B-GGUF/blob/main/LGAI-EXAONE_EXAONE-Deep-2.4B-Q3_K_S.gguf) | Q3_K_S | 1.14GB | false | Low quality, not recommended. |
| [EXAONE-Deep-2.4B-IQ3_XS.gguf](https://huggingface.co/bartowski/LGAI-EXAONE_EXAONE-Deep-2.4B-GGUF/blob/main/LGAI-EXAONE_EXAONE-Deep-2.4B-IQ3_XS.gguf) | IQ3_XS | 1.10GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [EXAONE-Deep-2.4B-Q2_K_L.gguf](https://huggingface.co/bartowski/LGAI-EXAONE_EXAONE-Deep-2.4B-GGUF/blob/main/LGAI-EXAONE_EXAONE-Deep-2.4B-Q2_K_L.gguf) | Q2_K_L | 1.07GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [EXAONE-Deep-2.4B-IQ3_XXS.gguf](https://huggingface.co/bartowski/LGAI-EXAONE_EXAONE-Deep-2.4B-GGUF/blob/main/LGAI-EXAONE_EXAONE-Deep-2.4B-IQ3_XXS.gguf) | IQ3_XXS | 1.01GB | false | Lower quality, new method with decent performance, comparable to Q3 quants. |
| [EXAONE-Deep-2.4B-Q2_K.gguf](https://huggingface.co/bartowski/LGAI-EXAONE_EXAONE-Deep-2.4B-GGUF/blob/main/LGAI-EXAONE_EXAONE-Deep-2.4B-Q2_K.gguf) | Q2_K | 1.01GB | false | Very low quality but surprisingly usable. |
| [EXAONE-Deep-2.4B-IQ2_M.gguf](https://huggingface.co/bartowski/LGAI-EXAONE_EXAONE-Deep-2.4B-GGUF/blob/main/LGAI-EXAONE_EXAONE-Deep-2.4B-IQ2_M.gguf) | IQ2_M | 0.91GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/LGAI-EXAONE_EXAONE-Deep-2.4B-GGUF --include "LGAI-EXAONE_EXAONE-Deep-2.4B-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/LGAI-EXAONE_EXAONE-Deep-2.4B-GGUF --include "LGAI-EXAONE_EXAONE-Deep-2.4B-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (LGAI-EXAONE_EXAONE-Deep-2.4B-Q8_0) or download them all in place (./)
</details>
## ARM/AVX information
Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.
Now, however, there is something called "online repacking" for weights. details in [this PR](https://github.com/ggerganov/llama.cpp/pull/9921). If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.
As of llama.cpp build [b4282](https://github.com/ggerganov/llama.cpp/releases/tag/b4282) you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.
Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to [this PR](https://github.com/ggerganov/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information (deprecated</summary>
I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Thank you to LM Studio for sponsoring my work.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}