初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Ichigo-llama3.1-s-instruct-v0.4-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-25 17:01:16 +08:00
commit 0a2d1a9002
28 changed files with 270 additions and 0 deletions

60
.gitattributes vendored Normal file
View File

@@ -0,0 +1,60 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Ichigo-llama3.1-s-instruct-v0.4-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Ichigo-llama3.1-s-instruct-v0.4-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Ichigo-llama3.1-s-instruct-v0.4-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Ichigo-llama3.1-s-instruct-v0.4-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Ichigo-llama3.1-s-instruct-v0.4-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Ichigo-llama3.1-s-instruct-v0.4-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Ichigo-llama3.1-s-instruct-v0.4-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Ichigo-llama3.1-s-instruct-v0.4-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Ichigo-llama3.1-s-instruct-v0.4-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Ichigo-llama3.1-s-instruct-v0.4-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
Ichigo-llama3.1-s-instruct-v0.4-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
Ichigo-llama3.1-s-instruct-v0.4-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
Ichigo-llama3.1-s-instruct-v0.4-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Ichigo-llama3.1-s-instruct-v0.4-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Ichigo-llama3.1-s-instruct-v0.4-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Ichigo-llama3.1-s-instruct-v0.4-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Ichigo-llama3.1-s-instruct-v0.4-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Ichigo-llama3.1-s-instruct-v0.4-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Ichigo-llama3.1-s-instruct-v0.4-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Ichigo-llama3.1-s-instruct-v0.4-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Ichigo-llama3.1-s-instruct-v0.4-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Ichigo-llama3.1-s-instruct-v0.4-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Ichigo-llama3.1-s-instruct-v0.4-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Ichigo-llama3.1-s-instruct-v0.4-f16.gguf filter=lfs diff=lfs merge=lfs -text
Ichigo-llama3.1-s-instruct-v0.4.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:64f3df4ea29efc7214e76ebdc254132b37d4f2ceaa23b76519c9043a9014b0af
size 2950656000

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:44118fab158ce39dfe3496a4c4542848b9f097c034c5dbb99ebc1da47189218f
size 3787478624

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6a07a3eb109ad4d81d321941afd367825ce5062c41166da1a2646e24f03f0845
size 3521402464

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f9ad192e2e5ea594d9387b5ec6e29b1fbd76f7fd97e54904c207e5ecb79a8ed7
size 4450532160

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:afc5615a3465092f33cdd374cea4ae73c99c0b22698dbd1b8d2a89815f7bf2dc
size 3181572480

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fd3a168c8ee4277fbc854934501f5584b0abb622d06a72d7949a9c54ed4e8b59
size 3696656480

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:586f6c31e236a8058ffe29af64cc964c50fb82ac3bd7aa3b129c56a176734147
size 4324611680

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:62824e054402cad48c52ebd56ab68fb99a0e36f464e1daf4200b33217eeb0230
size 4021573216

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c4745d75f34086665d616eacd918df037ab427453170135118efa94012a3e2a9
size 3667154528

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0b6cc834dd9d68f7efee31b9557e337888d00dba40334a633886b8cf6d705aba
size 4786126944

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6ef959b0fcc38488a8a91c38ef9c4a724aab5c1070e2044e394100e88aac859b
size 4678827200

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1fdd9e662e6a5fb7ae294238235918da304903536c2fbf8f44f7d32016db53db
size 4664147136

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:670a5901bc641ad077ee253da7f122960e3a9b57e9f0aefa7c06e8b35c2622c2
size 4664147136

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9fdf57be0f89c56cb0d471a0eb4a5f87b6d4bf81b8312eb19b3ddff3b8161bf0
size 4664147136

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:76d177852756b520d1a37782b0ccfdfd361bfedf4db9a07591a51cd515a17de9
size 5315133536

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:076cf6e0473e2ba11f9ea1e1f81a65b1b68600b3546daf231e791bb4387aee78
size 4923669696

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9051bedf3a12d8e52b6cab8b883926ee5683cda0bb780f0adda32d679671a53b
size 4695604416

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e4ce1daec06de05def39c1c08cccb8bd7b5102ed7cde20134dbc59792bafaa69
size 6061719648

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:31ac87e55f14ffc6f070b2aea26dc3a5f3357d23bd2853f423d48dd58b1b1aa9
size 5736186560

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:20c682885b479e7c059877bfac9e91ce50e0a102062b3b8e607067b42675736e
size 5602493120

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:32c4339993eaec9e412ea296c3e174b392f9e665a28e6ef3b7dd109aa0007913
size 6599485728

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a1070c2e18a0168c55390b4a901c4bb51f7ae3ee235c50173e01e9c99a2bde6e
size 6854967392

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7c1c9966a71b7ee3c4de1b5b2db49d396987103751d9b4d4f0735d05eec477a8
size 8545271904

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:438ba116df9ed8edd9c7001cfec5cc040abdd9dcb1034477df5518013ac9ae53
size 16077347104

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7e2ec6a5ee6bb2a20c864c8b65a85b53d11b3b80b2430b1f5d1c97a93b391a7e
size 4988170

134
README.md Normal file
View File

@@ -0,0 +1,134 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
datasets:
- homebrewltd/instruction-speech-whispervq-v2
language:
- en
tags:
- sound language model
base_model: homebrewltd/Ichigo-llama3.1-s-instruct-v0.4
license: apache-2.0
---
## Llamacpp imatrix Quantizations of Ichigo-llama3.1-s-instruct-v0.4
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4058">b4058</a> for quantization.
Original model: https://huggingface.co/homebrewltd/Ichigo-llama3.1-s-instruct-v0.4
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<|begin_of_text|><|start_header_id|>system<|end_header_id|>
Cutting Knowledge Date: December 2023
Today Date: 26 Jul 2024
{system_prompt}<|eot_id|><|start_header_id|>user<|end_header_id|>
{prompt}<|eot_id|><|start_header_id|>assistant<|end_header_id|>
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Ichigo-llama3.1-s-instruct-v0.4-f16.gguf](https://huggingface.co/bartowski/Ichigo-llama3.1-s-instruct-v0.4-GGUF/blob/main/Ichigo-llama3.1-s-instruct-v0.4-f16.gguf) | f16 | 16.08GB | false | Full F16 weights. |
| [Ichigo-llama3.1-s-instruct-v0.4-Q8_0.gguf](https://huggingface.co/bartowski/Ichigo-llama3.1-s-instruct-v0.4-GGUF/blob/main/Ichigo-llama3.1-s-instruct-v0.4-Q8_0.gguf) | Q8_0 | 8.55GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Ichigo-llama3.1-s-instruct-v0.4-Q6_K_L.gguf](https://huggingface.co/bartowski/Ichigo-llama3.1-s-instruct-v0.4-GGUF/blob/main/Ichigo-llama3.1-s-instruct-v0.4-Q6_K_L.gguf) | Q6_K_L | 6.85GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Ichigo-llama3.1-s-instruct-v0.4-Q6_K.gguf](https://huggingface.co/bartowski/Ichigo-llama3.1-s-instruct-v0.4-GGUF/blob/main/Ichigo-llama3.1-s-instruct-v0.4-Q6_K.gguf) | Q6_K | 6.60GB | false | Very high quality, near perfect, *recommended*. |
| [Ichigo-llama3.1-s-instruct-v0.4-Q5_K_L.gguf](https://huggingface.co/bartowski/Ichigo-llama3.1-s-instruct-v0.4-GGUF/blob/main/Ichigo-llama3.1-s-instruct-v0.4-Q5_K_L.gguf) | Q5_K_L | 6.06GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Ichigo-llama3.1-s-instruct-v0.4-Q5_K_M.gguf](https://huggingface.co/bartowski/Ichigo-llama3.1-s-instruct-v0.4-GGUF/blob/main/Ichigo-llama3.1-s-instruct-v0.4-Q5_K_M.gguf) | Q5_K_M | 5.74GB | false | High quality, *recommended*. |
| [Ichigo-llama3.1-s-instruct-v0.4-Q5_K_S.gguf](https://huggingface.co/bartowski/Ichigo-llama3.1-s-instruct-v0.4-GGUF/blob/main/Ichigo-llama3.1-s-instruct-v0.4-Q5_K_S.gguf) | Q5_K_S | 5.60GB | false | High quality, *recommended*. |
| [Ichigo-llama3.1-s-instruct-v0.4-Q4_K_L.gguf](https://huggingface.co/bartowski/Ichigo-llama3.1-s-instruct-v0.4-GGUF/blob/main/Ichigo-llama3.1-s-instruct-v0.4-Q4_K_L.gguf) | Q4_K_L | 5.32GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Ichigo-llama3.1-s-instruct-v0.4-Q4_K_M.gguf](https://huggingface.co/bartowski/Ichigo-llama3.1-s-instruct-v0.4-GGUF/blob/main/Ichigo-llama3.1-s-instruct-v0.4-Q4_K_M.gguf) | Q4_K_M | 4.92GB | false | Good quality, default size for most use cases, *recommended*. |
| [Ichigo-llama3.1-s-instruct-v0.4-Q3_K_XL.gguf](https://huggingface.co/bartowski/Ichigo-llama3.1-s-instruct-v0.4-GGUF/blob/main/Ichigo-llama3.1-s-instruct-v0.4-Q3_K_XL.gguf) | Q3_K_XL | 4.79GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Ichigo-llama3.1-s-instruct-v0.4-Q4_K_S.gguf](https://huggingface.co/bartowski/Ichigo-llama3.1-s-instruct-v0.4-GGUF/blob/main/Ichigo-llama3.1-s-instruct-v0.4-Q4_K_S.gguf) | Q4_K_S | 4.70GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Ichigo-llama3.1-s-instruct-v0.4-Q4_0.gguf](https://huggingface.co/bartowski/Ichigo-llama3.1-s-instruct-v0.4-GGUF/blob/main/Ichigo-llama3.1-s-instruct-v0.4-Q4_0.gguf) | Q4_0 | 4.68GB | false | Legacy format, generally not worth using over similarly sized formats |
| [Ichigo-llama3.1-s-instruct-v0.4-Q4_0_8_8.gguf](https://huggingface.co/bartowski/Ichigo-llama3.1-s-instruct-v0.4-GGUF/blob/main/Ichigo-llama3.1-s-instruct-v0.4-Q4_0_8_8.gguf) | Q4_0_8_8 | 4.66GB | false | Optimized for ARM inference. Requires 'sve' support (see link below). *Don't use on Mac or Windows*. |
| [Ichigo-llama3.1-s-instruct-v0.4-Q4_0_4_8.gguf](https://huggingface.co/bartowski/Ichigo-llama3.1-s-instruct-v0.4-GGUF/blob/main/Ichigo-llama3.1-s-instruct-v0.4-Q4_0_4_8.gguf) | Q4_0_4_8 | 4.66GB | false | Optimized for ARM inference. Requires 'i8mm' support (see link below). *Don't use on Mac or Windows*. |
| [Ichigo-llama3.1-s-instruct-v0.4-Q4_0_4_4.gguf](https://huggingface.co/bartowski/Ichigo-llama3.1-s-instruct-v0.4-GGUF/blob/main/Ichigo-llama3.1-s-instruct-v0.4-Q4_0_4_4.gguf) | Q4_0_4_4 | 4.66GB | false | Optimized for ARM inference. Should work well on all ARM chips, pick this if you're unsure. *Don't use on Mac or Windows*. |
| [Ichigo-llama3.1-s-instruct-v0.4-IQ4_XS.gguf](https://huggingface.co/bartowski/Ichigo-llama3.1-s-instruct-v0.4-GGUF/blob/main/Ichigo-llama3.1-s-instruct-v0.4-IQ4_XS.gguf) | IQ4_XS | 4.45GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Ichigo-llama3.1-s-instruct-v0.4-Q3_K_L.gguf](https://huggingface.co/bartowski/Ichigo-llama3.1-s-instruct-v0.4-GGUF/blob/main/Ichigo-llama3.1-s-instruct-v0.4-Q3_K_L.gguf) | Q3_K_L | 4.32GB | false | Lower quality but usable, good for low RAM availability. |
| [Ichigo-llama3.1-s-instruct-v0.4-Q3_K_M.gguf](https://huggingface.co/bartowski/Ichigo-llama3.1-s-instruct-v0.4-GGUF/blob/main/Ichigo-llama3.1-s-instruct-v0.4-Q3_K_M.gguf) | Q3_K_M | 4.02GB | false | Low quality. |
| [Ichigo-llama3.1-s-instruct-v0.4-IQ3_M.gguf](https://huggingface.co/bartowski/Ichigo-llama3.1-s-instruct-v0.4-GGUF/blob/main/Ichigo-llama3.1-s-instruct-v0.4-IQ3_M.gguf) | IQ3_M | 3.79GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Ichigo-llama3.1-s-instruct-v0.4-Q2_K_L.gguf](https://huggingface.co/bartowski/Ichigo-llama3.1-s-instruct-v0.4-GGUF/blob/main/Ichigo-llama3.1-s-instruct-v0.4-Q2_K_L.gguf) | Q2_K_L | 3.70GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Ichigo-llama3.1-s-instruct-v0.4-Q3_K_S.gguf](https://huggingface.co/bartowski/Ichigo-llama3.1-s-instruct-v0.4-GGUF/blob/main/Ichigo-llama3.1-s-instruct-v0.4-Q3_K_S.gguf) | Q3_K_S | 3.67GB | false | Low quality, not recommended. |
| [Ichigo-llama3.1-s-instruct-v0.4-IQ3_XS.gguf](https://huggingface.co/bartowski/Ichigo-llama3.1-s-instruct-v0.4-GGUF/blob/main/Ichigo-llama3.1-s-instruct-v0.4-IQ3_XS.gguf) | IQ3_XS | 3.52GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Ichigo-llama3.1-s-instruct-v0.4-Q2_K.gguf](https://huggingface.co/bartowski/Ichigo-llama3.1-s-instruct-v0.4-GGUF/blob/main/Ichigo-llama3.1-s-instruct-v0.4-Q2_K.gguf) | Q2_K | 3.18GB | false | Very low quality but surprisingly usable. |
| [Ichigo-llama3.1-s-instruct-v0.4-IQ2_M.gguf](https://huggingface.co/bartowski/Ichigo-llama3.1-s-instruct-v0.4-GGUF/blob/main/Ichigo-llama3.1-s-instruct-v0.4-IQ2_M.gguf) | IQ2_M | 2.95GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Ichigo-llama3.1-s-instruct-v0.4-GGUF --include "Ichigo-llama3.1-s-instruct-v0.4-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Ichigo-llama3.1-s-instruct-v0.4-GGUF --include "Ichigo-llama3.1-s-instruct-v0.4-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Ichigo-llama3.1-s-instruct-v0.4-Q8_0) or download them all in place (./)
## Q4_0_X_X
These are *NOT* for Metal (Apple) offloading, only ARM chips.
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}