初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Aura-4B-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-08 12:08:06 +08:00
commit 64f41c7850
27 changed files with 306 additions and 0 deletions

59
.gitattributes vendored Normal file
View File

@@ -0,0 +1,59 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Aura-4B-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Aura-4B-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Aura-4B-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Aura-4B-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Aura-4B-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Aura-4B-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Aura-4B-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Aura-4B-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Aura-4B-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Aura-4B-Q4_1.gguf filter=lfs diff=lfs merge=lfs -text
Aura-4B-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Aura-4B-IQ4_NL.gguf filter=lfs diff=lfs merge=lfs -text
Aura-4B-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Aura-4B-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Aura-4B-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Aura-4B-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Aura-4B-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Aura-4B-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Aura-4B-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Aura-4B-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Aura-4B-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Aura-4B-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Aura-4B-f16.gguf filter=lfs diff=lfs merge=lfs -text
Aura-4B.imatrix filter=lfs diff=lfs merge=lfs -text

3
Aura-4B-IQ2_M.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e74c5df245ccf4997198a8ead8caa06780a17309e058ddabc4fda3de5af365ce
size 1722634528

3
Aura-4B-IQ3_M.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:482b798a26f06f1fa6e492f0864bf52ecc5520ecc5b878ea22084a66efd47ab3
size 2183416096

3
Aura-4B-IQ3_XS.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f81677d714cd4b04482e07ea7536785ba4f1cf35c6bd00a27694ade9c505e6c3
size 2027604256

3
Aura-4B-IQ4_NL.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:42387744c58bbb3996dbf7dda46b9d9c457c9a68ce5e14301762e4e0f9d408ba
size 2661105952

3
Aura-4B-IQ4_XS.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c6ef0241ffaa6a0543500d8394a749f023b6f4f81eb2dc05cc3c276f1ad1acb5
size 2535547168

3
Aura-4B-Q2_K.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:138c7051a6e2bebcd84a3cd4eaa00c00b1a605f81b12f2f1e8d6db2a296e40ed
size 1839739168

3
Aura-4B-Q2_K_L.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f92fe73df310e1e91a7c34aaf256a3e45ecbd534fdd5faa5c1553e78c5cd1dad
size 2224507168

3
Aura-4B-Q3_K_L.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:64305acb11267b146a2c186454f3b4f10e3794d7c69ceb67e702c6c19bc7460e
size 2464860448

3
Aura-4B-Q3_K_M.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ef955b54327b338b72718433da10b7da5d95bd7b73f13cda03a7d9652187b0dd
size 2296564000

3
Aura-4B-Q3_K_S.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:47d6f755ba60b3785b68bea5c600cc2839352dd9b882d7965a67324d4b060ea9
size 2101528864

3
Aura-4B-Q3_K_XL.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:97c0a9b1fcf2935218fb85930ca2064dd5b96f25280d785f357917e522f7e08e
size 2809612576

3
Aura-4B-Q4_0.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ca619b01018b9701792be752fd5509ea5c96cbd609c40abfbe02268c56978c4d
size 2655600928

3
Aura-4B-Q4_1.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4027d0ba45845332b7314cbcaac128fc33b2ebf13e3801e4ed0ceed71def927f
size 2905932064

3
Aura-4B-Q4_K_L.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6c3c769747f74e6a5dca7862c5679b2c0d7128b833adda24e08061a338b92dcf
size 3070708000

3
Aura-4B-Q4_K_M.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e67175074ebbb2e25f7ee9c2c0d89e49edb564f9a56b194c7a02945e9d04823d
size 2778284320

3
Aura-4B-Q4_K_S.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b157007a31032d5c7e528fa7c212f299a381dd275dabd59c174540d2207ee270
size 2664251680

3
Aura-4B-Q5_K_L.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:918d8e82afe08c04b0cc60486feae9d62052435227fab5a2b306d50df8af5fbc
size 3473361184

3
Aura-4B-Q5_K_M.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6111a9ca0626bae8a28476982b050d03ae26a177f503f5a13cbcab2dff7ada33
size 3230187808

3
Aura-4B-Q5_K_S.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c01cf8a52944e274a02b32b60fd281f8f79a53f972315860dfd9cbf7da7be991
size 3163341088

3
Aura-4B-Q6_K.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:29b79caea8c44dd80320f41a82c326cd49579bd4ae5362734d77b9e7568b0760
size 3710335264

3
Aura-4B-Q6_K_L.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a9059dbc6bba3750b8bddd6c85682ce85cfe4ff4f2270f6280989fe0596ee94a
size 3901180192

3
Aura-4B-Q8_0.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c7ce2943a3a82340492a7d4e1dadaa543b02d67d04ff8e5c9a996980b8d39385
size 4803217696

3
Aura-4B-f16.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0648b2e3fa30ea5966f6aa4a24ccde0fd7059bbcd14d9ece317a2ed159eebb22
size 9033730080

3
Aura-4B.imatrix Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:45d2031fc6b99880c00bc151661105f58cbff4f76562d50e215b90820c301dd7
size 3677450

174
README.md Normal file
View File

@@ -0,0 +1,174 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
language:
- en
datasets:
- Mielikki/Erebus-87k
- FourOhFour/Instruct_Phase
- FourOhFour/RP_Phase
- anthracite-core/full-opus-chosen-hermes-rejected-kto-v1
base_model: AuraIndustries/Aura-4B
license: apache-2.0
---
## Llamacpp imatrix Quantizations of Aura-4B
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4381">b4381</a> for quantization.
Original model: https://huggingface.co/AuraIndustries/Aura-4B
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Aura-4B-f16.gguf](https://huggingface.co/bartowski/Aura-4B-GGUF/blob/main/Aura-4B-f16.gguf) | f16 | 9.03GB | false | Full F16 weights. |
| [Aura-4B-Q8_0.gguf](https://huggingface.co/bartowski/Aura-4B-GGUF/blob/main/Aura-4B-Q8_0.gguf) | Q8_0 | 4.80GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Aura-4B-Q6_K_L.gguf](https://huggingface.co/bartowski/Aura-4B-GGUF/blob/main/Aura-4B-Q6_K_L.gguf) | Q6_K_L | 3.90GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Aura-4B-Q6_K.gguf](https://huggingface.co/bartowski/Aura-4B-GGUF/blob/main/Aura-4B-Q6_K.gguf) | Q6_K | 3.71GB | false | Very high quality, near perfect, *recommended*. |
| [Aura-4B-Q5_K_L.gguf](https://huggingface.co/bartowski/Aura-4B-GGUF/blob/main/Aura-4B-Q5_K_L.gguf) | Q5_K_L | 3.47GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Aura-4B-Q5_K_M.gguf](https://huggingface.co/bartowski/Aura-4B-GGUF/blob/main/Aura-4B-Q5_K_M.gguf) | Q5_K_M | 3.23GB | false | High quality, *recommended*. |
| [Aura-4B-Q5_K_S.gguf](https://huggingface.co/bartowski/Aura-4B-GGUF/blob/main/Aura-4B-Q5_K_S.gguf) | Q5_K_S | 3.16GB | false | High quality, *recommended*. |
| [Aura-4B-Q4_K_L.gguf](https://huggingface.co/bartowski/Aura-4B-GGUF/blob/main/Aura-4B-Q4_K_L.gguf) | Q4_K_L | 3.07GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Aura-4B-Q4_1.gguf](https://huggingface.co/bartowski/Aura-4B-GGUF/blob/main/Aura-4B-Q4_1.gguf) | Q4_1 | 2.91GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| [Aura-4B-Q3_K_XL.gguf](https://huggingface.co/bartowski/Aura-4B-GGUF/blob/main/Aura-4B-Q3_K_XL.gguf) | Q3_K_XL | 2.81GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Aura-4B-Q4_K_M.gguf](https://huggingface.co/bartowski/Aura-4B-GGUF/blob/main/Aura-4B-Q4_K_M.gguf) | Q4_K_M | 2.78GB | false | Good quality, default size for most use cases, *recommended*. |
| [Aura-4B-Q4_K_S.gguf](https://huggingface.co/bartowski/Aura-4B-GGUF/blob/main/Aura-4B-Q4_K_S.gguf) | Q4_K_S | 2.66GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Aura-4B-Q4_0.gguf](https://huggingface.co/bartowski/Aura-4B-GGUF/blob/main/Aura-4B-Q4_0.gguf) | Q4_0 | 2.66GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| [Aura-4B-IQ4_NL.gguf](https://huggingface.co/bartowski/Aura-4B-GGUF/blob/main/Aura-4B-IQ4_NL.gguf) | IQ4_NL | 2.66GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [Aura-4B-IQ4_XS.gguf](https://huggingface.co/bartowski/Aura-4B-GGUF/blob/main/Aura-4B-IQ4_XS.gguf) | IQ4_XS | 2.54GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Aura-4B-Q3_K_L.gguf](https://huggingface.co/bartowski/Aura-4B-GGUF/blob/main/Aura-4B-Q3_K_L.gguf) | Q3_K_L | 2.46GB | false | Lower quality but usable, good for low RAM availability. |
| [Aura-4B-Q3_K_M.gguf](https://huggingface.co/bartowski/Aura-4B-GGUF/blob/main/Aura-4B-Q3_K_M.gguf) | Q3_K_M | 2.30GB | false | Low quality. |
| [Aura-4B-Q2_K_L.gguf](https://huggingface.co/bartowski/Aura-4B-GGUF/blob/main/Aura-4B-Q2_K_L.gguf) | Q2_K_L | 2.22GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Aura-4B-IQ3_M.gguf](https://huggingface.co/bartowski/Aura-4B-GGUF/blob/main/Aura-4B-IQ3_M.gguf) | IQ3_M | 2.18GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Aura-4B-Q3_K_S.gguf](https://huggingface.co/bartowski/Aura-4B-GGUF/blob/main/Aura-4B-Q3_K_S.gguf) | Q3_K_S | 2.10GB | false | Low quality, not recommended. |
| [Aura-4B-IQ3_XS.gguf](https://huggingface.co/bartowski/Aura-4B-GGUF/blob/main/Aura-4B-IQ3_XS.gguf) | IQ3_XS | 2.03GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Aura-4B-Q2_K.gguf](https://huggingface.co/bartowski/Aura-4B-GGUF/blob/main/Aura-4B-Q2_K.gguf) | Q2_K | 1.84GB | false | Very low quality but surprisingly usable. |
| [Aura-4B-IQ2_M.gguf](https://huggingface.co/bartowski/Aura-4B-GGUF/blob/main/Aura-4B-IQ2_M.gguf) | IQ2_M | 1.72GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Aura-4B-GGUF --include "Aura-4B-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Aura-4B-GGUF --include "Aura-4B-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Aura-4B-Q8_0) or download them all in place (./)
</details>
## ARM/AVX information
Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.
Now, however, there is something called "online repacking" for weights. details in [this PR](https://github.com/ggerganov/llama.cpp/pull/9921). If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.
As of llama.cpp build [b4282](https://github.com/ggerganov/llama.cpp/releases/tag/b4282) you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.
Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to [this PR](https://github.com/ggerganov/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information (deprecated</summary>
I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}