初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Aurore-Reveil_Koto-Small-7B-IT-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-21 19:55:06 +08:00
commit 5206281464
28 changed files with 287 additions and 0 deletions

47
.gitattributes vendored Normal file
View File

@@ -0,0 +1,47 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bin.* filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zstandard filter=lfs diff=lfs merge=lfs -text
*.tfevents* filter=lfs diff=lfs merge=lfs -text
*.db* filter=lfs diff=lfs merge=lfs -text
*.ark* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*data* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.meta filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.index filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.gguf* filter=lfs diff=lfs merge=lfs -text
*.ggml filter=lfs diff=lfs merge=lfs -text
*.llamafile* filter=lfs diff=lfs merge=lfs -text
*.pt2 filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a58edb1d38bf0fb22a07d7bace0ee08975564921b2317b9b434d245e1a4dfca0
size 2867920192

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f6455975e5ae3da2160aad918a7b1e6a51d526a8c310cdb4bb213209fd15cb32
size 3650063680

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d98b6e750d8cc031534e71b11a17a916c23cf3bfef6a8f38d3a9484b14fc4186
size 3396373824

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:322c6e264ebb755c19b988b0a40b38ae78439423445d4c102f4d1edbebedcca6
size 3152543040

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:92890d21792a098cc57717fd71e2cf27b8e215eb06e195349d0d6fdcd726e6f0
size 4474510656

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3ce23d1a9a3bf77b05250c76ab235533c5bd434f82797dceb1cc12fabf0648eb
size 4260453696

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3bbb393dff7a352dfe902200076050d3e4db21a0687193188a11e15d6c9ee67f
size 3076406592

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9da332a2e88eab5e6b9ec4664e7b1b7cbb8440e1d6f583422b2026207673743a
size 3683126592

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b349eb3d2be9427deb7c4041751d30e690000a5441de9f2631d1c950c535e04b
size 4138962240

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:796e550cafc3d8b454eb2400ea420841d4635c75b3f229d57475f4ebc724b14b
size 3854011712

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7095573c331c97f5ea13944c5b159449f09a48da82cc0dfce500f82ca8a79fe0
size 3525840192

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:305de438243018453a058b6788923275ecbcf418865e161eef6db11697d54c9f
size 4682583360

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:792ad810135795675e10265208386945ff5e4558da39dc480c18dafe5ab23aeb
size 4466908480

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3d57f85b9fe5ea2e09c04759178f3381b5ec818a5c2a9c047abbc77118f0afff
size 4893187392

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:581e9acf92e7e22f26720e3c36938ad77c971cdebbe8c0ab90ff9f584ed13e37
size 5145447744

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c5c38bfa5d8d5100e91a2e0050a0b2f3e082cd4bfd423cb527abc3b6f1ae180c
size 4684340544

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1b2005b3d55682b6f8c064af489019a1366456902c938fbc3a941e7a94011cc4
size 4480277824

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:931a09aa5ff4af927ffa4bfd2b569c1e54e8adc58d03e9413fd16391046bb5f5
size 5832002880

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f8c2ac4b0cfa2aed85a8d73454065e9ac25ae2a24cad038263221a982ab67af6
size 5448555840

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3a410e72c2e32c32f7b3ae804f87a42984e350e2c3921ef11e11488b12831aab
size 5330738496

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0da5296df352bc80efaee2b3f02a766032ed28288ccd5ba3df754884879388ca
size 6260534592

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ccc373fb8de4f505ecedd8d92ddafeb98777ea542fb32f0e6b5c8f582f59938e
size 6561467712

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a93554796485d52b4e998097a44b5cf760bea5ddc2f4b494132d2e3c10feb931
size 8106511680

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5f8350bf50b379015bbdcdb033f21e6a0971f737d5bf46dcb5b00331f6b591d6
size 15252229152

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:dd34d3fa952ce75f3a74a8bdf93662a108bb614f0ff5c7260dfca16618929f2b
size 5162880

164
README.md Normal file
View File

@@ -0,0 +1,164 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
base_model_relation: quantized
base_model: Aurore-Reveil/Koto-Small-7B-IT
---
## Llamacpp imatrix Quantizations of Koto-Small-7B-IT by Aurore-Reveil
Using <a href="https://github.com/ggml-org/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggml-org/llama.cpp/releases/tag/b6317">b6317</a> for quantization.
Original model: https://huggingface.co/Aurore-Reveil/Koto-Small-7B-IT
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8) combined with a subset of combined_all_small.parquet from Ed Addario [here](https://huggingface.co/datasets/eaddario/imatrix-calibration/blob/main/combined_all_small.parquet)
Run them in [LM Studio](https://lmstudio.ai/)
Run them directly with [llama.cpp](https://github.com/ggml-org/llama.cpp), or any other llama.cpp based project
## Prompt format
No prompt format found, check original model page
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Koto-Small-7B-IT-bf16.gguf](https://huggingface.co/bartowski/Aurore-Reveil_Koto-Small-7B-IT-GGUF/blob/main/Aurore-Reveil_Koto-Small-7B-IT-bf16.gguf) | bf16 | 15.25GB | false | Full BF16 weights. |
| [Koto-Small-7B-IT-Q8_0.gguf](https://huggingface.co/bartowski/Aurore-Reveil_Koto-Small-7B-IT-GGUF/blob/main/Aurore-Reveil_Koto-Small-7B-IT-Q8_0.gguf) | Q8_0 | 8.11GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Koto-Small-7B-IT-Q6_K_L.gguf](https://huggingface.co/bartowski/Aurore-Reveil_Koto-Small-7B-IT-GGUF/blob/main/Aurore-Reveil_Koto-Small-7B-IT-Q6_K_L.gguf) | Q6_K_L | 6.56GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Koto-Small-7B-IT-Q6_K.gguf](https://huggingface.co/bartowski/Aurore-Reveil_Koto-Small-7B-IT-GGUF/blob/main/Aurore-Reveil_Koto-Small-7B-IT-Q6_K.gguf) | Q6_K | 6.26GB | false | Very high quality, near perfect, *recommended*. |
| [Koto-Small-7B-IT-Q5_K_L.gguf](https://huggingface.co/bartowski/Aurore-Reveil_Koto-Small-7B-IT-GGUF/blob/main/Aurore-Reveil_Koto-Small-7B-IT-Q5_K_L.gguf) | Q5_K_L | 5.83GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Koto-Small-7B-IT-Q5_K_M.gguf](https://huggingface.co/bartowski/Aurore-Reveil_Koto-Small-7B-IT-GGUF/blob/main/Aurore-Reveil_Koto-Small-7B-IT-Q5_K_M.gguf) | Q5_K_M | 5.45GB | false | High quality, *recommended*. |
| [Koto-Small-7B-IT-Q5_K_S.gguf](https://huggingface.co/bartowski/Aurore-Reveil_Koto-Small-7B-IT-GGUF/blob/main/Aurore-Reveil_Koto-Small-7B-IT-Q5_K_S.gguf) | Q5_K_S | 5.33GB | false | High quality, *recommended*. |
| [Koto-Small-7B-IT-Q4_K_L.gguf](https://huggingface.co/bartowski/Aurore-Reveil_Koto-Small-7B-IT-GGUF/blob/main/Aurore-Reveil_Koto-Small-7B-IT-Q4_K_L.gguf) | Q4_K_L | 5.15GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Koto-Small-7B-IT-Q4_1.gguf](https://huggingface.co/bartowski/Aurore-Reveil_Koto-Small-7B-IT-GGUF/blob/main/Aurore-Reveil_Koto-Small-7B-IT-Q4_1.gguf) | Q4_1 | 4.89GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| [Koto-Small-7B-IT-Q4_K_M.gguf](https://huggingface.co/bartowski/Aurore-Reveil_Koto-Small-7B-IT-GGUF/blob/main/Aurore-Reveil_Koto-Small-7B-IT-Q4_K_M.gguf) | Q4_K_M | 4.68GB | false | Good quality, default size for most use cases, *recommended*. |
| [Koto-Small-7B-IT-Q3_K_XL.gguf](https://huggingface.co/bartowski/Aurore-Reveil_Koto-Small-7B-IT-GGUF/blob/main/Aurore-Reveil_Koto-Small-7B-IT-Q3_K_XL.gguf) | Q3_K_XL | 4.68GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Koto-Small-7B-IT-Q4_K_S.gguf](https://huggingface.co/bartowski/Aurore-Reveil_Koto-Small-7B-IT-GGUF/blob/main/Aurore-Reveil_Koto-Small-7B-IT-Q4_K_S.gguf) | Q4_K_S | 4.48GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Koto-Small-7B-IT-Q4_0.gguf](https://huggingface.co/bartowski/Aurore-Reveil_Koto-Small-7B-IT-GGUF/blob/main/Aurore-Reveil_Koto-Small-7B-IT-Q4_0.gguf) | Q4_0 | 4.47GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| [Koto-Small-7B-IT-IQ4_NL.gguf](https://huggingface.co/bartowski/Aurore-Reveil_Koto-Small-7B-IT-GGUF/blob/main/Aurore-Reveil_Koto-Small-7B-IT-IQ4_NL.gguf) | IQ4_NL | 4.47GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [Koto-Small-7B-IT-IQ4_XS.gguf](https://huggingface.co/bartowski/Aurore-Reveil_Koto-Small-7B-IT-GGUF/blob/main/Aurore-Reveil_Koto-Small-7B-IT-IQ4_XS.gguf) | IQ4_XS | 4.26GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Koto-Small-7B-IT-Q3_K_L.gguf](https://huggingface.co/bartowski/Aurore-Reveil_Koto-Small-7B-IT-GGUF/blob/main/Aurore-Reveil_Koto-Small-7B-IT-Q3_K_L.gguf) | Q3_K_L | 4.14GB | false | Lower quality but usable, good for low RAM availability. |
| [Koto-Small-7B-IT-Q3_K_M.gguf](https://huggingface.co/bartowski/Aurore-Reveil_Koto-Small-7B-IT-GGUF/blob/main/Aurore-Reveil_Koto-Small-7B-IT-Q3_K_M.gguf) | Q3_K_M | 3.85GB | false | Low quality. |
| [Koto-Small-7B-IT-Q2_K_L.gguf](https://huggingface.co/bartowski/Aurore-Reveil_Koto-Small-7B-IT-GGUF/blob/main/Aurore-Reveil_Koto-Small-7B-IT-Q2_K_L.gguf) | Q2_K_L | 3.68GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Koto-Small-7B-IT-IQ3_M.gguf](https://huggingface.co/bartowski/Aurore-Reveil_Koto-Small-7B-IT-GGUF/blob/main/Aurore-Reveil_Koto-Small-7B-IT-IQ3_M.gguf) | IQ3_M | 3.65GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Koto-Small-7B-IT-Q3_K_S.gguf](https://huggingface.co/bartowski/Aurore-Reveil_Koto-Small-7B-IT-GGUF/blob/main/Aurore-Reveil_Koto-Small-7B-IT-Q3_K_S.gguf) | Q3_K_S | 3.53GB | false | Low quality, not recommended. |
| [Koto-Small-7B-IT-IQ3_XS.gguf](https://huggingface.co/bartowski/Aurore-Reveil_Koto-Small-7B-IT-GGUF/blob/main/Aurore-Reveil_Koto-Small-7B-IT-IQ3_XS.gguf) | IQ3_XS | 3.40GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Koto-Small-7B-IT-IQ3_XXS.gguf](https://huggingface.co/bartowski/Aurore-Reveil_Koto-Small-7B-IT-GGUF/blob/main/Aurore-Reveil_Koto-Small-7B-IT-IQ3_XXS.gguf) | IQ3_XXS | 3.15GB | false | Lower quality, new method with decent performance, comparable to Q3 quants. |
| [Koto-Small-7B-IT-Q2_K.gguf](https://huggingface.co/bartowski/Aurore-Reveil_Koto-Small-7B-IT-GGUF/blob/main/Aurore-Reveil_Koto-Small-7B-IT-Q2_K.gguf) | Q2_K | 3.08GB | false | Very low quality but surprisingly usable. |
| [Koto-Small-7B-IT-IQ2_M.gguf](https://huggingface.co/bartowski/Aurore-Reveil_Koto-Small-7B-IT-GGUF/blob/main/Aurore-Reveil_Koto-Small-7B-IT-IQ2_M.gguf) | IQ2_M | 2.87GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Aurore-Reveil_Koto-Small-7B-IT-GGUF --include "Aurore-Reveil_Koto-Small-7B-IT-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Aurore-Reveil_Koto-Small-7B-IT-GGUF --include "Aurore-Reveil_Koto-Small-7B-IT-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Aurore-Reveil_Koto-Small-7B-IT-Q8_0) or download them all in place (./)
</details>
## ARM/AVX information
Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.
Now, however, there is something called "online repacking" for weights. details in [this PR](https://github.com/ggml-org/llama.cpp/pull/9921). If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.
As of llama.cpp build [b4282](https://github.com/ggml-org/llama.cpp/releases/tag/b4282) you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.
Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to [this PR](https://github.com/ggml-org/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information (deprecated</summary>
I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggml-org/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Thank you to LM Studio for sponsoring my work.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}