初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Captain-Eris-Diogenes_Twilight-V0.420-12B-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-21 18:31:07 +08:00
commit 9ef334c841
28 changed files with 302 additions and 0 deletions

60
.gitattributes vendored Normal file
View File

@@ -0,0 +1,60 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Captain-Eris-Diogenes_Twilight-V0.420-12B-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Captain-Eris-Diogenes_Twilight-V0.420-12B-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Captain-Eris-Diogenes_Twilight-V0.420-12B-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Captain-Eris-Diogenes_Twilight-V0.420-12B-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Captain-Eris-Diogenes_Twilight-V0.420-12B-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Captain-Eris-Diogenes_Twilight-V0.420-12B-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Captain-Eris-Diogenes_Twilight-V0.420-12B-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Captain-Eris-Diogenes_Twilight-V0.420-12B-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Captain-Eris-Diogenes_Twilight-V0.420-12B-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Captain-Eris-Diogenes_Twilight-V0.420-12B-Q4_1.gguf filter=lfs diff=lfs merge=lfs -text
Captain-Eris-Diogenes_Twilight-V0.420-12B-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Captain-Eris-Diogenes_Twilight-V0.420-12B-IQ4_NL.gguf filter=lfs diff=lfs merge=lfs -text
Captain-Eris-Diogenes_Twilight-V0.420-12B-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Captain-Eris-Diogenes_Twilight-V0.420-12B-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Captain-Eris-Diogenes_Twilight-V0.420-12B-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Captain-Eris-Diogenes_Twilight-V0.420-12B-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Captain-Eris-Diogenes_Twilight-V0.420-12B-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Captain-Eris-Diogenes_Twilight-V0.420-12B-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Captain-Eris-Diogenes_Twilight-V0.420-12B-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Captain-Eris-Diogenes_Twilight-V0.420-12B-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Captain-Eris-Diogenes_Twilight-V0.420-12B-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Captain-Eris-Diogenes_Twilight-V0.420-12B-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Captain-Eris-Diogenes_Twilight-V0.420-12B-IQ2_S.gguf filter=lfs diff=lfs merge=lfs -text
Captain-Eris-Diogenes_Twilight-V0.420-12B-f16.gguf filter=lfs diff=lfs merge=lfs -text
Captain-Eris-Diogenes_Twilight-V0.420-12B.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0322e1a4315a99ed18ca1afe8bcb3b1856dac2a7eb688609fbac494fff5bec2b
size 4435027072

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c3ad5c025d656b190c9f326e6d64949fb71dd4bc5dba6bba37a4471e57d55c56
size 4138476672

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:992492e73327b83bb574fe8ca354833359904b02f14947eaa8d02126f2c1597b
size 5722236032

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5ba77019f9e7ec55b0bcba38a4419a214cca2bab8e2124e1ee256ce116fc3fee
size 5306492032

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:194c1aadd3d5f643157177334710e487a17ab499f9e80ab101415f8446489b18
size 7097918592

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f3db5ced5ec01d846c4a0bc5fc48235786f4688cc5bdc444679732fc7347c5f5
size 6742713472

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0cb351f1887f2b80da1331d186cec94007871e92b3acdf467f82e5bf74947ef9
size 4791051392

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5392591a92991c6d97b4d24cbb46c707d204a57007695e1cbb106463a02ab249
size 5446411392

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fe1889331f98987c03b973b530fb0048d8ed252fc86d09d8fd48db0e16e4da22
size 6561506432

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:237c253d3c860807968fa3ca70c9af0cf5c8631dfd1f4bc39ee71ede32bea07c
size 6083093632

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:126d7c7ca65ee7183206777dec8c3e9fbb83a1ea9e362f0b6ade68d25680621e
size 5534229632

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a85fb102fa54b5ec418623ba9c9ad0ef0431218153d6f8bab1a07fb17f150021
size 7148708992

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cb257f23e8a1601e0ed034659951f00bb5d4c88f4fc99fde9a2f0c4aaffbc03c
size 7094641792

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0bc211bbef0c8f618a73d4930a58b8e2f3a59d9ca24a3d8759a632c4323ea4f9
size 7795221632

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:181144e06957d32dc308c95710d0b21af8b1fb01a7af64ebb4bace333fce843e
size 7975281792

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4390b2ac644d69546f48fb817ecce293543afafcb754abd3e6d17ed9865dc40e
size 7477208192

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f5bab813830413cf9a4acc1b3ce7dcede5de4f1781917d2ad1872d6f0fc109ae
size 7120200832

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8ad951dc7e984fdb0a023672558c24f494f8b294ddc545bfa09180c300021a6d
size 9141822592

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:07aa2f8811992db4b340e50dd64109a0fc4dc4e3db8695e1996b9efc744d3385
size 8727635072

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:09ad40552342b6f89c57c3ab2996f9ee7d7a9683d645b27720a86e32d8498220
size 8518739072

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9d7b9b0301fe252e0aacb559df14ca87233b48f6456bede1c204d72f4533a9f3
size 10056213632

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d1b24584e2bb40b86c6155af8b7f33f1483b2cda84c7770d7f8bcf6983dac20c
size 10381272192

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f882d2405d53638ea796ee6e76ce530bc29d292a4cff5840711807863dfab72c
size 13022372992

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ac5ee50436dd46b12de74f60647f73b701dc37a9fcbbff3c5fc2580764cc9345
size 24504279872

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6fccb9eb99ff262906480465bdc01e4f628428c0974097d70ff886f4e56a7502
size 7054418

166
README.md Normal file
View File

@@ -0,0 +1,166 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
base_model: Nitral-AI/Captain-Eris-Diogenes_Twilight-V0.420-12B
tags:
- mergekit
- merge
---
## Llamacpp imatrix Quantizations of Captain-Eris-Diogenes_Twilight-V0.420-12B
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4404">b4404</a> for quantization.
Original model: https://huggingface.co/Nitral-AI/Captain-Eris-Diogenes_Twilight-V0.420-12B
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<s>[INST]{prompt}[/INST]
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Captain-Eris-Diogenes_Twilight-V0.420-12B-f16.gguf](https://huggingface.co/bartowski/Captain-Eris-Diogenes_Twilight-V0.420-12B-GGUF/blob/main/Captain-Eris-Diogenes_Twilight-V0.420-12B-f16.gguf) | f16 | 24.50GB | false | Full F16 weights. |
| [Captain-Eris-Diogenes_Twilight-V0.420-12B-Q8_0.gguf](https://huggingface.co/bartowski/Captain-Eris-Diogenes_Twilight-V0.420-12B-GGUF/blob/main/Captain-Eris-Diogenes_Twilight-V0.420-12B-Q8_0.gguf) | Q8_0 | 13.02GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Captain-Eris-Diogenes_Twilight-V0.420-12B-Q6_K_L.gguf](https://huggingface.co/bartowski/Captain-Eris-Diogenes_Twilight-V0.420-12B-GGUF/blob/main/Captain-Eris-Diogenes_Twilight-V0.420-12B-Q6_K_L.gguf) | Q6_K_L | 10.38GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Captain-Eris-Diogenes_Twilight-V0.420-12B-Q6_K.gguf](https://huggingface.co/bartowski/Captain-Eris-Diogenes_Twilight-V0.420-12B-GGUF/blob/main/Captain-Eris-Diogenes_Twilight-V0.420-12B-Q6_K.gguf) | Q6_K | 10.06GB | false | Very high quality, near perfect, *recommended*. |
| [Captain-Eris-Diogenes_Twilight-V0.420-12B-Q5_K_L.gguf](https://huggingface.co/bartowski/Captain-Eris-Diogenes_Twilight-V0.420-12B-GGUF/blob/main/Captain-Eris-Diogenes_Twilight-V0.420-12B-Q5_K_L.gguf) | Q5_K_L | 9.14GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Captain-Eris-Diogenes_Twilight-V0.420-12B-Q5_K_M.gguf](https://huggingface.co/bartowski/Captain-Eris-Diogenes_Twilight-V0.420-12B-GGUF/blob/main/Captain-Eris-Diogenes_Twilight-V0.420-12B-Q5_K_M.gguf) | Q5_K_M | 8.73GB | false | High quality, *recommended*. |
| [Captain-Eris-Diogenes_Twilight-V0.420-12B-Q5_K_S.gguf](https://huggingface.co/bartowski/Captain-Eris-Diogenes_Twilight-V0.420-12B-GGUF/blob/main/Captain-Eris-Diogenes_Twilight-V0.420-12B-Q5_K_S.gguf) | Q5_K_S | 8.52GB | false | High quality, *recommended*. |
| [Captain-Eris-Diogenes_Twilight-V0.420-12B-Q4_K_L.gguf](https://huggingface.co/bartowski/Captain-Eris-Diogenes_Twilight-V0.420-12B-GGUF/blob/main/Captain-Eris-Diogenes_Twilight-V0.420-12B-Q4_K_L.gguf) | Q4_K_L | 7.98GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Captain-Eris-Diogenes_Twilight-V0.420-12B-Q4_1.gguf](https://huggingface.co/bartowski/Captain-Eris-Diogenes_Twilight-V0.420-12B-GGUF/blob/main/Captain-Eris-Diogenes_Twilight-V0.420-12B-Q4_1.gguf) | Q4_1 | 7.80GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| [Captain-Eris-Diogenes_Twilight-V0.420-12B-Q4_K_M.gguf](https://huggingface.co/bartowski/Captain-Eris-Diogenes_Twilight-V0.420-12B-GGUF/blob/main/Captain-Eris-Diogenes_Twilight-V0.420-12B-Q4_K_M.gguf) | Q4_K_M | 7.48GB | false | Good quality, default size for most use cases, *recommended*. |
| [Captain-Eris-Diogenes_Twilight-V0.420-12B-Q3_K_XL.gguf](https://huggingface.co/bartowski/Captain-Eris-Diogenes_Twilight-V0.420-12B-GGUF/blob/main/Captain-Eris-Diogenes_Twilight-V0.420-12B-Q3_K_XL.gguf) | Q3_K_XL | 7.15GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Captain-Eris-Diogenes_Twilight-V0.420-12B-Q4_K_S.gguf](https://huggingface.co/bartowski/Captain-Eris-Diogenes_Twilight-V0.420-12B-GGUF/blob/main/Captain-Eris-Diogenes_Twilight-V0.420-12B-Q4_K_S.gguf) | Q4_K_S | 7.12GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Captain-Eris-Diogenes_Twilight-V0.420-12B-IQ4_NL.gguf](https://huggingface.co/bartowski/Captain-Eris-Diogenes_Twilight-V0.420-12B-GGUF/blob/main/Captain-Eris-Diogenes_Twilight-V0.420-12B-IQ4_NL.gguf) | IQ4_NL | 7.10GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [Captain-Eris-Diogenes_Twilight-V0.420-12B-Q4_0.gguf](https://huggingface.co/bartowski/Captain-Eris-Diogenes_Twilight-V0.420-12B-GGUF/blob/main/Captain-Eris-Diogenes_Twilight-V0.420-12B-Q4_0.gguf) | Q4_0 | 7.09GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| [Captain-Eris-Diogenes_Twilight-V0.420-12B-IQ4_XS.gguf](https://huggingface.co/bartowski/Captain-Eris-Diogenes_Twilight-V0.420-12B-GGUF/blob/main/Captain-Eris-Diogenes_Twilight-V0.420-12B-IQ4_XS.gguf) | IQ4_XS | 6.74GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Captain-Eris-Diogenes_Twilight-V0.420-12B-Q3_K_L.gguf](https://huggingface.co/bartowski/Captain-Eris-Diogenes_Twilight-V0.420-12B-GGUF/blob/main/Captain-Eris-Diogenes_Twilight-V0.420-12B-Q3_K_L.gguf) | Q3_K_L | 6.56GB | false | Lower quality but usable, good for low RAM availability. |
| [Captain-Eris-Diogenes_Twilight-V0.420-12B-Q3_K_M.gguf](https://huggingface.co/bartowski/Captain-Eris-Diogenes_Twilight-V0.420-12B-GGUF/blob/main/Captain-Eris-Diogenes_Twilight-V0.420-12B-Q3_K_M.gguf) | Q3_K_M | 6.08GB | false | Low quality. |
| [Captain-Eris-Diogenes_Twilight-V0.420-12B-IQ3_M.gguf](https://huggingface.co/bartowski/Captain-Eris-Diogenes_Twilight-V0.420-12B-GGUF/blob/main/Captain-Eris-Diogenes_Twilight-V0.420-12B-IQ3_M.gguf) | IQ3_M | 5.72GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Captain-Eris-Diogenes_Twilight-V0.420-12B-Q3_K_S.gguf](https://huggingface.co/bartowski/Captain-Eris-Diogenes_Twilight-V0.420-12B-GGUF/blob/main/Captain-Eris-Diogenes_Twilight-V0.420-12B-Q3_K_S.gguf) | Q3_K_S | 5.53GB | false | Low quality, not recommended. |
| [Captain-Eris-Diogenes_Twilight-V0.420-12B-Q2_K_L.gguf](https://huggingface.co/bartowski/Captain-Eris-Diogenes_Twilight-V0.420-12B-GGUF/blob/main/Captain-Eris-Diogenes_Twilight-V0.420-12B-Q2_K_L.gguf) | Q2_K_L | 5.45GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Captain-Eris-Diogenes_Twilight-V0.420-12B-IQ3_XS.gguf](https://huggingface.co/bartowski/Captain-Eris-Diogenes_Twilight-V0.420-12B-GGUF/blob/main/Captain-Eris-Diogenes_Twilight-V0.420-12B-IQ3_XS.gguf) | IQ3_XS | 5.31GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Captain-Eris-Diogenes_Twilight-V0.420-12B-Q2_K.gguf](https://huggingface.co/bartowski/Captain-Eris-Diogenes_Twilight-V0.420-12B-GGUF/blob/main/Captain-Eris-Diogenes_Twilight-V0.420-12B-Q2_K.gguf) | Q2_K | 4.79GB | false | Very low quality but surprisingly usable. |
| [Captain-Eris-Diogenes_Twilight-V0.420-12B-IQ2_M.gguf](https://huggingface.co/bartowski/Captain-Eris-Diogenes_Twilight-V0.420-12B-GGUF/blob/main/Captain-Eris-Diogenes_Twilight-V0.420-12B-IQ2_M.gguf) | IQ2_M | 4.44GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
| [Captain-Eris-Diogenes_Twilight-V0.420-12B-IQ2_S.gguf](https://huggingface.co/bartowski/Captain-Eris-Diogenes_Twilight-V0.420-12B-GGUF/blob/main/Captain-Eris-Diogenes_Twilight-V0.420-12B-IQ2_S.gguf) | IQ2_S | 4.14GB | false | Low quality, uses SOTA techniques to be usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Captain-Eris-Diogenes_Twilight-V0.420-12B-GGUF --include "Captain-Eris-Diogenes_Twilight-V0.420-12B-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Captain-Eris-Diogenes_Twilight-V0.420-12B-GGUF --include "Captain-Eris-Diogenes_Twilight-V0.420-12B-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Captain-Eris-Diogenes_Twilight-V0.420-12B-Q8_0) or download them all in place (./)
</details>
## ARM/AVX information
Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.
Now, however, there is something called "online repacking" for weights. details in [this PR](https://github.com/ggerganov/llama.cpp/pull/9921). If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.
As of llama.cpp build [b4282](https://github.com/ggerganov/llama.cpp/releases/tag/b4282) you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.
Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to [this PR](https://github.com/ggerganov/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information (deprecated</summary>
I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}