初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Dans-DangerousWinds-V1.1.0-12b-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-30 02:08:13 +08:00
commit 25ebdecd94
29 changed files with 297 additions and 0 deletions

49
.gitattributes vendored Normal file
View File

@@ -0,0 +1,49 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bin.* filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zstandard filter=lfs diff=lfs merge=lfs -text
*.tfevents* filter=lfs diff=lfs merge=lfs -text
*.db* filter=lfs diff=lfs merge=lfs -text
*.ark* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*data* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.meta filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.index filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.gguf* filter=lfs diff=lfs merge=lfs -text
*.ggml filter=lfs diff=lfs merge=lfs -text
*.llamafile* filter=lfs diff=lfs merge=lfs -text
*.pt2 filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Dans-DangerousWinds-V1.1.0-12b.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:885f85473bc28e0ed554cda004f4b2864d6dc99a958d94719b77a0b31c475696
size 4435023776

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a6d18e043b126d5a2cbd80494b5eb61a683d18fa75b2dc30fef89643f8791bad
size 4138473376

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a9cf73a49b5f0b2453b4f91397bb534ce0a416c8df01dd85f53793018358ce84
size 3915077536

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e5f2056abc9b294b78eae9b677ee6d47e71a4bd963460f78d83984a858ebb984
size 5722232736

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:92408fc3da053a9bfb514c9ba67e026aa5456880e4770d3c43acb6c727ad7314
size 5306488736

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6a15172780b8a857bec6ac4dcb1818bd24a3280ea7e3f0d1a860b038bee2e789
size 7097915296

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:dfdcfa332defadc4fcddd57da2f2a8d37ae9eb043bb77904ef33c5b0461ea3d9
size 6742710176

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b997cb8cd46ab225ab1394783b86e8a4f876147e7c0ae18d24d1ad25b447430a
size 4791048096

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3cbabf197786b97b8e917120ae1909aa684c885824f26dfa7673492ba17d957b
size 5446408096

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:569c545587a84b420148fc94129531b11e213018101357a938516c0ac8311ebb
size 6561503136

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0be3ed61f99a1b33e88cbeaf1a531ccae5371933d9cc5f995ce8cf9aeb40d937
size 6083090336

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ba3b7890c2fe92967e33c17a5a2d4320196d04a3ffafed96cb8485b6e3a7db55
size 5534226336

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9e92783ea4038f89865996d1602b83fc2362f1bde5fa71729c55058fa9807bf7
size 7148705696

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:603b3d44cc86ea6d06bf95351a767483ad622bbc80c6ba1f38817e85a1b8d259
size 7094638496

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b57612438f43aabe13dc34c8f1eb02ecb098cbe71a25ddc35040ff125eb7b32a
size 7795218336

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:71500db31c2fcb1463313725532c144ecba6cc57724bc7f3389025c299a31d09
size 7975278496

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ab4620eff87fccd5a7932fc35591e0a1395d89bb7bcff19691126dbad44eb529
size 7477204896

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8e236d6cb5bc851af8a23e071cc760b22ec5117dab11b466ae7e9ad106a4b5f8
size 7120197536

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c615b59c892c5a44aeb1a54a3b4ad426331944f205d8b8590731d5cc9ba2ae27
size 9141819296

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:bd6b2bc337ee51b8f56fc07de2e609daa8dafb80d0d789e837b5b02233765b63
size 8727631776

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f7c2821a33cb97eb8de42b3f4836fa7195ce1a794c2c8863a7662d6d0fd7c765
size 8518735776

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3cd7f29bce8f375c644aaa28d3d68efffacc289d29c5440539882d8984d40743
size 10056210336

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c3285ba55c6685d50439e45b353dfc6bfffc086665eec8a2c275589acbf6acd9
size 10381268896

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e4474df863692394704a958cfcc15e60d1a314ac84b2faf5d4d606a9e32df493
size 13022369696

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e54f5317d7bf5b165d277553de66e69680aa82005cbd6489f45fcc2ea1475547
size 24504276608

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f8a3d0a2fb891e8608f42f1e46e2c719f323460f17b2c91795c0172ff483f4fc
size 7054418

169
README.md Normal file
View File

@@ -0,0 +1,169 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
license: apache-2.0
base_model: PocketDoc/Dans-DangerousWinds-V1.1.0-12b
datasets:
- PocketDoc/Dans-Prosemaxx-Adventure
- PocketDoc/Dans-Failuremaxx-Adventure
- PocketDoc/Dans-Prosemaxx-Cowriter-2-S
language:
- en
---
## Llamacpp imatrix Quantizations of Dans-DangerousWinds-V1.1.0-12b
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4497">b4497</a> for quantization.
Original model: https://huggingface.co/PocketDoc/Dans-DangerousWinds-V1.1.0-12b
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
No prompt format found, check original model page
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Dans-DangerousWinds-V1.1.0-12b-f16.gguf](https://huggingface.co/bartowski/Dans-DangerousWinds-V1.1.0-12b-GGUF/blob/main/Dans-DangerousWinds-V1.1.0-12b-f16.gguf) | f16 | 24.50GB | false | Full F16 weights. |
| [Dans-DangerousWinds-V1.1.0-12b-Q8_0.gguf](https://huggingface.co/bartowski/Dans-DangerousWinds-V1.1.0-12b-GGUF/blob/main/Dans-DangerousWinds-V1.1.0-12b-Q8_0.gguf) | Q8_0 | 13.02GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Dans-DangerousWinds-V1.1.0-12b-Q6_K_L.gguf](https://huggingface.co/bartowski/Dans-DangerousWinds-V1.1.0-12b-GGUF/blob/main/Dans-DangerousWinds-V1.1.0-12b-Q6_K_L.gguf) | Q6_K_L | 10.38GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Dans-DangerousWinds-V1.1.0-12b-Q6_K.gguf](https://huggingface.co/bartowski/Dans-DangerousWinds-V1.1.0-12b-GGUF/blob/main/Dans-DangerousWinds-V1.1.0-12b-Q6_K.gguf) | Q6_K | 10.06GB | false | Very high quality, near perfect, *recommended*. |
| [Dans-DangerousWinds-V1.1.0-12b-Q5_K_L.gguf](https://huggingface.co/bartowski/Dans-DangerousWinds-V1.1.0-12b-GGUF/blob/main/Dans-DangerousWinds-V1.1.0-12b-Q5_K_L.gguf) | Q5_K_L | 9.14GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Dans-DangerousWinds-V1.1.0-12b-Q5_K_M.gguf](https://huggingface.co/bartowski/Dans-DangerousWinds-V1.1.0-12b-GGUF/blob/main/Dans-DangerousWinds-V1.1.0-12b-Q5_K_M.gguf) | Q5_K_M | 8.73GB | false | High quality, *recommended*. |
| [Dans-DangerousWinds-V1.1.0-12b-Q5_K_S.gguf](https://huggingface.co/bartowski/Dans-DangerousWinds-V1.1.0-12b-GGUF/blob/main/Dans-DangerousWinds-V1.1.0-12b-Q5_K_S.gguf) | Q5_K_S | 8.52GB | false | High quality, *recommended*. |
| [Dans-DangerousWinds-V1.1.0-12b-Q4_K_L.gguf](https://huggingface.co/bartowski/Dans-DangerousWinds-V1.1.0-12b-GGUF/blob/main/Dans-DangerousWinds-V1.1.0-12b-Q4_K_L.gguf) | Q4_K_L | 7.98GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Dans-DangerousWinds-V1.1.0-12b-Q4_1.gguf](https://huggingface.co/bartowski/Dans-DangerousWinds-V1.1.0-12b-GGUF/blob/main/Dans-DangerousWinds-V1.1.0-12b-Q4_1.gguf) | Q4_1 | 7.80GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| [Dans-DangerousWinds-V1.1.0-12b-Q4_K_M.gguf](https://huggingface.co/bartowski/Dans-DangerousWinds-V1.1.0-12b-GGUF/blob/main/Dans-DangerousWinds-V1.1.0-12b-Q4_K_M.gguf) | Q4_K_M | 7.48GB | false | Good quality, default size for most use cases, *recommended*. |
| [Dans-DangerousWinds-V1.1.0-12b-Q3_K_XL.gguf](https://huggingface.co/bartowski/Dans-DangerousWinds-V1.1.0-12b-GGUF/blob/main/Dans-DangerousWinds-V1.1.0-12b-Q3_K_XL.gguf) | Q3_K_XL | 7.15GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Dans-DangerousWinds-V1.1.0-12b-Q4_K_S.gguf](https://huggingface.co/bartowski/Dans-DangerousWinds-V1.1.0-12b-GGUF/blob/main/Dans-DangerousWinds-V1.1.0-12b-Q4_K_S.gguf) | Q4_K_S | 7.12GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Dans-DangerousWinds-V1.1.0-12b-IQ4_NL.gguf](https://huggingface.co/bartowski/Dans-DangerousWinds-V1.1.0-12b-GGUF/blob/main/Dans-DangerousWinds-V1.1.0-12b-IQ4_NL.gguf) | IQ4_NL | 7.10GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [Dans-DangerousWinds-V1.1.0-12b-Q4_0.gguf](https://huggingface.co/bartowski/Dans-DangerousWinds-V1.1.0-12b-GGUF/blob/main/Dans-DangerousWinds-V1.1.0-12b-Q4_0.gguf) | Q4_0 | 7.09GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| [Dans-DangerousWinds-V1.1.0-12b-IQ4_XS.gguf](https://huggingface.co/bartowski/Dans-DangerousWinds-V1.1.0-12b-GGUF/blob/main/Dans-DangerousWinds-V1.1.0-12b-IQ4_XS.gguf) | IQ4_XS | 6.74GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Dans-DangerousWinds-V1.1.0-12b-Q3_K_L.gguf](https://huggingface.co/bartowski/Dans-DangerousWinds-V1.1.0-12b-GGUF/blob/main/Dans-DangerousWinds-V1.1.0-12b-Q3_K_L.gguf) | Q3_K_L | 6.56GB | false | Lower quality but usable, good for low RAM availability. |
| [Dans-DangerousWinds-V1.1.0-12b-Q3_K_M.gguf](https://huggingface.co/bartowski/Dans-DangerousWinds-V1.1.0-12b-GGUF/blob/main/Dans-DangerousWinds-V1.1.0-12b-Q3_K_M.gguf) | Q3_K_M | 6.08GB | false | Low quality. |
| [Dans-DangerousWinds-V1.1.0-12b-IQ3_M.gguf](https://huggingface.co/bartowski/Dans-DangerousWinds-V1.1.0-12b-GGUF/blob/main/Dans-DangerousWinds-V1.1.0-12b-IQ3_M.gguf) | IQ3_M | 5.72GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Dans-DangerousWinds-V1.1.0-12b-Q3_K_S.gguf](https://huggingface.co/bartowski/Dans-DangerousWinds-V1.1.0-12b-GGUF/blob/main/Dans-DangerousWinds-V1.1.0-12b-Q3_K_S.gguf) | Q3_K_S | 5.53GB | false | Low quality, not recommended. |
| [Dans-DangerousWinds-V1.1.0-12b-Q2_K_L.gguf](https://huggingface.co/bartowski/Dans-DangerousWinds-V1.1.0-12b-GGUF/blob/main/Dans-DangerousWinds-V1.1.0-12b-Q2_K_L.gguf) | Q2_K_L | 5.45GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Dans-DangerousWinds-V1.1.0-12b-IQ3_XS.gguf](https://huggingface.co/bartowski/Dans-DangerousWinds-V1.1.0-12b-GGUF/blob/main/Dans-DangerousWinds-V1.1.0-12b-IQ3_XS.gguf) | IQ3_XS | 5.31GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Dans-DangerousWinds-V1.1.0-12b-Q2_K.gguf](https://huggingface.co/bartowski/Dans-DangerousWinds-V1.1.0-12b-GGUF/blob/main/Dans-DangerousWinds-V1.1.0-12b-Q2_K.gguf) | Q2_K | 4.79GB | false | Very low quality but surprisingly usable. |
| [Dans-DangerousWinds-V1.1.0-12b-IQ2_M.gguf](https://huggingface.co/bartowski/Dans-DangerousWinds-V1.1.0-12b-GGUF/blob/main/Dans-DangerousWinds-V1.1.0-12b-IQ2_M.gguf) | IQ2_M | 4.44GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
| [Dans-DangerousWinds-V1.1.0-12b-IQ2_S.gguf](https://huggingface.co/bartowski/Dans-DangerousWinds-V1.1.0-12b-GGUF/blob/main/Dans-DangerousWinds-V1.1.0-12b-IQ2_S.gguf) | IQ2_S | 4.14GB | false | Low quality, uses SOTA techniques to be usable. |
| [Dans-DangerousWinds-V1.1.0-12b-IQ2_XS.gguf](https://huggingface.co/bartowski/Dans-DangerousWinds-V1.1.0-12b-GGUF/blob/main/Dans-DangerousWinds-V1.1.0-12b-IQ2_XS.gguf) | IQ2_XS | 3.92GB | false | Low quality, uses SOTA techniques to be usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Dans-DangerousWinds-V1.1.0-12b-GGUF --include "Dans-DangerousWinds-V1.1.0-12b-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Dans-DangerousWinds-V1.1.0-12b-GGUF --include "Dans-DangerousWinds-V1.1.0-12b-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Dans-DangerousWinds-V1.1.0-12b-Q8_0) or download them all in place (./)
</details>
## ARM/AVX information
Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.
Now, however, there is something called "online repacking" for weights. details in [this PR](https://github.com/ggerganov/llama.cpp/pull/9921). If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.
As of llama.cpp build [b4282](https://github.com/ggerganov/llama.cpp/releases/tag/b4282) you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.
Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to [this PR](https://github.com/ggerganov/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information (deprecated</summary>
I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}