初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-27 17:33:07 +08:00
commit c0a13901ee
29 changed files with 326 additions and 0 deletions

61
.gitattributes vendored Normal file
View File

@@ -0,0 +1,61 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q4_1.gguf filter=lfs diff=lfs merge=lfs -text
Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-IQ4_NL.gguf filter=lfs diff=lfs merge=lfs -text
Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-IQ3_XXS.gguf filter=lfs diff=lfs merge=lfs -text
Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-IQ2_S.gguf filter=lfs diff=lfs merge=lfs -text
Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-f16.gguf filter=lfs diff=lfs merge=lfs -text
Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a62c289a396f54388947e283069e7379203d02f2db72359315b1f44d317d1310
size 5353855872

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:81e0efbb1b50e779ef42d3f8c7e7c5d08583a6187a93c6846f7a63c89d08400f
size 5001436032

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:bd53384d640da9bb022ef126ab7817657c72ed600e20c7abddda358e7f6b4a53
size 6913976256

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c0371e853f97928feb2819823b61c1eee35a8bfb35c6562a630c1f0d8dd206dc
size 6380799936

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:73d3cec2c50c4667809208f57558030a602a78576b3aa704779b5a082de77168
size 5944417152

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:34c3de1e6fe33101277ae2b453f9d301b958680576fa74f00b4f99af41e71183
size 8546350048

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:abfb7c49435beafbc1f323614fdf81677678de6806b02326f20581106bd976d7
size 8117071168

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ed3b42e3f7214df15848d8e9533e93b8aeab47b58ecf2faaf936c02297503c69
size 5768143424

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7d4d59678293bf5a0e5eece60368ee2078e3011bb5c98792ce749284cb61cb83
size 6526468384

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:61500aafab357c69eed66fa2d4a20f7966c66ff7dd5705b051cc442030520a62
size 7922206656

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9ef0e8d8e9600d504cf44a52f8e6df50d13c1b69c7634ea75e90455452cf1692
size 7336642496

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ca6f368add0a244fa37d2f779b06f982f136d9fbb5d757681207bd1a9b3aa017
size 6657034176

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c0a21148d37b7dac1a594cfb2b3de395790b81d525b9c5e77c154d7a26833da1
size 8601665824

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d1871b8cf29a2a37c808246be326221f19117f9009d5603c1f86516913d8053f
size 8541434848

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:459ad0a0f625c684a2baece1e609685593b2ac01d8db995daeb4629ccab81aaf
size 9389179168

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:42a8f086ca9a6e2adffea8195ae7aec8a163205ac2795dc86668669a1d8e8e3e
size 9561604384

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:dde3451fb0bb07c90849d9a7eeb4a9b3cf74ffc1b3d32e07c9a480a5e5a8c582
size 8985277408

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d2bdd7d8860b6461865be53708a05f4edc423f2464e6fddd03d7c5a5fb9f7823
size 8570598368

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ec42e66306ce677b007b838335ff2af08c1c9fa0ed83afad0a215afd263b0b39
size 10985046304

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6dbeb0d8c94a6b3bf2816f365c9da68d50b83ce3f9723a56e0084737e9546566
size 10505784928

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6d74f121a9222e01fb58f2335b9f9363535642d46021ecf168039b571bcc6ef8
size 10263465568

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3a2f5ef671b6c2fb34d2d82b9de581b3fe477679b1a7d516319fd6d10cfa85d7
size 12121324192

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d3310ee86706e387bbefefd6c6c1d83ea345fdeeacdcaac2ce0cb2da35ce7fab
size 12497453344

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:45ceabc7005fd1dc8d5ea4e81b3bd53bf0572f74f1d00ac687066fd2d74f2bed
size 15697248544

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d7322fc3ee477873c52a26ff5aeef12fd54e282b63cea710136cf07921b53be3
size 29539536192

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f68609c4218997b0707ca94128e49c911db9d9dac1e2e0444c05ac4c25d825a0
size 8563610

186
README.md Normal file
View File

@@ -0,0 +1,186 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
tags:
- text-generation-inference
- transformers
- unsloth
- qwen2
- trl
- sft
- chat
base_model: Cran-May/NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate
language:
- en
- zh
datasets:
- Lunzima/Semiboobs
license: apache-2.0
---
## Llamacpp imatrix Quantizations of NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate by Cran-May
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4738">b4738</a> for quantization.
Original model: https://huggingface.co/Cran-May/NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
Run them directly with [llama.cpp](https://github.com/ggerganov/llama.cpp), or any other llama.cpp based project
## Prompt format
```
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-f16.gguf](https://huggingface.co/bartowski/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-GGUF/blob/main/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-f16.gguf) | f16 | 29.54GB | false | Full F16 weights. |
| [NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q8_0.gguf](https://huggingface.co/bartowski/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-GGUF/blob/main/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q8_0.gguf) | Q8_0 | 15.70GB | false | Extremely high quality, generally unneeded but max available quant. |
| [NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q6_K_L.gguf](https://huggingface.co/bartowski/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-GGUF/blob/main/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q6_K_L.gguf) | Q6_K_L | 12.50GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q6_K.gguf](https://huggingface.co/bartowski/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-GGUF/blob/main/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q6_K.gguf) | Q6_K | 12.12GB | false | Very high quality, near perfect, *recommended*. |
| [NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q5_K_L.gguf](https://huggingface.co/bartowski/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-GGUF/blob/main/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q5_K_L.gguf) | Q5_K_L | 10.99GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q5_K_M.gguf](https://huggingface.co/bartowski/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-GGUF/blob/main/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q5_K_M.gguf) | Q5_K_M | 10.51GB | false | High quality, *recommended*. |
| [NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q5_K_S.gguf](https://huggingface.co/bartowski/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-GGUF/blob/main/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q5_K_S.gguf) | Q5_K_S | 10.26GB | false | High quality, *recommended*. |
| [NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q4_K_L.gguf](https://huggingface.co/bartowski/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-GGUF/blob/main/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q4_K_L.gguf) | Q4_K_L | 9.56GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q4_1.gguf](https://huggingface.co/bartowski/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-GGUF/blob/main/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q4_1.gguf) | Q4_1 | 9.39GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| [NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q4_K_M.gguf](https://huggingface.co/bartowski/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-GGUF/blob/main/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q4_K_M.gguf) | Q4_K_M | 8.99GB | false | Good quality, default size for most use cases, *recommended*. |
| [NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q3_K_XL.gguf](https://huggingface.co/bartowski/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-GGUF/blob/main/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q3_K_XL.gguf) | Q3_K_XL | 8.60GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q4_K_S.gguf](https://huggingface.co/bartowski/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-GGUF/blob/main/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q4_K_S.gguf) | Q4_K_S | 8.57GB | false | Slightly lower quality with more space savings, *recommended*. |
| [NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-IQ4_NL.gguf](https://huggingface.co/bartowski/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-GGUF/blob/main/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-IQ4_NL.gguf) | IQ4_NL | 8.55GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q4_0.gguf](https://huggingface.co/bartowski/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-GGUF/blob/main/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q4_0.gguf) | Q4_0 | 8.54GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| [NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-IQ4_XS.gguf](https://huggingface.co/bartowski/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-GGUF/blob/main/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-IQ4_XS.gguf) | IQ4_XS | 8.12GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q3_K_L.gguf](https://huggingface.co/bartowski/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-GGUF/blob/main/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q3_K_L.gguf) | Q3_K_L | 7.92GB | false | Lower quality but usable, good for low RAM availability. |
| [NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q3_K_M.gguf](https://huggingface.co/bartowski/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-GGUF/blob/main/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q3_K_M.gguf) | Q3_K_M | 7.34GB | false | Low quality. |
| [NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-IQ3_M.gguf](https://huggingface.co/bartowski/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-GGUF/blob/main/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-IQ3_M.gguf) | IQ3_M | 6.91GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q3_K_S.gguf](https://huggingface.co/bartowski/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-GGUF/blob/main/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q3_K_S.gguf) | Q3_K_S | 6.66GB | false | Low quality, not recommended. |
| [NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q2_K_L.gguf](https://huggingface.co/bartowski/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-GGUF/blob/main/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q2_K_L.gguf) | Q2_K_L | 6.53GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-IQ3_XS.gguf](https://huggingface.co/bartowski/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-GGUF/blob/main/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-IQ3_XS.gguf) | IQ3_XS | 6.38GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-IQ3_XXS.gguf](https://huggingface.co/bartowski/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-GGUF/blob/main/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-IQ3_XXS.gguf) | IQ3_XXS | 5.94GB | false | Lower quality, new method with decent performance, comparable to Q3 quants. |
| [NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q2_K.gguf](https://huggingface.co/bartowski/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-GGUF/blob/main/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q2_K.gguf) | Q2_K | 5.77GB | false | Very low quality but surprisingly usable. |
| [NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-IQ2_M.gguf](https://huggingface.co/bartowski/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-GGUF/blob/main/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-IQ2_M.gguf) | IQ2_M | 5.35GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
| [NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-IQ2_S.gguf](https://huggingface.co/bartowski/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-GGUF/blob/main/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-IQ2_S.gguf) | IQ2_S | 5.00GB | false | Low quality, uses SOTA techniques to be usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-GGUF --include "Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-GGUF --include "Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Cran-May_NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay-duplicate-Q8_0) or download them all in place (./)
</details>
## ARM/AVX information
Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.
Now, however, there is something called "online repacking" for weights. details in [this PR](https://github.com/ggerganov/llama.cpp/pull/9921). If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.
As of llama.cpp build [b4282](https://github.com/ggerganov/llama.cpp/releases/tag/b4282) you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.
Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to [this PR](https://github.com/ggerganov/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information (deprecated</summary>
I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Thank you to LM Studio for sponsoring my work.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}