初始化项目,由ModelHub XC社区提供模型

Model: bartowski/NightyGurps-14b-v1.1-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-09-01 13:28:13 +08:00
commit fd9e4f372b
28 changed files with 263 additions and 0 deletions

60
.gitattributes vendored Normal file
View File

@@ -0,0 +1,60 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
NightyGurps-14b-v1.1-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
NightyGurps-14b-v1.1-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
NightyGurps-14b-v1.1-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
NightyGurps-14b-v1.1-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
NightyGurps-14b-v1.1-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
NightyGurps-14b-v1.1-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
NightyGurps-14b-v1.1-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
NightyGurps-14b-v1.1-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
NightyGurps-14b-v1.1-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
NightyGurps-14b-v1.1-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
NightyGurps-14b-v1.1-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
NightyGurps-14b-v1.1-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
NightyGurps-14b-v1.1-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
NightyGurps-14b-v1.1-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
NightyGurps-14b-v1.1-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
NightyGurps-14b-v1.1-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
NightyGurps-14b-v1.1-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
NightyGurps-14b-v1.1-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
NightyGurps-14b-v1.1-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
NightyGurps-14b-v1.1-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
NightyGurps-14b-v1.1-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
NightyGurps-14b-v1.1-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
NightyGurps-14b-v1.1-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
NightyGurps-14b-v1.1-f16.gguf filter=lfs diff=lfs merge=lfs -text
NightyGurps-14b-v1.1.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7753b4acc5b30715d0b75089739f74931373b5742e76b6c792121a6f93a1f89d
size 5356146688

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:aee30af7eb5fa02b05b7d2103095fad5402bdce28cc252c60f7444be4afe4ecc
size 6916538368

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cb55ce9a9a040dd80fc902a389c74ebd4957198a9bc99090b0fdf14650f84cbb
size 6383362048

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6f8bdf45cc14fd76f95c0c46c96f5364d926aa8399323f4d671a1f4f23aea84d
size 8119840768

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1d001befaeb15502817c33bca39db35fd7598f08d8d407a31e335bc2721b371a
size 5770498048

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:255095896f8785b04052a4ac80cc9da6b9747665a8de6243e974c28ddedf203b
size 6530818048

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:19131bf4972bd2fe4adce2ddd1ad9c2a8bc8417b2576a0bbb2f8798efc13f36e
size 7924768768

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e1165b7fce255a497a5f43c244f81341552d8cc12b80b054908d133598dd6f72
size 7339204608

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8603785709ebb325074e487b62867a517cc63272d484bc776d9046dae5edd888
size 6659596288

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cbef78b24be89cad4170f9d56ca1e5c88e9c492a09af4b0090968995c32bfe19
size 8606015488

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e25e472fb4b8b16509eaf1695f69f6a8d25c4242357f620ccbac9a13b668b63a
size 8544268288

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:384ba1a8014356dea7a65b2f949260259d5c370764c8c270744f4674b2298bd2
size 8517726208

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e658b223c18a36326a4c0160fa9602bef1d55489a7c9c5ee0a63159266a773ce
size 8517726208

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:56db3ebf7e2e906e13b0f7ec6e22abf791ec7890375f03b9fd15c8a70e20559e
size 8517726208

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9e9d2b8dc87dae4b86fde3037e5234c03fdc6dbd7d17adfd127fec9f10f3324a
size 9565954048

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d09d53259ad2c0298150fa8c2db98fe42f11731af89fdc80ad0e255a19adc4b0
size 8988110848

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cca2769d801cef82cb9538564556ba698b745909b1531d81d92fa13d3f0d9e0a
size 8573431808

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8f30029d1a5e950419586b85a183a2995c14e1081ef696dc09172ec69ee34c9c
size 10989395968

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:18d97dab263e1a89ba0d5af19e424f41bdf4cfab546fccce6c518e0c0e1249bc
size 10508873728

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ad6a0a5e281cdbd8430368c4d1d5a1943b5fe06af4d8138afbc9dc24166deb81
size 10266554368

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:59b2cfb325ac7954a0d4c82d8034a1b129c6662c53df2c43cc0d1ef55080de79
size 12124684288

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a17390174429a9970e1274f7b7da6e2070d5185be4be43ae7fdae4cb23eb7c0e
size 12501803008

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c0f6c277eb5b128fcaaa4fc19ca895a846856a9410170aa53f1ac697ed2d386c
size 15701598208

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:bea5ea8b4421a08fdff334d93c26b9b4429424d2e775921d178c92a197c34e9d
size 29547716352

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a384e3185f2107ba0a5ed4da63a014509012e790d5529e1448792d6010e76318
size 8563610

127
README.md Normal file
View File

@@ -0,0 +1,127 @@
---
base_model: AlexBefest/NightyGurps-14b-v1.1
language:
- ru
license: apache-2.0
pipeline_tag: text-generation
quantized_by: bartowski
---
## Llamacpp imatrix Quantizations of NightyGurps-14b-v1.1
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b3787">b3787</a> for quantization.
Original model: https://huggingface.co/AlexBefest/NightyGurps-14b-v1.1
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [NightyGurps-14b-v1.1-f16.gguf](https://huggingface.co/bartowski/NightyGurps-14b-v1.1-GGUF/blob/main/NightyGurps-14b-v1.1-f16.gguf) | f16 | 29.55GB | false | Full F16 weights. |
| [NightyGurps-14b-v1.1-Q8_0.gguf](https://huggingface.co/bartowski/NightyGurps-14b-v1.1-GGUF/blob/main/NightyGurps-14b-v1.1-Q8_0.gguf) | Q8_0 | 15.70GB | false | Extremely high quality, generally unneeded but max available quant. |
| [NightyGurps-14b-v1.1-Q6_K_L.gguf](https://huggingface.co/bartowski/NightyGurps-14b-v1.1-GGUF/blob/main/NightyGurps-14b-v1.1-Q6_K_L.gguf) | Q6_K_L | 12.50GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [NightyGurps-14b-v1.1-Q6_K.gguf](https://huggingface.co/bartowski/NightyGurps-14b-v1.1-GGUF/blob/main/NightyGurps-14b-v1.1-Q6_K.gguf) | Q6_K | 12.12GB | false | Very high quality, near perfect, *recommended*. |
| [NightyGurps-14b-v1.1-Q5_K_L.gguf](https://huggingface.co/bartowski/NightyGurps-14b-v1.1-GGUF/blob/main/NightyGurps-14b-v1.1-Q5_K_L.gguf) | Q5_K_L | 10.99GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [NightyGurps-14b-v1.1-Q5_K_M.gguf](https://huggingface.co/bartowski/NightyGurps-14b-v1.1-GGUF/blob/main/NightyGurps-14b-v1.1-Q5_K_M.gguf) | Q5_K_M | 10.51GB | false | High quality, *recommended*. |
| [NightyGurps-14b-v1.1-Q5_K_S.gguf](https://huggingface.co/bartowski/NightyGurps-14b-v1.1-GGUF/blob/main/NightyGurps-14b-v1.1-Q5_K_S.gguf) | Q5_K_S | 10.27GB | false | High quality, *recommended*. |
| [NightyGurps-14b-v1.1-Q4_K_L.gguf](https://huggingface.co/bartowski/NightyGurps-14b-v1.1-GGUF/blob/main/NightyGurps-14b-v1.1-Q4_K_L.gguf) | Q4_K_L | 9.57GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [NightyGurps-14b-v1.1-Q4_K_M.gguf](https://huggingface.co/bartowski/NightyGurps-14b-v1.1-GGUF/blob/main/NightyGurps-14b-v1.1-Q4_K_M.gguf) | Q4_K_M | 8.99GB | false | Good quality, default size for must use cases, *recommended*. |
| [NightyGurps-14b-v1.1-Q3_K_XL.gguf](https://huggingface.co/bartowski/NightyGurps-14b-v1.1-GGUF/blob/main/NightyGurps-14b-v1.1-Q3_K_XL.gguf) | Q3_K_XL | 8.61GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [NightyGurps-14b-v1.1-Q4_K_S.gguf](https://huggingface.co/bartowski/NightyGurps-14b-v1.1-GGUF/blob/main/NightyGurps-14b-v1.1-Q4_K_S.gguf) | Q4_K_S | 8.57GB | false | Slightly lower quality with more space savings, *recommended*. |
| [NightyGurps-14b-v1.1-Q4_0.gguf](https://huggingface.co/bartowski/NightyGurps-14b-v1.1-GGUF/blob/main/NightyGurps-14b-v1.1-Q4_0.gguf) | Q4_0 | 8.54GB | false | Legacy format, generally not worth using over similarly sized formats |
| [NightyGurps-14b-v1.1-Q4_0_8_8.gguf](https://huggingface.co/bartowski/NightyGurps-14b-v1.1-GGUF/blob/main/NightyGurps-14b-v1.1-Q4_0_8_8.gguf) | Q4_0_8_8 | 8.52GB | false | Optimized for ARM inference. Requires 'sve' support (see link below). |
| [NightyGurps-14b-v1.1-Q4_0_4_8.gguf](https://huggingface.co/bartowski/NightyGurps-14b-v1.1-GGUF/blob/main/NightyGurps-14b-v1.1-Q4_0_4_8.gguf) | Q4_0_4_8 | 8.52GB | false | Optimized for ARM inference. Requires 'i8mm' support (see link below). |
| [NightyGurps-14b-v1.1-Q4_0_4_4.gguf](https://huggingface.co/bartowski/NightyGurps-14b-v1.1-GGUF/blob/main/NightyGurps-14b-v1.1-Q4_0_4_4.gguf) | Q4_0_4_4 | 8.52GB | false | Optimized for ARM inference. Should work well on all ARM chips, pick this if you're unsure. |
| [NightyGurps-14b-v1.1-IQ4_XS.gguf](https://huggingface.co/bartowski/NightyGurps-14b-v1.1-GGUF/blob/main/NightyGurps-14b-v1.1-IQ4_XS.gguf) | IQ4_XS | 8.12GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [NightyGurps-14b-v1.1-Q3_K_L.gguf](https://huggingface.co/bartowski/NightyGurps-14b-v1.1-GGUF/blob/main/NightyGurps-14b-v1.1-Q3_K_L.gguf) | Q3_K_L | 7.92GB | false | Lower quality but usable, good for low RAM availability. |
| [NightyGurps-14b-v1.1-Q3_K_M.gguf](https://huggingface.co/bartowski/NightyGurps-14b-v1.1-GGUF/blob/main/NightyGurps-14b-v1.1-Q3_K_M.gguf) | Q3_K_M | 7.34GB | false | Low quality. |
| [NightyGurps-14b-v1.1-IQ3_M.gguf](https://huggingface.co/bartowski/NightyGurps-14b-v1.1-GGUF/blob/main/NightyGurps-14b-v1.1-IQ3_M.gguf) | IQ3_M | 6.92GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [NightyGurps-14b-v1.1-Q3_K_S.gguf](https://huggingface.co/bartowski/NightyGurps-14b-v1.1-GGUF/blob/main/NightyGurps-14b-v1.1-Q3_K_S.gguf) | Q3_K_S | 6.66GB | false | Low quality, not recommended. |
| [NightyGurps-14b-v1.1-Q2_K_L.gguf](https://huggingface.co/bartowski/NightyGurps-14b-v1.1-GGUF/blob/main/NightyGurps-14b-v1.1-Q2_K_L.gguf) | Q2_K_L | 6.53GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [NightyGurps-14b-v1.1-IQ3_XS.gguf](https://huggingface.co/bartowski/NightyGurps-14b-v1.1-GGUF/blob/main/NightyGurps-14b-v1.1-IQ3_XS.gguf) | IQ3_XS | 6.38GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [NightyGurps-14b-v1.1-Q2_K.gguf](https://huggingface.co/bartowski/NightyGurps-14b-v1.1-GGUF/blob/main/NightyGurps-14b-v1.1-Q2_K.gguf) | Q2_K | 5.77GB | false | Very low quality but surprisingly usable. |
| [NightyGurps-14b-v1.1-IQ2_M.gguf](https://huggingface.co/bartowski/NightyGurps-14b-v1.1-GGUF/blob/main/NightyGurps-14b-v1.1-IQ2_M.gguf) | IQ2_M | 5.36GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/NightyGurps-14b-v1.1-GGUF --include "NightyGurps-14b-v1.1-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/NightyGurps-14b-v1.1-GGUF --include "NightyGurps-14b-v1.1-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (NightyGurps-14b-v1.1-Q8_0) or download them all in place (./)
## Q4_0_X_X
These are *NOT* for Metal (Apple) offloading, only ARM chips.
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset
Thank you ZeroWw for the inspiration to experiment with embed/output
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}