初始化项目,由ModelHub XC社区提供模型

Model: bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-09-03 23:56:13 +08:00
commit 439e863af2
30 changed files with 341 additions and 0 deletions

75
.gitattributes vendored Normal file
View File

@@ -0,0 +1,75 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bin.* filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zstandard filter=lfs diff=lfs merge=lfs -text
*.tfevents* filter=lfs diff=lfs merge=lfs -text
*.db* filter=lfs diff=lfs merge=lfs -text
*.ark* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*data* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.meta filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.index filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ggml filter=lfs diff=lfs merge=lfs -text
*.llamafile* filter=lfs diff=lfs merge=lfs -text
*.pt2 filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
nvidia_Nemotron-Cascade-14B-Thinking-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
nvidia_Nemotron-Cascade-14B-Thinking-IQ2_S.gguf filter=lfs diff=lfs merge=lfs -text
nvidia_Nemotron-Cascade-14B-Thinking-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
nvidia_Nemotron-Cascade-14B-Thinking-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
nvidia_Nemotron-Cascade-14B-Thinking-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
nvidia_Nemotron-Cascade-14B-Thinking-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
nvidia_Nemotron-Cascade-14B-Thinking-bf16.gguf filter=lfs diff=lfs merge=lfs -text
nvidia_Nemotron-Cascade-14B-Thinking-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
nvidia_Nemotron-Cascade-14B-Thinking-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
nvidia_Nemotron-Cascade-14B-Thinking-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
nvidia_Nemotron-Cascade-14B-Thinking-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
nvidia_Nemotron-Cascade-14B-Thinking-IQ2_XS.gguf filter=lfs diff=lfs merge=lfs -text
nvidia_Nemotron-Cascade-14B-Thinking-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
nvidia_Nemotron-Cascade-14B-Thinking-IQ4_NL.gguf filter=lfs diff=lfs merge=lfs -text
nvidia_Nemotron-Cascade-14B-Thinking-IQ3_XXS.gguf filter=lfs diff=lfs merge=lfs -text
nvidia_Nemotron-Cascade-14B-Thinking-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
nvidia_Nemotron-Cascade-14B-Thinking-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
nvidia_Nemotron-Cascade-14B-Thinking-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
nvidia_Nemotron-Cascade-14B-Thinking-Q4_1.gguf filter=lfs diff=lfs merge=lfs -text
nvidia_Nemotron-Cascade-14B-Thinking-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
nvidia_Nemotron-Cascade-14B-Thinking-imatrix.gguf filter=lfs diff=lfs merge=lfs -text
nvidia_Nemotron-Cascade-14B-Thinking-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
nvidia_Nemotron-Cascade-14B-Thinking-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
nvidia_Nemotron-Cascade-14B-Thinking-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
nvidia_Nemotron-Cascade-14B-Thinking-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
nvidia_Nemotron-Cascade-14B-Thinking-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
nvidia_Nemotron-Cascade-14B-Thinking-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text

184
README.md Normal file
View File

@@ -0,0 +1,184 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
tags:
- nvidia
- nemotron-cascade
- reasoning
- general-purpose
- SFT
- RL
license: other
license_name: nvidia-open-model-license
language:
- en
base_model: nvidia/Nemotron-Cascade-14B-Thinking
base_model_relation: quantized
license_link: https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/
---
## Llamacpp imatrix Quantizations of Nemotron-Cascade-14B-Thinking by nvidia
Using <a href="https://github.com/ggml-org/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggml-org/llama.cpp/releases/tag/b7423">b7423</a> for quantization.
Original model: https://huggingface.co/nvidia/Nemotron-Cascade-14B-Thinking
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8) combined with a subset of combined_all_small.parquet from Ed Addario [here](https://huggingface.co/datasets/eaddario/imatrix-calibration/blob/main/combined_all_small.parquet)
Run them in [LM Studio](https://lmstudio.ai/)
Run them directly with [llama.cpp](https://github.com/ggml-org/llama.cpp), or any other llama.cpp based project
## Prompt format
```
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt} /think<|im_end|>
<|im_start|>assistant
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Nemotron-Cascade-14B-Thinking-bf16.gguf](https://huggingface.co/bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF/blob/main/nvidia_Nemotron-Cascade-14B-Thinking-bf16.gguf) | bf16 | 29.54GB | false | Full BF16 weights. |
| [Nemotron-Cascade-14B-Thinking-Q8_0.gguf](https://huggingface.co/bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF/blob/main/nvidia_Nemotron-Cascade-14B-Thinking-Q8_0.gguf) | Q8_0 | 15.70GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Nemotron-Cascade-14B-Thinking-Q6_K_L.gguf](https://huggingface.co/bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF/blob/main/nvidia_Nemotron-Cascade-14B-Thinking-Q6_K_L.gguf) | Q6_K_L | 12.50GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Nemotron-Cascade-14B-Thinking-Q6_K.gguf](https://huggingface.co/bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF/blob/main/nvidia_Nemotron-Cascade-14B-Thinking-Q6_K.gguf) | Q6_K | 12.12GB | false | Very high quality, near perfect, *recommended*. |
| [Nemotron-Cascade-14B-Thinking-Q5_K_L.gguf](https://huggingface.co/bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF/blob/main/nvidia_Nemotron-Cascade-14B-Thinking-Q5_K_L.gguf) | Q5_K_L | 10.99GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Nemotron-Cascade-14B-Thinking-Q5_K_M.gguf](https://huggingface.co/bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF/blob/main/nvidia_Nemotron-Cascade-14B-Thinking-Q5_K_M.gguf) | Q5_K_M | 10.51GB | false | High quality, *recommended*. |
| [Nemotron-Cascade-14B-Thinking-Q5_K_S.gguf](https://huggingface.co/bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF/blob/main/nvidia_Nemotron-Cascade-14B-Thinking-Q5_K_S.gguf) | Q5_K_S | 10.26GB | false | High quality, *recommended*. |
| [Nemotron-Cascade-14B-Thinking-Q4_K_L.gguf](https://huggingface.co/bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF/blob/main/nvidia_Nemotron-Cascade-14B-Thinking-Q4_K_L.gguf) | Q4_K_L | 9.58GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Nemotron-Cascade-14B-Thinking-Q4_1.gguf](https://huggingface.co/bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF/blob/main/nvidia_Nemotron-Cascade-14B-Thinking-Q4_1.gguf) | Q4_1 | 9.39GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| [Nemotron-Cascade-14B-Thinking-Q4_K_M.gguf](https://huggingface.co/bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF/blob/main/nvidia_Nemotron-Cascade-14B-Thinking-Q4_K_M.gguf) | Q4_K_M | 9.00GB | false | Good quality, default size for most use cases, *recommended*. |
| [Nemotron-Cascade-14B-Thinking-Q3_K_XL.gguf](https://huggingface.co/bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF/blob/main/nvidia_Nemotron-Cascade-14B-Thinking-Q3_K_XL.gguf) | Q3_K_XL | 8.58GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Nemotron-Cascade-14B-Thinking-Q4_K_S.gguf](https://huggingface.co/bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF/blob/main/nvidia_Nemotron-Cascade-14B-Thinking-Q4_K_S.gguf) | Q4_K_S | 8.57GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Nemotron-Cascade-14B-Thinking-Q4_0.gguf](https://huggingface.co/bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF/blob/main/nvidia_Nemotron-Cascade-14B-Thinking-Q4_0.gguf) | Q4_0 | 8.54GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| [Nemotron-Cascade-14B-Thinking-IQ4_NL.gguf](https://huggingface.co/bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF/blob/main/nvidia_Nemotron-Cascade-14B-Thinking-IQ4_NL.gguf) | IQ4_NL | 8.54GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [Nemotron-Cascade-14B-Thinking-IQ4_XS.gguf](https://huggingface.co/bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF/blob/main/nvidia_Nemotron-Cascade-14B-Thinking-IQ4_XS.gguf) | IQ4_XS | 8.11GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Nemotron-Cascade-14B-Thinking-Q3_K_L.gguf](https://huggingface.co/bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF/blob/main/nvidia_Nemotron-Cascade-14B-Thinking-Q3_K_L.gguf) | Q3_K_L | 7.90GB | false | Lower quality but usable, good for low RAM availability. |
| [Nemotron-Cascade-14B-Thinking-Q3_K_M.gguf](https://huggingface.co/bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF/blob/main/nvidia_Nemotron-Cascade-14B-Thinking-Q3_K_M.gguf) | Q3_K_M | 7.32GB | false | Low quality. |
| [Nemotron-Cascade-14B-Thinking-IQ3_M.gguf](https://huggingface.co/bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF/blob/main/nvidia_Nemotron-Cascade-14B-Thinking-IQ3_M.gguf) | IQ3_M | 6.88GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Nemotron-Cascade-14B-Thinking-Q3_K_S.gguf](https://huggingface.co/bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF/blob/main/nvidia_Nemotron-Cascade-14B-Thinking-Q3_K_S.gguf) | Q3_K_S | 6.66GB | false | Low quality, not recommended. |
| [Nemotron-Cascade-14B-Thinking-Q2_K_L.gguf](https://huggingface.co/bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF/blob/main/nvidia_Nemotron-Cascade-14B-Thinking-Q2_K_L.gguf) | Q2_K_L | 6.51GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Nemotron-Cascade-14B-Thinking-IQ3_XS.gguf](https://huggingface.co/bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF/blob/main/nvidia_Nemotron-Cascade-14B-Thinking-IQ3_XS.gguf) | IQ3_XS | 6.38GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Nemotron-Cascade-14B-Thinking-IQ3_XXS.gguf](https://huggingface.co/bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF/blob/main/nvidia_Nemotron-Cascade-14B-Thinking-IQ3_XXS.gguf) | IQ3_XXS | 5.94GB | false | Lower quality, new method with decent performance, comparable to Q3 quants. |
| [Nemotron-Cascade-14B-Thinking-Q2_K.gguf](https://huggingface.co/bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF/blob/main/nvidia_Nemotron-Cascade-14B-Thinking-Q2_K.gguf) | Q2_K | 5.75GB | false | Very low quality but surprisingly usable. |
| [Nemotron-Cascade-14B-Thinking-IQ2_M.gguf](https://huggingface.co/bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF/blob/main/nvidia_Nemotron-Cascade-14B-Thinking-IQ2_M.gguf) | IQ2_M | 5.32GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
| [Nemotron-Cascade-14B-Thinking-IQ2_S.gguf](https://huggingface.co/bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF/blob/main/nvidia_Nemotron-Cascade-14B-Thinking-IQ2_S.gguf) | IQ2_S | 4.96GB | false | Low quality, uses SOTA techniques to be usable. |
| [Nemotron-Cascade-14B-Thinking-IQ2_XS.gguf](https://huggingface.co/bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF/blob/main/nvidia_Nemotron-Cascade-14B-Thinking-IQ2_XS.gguf) | IQ2_XS | 4.69GB | false | Low quality, uses SOTA techniques to be usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF --include "nvidia_Nemotron-Cascade-14B-Thinking-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/nvidia_Nemotron-Cascade-14B-Thinking-GGUF --include "nvidia_Nemotron-Cascade-14B-Thinking-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (nvidia_Nemotron-Cascade-14B-Thinking-Q8_0) or download them all in place (./)
</details>
## ARM/AVX information
Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.
Now, however, there is something called "online repacking" for weights. details in [this PR](https://github.com/ggml-org/llama.cpp/pull/9921). If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.
As of llama.cpp build [b4282](https://github.com/ggml-org/llama.cpp/releases/tag/b4282) you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.
Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to [this PR](https://github.com/ggml-org/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information (deprecated</summary>
I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggml-org/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Thank you to LM Studio for sponsoring my work.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3e7d71678abb9cc24ea049b13a87e40533649f9670a81563cc0a42f3bfd5f1a5
size 5322942080

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0365e192fb522fb14d95789ec3502596d58bbb642abf1d15777c09866baaa1dc
size 4963313280

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d06e90b472f2fb9840db9318e83382c5fcb927c8efb7a0402a3f997a54336a65
size 4691589760

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:025f78b0718f7b3df0a097d57ebdeb681b20a954491d762b043f060f3eecde1f
size 6883410560

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e66d0ffd3c16fc08fcda635350d60c76df1bf7bc3b97dfa058cfe000199468d5
size 6375301760

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:869b931304d697a0aa7e6d134816aba8d850204f58db0c772a2384e556a02e1c
size 5942666880

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:66bd78af480e8830ad148b3051c989a04b8d6d6c3fd049cf99b8e8e225925ad8
size 8541363840

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a5faccd9f87c3766e58e4afa0a49cde647981272f33f6219ec48fc05364a76c2
size 8110730880

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:89b23807c2a0022dfd3f32a72be9d7b683ba56277bebe67ddb7f1757d3a3ce6b
size 5753984640

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:dfb63fb302a42b785bf6e62dbf3832dbb4219d4a16f7541f3d5dac7cdd5d5343
size 6513664640

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1e72d106b5fcf78ed49ae8368a352622849bf06feee3aad132b14a300b3482db
size 7900652160

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3266092c1cb5ca3655e66dfb6ecde0ab879292f3a401e3b9fed74d3383f17a08
size 7321313920

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fd9eccebd8997e2a37a431b552eec95c6914e3f5c112301624a702537e6eae97
size 6657106560

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:36cc41cf93ae4f18970ea3a670377400276ce1886a972a6f15a4417550f7fc9f
size 8581325440

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2ca9fef0378e94a53da4611e3712e3915c3858c5b814c0e60dd5c6c7bec30957
size 8543002240

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6d4b9d88e501c0e6d64887970b736197a13d5e96f147b9c93685de682be22af7
size 9389522560

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4381f1936520e1fe3e28b8f0a1f69874fe3be824cd5941f7c081a8643e0c87e0
size 9579111040

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:af36aa06103ebf068ddd36d00aba46b3a2a8d448da08f8824fdf8ea4c76a9522
size 9001754240

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f7b61e2698036d8028f868fe06e78712134f6416fb55b2f8b787f8a1f03fa5bc
size 8573476480

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0b0cf91f28257d9c771b4bbb753dd31ce573fea5f53c3139e7df9c349ecb8031
size 10994688640

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:964f654b6f573b7552decad7806b7e78b2069874cec47344679c46c2774e1cb7
size 10514570880

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:faed2ee2644539ae728515c4e4ee31452ec96ce48885aac44a85b64eb7f14db0
size 10263895680

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:edcba986796a702287479885c4bc14b59ef52c421bccf149e2bc63f6a397f0de
size 12121938560

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:bfaaf894b67f0d0d8464369e994e9a08b21bddb109370180e2ec793fe9c14575
size 12498739840

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:99c533a14f77d214cd90e1453ecdf0ff7fbf61d189afdf6eb3aa72252258af03
size 15698535040

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:694680cdfd023b11df8665a6b8ba7cf62bfac49ba2d6142bc66933cda6d5b7e8
size 29543424320

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7cbd679c8f9c5d353205a05bf0cf983b317f3589b15b2706258da007c8e58265
size 7743584