commit bb3d2b677a7f611d8a8bbe8634dd0b7532454c93 Author: ModelHub XC Date: Tue Jul 7 07:56:16 2026 +0800 初始化项目,由ModelHub XC社区提供模型 Model: mradermacher/Nemotron-Cascade-14B-Thinking-i1-GGUF Source: Original Platform diff --git a/.gitattributes b/.gitattributes new file mode 100644 index 0000000..552654a --- /dev/null +++ b/.gitattributes @@ -0,0 +1,60 @@ +*.7z filter=lfs diff=lfs merge=lfs -text +*.arrow filter=lfs diff=lfs merge=lfs -text +*.bin filter=lfs diff=lfs merge=lfs -text +*.bz2 filter=lfs diff=lfs merge=lfs -text +*.ckpt filter=lfs diff=lfs merge=lfs -text +*.ftz filter=lfs diff=lfs merge=lfs -text +*.gz filter=lfs diff=lfs merge=lfs -text +*.h5 filter=lfs diff=lfs merge=lfs -text +*.joblib filter=lfs diff=lfs merge=lfs -text +*.lfs.* filter=lfs diff=lfs merge=lfs -text +*.mlmodel filter=lfs diff=lfs merge=lfs -text +*.model filter=lfs diff=lfs merge=lfs -text +*.msgpack filter=lfs diff=lfs merge=lfs -text +*.npy filter=lfs diff=lfs merge=lfs -text +*.npz filter=lfs diff=lfs merge=lfs -text +*.onnx filter=lfs diff=lfs merge=lfs -text +*.ot filter=lfs diff=lfs merge=lfs -text +*.parquet filter=lfs diff=lfs merge=lfs -text +*.pb filter=lfs diff=lfs merge=lfs -text +*.pickle filter=lfs diff=lfs merge=lfs -text +*.pkl filter=lfs diff=lfs merge=lfs -text +*.pt filter=lfs diff=lfs merge=lfs -text +*.pth filter=lfs diff=lfs merge=lfs -text +*.rar filter=lfs diff=lfs merge=lfs -text +*.safetensors filter=lfs diff=lfs merge=lfs -text +saved_model/**/* filter=lfs diff=lfs merge=lfs -text +*.tar.* filter=lfs diff=lfs merge=lfs -text +*.tar filter=lfs diff=lfs merge=lfs -text +*.tflite filter=lfs diff=lfs merge=lfs -text +*.tgz filter=lfs diff=lfs merge=lfs -text +*.wasm filter=lfs diff=lfs merge=lfs -text +*.xz filter=lfs diff=lfs merge=lfs -text +*.zip filter=lfs diff=lfs merge=lfs -text +*.zst filter=lfs diff=lfs merge=lfs -text +*tfevents* filter=lfs diff=lfs merge=lfs -text +Nemotron-Cascade-14B-Thinking.imatrix.gguf filter=lfs diff=lfs merge=lfs -text +Nemotron-Cascade-14B-Thinking.i1-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text +Nemotron-Cascade-14B-Thinking.i1-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text +Nemotron-Cascade-14B-Thinking.i1-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text +Nemotron-Cascade-14B-Thinking.i1-IQ3_XXS.gguf filter=lfs diff=lfs merge=lfs -text +Nemotron-Cascade-14B-Thinking.i1-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text +Nemotron-Cascade-14B-Thinking.i1-IQ4_NL.gguf filter=lfs diff=lfs merge=lfs -text +Nemotron-Cascade-14B-Thinking.i1-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text +Nemotron-Cascade-14B-Thinking.i1-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text +Nemotron-Cascade-14B-Thinking.i1-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text +Nemotron-Cascade-14B-Thinking.i1-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text +Nemotron-Cascade-14B-Thinking.i1-Q2_K_S.gguf filter=lfs diff=lfs merge=lfs -text +Nemotron-Cascade-14B-Thinking.i1-IQ1_M.gguf filter=lfs diff=lfs merge=lfs -text +Nemotron-Cascade-14B-Thinking.i1-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text +Nemotron-Cascade-14B-Thinking.i1-IQ2_XXS.gguf filter=lfs diff=lfs merge=lfs -text +Nemotron-Cascade-14B-Thinking.i1-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text +Nemotron-Cascade-14B-Thinking.i1-IQ2_XS.gguf filter=lfs diff=lfs merge=lfs -text +Nemotron-Cascade-14B-Thinking.i1-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text +Nemotron-Cascade-14B-Thinking.i1-IQ2_S.gguf filter=lfs diff=lfs merge=lfs -text +Nemotron-Cascade-14B-Thinking.i1-IQ1_S.gguf filter=lfs diff=lfs merge=lfs -text +Nemotron-Cascade-14B-Thinking.i1-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text +Nemotron-Cascade-14B-Thinking.i1-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text +Nemotron-Cascade-14B-Thinking.i1-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text +Nemotron-Cascade-14B-Thinking.i1-Q4_1.gguf filter=lfs diff=lfs merge=lfs -text +Nemotron-Cascade-14B-Thinking.i1-IQ3_S.gguf filter=lfs diff=lfs merge=lfs -text diff --git a/Nemotron-Cascade-14B-Thinking.i1-IQ1_M.gguf b/Nemotron-Cascade-14B-Thinking.i1-IQ1_M.gguf new file mode 100644 index 0000000..dd0d65c --- /dev/null +++ b/Nemotron-Cascade-14B-Thinking.i1-IQ1_M.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:cab68fe305aa35b9fb4b470abee215e927ed3fb4107d9c4beacf6357a8037056 +size 3849657376 diff --git a/Nemotron-Cascade-14B-Thinking.i1-IQ1_S.gguf b/Nemotron-Cascade-14B-Thinking.i1-IQ1_S.gguf new file mode 100644 index 0000000..3b9c79b --- /dev/null +++ b/Nemotron-Cascade-14B-Thinking.i1-IQ1_S.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:ad76a6810e9131d41d3e01076d83e24e853d14f710569add466ddce4a3c2e4e9 +size 3579935776 diff --git a/Nemotron-Cascade-14B-Thinking.i1-IQ2_M.gguf b/Nemotron-Cascade-14B-Thinking.i1-IQ2_M.gguf new file mode 100644 index 0000000..101fd81 --- /dev/null +++ b/Nemotron-Cascade-14B-Thinking.i1-IQ2_M.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:9b919630eea4596fd8a7ce3368b3eebb1d5ef3b597020b35572fceabca612acf +size 5322942496 diff --git a/Nemotron-Cascade-14B-Thinking.i1-IQ2_S.gguf b/Nemotron-Cascade-14B-Thinking.i1-IQ2_S.gguf new file mode 100644 index 0000000..c1e32af --- /dev/null +++ b/Nemotron-Cascade-14B-Thinking.i1-IQ2_S.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:1f2403d09d313a3c5f7d4edf9ea9dd689aa4f5acbf109d05671d464d9382eed7 +size 4963313696 diff --git a/Nemotron-Cascade-14B-Thinking.i1-IQ2_XS.gguf b/Nemotron-Cascade-14B-Thinking.i1-IQ2_XS.gguf new file mode 100644 index 0000000..d802b0b --- /dev/null +++ b/Nemotron-Cascade-14B-Thinking.i1-IQ2_XS.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:486a72eb4a7852aabb30cf8d3e1e24a8ad81290855f9bbbbef19e26f72388cd5 +size 4691590176 diff --git a/Nemotron-Cascade-14B-Thinking.i1-IQ2_XXS.gguf b/Nemotron-Cascade-14B-Thinking.i1-IQ2_XXS.gguf new file mode 100644 index 0000000..924b9c0 --- /dev/null +++ b/Nemotron-Cascade-14B-Thinking.i1-IQ2_XXS.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:f2b8145581b700d47fefa0ebc5a42b5a73191971b383885972a0af630e17b211 +size 4299193376 diff --git a/Nemotron-Cascade-14B-Thinking.i1-IQ3_M.gguf b/Nemotron-Cascade-14B-Thinking.i1-IQ3_M.gguf new file mode 100644 index 0000000..f91c247 --- /dev/null +++ b/Nemotron-Cascade-14B-Thinking.i1-IQ3_M.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:f4c6e1cfa50ccac06d22a9b133d9b2e61654f35c63f69619d01619e2bc268407 +size 6883410976 diff --git a/Nemotron-Cascade-14B-Thinking.i1-IQ3_S.gguf b/Nemotron-Cascade-14B-Thinking.i1-IQ3_S.gguf new file mode 100644 index 0000000..23ac643 --- /dev/null +++ b/Nemotron-Cascade-14B-Thinking.i1-IQ3_S.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:9ed3c063680c3d6d1ef04d5caa6df77ffe3e7f8acd6f24fce6f1344965d821ec +size 6684959776 diff --git a/Nemotron-Cascade-14B-Thinking.i1-IQ3_XS.gguf b/Nemotron-Cascade-14B-Thinking.i1-IQ3_XS.gguf new file mode 100644 index 0000000..0cd5501 --- /dev/null +++ b/Nemotron-Cascade-14B-Thinking.i1-IQ3_XS.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:4526c17176ef13fa3ae33e79023db32f90fbd8ef1597a433ad715753601451f3 +size 6375302176 diff --git a/Nemotron-Cascade-14B-Thinking.i1-IQ3_XXS.gguf b/Nemotron-Cascade-14B-Thinking.i1-IQ3_XXS.gguf new file mode 100644 index 0000000..e183e63 --- /dev/null +++ b/Nemotron-Cascade-14B-Thinking.i1-IQ3_XXS.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:046c9e0c9bcfa0606f617ee9dad82b040b30954eaae5b10573da7eb34d389776 +size 5942667296 diff --git a/Nemotron-Cascade-14B-Thinking.i1-IQ4_NL.gguf b/Nemotron-Cascade-14B-Thinking.i1-IQ4_NL.gguf new file mode 100644 index 0000000..0960011 --- /dev/null +++ b/Nemotron-Cascade-14B-Thinking.i1-IQ4_NL.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:547d37697a0d062d01aab46e271652b131f607d3fd2c1e80bab2e39fddb7f4bc +size 8541364256 diff --git a/Nemotron-Cascade-14B-Thinking.i1-IQ4_XS.gguf b/Nemotron-Cascade-14B-Thinking.i1-IQ4_XS.gguf new file mode 100644 index 0000000..55df56c --- /dev/null +++ b/Nemotron-Cascade-14B-Thinking.i1-IQ4_XS.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:8ca7e4998ae9bbc2e96e2dd88df42dbe68992bef66f20cde6e5a802e982727d0 +size 8110731296 diff --git a/Nemotron-Cascade-14B-Thinking.i1-Q2_K.gguf b/Nemotron-Cascade-14B-Thinking.i1-Q2_K.gguf new file mode 100644 index 0000000..4d171ce --- /dev/null +++ b/Nemotron-Cascade-14B-Thinking.i1-Q2_K.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:d66f213258fa3d04f185b562b4a30c49e93fcfaa69537995df7938d0e6bf98ad +size 5753985056 diff --git a/Nemotron-Cascade-14B-Thinking.i1-Q2_K_S.gguf b/Nemotron-Cascade-14B-Thinking.i1-Q2_K_S.gguf new file mode 100644 index 0000000..3748de6 --- /dev/null +++ b/Nemotron-Cascade-14B-Thinking.i1-Q2_K_S.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:3c87c4d9f76a48352f24b88e31214d6ba76c3aefd741bb19e58db2ec3e5314c1 +size 5389850656 diff --git a/Nemotron-Cascade-14B-Thinking.i1-Q3_K_L.gguf b/Nemotron-Cascade-14B-Thinking.i1-Q3_K_L.gguf new file mode 100644 index 0000000..d5da39b --- /dev/null +++ b/Nemotron-Cascade-14B-Thinking.i1-Q3_K_L.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:7d9425fae4716edb132982f73dbaa6c9188621b381d04e61a5da641400d1d66c +size 7900652576 diff --git a/Nemotron-Cascade-14B-Thinking.i1-Q3_K_M.gguf b/Nemotron-Cascade-14B-Thinking.i1-Q3_K_M.gguf new file mode 100644 index 0000000..d797060 --- /dev/null +++ b/Nemotron-Cascade-14B-Thinking.i1-Q3_K_M.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:d8c74bb1eee97f643aac70c438dbfcebe905562bf3a003199d08fc653d196046 +size 7321314336 diff --git a/Nemotron-Cascade-14B-Thinking.i1-Q3_K_S.gguf b/Nemotron-Cascade-14B-Thinking.i1-Q3_K_S.gguf new file mode 100644 index 0000000..c329a29 --- /dev/null +++ b/Nemotron-Cascade-14B-Thinking.i1-Q3_K_S.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:70ec3317ff0fac1a4cdc36426eede063987f60f9212fc8891938d17713da011a +size 6657106976 diff --git a/Nemotron-Cascade-14B-Thinking.i1-Q4_0.gguf b/Nemotron-Cascade-14B-Thinking.i1-Q4_0.gguf new file mode 100644 index 0000000..e2fc61c --- /dev/null +++ b/Nemotron-Cascade-14B-Thinking.i1-Q4_0.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:05de601a2520cae9a187ecfcf912a631841b211c0021db0df2d158dbfae15396 +size 8543002656 diff --git a/Nemotron-Cascade-14B-Thinking.i1-Q4_1.gguf b/Nemotron-Cascade-14B-Thinking.i1-Q4_1.gguf new file mode 100644 index 0000000..0f12adb --- /dev/null +++ b/Nemotron-Cascade-14B-Thinking.i1-Q4_1.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:47c7fdfcc3f410875e7c1b84ecaa1b126205d10464aba816d9bf9c4ef8ebaa9e +size 9389522976 diff --git a/Nemotron-Cascade-14B-Thinking.i1-Q4_K_M.gguf b/Nemotron-Cascade-14B-Thinking.i1-Q4_K_M.gguf new file mode 100644 index 0000000..5c5e88d --- /dev/null +++ b/Nemotron-Cascade-14B-Thinking.i1-Q4_K_M.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:d9f170d7ce62ea3d87b545943f90d6e7985af6bc50b9317d259ff5608465e523 +size 9001754656 diff --git a/Nemotron-Cascade-14B-Thinking.i1-Q4_K_S.gguf b/Nemotron-Cascade-14B-Thinking.i1-Q4_K_S.gguf new file mode 100644 index 0000000..124b483 --- /dev/null +++ b/Nemotron-Cascade-14B-Thinking.i1-Q4_K_S.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:f8e54a110f05408721eeedc1f67ebbbf4e3e6e1c662379e558645872ebdf9702 +size 8573476896 diff --git a/Nemotron-Cascade-14B-Thinking.i1-Q5_K_M.gguf b/Nemotron-Cascade-14B-Thinking.i1-Q5_K_M.gguf new file mode 100644 index 0000000..efc5c4f --- /dev/null +++ b/Nemotron-Cascade-14B-Thinking.i1-Q5_K_M.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:cf006d06fe86280138448e85d23edf4f54445c1a92251c52176f747a1e67bf81 +size 10514571296 diff --git a/Nemotron-Cascade-14B-Thinking.i1-Q5_K_S.gguf b/Nemotron-Cascade-14B-Thinking.i1-Q5_K_S.gguf new file mode 100644 index 0000000..e2de61e --- /dev/null +++ b/Nemotron-Cascade-14B-Thinking.i1-Q5_K_S.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:cfe289c2ebcaf31f4a4fb8d7d0ebe8aa3dc39895972177405615dcaabba8ad49 +size 10263896096 diff --git a/Nemotron-Cascade-14B-Thinking.i1-Q6_K.gguf b/Nemotron-Cascade-14B-Thinking.i1-Q6_K.gguf new file mode 100644 index 0000000..8c13df6 --- /dev/null +++ b/Nemotron-Cascade-14B-Thinking.i1-Q6_K.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:b523a9d5dfc92948bfa88886fcde58846e035a24892942a799164f2f317073af +size 12121938976 diff --git a/Nemotron-Cascade-14B-Thinking.imatrix.gguf b/Nemotron-Cascade-14B-Thinking.imatrix.gguf new file mode 100644 index 0000000..63d64f4 --- /dev/null +++ b/Nemotron-Cascade-14B-Thinking.imatrix.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:d3b2c403768851bcf36e07fad99d74dc7956115b84ae4d0a6c7523dc237df3d0 +size 7743552 diff --git a/README.md b/README.md new file mode 100644 index 0000000..0503bd8 --- /dev/null +++ b/README.md @@ -0,0 +1,95 @@ +--- +base_model: nvidia/Nemotron-Cascade-14B-Thinking +language: +- en +library_name: transformers +license: other +license_link: https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/ +license_name: nvidia-open-model-license +mradermacher: + readme_rev: 1 +quantized_by: mradermacher +tags: +- nvidia +- nemotron-cascade +- reasoning +- general-purpose +- SFT +- RL +- pytorch +--- +## About + + + + + + + + + +weighted/imatrix quants of https://huggingface.co/nvidia/Nemotron-Cascade-14B-Thinking + + + +***For a convenient overview and download list, visit our [model page for this model](https://hf.tst.eu/model#Nemotron-Cascade-14B-Thinking-i1-GGUF).*** + +static quants are available at https://huggingface.co/mradermacher/Nemotron-Cascade-14B-Thinking-GGUF +## Usage + +If you are unsure how to use GGUF files, refer to one of [TheBloke's +READMEs](https://huggingface.co/TheBloke/KafkaLM-70B-German-V0.1-GGUF) for +more details, including on how to concatenate multi-part files. + +## Provided Quants + +(sorted by size, not necessarily quality. IQ-quants are often preferable over similar sized non-IQ quants) + +| Link | Type | Size/GB | Notes | +|:-----|:-----|--------:|:------| +| [GGUF](https://huggingface.co/mradermacher/Nemotron-Cascade-14B-Thinking-i1-GGUF/resolve/main/Nemotron-Cascade-14B-Thinking.imatrix.gguf) | imatrix | 0.1 | imatrix file (for creating your own quants) | +| [GGUF](https://huggingface.co/mradermacher/Nemotron-Cascade-14B-Thinking-i1-GGUF/resolve/main/Nemotron-Cascade-14B-Thinking.i1-IQ1_S.gguf) | i1-IQ1_S | 3.7 | for the desperate | +| [GGUF](https://huggingface.co/mradermacher/Nemotron-Cascade-14B-Thinking-i1-GGUF/resolve/main/Nemotron-Cascade-14B-Thinking.i1-IQ1_M.gguf) | i1-IQ1_M | 3.9 | mostly desperate | +| [GGUF](https://huggingface.co/mradermacher/Nemotron-Cascade-14B-Thinking-i1-GGUF/resolve/main/Nemotron-Cascade-14B-Thinking.i1-IQ2_XXS.gguf) | i1-IQ2_XXS | 4.4 | | +| [GGUF](https://huggingface.co/mradermacher/Nemotron-Cascade-14B-Thinking-i1-GGUF/resolve/main/Nemotron-Cascade-14B-Thinking.i1-IQ2_XS.gguf) | i1-IQ2_XS | 4.8 | | +| [GGUF](https://huggingface.co/mradermacher/Nemotron-Cascade-14B-Thinking-i1-GGUF/resolve/main/Nemotron-Cascade-14B-Thinking.i1-IQ2_S.gguf) | i1-IQ2_S | 5.1 | | +| [GGUF](https://huggingface.co/mradermacher/Nemotron-Cascade-14B-Thinking-i1-GGUF/resolve/main/Nemotron-Cascade-14B-Thinking.i1-IQ2_M.gguf) | i1-IQ2_M | 5.4 | | +| [GGUF](https://huggingface.co/mradermacher/Nemotron-Cascade-14B-Thinking-i1-GGUF/resolve/main/Nemotron-Cascade-14B-Thinking.i1-Q2_K_S.gguf) | i1-Q2_K_S | 5.5 | very low quality | +| [GGUF](https://huggingface.co/mradermacher/Nemotron-Cascade-14B-Thinking-i1-GGUF/resolve/main/Nemotron-Cascade-14B-Thinking.i1-Q2_K.gguf) | i1-Q2_K | 5.9 | IQ3_XXS probably better | +| [GGUF](https://huggingface.co/mradermacher/Nemotron-Cascade-14B-Thinking-i1-GGUF/resolve/main/Nemotron-Cascade-14B-Thinking.i1-IQ3_XXS.gguf) | i1-IQ3_XXS | 6.0 | lower quality | +| [GGUF](https://huggingface.co/mradermacher/Nemotron-Cascade-14B-Thinking-i1-GGUF/resolve/main/Nemotron-Cascade-14B-Thinking.i1-IQ3_XS.gguf) | i1-IQ3_XS | 6.5 | | +| [GGUF](https://huggingface.co/mradermacher/Nemotron-Cascade-14B-Thinking-i1-GGUF/resolve/main/Nemotron-Cascade-14B-Thinking.i1-Q3_K_S.gguf) | i1-Q3_K_S | 6.8 | IQ3_XS probably better | +| [GGUF](https://huggingface.co/mradermacher/Nemotron-Cascade-14B-Thinking-i1-GGUF/resolve/main/Nemotron-Cascade-14B-Thinking.i1-IQ3_S.gguf) | i1-IQ3_S | 6.8 | beats Q3_K* | +| [GGUF](https://huggingface.co/mradermacher/Nemotron-Cascade-14B-Thinking-i1-GGUF/resolve/main/Nemotron-Cascade-14B-Thinking.i1-IQ3_M.gguf) | i1-IQ3_M | 7.0 | | +| [GGUF](https://huggingface.co/mradermacher/Nemotron-Cascade-14B-Thinking-i1-GGUF/resolve/main/Nemotron-Cascade-14B-Thinking.i1-Q3_K_M.gguf) | i1-Q3_K_M | 7.4 | IQ3_S probably better | +| [GGUF](https://huggingface.co/mradermacher/Nemotron-Cascade-14B-Thinking-i1-GGUF/resolve/main/Nemotron-Cascade-14B-Thinking.i1-Q3_K_L.gguf) | i1-Q3_K_L | 8.0 | IQ3_M probably better | +| [GGUF](https://huggingface.co/mradermacher/Nemotron-Cascade-14B-Thinking-i1-GGUF/resolve/main/Nemotron-Cascade-14B-Thinking.i1-IQ4_XS.gguf) | i1-IQ4_XS | 8.2 | | +| [GGUF](https://huggingface.co/mradermacher/Nemotron-Cascade-14B-Thinking-i1-GGUF/resolve/main/Nemotron-Cascade-14B-Thinking.i1-IQ4_NL.gguf) | i1-IQ4_NL | 8.6 | prefer IQ4_XS | +| [GGUF](https://huggingface.co/mradermacher/Nemotron-Cascade-14B-Thinking-i1-GGUF/resolve/main/Nemotron-Cascade-14B-Thinking.i1-Q4_0.gguf) | i1-Q4_0 | 8.6 | fast, low quality | +| [GGUF](https://huggingface.co/mradermacher/Nemotron-Cascade-14B-Thinking-i1-GGUF/resolve/main/Nemotron-Cascade-14B-Thinking.i1-Q4_K_S.gguf) | i1-Q4_K_S | 8.7 | optimal size/speed/quality | +| [GGUF](https://huggingface.co/mradermacher/Nemotron-Cascade-14B-Thinking-i1-GGUF/resolve/main/Nemotron-Cascade-14B-Thinking.i1-Q4_K_M.gguf) | i1-Q4_K_M | 9.1 | fast, recommended | +| [GGUF](https://huggingface.co/mradermacher/Nemotron-Cascade-14B-Thinking-i1-GGUF/resolve/main/Nemotron-Cascade-14B-Thinking.i1-Q4_1.gguf) | i1-Q4_1 | 9.5 | | +| [GGUF](https://huggingface.co/mradermacher/Nemotron-Cascade-14B-Thinking-i1-GGUF/resolve/main/Nemotron-Cascade-14B-Thinking.i1-Q5_K_S.gguf) | i1-Q5_K_S | 10.4 | | +| [GGUF](https://huggingface.co/mradermacher/Nemotron-Cascade-14B-Thinking-i1-GGUF/resolve/main/Nemotron-Cascade-14B-Thinking.i1-Q5_K_M.gguf) | i1-Q5_K_M | 10.6 | | +| [GGUF](https://huggingface.co/mradermacher/Nemotron-Cascade-14B-Thinking-i1-GGUF/resolve/main/Nemotron-Cascade-14B-Thinking.i1-Q6_K.gguf) | i1-Q6_K | 12.2 | practically like static Q6_K | + +Here is a handy graph by ikawrakow comparing some lower-quality quant +types (lower is better): + +![image.png](https://www.nethype.de/huggingface_embed/quantpplgraph.png) + +And here are Artefact2's thoughts on the matter: +https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9 + +## FAQ / Model Request + +See https://huggingface.co/mradermacher/model_requests for some answers to +questions you might have and/or if you want some other model quantized. + +## Thanks + +I thank my company, [nethype GmbH](https://www.nethype.de/), for letting +me use its servers and providing upgrades to my workstation to enable +this work in my free time. Additional thanks to [@nicoboss](https://huggingface.co/nicoboss) for giving me access to his private supercomputer, enabling me to provide many more imatrix quants, at much higher quality, than I would otherwise be able to. + +