初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Nemotron-Mini-4B-Instruct-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-22 02:10:06 +08:00
commit 64c20d73c3
22 changed files with 239 additions and 0 deletions

54
.gitattributes vendored Normal file
View File

@@ -0,0 +1,54 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Nemotron-Mini-4B-Instruct-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Nemotron-Mini-4B-Instruct-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Nemotron-Mini-4B-Instruct-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Nemotron-Mini-4B-Instruct-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Nemotron-Mini-4B-Instruct-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Nemotron-Mini-4B-Instruct-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Nemotron-Mini-4B-Instruct-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Nemotron-Mini-4B-Instruct-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Nemotron-Mini-4B-Instruct-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Nemotron-Mini-4B-Instruct-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
Nemotron-Mini-4B-Instruct-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
Nemotron-Mini-4B-Instruct-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
Nemotron-Mini-4B-Instruct-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Nemotron-Mini-4B-Instruct-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Nemotron-Mini-4B-Instruct-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Nemotron-Mini-4B-Instruct-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Nemotron-Mini-4B-Instruct-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Nemotron-Mini-4B-Instruct-f16.gguf filter=lfs diff=lfs merge=lfs -text
Nemotron-Mini-4B-Instruct.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:603beac8e7a39c74a16770c7f6bcee15b5010be0d152765e5d43075a1f886d94
size 2184092736

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9bcaa30cb040d54a4daa2c42d22750b3491136b94274948a0b124ad17cc9c470
size 2461260864

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f0d8719ccce6f1db4a4db64530ec6536b62bb0c433e2dc363ae5ed4c42d49ad4
size 2452954176

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:184962ffa73badcba10c6557e656f42ceeacc9ef93f8bbd6967f6926ed0148d0
size 3141082176

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ef4af9471ef1e2c538f05063ecc41387ef8c720ccad30f93e9caf8c0af5c6696
size 2574703680

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5987bf8d5ef5a9756a471edf9a64c735954d363d5b77be8dd0156802325d3504
size 2567625792

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7c64d479b7d4988a3450cb63c8ed0c2cea07fa283af147251d7729ecc2acc1ba
size 2567625792

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f258902b5425d0fa31284bc917f82b875eb246bed126c94595b61d4b8da49c16
size 2567625792

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:051edb8061a9ae18916e59b5ea575668caee3a11c23058eb2d4698ad049e831d
size 3281067072

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2bf02846dbd45e9580b9338b90505fb54e00d185713d4b7f06699e2d5d298e48
size 2697387072

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2321664dc50623f8c1a7ef800b515038392da492dcc67c10518d307265f97ad5
size 2583354432

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4662f7e017d61de3a011d6b8420ce75046997c90f99fe392f7e1478e59302e6c
size 3545308224

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6230395444f14a10e4675e4d625183ae3c094567ddfe757991c0909f0f21c2d8
size 3059932224

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:47b8be9f954aa13e9cbc96096308e56dd90243bf28568f0da43344609dfda922
size 2993085504

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3232032ab9c700437447441893be0735d523db659412fc0f6dc24d9a2a293a4d
size 3445136448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:78c70e1cd078576b57f8ba1dda6c626a24fcaf2a85331680eb1e7e7fbf2ec42b
size 3826064448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8ed6d148ded733d4401495e44dc834ffa79e8cfee6f44d8fb60808dddd78b8eb
size 4459928640

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a83fbc63c044779ed75ee1f0492e9507da83d470bc9f7953b84901b24c7b341d
size 8388156192

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:38f3594e99beb4bb6516a2e8efed23111a9e170ac872f183129d93730dd0cd5d
size 3152084

127
README.md Normal file
View File

@@ -0,0 +1,127 @@
---
base_model: nvidia/Nemotron-Mini-4B-Instruct
language:
- en
library_name: nemo
license: other
license_name: nvidia-community-model-license
license_link: https://huggingface.co/nvidia/Nemotron-Mini-4B-Instruct/blob/main/nvidia-community-model-license-aug2024.pdf
pipeline_tag: text-generation
quantized_by: bartowski
---
## Llamacpp imatrix Quantizations of Nemotron-Mini-4B-Instruct
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b3715">b3715</a> for quantization.
Original model: https://huggingface.co/nvidia/Nemotron-Mini-4B-Instruct
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<extra_id_0>System
{system_prompt}
<extra_id_1>User
{prompt}
<extra_id_1>Assistant
<extra_id_1>Assistant
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Nemotron-Mini-4B-Instruct-f16.gguf](https://huggingface.co/bartowski/Nemotron-Mini-4B-Instruct-GGUF/blob/main/Nemotron-Mini-4B-Instruct-f16.gguf) | f16 | 8.39GB | false | Full F16 weights. |
| [Nemotron-Mini-4B-Instruct-Q8_0.gguf](https://huggingface.co/bartowski/Nemotron-Mini-4B-Instruct-GGUF/blob/main/Nemotron-Mini-4B-Instruct-Q8_0.gguf) | Q8_0 | 4.46GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Nemotron-Mini-4B-Instruct-Q6_K_L.gguf](https://huggingface.co/bartowski/Nemotron-Mini-4B-Instruct-GGUF/blob/main/Nemotron-Mini-4B-Instruct-Q6_K_L.gguf) | Q6_K_L | 3.83GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Nemotron-Mini-4B-Instruct-Q5_K_L.gguf](https://huggingface.co/bartowski/Nemotron-Mini-4B-Instruct-GGUF/blob/main/Nemotron-Mini-4B-Instruct-Q5_K_L.gguf) | Q5_K_L | 3.55GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Nemotron-Mini-4B-Instruct-Q6_K.gguf](https://huggingface.co/bartowski/Nemotron-Mini-4B-Instruct-GGUF/blob/main/Nemotron-Mini-4B-Instruct-Q6_K.gguf) | Q6_K | 3.45GB | false | Very high quality, near perfect, *recommended*. |
| [Nemotron-Mini-4B-Instruct-Q4_K_L.gguf](https://huggingface.co/bartowski/Nemotron-Mini-4B-Instruct-GGUF/blob/main/Nemotron-Mini-4B-Instruct-Q4_K_L.gguf) | Q4_K_L | 3.28GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Nemotron-Mini-4B-Instruct-Q3_K_XL.gguf](https://huggingface.co/bartowski/Nemotron-Mini-4B-Instruct-GGUF/blob/main/Nemotron-Mini-4B-Instruct-Q3_K_XL.gguf) | Q3_K_XL | 3.14GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Nemotron-Mini-4B-Instruct-Q5_K_M.gguf](https://huggingface.co/bartowski/Nemotron-Mini-4B-Instruct-GGUF/blob/main/Nemotron-Mini-4B-Instruct-Q5_K_M.gguf) | Q5_K_M | 3.06GB | false | High quality, *recommended*. |
| [Nemotron-Mini-4B-Instruct-Q5_K_S.gguf](https://huggingface.co/bartowski/Nemotron-Mini-4B-Instruct-GGUF/blob/main/Nemotron-Mini-4B-Instruct-Q5_K_S.gguf) | Q5_K_S | 2.99GB | false | High quality, *recommended*. |
| [Nemotron-Mini-4B-Instruct-Q4_K_M.gguf](https://huggingface.co/bartowski/Nemotron-Mini-4B-Instruct-GGUF/blob/main/Nemotron-Mini-4B-Instruct-Q4_K_M.gguf) | Q4_K_M | 2.70GB | false | Good quality, default size for must use cases, *recommended*. |
| [Nemotron-Mini-4B-Instruct-Q4_K_S.gguf](https://huggingface.co/bartowski/Nemotron-Mini-4B-Instruct-GGUF/blob/main/Nemotron-Mini-4B-Instruct-Q4_K_S.gguf) | Q4_K_S | 2.58GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Nemotron-Mini-4B-Instruct-Q4_0_8_8.gguf](https://huggingface.co/bartowski/Nemotron-Mini-4B-Instruct-GGUF/blob/main/Nemotron-Mini-4B-Instruct-Q4_0_8_8.gguf) | Q4_0_8_8 | 2.57GB | false | Optimized for ARM inference. Requires 'sve' support (see link below). |
| [Nemotron-Mini-4B-Instruct-Q4_0_4_8.gguf](https://huggingface.co/bartowski/Nemotron-Mini-4B-Instruct-GGUF/blob/main/Nemotron-Mini-4B-Instruct-Q4_0_4_8.gguf) | Q4_0_4_8 | 2.57GB | false | Optimized for ARM inference. Requires 'i8mm' support (see link below). |
| [Nemotron-Mini-4B-Instruct-Q4_0_4_4.gguf](https://huggingface.co/bartowski/Nemotron-Mini-4B-Instruct-GGUF/blob/main/Nemotron-Mini-4B-Instruct-Q4_0_4_4.gguf) | Q4_0_4_4 | 2.57GB | false | Optimized for ARM inference. Should work well on all ARM chips, pick this if you're unsure. |
| [Nemotron-Mini-4B-Instruct-Q4_0.gguf](https://huggingface.co/bartowski/Nemotron-Mini-4B-Instruct-GGUF/blob/main/Nemotron-Mini-4B-Instruct-Q4_0.gguf) | Q4_0 | 2.57GB | false | Legacy format, generally not worth using over similarly sized formats |
| [Nemotron-Mini-4B-Instruct-IQ4_XS.gguf](https://huggingface.co/bartowski/Nemotron-Mini-4B-Instruct-GGUF/blob/main/Nemotron-Mini-4B-Instruct-IQ4_XS.gguf) | IQ4_XS | 2.46GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Nemotron-Mini-4B-Instruct-Q3_K_L.gguf](https://huggingface.co/bartowski/Nemotron-Mini-4B-Instruct-GGUF/blob/main/Nemotron-Mini-4B-Instruct-Q3_K_L.gguf) | Q3_K_L | 2.45GB | false | Lower quality but usable, good for low RAM availability. |
| [Nemotron-Mini-4B-Instruct-IQ3_M.gguf](https://huggingface.co/bartowski/Nemotron-Mini-4B-Instruct-GGUF/blob/main/Nemotron-Mini-4B-Instruct-IQ3_M.gguf) | IQ3_M | 2.18GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Nemotron-Mini-4B-Instruct-GGUF --include "Nemotron-Mini-4B-Instruct-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Nemotron-Mini-4B-Instruct-GGUF --include "Nemotron-Mini-4B-Instruct-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Nemotron-Mini-4B-Instruct-Q8_0) or download them all in place (./)
## Q4_0_X_X
These are *NOT* for Metal (Apple) offloading, only ARM chips.
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset
Thank you ZeroWw for the inspiration to experiment with embed/output
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}