初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Ministral-8B-Instruct-2410-HF-GGUF-TEST
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-18 15:30:07 +08:00
commit 96f8c9f1b7
28 changed files with 259 additions and 0 deletions

60
.gitattributes vendored Normal file
View File

@@ -0,0 +1,60 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Ministral-8B-Instruct-2410-HF-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Ministral-8B-Instruct-2410-HF-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Ministral-8B-Instruct-2410-HF-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Ministral-8B-Instruct-2410-HF-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Ministral-8B-Instruct-2410-HF-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Ministral-8B-Instruct-2410-HF-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Ministral-8B-Instruct-2410-HF-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Ministral-8B-Instruct-2410-HF-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Ministral-8B-Instruct-2410-HF-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Ministral-8B-Instruct-2410-HF-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
Ministral-8B-Instruct-2410-HF-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
Ministral-8B-Instruct-2410-HF-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
Ministral-8B-Instruct-2410-HF-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Ministral-8B-Instruct-2410-HF-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Ministral-8B-Instruct-2410-HF-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Ministral-8B-Instruct-2410-HF-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Ministral-8B-Instruct-2410-HF-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Ministral-8B-Instruct-2410-HF-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Ministral-8B-Instruct-2410-HF-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Ministral-8B-Instruct-2410-HF-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Ministral-8B-Instruct-2410-HF-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Ministral-8B-Instruct-2410-HF-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Ministral-8B-Instruct-2410-HF-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Ministral-8B-Instruct-2410-HF-f16.gguf filter=lfs diff=lfs merge=lfs -text
Ministral-8B-Instruct-2410-HF.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0a6b6f689424e9abc062ae2e314fc698d8b9132a2e0d2bb7958a0c199b7ec502
size 2958326400

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:91e22bb67b1baa13e19f8a81f9acba24daa44ea2f1656a3f70e87ff334422d22
size 3791682176

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:63d09efdc9f32282b41ad2903e384dc37c3af30f33b231843e94e0672113acb9
size 3521936000

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ec07dba75d0ad447967afa413a57a3111a5f74b82517018a4e5424d3edde811b
size 4448221824

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2cd8a8cde5b27103c3b196cec3dd11a307e7cc68ad730ffca895e4424b72b416
size 3185474176

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:dad8ce3efd4f60e292204b26d83f2acac0f82927f5a89a949c7afbec8fd2af1e
size 3709762176

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:823dc2f68ca09652a917e76999a083e44b3c25d55eb261cac2e7c9ab7c6fd214
size 4326455936

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b5b2a655634faa13c1be514317ee1ae7548f1a4881192f842355652be7fcbd9a
size 4019223168

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d9845e38c6756704700e3542c3ccc2192ae3961650227ced40074dd8da3a0fd1
size 3664673408

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:026af07eabc931bf0c9412d06edf31c7cdf4c2dfc4ff240265ac396f342ba76b
size 4796217984

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cd2be2c865893447e4265162e442e1451ed5944a4eaf7f7a2387ab4b5a5e35da
size 4671044224

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0cdcf89653d808b31eea86913916ad9fd7c0899f9dbdd3fec97f127477148d6a
size 4658461312

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f8654c62b84525d3ee3997bf609c98fa0bff724ef13321ff40c7a4d6c22a40a2
size 4658461312

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7cb0c106ae0a27ecea9fca670c3918b42984be42cab98aa0be36c21a0bc35f02
size 4658461312

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2d3746aa3175ad70902b4f885eb6a6a475eec1bcd60048dc7820bc9c70ba88cf
size 5309954688

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4968abc6acadc8eeb15577c90d9a80aa959ea6e84e7db3d8d3dcacea17e5da6e
size 4911495808

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:05fb8467e9ec82f4c55d4eca57eb96ded605ec93822f2e9ab1f0e88d67c33d4b
size 4685724288

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:48229a1e0341594e64bdea9445f05ae6b25b33cecf8a2d08ef6198f82719c238
size 6055492224

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7df4404411055f30922c09af49e22c96503c4ec073758d696a9ee8d59bf98656
size 5724142208

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b2c2ab20a4cf33395f72a9b11c2eabe867b927417b0ecc2b5abc7ea581be98b2
size 5593791104

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:66b38cbcbe8ff57488c4b671750f548d5f176784c43842a5c06d838c880a9c9b
size 6587579008

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f11d6026cac8f45689773185c8ca7e9d08a5e93eac74082dba3a1558def5b44b
size 6847625856

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3ce9bd95fdbd6eca354b9eee1519cfbe2e5fd7f49b410d9a5833bcc28eb56a71
size 8529803904

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:744526a87fe16da254edc46bd1dd1aaddde51d70d3966276ebc8ffdb12afe641
size 16048093536

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:711041574b3260831ebdb06fe43e5d4bfc8234611fae33364135906382a00d11
size 5316782

123
README.md Normal file
View File

@@ -0,0 +1,123 @@
---
base_model: prince-canuma/Ministral-8B-Instruct-2410-HF
pipeline_tag: text-generation
tags: []
quantized_by: bartowski
---
## Llamacpp imatrix Quantizations of Ministral-8B-Instruct-2410-HF
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b3901">b3901</a> for quantization.
Original model: https://huggingface.co/prince-canuma/Ministral-8B-Instruct-2410-HF
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
### Warning: These are based on an unverified conversion and before finalized llama.cpp support. If you still see this message, know that you may have issues.
## Prompt format
```
<s>[INST] {prompt}[/INST] </s>
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Ministral-8B-Instruct-2410-HF-f16.gguf](https://huggingface.co/bartowski/Ministral-8B-Instruct-2410-HF-GGUF/blob/main/Ministral-8B-Instruct-2410-HF-f16.gguf) | f16 | 16.05GB | false | Full F16 weights. |
| [Ministral-8B-Instruct-2410-HF-Q8_0.gguf](https://huggingface.co/bartowski/Ministral-8B-Instruct-2410-HF-GGUF/blob/main/Ministral-8B-Instruct-2410-HF-Q8_0.gguf) | Q8_0 | 8.53GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Ministral-8B-Instruct-2410-HF-Q6_K_L.gguf](https://huggingface.co/bartowski/Ministral-8B-Instruct-2410-HF-GGUF/blob/main/Ministral-8B-Instruct-2410-HF-Q6_K_L.gguf) | Q6_K_L | 6.85GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Ministral-8B-Instruct-2410-HF-Q6_K.gguf](https://huggingface.co/bartowski/Ministral-8B-Instruct-2410-HF-GGUF/blob/main/Ministral-8B-Instruct-2410-HF-Q6_K.gguf) | Q6_K | 6.59GB | false | Very high quality, near perfect, *recommended*. |
| [Ministral-8B-Instruct-2410-HF-Q5_K_L.gguf](https://huggingface.co/bartowski/Ministral-8B-Instruct-2410-HF-GGUF/blob/main/Ministral-8B-Instruct-2410-HF-Q5_K_L.gguf) | Q5_K_L | 6.06GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Ministral-8B-Instruct-2410-HF-Q5_K_M.gguf](https://huggingface.co/bartowski/Ministral-8B-Instruct-2410-HF-GGUF/blob/main/Ministral-8B-Instruct-2410-HF-Q5_K_M.gguf) | Q5_K_M | 5.72GB | false | High quality, *recommended*. |
| [Ministral-8B-Instruct-2410-HF-Q5_K_S.gguf](https://huggingface.co/bartowski/Ministral-8B-Instruct-2410-HF-GGUF/blob/main/Ministral-8B-Instruct-2410-HF-Q5_K_S.gguf) | Q5_K_S | 5.59GB | false | High quality, *recommended*. |
| [Ministral-8B-Instruct-2410-HF-Q4_K_L.gguf](https://huggingface.co/bartowski/Ministral-8B-Instruct-2410-HF-GGUF/blob/main/Ministral-8B-Instruct-2410-HF-Q4_K_L.gguf) | Q4_K_L | 5.31GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Ministral-8B-Instruct-2410-HF-Q4_K_M.gguf](https://huggingface.co/bartowski/Ministral-8B-Instruct-2410-HF-GGUF/blob/main/Ministral-8B-Instruct-2410-HF-Q4_K_M.gguf) | Q4_K_M | 4.91GB | false | Good quality, default size for must use cases, *recommended*. |
| [Ministral-8B-Instruct-2410-HF-Q3_K_XL.gguf](https://huggingface.co/bartowski/Ministral-8B-Instruct-2410-HF-GGUF/blob/main/Ministral-8B-Instruct-2410-HF-Q3_K_XL.gguf) | Q3_K_XL | 4.80GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Ministral-8B-Instruct-2410-HF-Q4_K_S.gguf](https://huggingface.co/bartowski/Ministral-8B-Instruct-2410-HF-GGUF/blob/main/Ministral-8B-Instruct-2410-HF-Q4_K_S.gguf) | Q4_K_S | 4.69GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Ministral-8B-Instruct-2410-HF-Q4_0.gguf](https://huggingface.co/bartowski/Ministral-8B-Instruct-2410-HF-GGUF/blob/main/Ministral-8B-Instruct-2410-HF-Q4_0.gguf) | Q4_0 | 4.67GB | false | Legacy format, generally not worth using over similarly sized formats |
| [Ministral-8B-Instruct-2410-HF-Q4_0_8_8.gguf](https://huggingface.co/bartowski/Ministral-8B-Instruct-2410-HF-GGUF/blob/main/Ministral-8B-Instruct-2410-HF-Q4_0_8_8.gguf) | Q4_0_8_8 | 4.66GB | false | Optimized for ARM inference. Requires 'sve' support (see link below). *Don't use on Mac or Windows*. |
| [Ministral-8B-Instruct-2410-HF-Q4_0_4_8.gguf](https://huggingface.co/bartowski/Ministral-8B-Instruct-2410-HF-GGUF/blob/main/Ministral-8B-Instruct-2410-HF-Q4_0_4_8.gguf) | Q4_0_4_8 | 4.66GB | false | Optimized for ARM inference. Requires 'i8mm' support (see link below). *Don't use on Mac or Windows*. |
| [Ministral-8B-Instruct-2410-HF-Q4_0_4_4.gguf](https://huggingface.co/bartowski/Ministral-8B-Instruct-2410-HF-GGUF/blob/main/Ministral-8B-Instruct-2410-HF-Q4_0_4_4.gguf) | Q4_0_4_4 | 4.66GB | false | Optimized for ARM inference. Should work well on all ARM chips, pick this if you're unsure. *Don't use on Mac or Windows*. |
| [Ministral-8B-Instruct-2410-HF-IQ4_XS.gguf](https://huggingface.co/bartowski/Ministral-8B-Instruct-2410-HF-GGUF/blob/main/Ministral-8B-Instruct-2410-HF-IQ4_XS.gguf) | IQ4_XS | 4.45GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Ministral-8B-Instruct-2410-HF-Q3_K_L.gguf](https://huggingface.co/bartowski/Ministral-8B-Instruct-2410-HF-GGUF/blob/main/Ministral-8B-Instruct-2410-HF-Q3_K_L.gguf) | Q3_K_L | 4.33GB | false | Lower quality but usable, good for low RAM availability. |
| [Ministral-8B-Instruct-2410-HF-Q3_K_M.gguf](https://huggingface.co/bartowski/Ministral-8B-Instruct-2410-HF-GGUF/blob/main/Ministral-8B-Instruct-2410-HF-Q3_K_M.gguf) | Q3_K_M | 4.02GB | false | Low quality. |
| [Ministral-8B-Instruct-2410-HF-IQ3_M.gguf](https://huggingface.co/bartowski/Ministral-8B-Instruct-2410-HF-GGUF/blob/main/Ministral-8B-Instruct-2410-HF-IQ3_M.gguf) | IQ3_M | 3.79GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Ministral-8B-Instruct-2410-HF-Q2_K_L.gguf](https://huggingface.co/bartowski/Ministral-8B-Instruct-2410-HF-GGUF/blob/main/Ministral-8B-Instruct-2410-HF-Q2_K_L.gguf) | Q2_K_L | 3.71GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Ministral-8B-Instruct-2410-HF-Q3_K_S.gguf](https://huggingface.co/bartowski/Ministral-8B-Instruct-2410-HF-GGUF/blob/main/Ministral-8B-Instruct-2410-HF-Q3_K_S.gguf) | Q3_K_S | 3.66GB | false | Low quality, not recommended. |
| [Ministral-8B-Instruct-2410-HF-IQ3_XS.gguf](https://huggingface.co/bartowski/Ministral-8B-Instruct-2410-HF-GGUF/blob/main/Ministral-8B-Instruct-2410-HF-IQ3_XS.gguf) | IQ3_XS | 3.52GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Ministral-8B-Instruct-2410-HF-Q2_K.gguf](https://huggingface.co/bartowski/Ministral-8B-Instruct-2410-HF-GGUF/blob/main/Ministral-8B-Instruct-2410-HF-Q2_K.gguf) | Q2_K | 3.19GB | false | Very low quality but surprisingly usable. |
| [Ministral-8B-Instruct-2410-HF-IQ2_M.gguf](https://huggingface.co/bartowski/Ministral-8B-Instruct-2410-HF-GGUF/blob/main/Ministral-8B-Instruct-2410-HF-IQ2_M.gguf) | IQ2_M | 2.96GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Ministral-8B-Instruct-2410-HF-GGUF --include "Ministral-8B-Instruct-2410-HF-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Ministral-8B-Instruct-2410-HF-GGUF --include "Ministral-8B-Instruct-2410-HF-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Ministral-8B-Instruct-2410-HF-Q8_0) or download them all in place (./)
## Q4_0_X_X
These are *NOT* for Metal (Apple) offloading, only ARM chips.
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset
Thank you ZeroWw for the inspiration to experiment with embed/output
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}