初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Ministrations-8B-v1-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-18 03:36:09 +08:00
commit 54c2e5e995
30 changed files with 267 additions and 0 deletions

62
.gitattributes vendored Normal file
View File

@@ -0,0 +1,62 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Ministrations-8B-v1-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Ministrations-8B-v1-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Ministrations-8B-v1-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Ministrations-8B-v1-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Ministrations-8B-v1-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Ministrations-8B-v1-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Ministrations-8B-v1-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Ministrations-8B-v1-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Ministrations-8B-v1-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Ministrations-8B-v1-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
Ministrations-8B-v1-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
Ministrations-8B-v1-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
Ministrations-8B-v1-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Ministrations-8B-v1-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Ministrations-8B-v1-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Ministrations-8B-v1-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Ministrations-8B-v1-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Ministrations-8B-v1-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Ministrations-8B-v1-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Ministrations-8B-v1-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Ministrations-8B-v1-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Ministrations-8B-v1-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Ministrations-8B-v1-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Ministrations-8B-v1-IQ2_S.gguf filter=lfs diff=lfs merge=lfs -text
Ministrations-8B-v1-IQ2_XS.gguf filter=lfs diff=lfs merge=lfs -text
Ministrations-8B-v1-f16.gguf filter=lfs diff=lfs merge=lfs -text
Ministrations-8B-v1.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f004c7be9b01e70cc7733043a97a16862c901cdbc63dd27974cabd03fe997fbc
size 2958330496

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:50087005a64f353b90731915bdcf372814d90a04405d22fc495eb6567f241d4b
size 2771159680

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3216dca82e06caf6ebae40d58f826ea26de44fc2c5ae7f709121fb6e36079d22
size 2611251840

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6e3b52490fdf48c9a5b2df72958086c8bfa77638b45cbc0f28463b2d42f164c3
size 3791686272

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:23d0c3f5573b890d9cd0449d36901cfca8f0f5b7e8ff65289ddd91547727a836
size 3521940096

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ed6b0e31977b06a0fb9c9fcf073516c8194ac31c4517e8785b89fe3e12070fd8
size 4448225920

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5c56053b31c0936b37a2be6b6b3b07a3b44f73102f597d7f676a3652bb3b5ebe
size 3185478272

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4f101ea9ce6c4a97c0618e1ef0cf82bec273705e3b107a83a17b07603768b664
size 3709766272

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4f849e2c5fbbec2e16f43fff9d4339de8dc4a0118493ff43d8bebf2b1b6bfb70
size 4326460032

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ee8d3ded0c03181e6cff009e04e1d6abae050648d14984e3334ba8ec4c2ae6cf
size 4019227264

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:651c0b9341be1ed6f7e74e75a21841cc4a9eea25963a663490257394cff3506c
size 3664677504

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:165db94c263b89c75c9f6b198696f0ffcc8c126aaffdb342ce388f2584069232
size 4796222080

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:291b11baf5e18f1bc3e6f563053b86942608989cb9f283d382b314b980ac51de
size 4671048320

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b00161afbc3f44e2acfd6ccf4107dcd7aa1329a92dae38a96ceedcb60d97ba79
size 4658465408

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3c9355804e7d9150dfd226b94a1ede64574ab6839d9b178ef3459e86bab60db9
size 4658465408

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9f470642cd5b6d68b014d286cc6179edc7f012e9d9461ca360fcefe0aebb5eb6
size 4658465408

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1949c9bc233846a30b2d4429b81115fcc0e55851538c4c7302d81b45579074a7
size 5309958784

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:172adc75000a14cdc85e0b3f9ea34fb3f26ddc3ace9a397ce7957cc8cdc2b1a7
size 4911499904

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:16b312d691b1aaeed1de2a947900af6921cad4f3de21396213021b726082fb7a
size 4685728384

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:86492d3d91da19bf7aeb930849699874571071ff51a2394266ec14a923e86491
size 6055496320

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1d0d9a83e16bc476839c3900e88df58dbd3437d5317dfe5d8c357204046a8040
size 5724146304

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5ce8c819a2ee971835847eab1c7866e09b4405836f22b515308f95464baed273
size 5593795200

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0cad9c66890a32a29f42b6f59af19f4d19811d2baab60c9a4bdb57836ffa4c1c
size 6587583104

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2c1e4d327902982c446e23db8295da8dee55ba8cd27459105767128e1f7f261a
size 6847629952

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cbd9d4edc21d014971c0e76524e408ee671b70341e3eedd81f40772980055f8f
size 8529808000

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9d2ec59bbea16402e68534bec7f6795c22407bb17cc60306d59a71732acdf805
size 16048097632

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d1147228287650d101380b6d2735d10b92e0bd4a6293ad01d818f3b11c284638
size 5316782

123
README.md Normal file
View File

@@ -0,0 +1,123 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
license: other
base_model: TheDrummer/Ministrations-8B-v1
---
## Llamacpp imatrix Quantizations of Ministrations-8B-v1
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4014">b4014</a> for quantization.
Original model: https://huggingface.co/TheDrummer/Ministrations-8B-v1
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<s>[INST]{prompt}[/INST]
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Ministrations-8B-v1-f16.gguf](https://huggingface.co/bartowski/Ministrations-8B-v1-GGUF/blob/main/Ministrations-8B-v1-f16.gguf) | f16 | 16.05GB | false | Full F16 weights. |
| [Ministrations-8B-v1-Q8_0.gguf](https://huggingface.co/bartowski/Ministrations-8B-v1-GGUF/blob/main/Ministrations-8B-v1-Q8_0.gguf) | Q8_0 | 8.53GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Ministrations-8B-v1-Q6_K_L.gguf](https://huggingface.co/bartowski/Ministrations-8B-v1-GGUF/blob/main/Ministrations-8B-v1-Q6_K_L.gguf) | Q6_K_L | 6.85GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Ministrations-8B-v1-Q6_K.gguf](https://huggingface.co/bartowski/Ministrations-8B-v1-GGUF/blob/main/Ministrations-8B-v1-Q6_K.gguf) | Q6_K | 6.59GB | false | Very high quality, near perfect, *recommended*. |
| [Ministrations-8B-v1-Q5_K_L.gguf](https://huggingface.co/bartowski/Ministrations-8B-v1-GGUF/blob/main/Ministrations-8B-v1-Q5_K_L.gguf) | Q5_K_L | 6.06GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Ministrations-8B-v1-Q5_K_M.gguf](https://huggingface.co/bartowski/Ministrations-8B-v1-GGUF/blob/main/Ministrations-8B-v1-Q5_K_M.gguf) | Q5_K_M | 5.72GB | false | High quality, *recommended*. |
| [Ministrations-8B-v1-Q5_K_S.gguf](https://huggingface.co/bartowski/Ministrations-8B-v1-GGUF/blob/main/Ministrations-8B-v1-Q5_K_S.gguf) | Q5_K_S | 5.59GB | false | High quality, *recommended*. |
| [Ministrations-8B-v1-Q4_K_L.gguf](https://huggingface.co/bartowski/Ministrations-8B-v1-GGUF/blob/main/Ministrations-8B-v1-Q4_K_L.gguf) | Q4_K_L | 5.31GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Ministrations-8B-v1-Q4_K_M.gguf](https://huggingface.co/bartowski/Ministrations-8B-v1-GGUF/blob/main/Ministrations-8B-v1-Q4_K_M.gguf) | Q4_K_M | 4.91GB | false | Good quality, default size for must use cases, *recommended*. |
| [Ministrations-8B-v1-Q3_K_XL.gguf](https://huggingface.co/bartowski/Ministrations-8B-v1-GGUF/blob/main/Ministrations-8B-v1-Q3_K_XL.gguf) | Q3_K_XL | 4.80GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Ministrations-8B-v1-Q4_K_S.gguf](https://huggingface.co/bartowski/Ministrations-8B-v1-GGUF/blob/main/Ministrations-8B-v1-Q4_K_S.gguf) | Q4_K_S | 4.69GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Ministrations-8B-v1-Q4_0.gguf](https://huggingface.co/bartowski/Ministrations-8B-v1-GGUF/blob/main/Ministrations-8B-v1-Q4_0.gguf) | Q4_0 | 4.67GB | false | Legacy format, generally not worth using over similarly sized formats |
| [Ministrations-8B-v1-Q4_0_8_8.gguf](https://huggingface.co/bartowski/Ministrations-8B-v1-GGUF/blob/main/Ministrations-8B-v1-Q4_0_8_8.gguf) | Q4_0_8_8 | 4.66GB | false | Optimized for ARM inference. Requires 'sve' support (see link below). *Don't use on Mac or Windows*. |
| [Ministrations-8B-v1-Q4_0_4_8.gguf](https://huggingface.co/bartowski/Ministrations-8B-v1-GGUF/blob/main/Ministrations-8B-v1-Q4_0_4_8.gguf) | Q4_0_4_8 | 4.66GB | false | Optimized for ARM inference. Requires 'i8mm' support (see link below). *Don't use on Mac or Windows*. |
| [Ministrations-8B-v1-Q4_0_4_4.gguf](https://huggingface.co/bartowski/Ministrations-8B-v1-GGUF/blob/main/Ministrations-8B-v1-Q4_0_4_4.gguf) | Q4_0_4_4 | 4.66GB | false | Optimized for ARM inference. Should work well on all ARM chips, pick this if you're unsure. *Don't use on Mac or Windows*. |
| [Ministrations-8B-v1-IQ4_XS.gguf](https://huggingface.co/bartowski/Ministrations-8B-v1-GGUF/blob/main/Ministrations-8B-v1-IQ4_XS.gguf) | IQ4_XS | 4.45GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Ministrations-8B-v1-Q3_K_L.gguf](https://huggingface.co/bartowski/Ministrations-8B-v1-GGUF/blob/main/Ministrations-8B-v1-Q3_K_L.gguf) | Q3_K_L | 4.33GB | false | Lower quality but usable, good for low RAM availability. |
| [Ministrations-8B-v1-Q3_K_M.gguf](https://huggingface.co/bartowski/Ministrations-8B-v1-GGUF/blob/main/Ministrations-8B-v1-Q3_K_M.gguf) | Q3_K_M | 4.02GB | false | Low quality. |
| [Ministrations-8B-v1-IQ3_M.gguf](https://huggingface.co/bartowski/Ministrations-8B-v1-GGUF/blob/main/Ministrations-8B-v1-IQ3_M.gguf) | IQ3_M | 3.79GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Ministrations-8B-v1-Q2_K_L.gguf](https://huggingface.co/bartowski/Ministrations-8B-v1-GGUF/blob/main/Ministrations-8B-v1-Q2_K_L.gguf) | Q2_K_L | 3.71GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Ministrations-8B-v1-Q3_K_S.gguf](https://huggingface.co/bartowski/Ministrations-8B-v1-GGUF/blob/main/Ministrations-8B-v1-Q3_K_S.gguf) | Q3_K_S | 3.66GB | false | Low quality, not recommended. |
| [Ministrations-8B-v1-IQ3_XS.gguf](https://huggingface.co/bartowski/Ministrations-8B-v1-GGUF/blob/main/Ministrations-8B-v1-IQ3_XS.gguf) | IQ3_XS | 3.52GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Ministrations-8B-v1-Q2_K.gguf](https://huggingface.co/bartowski/Ministrations-8B-v1-GGUF/blob/main/Ministrations-8B-v1-Q2_K.gguf) | Q2_K | 3.19GB | false | Very low quality but surprisingly usable. |
| [Ministrations-8B-v1-IQ2_M.gguf](https://huggingface.co/bartowski/Ministrations-8B-v1-GGUF/blob/main/Ministrations-8B-v1-IQ2_M.gguf) | IQ2_M | 2.96GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
| [Ministrations-8B-v1-IQ2_S.gguf](https://huggingface.co/bartowski/Ministrations-8B-v1-GGUF/blob/main/Ministrations-8B-v1-IQ2_S.gguf) | IQ2_S | 2.77GB | false | Low quality, uses SOTA techniques to be usable. |
| [Ministrations-8B-v1-IQ2_XS.gguf](https://huggingface.co/bartowski/Ministrations-8B-v1-GGUF/blob/main/Ministrations-8B-v1-IQ2_XS.gguf) | IQ2_XS | 2.61GB | false | Low quality, uses SOTA techniques to be usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Ministrations-8B-v1-GGUF --include "Ministrations-8B-v1-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Ministrations-8B-v1-GGUF --include "Ministrations-8B-v1-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Ministrations-8B-v1-Q8_0) or download them all in place (./)
## Q4_0_X_X
These are *NOT* for Metal (Apple) offloading, only ARM chips.
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset
Thank you ZeroWw for the inspiration to experiment with embed/output
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}