初始化项目,由ModelHub XC社区提供模型

Model: bartowski/magnum-v4-12b-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-08 11:18:06 +08:00
commit 398dac6229
29 changed files with 267 additions and 0 deletions

61
.gitattributes vendored Normal file
View File

@@ -0,0 +1,61 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
magnum-v4-12b-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
magnum-v4-12b-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
magnum-v4-12b-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
magnum-v4-12b-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
magnum-v4-12b-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
magnum-v4-12b-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
magnum-v4-12b-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
magnum-v4-12b-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
magnum-v4-12b-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
magnum-v4-12b-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
magnum-v4-12b-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
magnum-v4-12b-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
magnum-v4-12b-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
magnum-v4-12b-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
magnum-v4-12b-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
magnum-v4-12b-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
magnum-v4-12b-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
magnum-v4-12b-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
magnum-v4-12b-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
magnum-v4-12b-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
magnum-v4-12b-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
magnum-v4-12b-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
magnum-v4-12b-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
magnum-v4-12b-IQ2_S.gguf filter=lfs diff=lfs merge=lfs -text
magnum-v4-12b-f16.gguf filter=lfs diff=lfs merge=lfs -text
magnum-v4-12b.imatrix filter=lfs diff=lfs merge=lfs -text

127
README.md Normal file
View File

@@ -0,0 +1,127 @@
---
base_model: anthracite-org/magnum-v4-12b
language:
- en
license: other
license_name: mrl
pipeline_tag: text-generation
tags:
- chat
quantized_by: bartowski
---
## Llamacpp imatrix Quantizations of magnum-v4-12b
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b3930">b3930</a> for quantization.
Original model: https://huggingface.co/anthracite-org/magnum-v4-12b
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<s>[INST]{prompt}[/INST]
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [magnum-v4-12b-f16.gguf](https://huggingface.co/bartowski/magnum-v4-12b-GGUF/blob/main/magnum-v4-12b-f16.gguf) | f16 | 24.50GB | false | Full F16 weights. |
| [magnum-v4-12b-Q8_0.gguf](https://huggingface.co/bartowski/magnum-v4-12b-GGUF/blob/main/magnum-v4-12b-Q8_0.gguf) | Q8_0 | 13.02GB | false | Extremely high quality, generally unneeded but max available quant. |
| [magnum-v4-12b-Q6_K_L.gguf](https://huggingface.co/bartowski/magnum-v4-12b-GGUF/blob/main/magnum-v4-12b-Q6_K_L.gguf) | Q6_K_L | 10.38GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [magnum-v4-12b-Q6_K.gguf](https://huggingface.co/bartowski/magnum-v4-12b-GGUF/blob/main/magnum-v4-12b-Q6_K.gguf) | Q6_K | 10.06GB | false | Very high quality, near perfect, *recommended*. |
| [magnum-v4-12b-Q5_K_L.gguf](https://huggingface.co/bartowski/magnum-v4-12b-GGUF/blob/main/magnum-v4-12b-Q5_K_L.gguf) | Q5_K_L | 9.14GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [magnum-v4-12b-Q5_K_M.gguf](https://huggingface.co/bartowski/magnum-v4-12b-GGUF/blob/main/magnum-v4-12b-Q5_K_M.gguf) | Q5_K_M | 8.73GB | false | High quality, *recommended*. |
| [magnum-v4-12b-Q5_K_S.gguf](https://huggingface.co/bartowski/magnum-v4-12b-GGUF/blob/main/magnum-v4-12b-Q5_K_S.gguf) | Q5_K_S | 8.52GB | false | High quality, *recommended*. |
| [magnum-v4-12b-Q4_K_L.gguf](https://huggingface.co/bartowski/magnum-v4-12b-GGUF/blob/main/magnum-v4-12b-Q4_K_L.gguf) | Q4_K_L | 7.98GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [magnum-v4-12b-Q4_K_M.gguf](https://huggingface.co/bartowski/magnum-v4-12b-GGUF/blob/main/magnum-v4-12b-Q4_K_M.gguf) | Q4_K_M | 7.48GB | false | Good quality, default size for must use cases, *recommended*. |
| [magnum-v4-12b-Q3_K_XL.gguf](https://huggingface.co/bartowski/magnum-v4-12b-GGUF/blob/main/magnum-v4-12b-Q3_K_XL.gguf) | Q3_K_XL | 7.15GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [magnum-v4-12b-Q4_K_S.gguf](https://huggingface.co/bartowski/magnum-v4-12b-GGUF/blob/main/magnum-v4-12b-Q4_K_S.gguf) | Q4_K_S | 7.12GB | false | Slightly lower quality with more space savings, *recommended*. |
| [magnum-v4-12b-Q4_0.gguf](https://huggingface.co/bartowski/magnum-v4-12b-GGUF/blob/main/magnum-v4-12b-Q4_0.gguf) | Q4_0 | 7.09GB | false | Legacy format, generally not worth using over similarly sized formats |
| [magnum-v4-12b-Q4_0_8_8.gguf](https://huggingface.co/bartowski/magnum-v4-12b-GGUF/blob/main/magnum-v4-12b-Q4_0_8_8.gguf) | Q4_0_8_8 | 7.07GB | false | Optimized for ARM inference. Requires 'sve' support (see link below). *Don't use on Mac or Windows*. |
| [magnum-v4-12b-Q4_0_4_8.gguf](https://huggingface.co/bartowski/magnum-v4-12b-GGUF/blob/main/magnum-v4-12b-Q4_0_4_8.gguf) | Q4_0_4_8 | 7.07GB | false | Optimized for ARM inference. Requires 'i8mm' support (see link below). *Don't use on Mac or Windows*. |
| [magnum-v4-12b-Q4_0_4_4.gguf](https://huggingface.co/bartowski/magnum-v4-12b-GGUF/blob/main/magnum-v4-12b-Q4_0_4_4.gguf) | Q4_0_4_4 | 7.07GB | false | Optimized for ARM inference. Should work well on all ARM chips, pick this if you're unsure. *Don't use on Mac or Windows*. |
| [magnum-v4-12b-IQ4_XS.gguf](https://huggingface.co/bartowski/magnum-v4-12b-GGUF/blob/main/magnum-v4-12b-IQ4_XS.gguf) | IQ4_XS | 6.74GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [magnum-v4-12b-Q3_K_L.gguf](https://huggingface.co/bartowski/magnum-v4-12b-GGUF/blob/main/magnum-v4-12b-Q3_K_L.gguf) | Q3_K_L | 6.56GB | false | Lower quality but usable, good for low RAM availability. |
| [magnum-v4-12b-Q3_K_M.gguf](https://huggingface.co/bartowski/magnum-v4-12b-GGUF/blob/main/magnum-v4-12b-Q3_K_M.gguf) | Q3_K_M | 6.08GB | false | Low quality. |
| [magnum-v4-12b-IQ3_M.gguf](https://huggingface.co/bartowski/magnum-v4-12b-GGUF/blob/main/magnum-v4-12b-IQ3_M.gguf) | IQ3_M | 5.72GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [magnum-v4-12b-Q3_K_S.gguf](https://huggingface.co/bartowski/magnum-v4-12b-GGUF/blob/main/magnum-v4-12b-Q3_K_S.gguf) | Q3_K_S | 5.53GB | false | Low quality, not recommended. |
| [magnum-v4-12b-Q2_K_L.gguf](https://huggingface.co/bartowski/magnum-v4-12b-GGUF/blob/main/magnum-v4-12b-Q2_K_L.gguf) | Q2_K_L | 5.45GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [magnum-v4-12b-IQ3_XS.gguf](https://huggingface.co/bartowski/magnum-v4-12b-GGUF/blob/main/magnum-v4-12b-IQ3_XS.gguf) | IQ3_XS | 5.31GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [magnum-v4-12b-Q2_K.gguf](https://huggingface.co/bartowski/magnum-v4-12b-GGUF/blob/main/magnum-v4-12b-Q2_K.gguf) | Q2_K | 4.79GB | false | Very low quality but surprisingly usable. |
| [magnum-v4-12b-IQ2_M.gguf](https://huggingface.co/bartowski/magnum-v4-12b-GGUF/blob/main/magnum-v4-12b-IQ2_M.gguf) | IQ2_M | 4.44GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
| [magnum-v4-12b-IQ2_S.gguf](https://huggingface.co/bartowski/magnum-v4-12b-GGUF/blob/main/magnum-v4-12b-IQ2_S.gguf) | IQ2_S | 4.14GB | false | Low quality, uses SOTA techniques to be usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/magnum-v4-12b-GGUF --include "magnum-v4-12b-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/magnum-v4-12b-GGUF --include "magnum-v4-12b-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (magnum-v4-12b-Q8_0) or download them all in place (./)
## Q4_0_X_X
These are *NOT* for Metal (Apple) offloading, only ARM chips.
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset
Thank you ZeroWw for the inspiration to experiment with embed/output
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}

3
magnum-v4-12b-IQ2_M.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8a11aefc53cb922ee42b79132d7fadae3b9a087b8c098af4f2ae3a2f246b7779
size 4435026816

3
magnum-v4-12b-IQ2_S.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7e667460accc80c006ff9abdabb1d2254bcc696e043d7b87cc732928892ec9fd
size 4138476416

3
magnum-v4-12b-IQ3_M.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:379030b7dde2baecae3889a988bb627bab5a0a84caf16823d296b5e79a67a53f
size 5722235776

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4c1a16dbccd0e3d8eb1dc8e152bd93bb50db1eed62139376312a8bb8a04185f8
size 5306491776

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:666bf759b2395cbf8feaba08dacb4325827c0c3c5b06cb2f764c7dbb659e2756
size 6742713216

3
magnum-v4-12b-Q2_K.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:986c73f7415ae8a2cc9abaa3804789cc815107e7232d3f16d14b31490003b16d
size 4791051136

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d513dd17bf9395645a8c0abd373f7f8e6bbed403f4132d315bc70697046d322e
size 5446411136

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:244bb8420ecbac574beb2178d445cbde5b200e830bb3a20e398f48c16d2d97aa
size 6561506176

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ee7012920bbcf94516092f0a73353e38c97605502f2121bf927cbbaa5929cd98
size 6083093376

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7c7278e8c31e55fe170f96e679c51ae97973b3b3caa12a92d16f110538663f67
size 5534229376

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a4b5e1b9c340a0839cf2be74249dd2bb71f615b0b7956362de1b566500116908
size 7148708736

3
magnum-v4-12b-Q4_0.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:29b5dc93c378f098e3ca5e6a8512d45f1177e5a30a7d62ad07ef23085fef4751
size 7094641536

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:57b7caa67026217a28a815912a079c2f6c94898210d92abd4e21f65f6fb3a427
size 7071703936

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c6a1fcbab2609a9cf2f6405ac7bc48d355bb08a5e365bf88633ee76678535d7a
size 7071703936

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6a83c3b0b933cc568bdc607ed70bdf05b56c5d2114a373be475529746fa49084
size 7071703936

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3fa092ba0685a741539884f0fb784b1e9dbea2652d62d0f60dfd9499b1ca3fe0
size 7975281536

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:08b1d35ff06c5033be8818759d4db51195017fc6b2e0b57cae77ccd23a2649cc
size 7477207936

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1e5037b55a481129f73531fb72ecf36e7d176a5c42372b895323ce81cdc2e6cd
size 7120200576

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d46583b873a13e8519261768958579e369f8fefa8add890b3803d8076ce66bbd
size 9141822336

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fcbb958ddbdcd942a7597ada6cb806ff547f281d85ad0e59e47d8911079e87ef
size 8727634816

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5a541491f698ae759d5a0636eea6d592c2a39907379bbe5eb6340b26f050c7df
size 8518738816

3
magnum-v4-12b-Q6_K.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a7281f0c00ecd42a248e2b1e17adcc6c9bf81abc9862490e947bc4c846308ba1
size 10056213376

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:770f761d9e4c0849d2f6ea0b23213c6dc305fcbec286c1b6bcda7abf7d1a7547
size 10381271936

3
magnum-v4-12b-Q8_0.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:32ce4a85f7d4e9bec073ac707c3f6a49df591b09fc6889dec0fad1c731b4abba
size 13022372736

3
magnum-v4-12b-f16.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9f5998c5f2e6c192c10c9ee944df5e302ec9e781ba5df29e8307cd112d9bc648
size 24504279680

3
magnum-v4-12b.imatrix Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b7cca5e5906baa8aa05899a71a4fd913edef2de21c53fa31cee8716ab722a81d
size 7054418