初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Prismatic-12b-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-23 20:57:13 +08:00
commit f12d566d73
29 changed files with 262 additions and 0 deletions

61
.gitattributes vendored Normal file
View File

@@ -0,0 +1,61 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Prismatic-12b-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Prismatic-12b-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Prismatic-12b-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Prismatic-12b-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Prismatic-12b-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Prismatic-12b-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Prismatic-12b-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Prismatic-12b-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Prismatic-12b-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Prismatic-12b-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
Prismatic-12b-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
Prismatic-12b-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
Prismatic-12b-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Prismatic-12b-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Prismatic-12b-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Prismatic-12b-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Prismatic-12b-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Prismatic-12b-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Prismatic-12b-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Prismatic-12b-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Prismatic-12b-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Prismatic-12b-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Prismatic-12b-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Prismatic-12b-IQ2_S.gguf filter=lfs diff=lfs merge=lfs -text
Prismatic-12b-f16.gguf filter=lfs diff=lfs merge=lfs -text
Prismatic-12b.imatrix filter=lfs diff=lfs merge=lfs -text

3
Prismatic-12b-IQ2_M.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3d69c902e22454c4d89b50fc5622858ef28bf4367feb2d3186df51eef350e184
size 4435022624

3
Prismatic-12b-IQ2_S.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f3fe0845e76e64b7a5942cadb69f97aa973533a959012583e34b97537c1c1f4f
size 4138472224

3
Prismatic-12b-IQ3_M.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4ca3d1742a9c5ce2727b75684cde34fd21ff2feab554ff146d42d38b736f0c12
size 5722231584

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8f596659ef64c88336a757b38c85733ab178238312ba9f35dbe8f95a1708a232
size 5306487584

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4c025caefec903fa4e9c4dac365087972649bb16c34cbd5b81b9601d592c5aec
size 6742709024

3
Prismatic-12b-Q2_K.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b3a5716c6b952d477ff2246dc6d36abee794df888f5613acaac5a5cc2c3d720c
size 4791046944

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2cf8563057076e0fb778066326bffaae460516cd8e05a646b83af450a4e0d46a
size 5446406944

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ed6d2433ad6c14e479cc432f475852bb71c03f64691d0897301c979d395755e3
size 6561501984

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5fc178465d3785c2c2702f4daca8b25cb3d3a8cb4a50baaa2eb5b29f002c2a7c
size 6083089184

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1a0333d062d1819f215603a6d43dd52a87d6228c49eb2ef51cc7acee86a1bf5c
size 5534225184

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d6f91d5d4f28de23c3cdee56aa8fb1f1b88a818333775096a619118854bbe3fc
size 7148704544

3
Prismatic-12b-Q4_0.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:064cafb73a114b1f465d8a61b70cbd402be622d4de2cbe5f1b3e4aa52a1dc1a2
size 7094637344

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:dc9f641fb7e45e521049cde0af672db15734ac2b9b5ad8f6ac2ad7c2fb6f3586
size 7071699744

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cc8e89ee2ca980f0a5c090cb3ab9cd079020ee172d0656c27dbc7cc21343e7d8
size 7071699744

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3f95d213370a0d72f30c8f6dd0488e118f449440069b4d644b5ac7bd18f73a69
size 7071699744

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c87743378fcbb197c94444e6314b5cdee2ed50a6a350727ad3cfcd279e85ab80
size 7975277344

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6211f5e0ebaacc2311bf16435c4a0471ebc6102903f1893c9ef0fb5b938905f4
size 7477203744

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:313948528af812f5812db2ad34fd447650df6865876d21e256a7d71ce76deaa6
size 7120196384

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fc462e74efe0fe529f3bfe84e5fcc93e297259cf6c19be50c41625a3c7625c3c
size 9141818144

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:741afb0678caadd2c223836376516e63b7450956bdb923b8a618d413e0bc931e
size 8727630624

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f837e9f84f69fce46a453292d75e5b0f9e8fcba3855921c39680d272af144c56
size 8518734624

3
Prismatic-12b-Q6_K.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b83bb69155f701ef56998a37188b8e9f0d77237246479b620042ee85ee7df04d
size 10056209184

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9246fbf2033553cb39cdb53611584abd1f54ca928e8a8e806e190ea58642ebf5
size 10381267744

3
Prismatic-12b-Q8_0.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f31fe3e8b3e083b610d682be07e52937bb3d68786544ca8f07ba04cf7e907157
size 13022368544

3
Prismatic-12b-f16.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:95a95734e83b3fd03b86a129a6beae12684709d4146b98e60065dd5de68b987d
size 24504275488

3
Prismatic-12b.imatrix Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3c8ea6986f27f0bc09cfc4fe7d71a7ac99f60092e6b353d523b18cb71f756274
size 7054418

122
README.md Normal file
View File

@@ -0,0 +1,122 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
tags:
- mergekit
- merge
base_model: ProdeusUnity/Prismatic-12b
---
## Llamacpp imatrix Quantizations of Prismatic-12b
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4058">b4058</a> for quantization.
Original model: https://huggingface.co/ProdeusUnity/Prismatic-12b
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
No prompt format found, check original model page
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Prismatic-12b-f16.gguf](https://huggingface.co/bartowski/Prismatic-12b-GGUF/blob/main/Prismatic-12b-f16.gguf) | f16 | 24.50GB | false | Full F16 weights. |
| [Prismatic-12b-Q8_0.gguf](https://huggingface.co/bartowski/Prismatic-12b-GGUF/blob/main/Prismatic-12b-Q8_0.gguf) | Q8_0 | 13.02GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Prismatic-12b-Q6_K_L.gguf](https://huggingface.co/bartowski/Prismatic-12b-GGUF/blob/main/Prismatic-12b-Q6_K_L.gguf) | Q6_K_L | 10.38GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Prismatic-12b-Q6_K.gguf](https://huggingface.co/bartowski/Prismatic-12b-GGUF/blob/main/Prismatic-12b-Q6_K.gguf) | Q6_K | 10.06GB | false | Very high quality, near perfect, *recommended*. |
| [Prismatic-12b-Q5_K_L.gguf](https://huggingface.co/bartowski/Prismatic-12b-GGUF/blob/main/Prismatic-12b-Q5_K_L.gguf) | Q5_K_L | 9.14GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Prismatic-12b-Q5_K_M.gguf](https://huggingface.co/bartowski/Prismatic-12b-GGUF/blob/main/Prismatic-12b-Q5_K_M.gguf) | Q5_K_M | 8.73GB | false | High quality, *recommended*. |
| [Prismatic-12b-Q5_K_S.gguf](https://huggingface.co/bartowski/Prismatic-12b-GGUF/blob/main/Prismatic-12b-Q5_K_S.gguf) | Q5_K_S | 8.52GB | false | High quality, *recommended*. |
| [Prismatic-12b-Q4_K_L.gguf](https://huggingface.co/bartowski/Prismatic-12b-GGUF/blob/main/Prismatic-12b-Q4_K_L.gguf) | Q4_K_L | 7.98GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Prismatic-12b-Q4_K_M.gguf](https://huggingface.co/bartowski/Prismatic-12b-GGUF/blob/main/Prismatic-12b-Q4_K_M.gguf) | Q4_K_M | 7.48GB | false | Good quality, default size for most use cases, *recommended*. |
| [Prismatic-12b-Q3_K_XL.gguf](https://huggingface.co/bartowski/Prismatic-12b-GGUF/blob/main/Prismatic-12b-Q3_K_XL.gguf) | Q3_K_XL | 7.15GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Prismatic-12b-Q4_K_S.gguf](https://huggingface.co/bartowski/Prismatic-12b-GGUF/blob/main/Prismatic-12b-Q4_K_S.gguf) | Q4_K_S | 7.12GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Prismatic-12b-Q4_0.gguf](https://huggingface.co/bartowski/Prismatic-12b-GGUF/blob/main/Prismatic-12b-Q4_0.gguf) | Q4_0 | 7.09GB | false | Legacy format, generally not worth using over similarly sized formats |
| [Prismatic-12b-Q4_0_8_8.gguf](https://huggingface.co/bartowski/Prismatic-12b-GGUF/blob/main/Prismatic-12b-Q4_0_8_8.gguf) | Q4_0_8_8 | 7.07GB | false | Optimized for ARM inference. Requires 'sve' support (see link below). *Don't use on Mac or Windows*. |
| [Prismatic-12b-Q4_0_4_8.gguf](https://huggingface.co/bartowski/Prismatic-12b-GGUF/blob/main/Prismatic-12b-Q4_0_4_8.gguf) | Q4_0_4_8 | 7.07GB | false | Optimized for ARM inference. Requires 'i8mm' support (see link below). *Don't use on Mac or Windows*. |
| [Prismatic-12b-Q4_0_4_4.gguf](https://huggingface.co/bartowski/Prismatic-12b-GGUF/blob/main/Prismatic-12b-Q4_0_4_4.gguf) | Q4_0_4_4 | 7.07GB | false | Optimized for ARM inference. Should work well on all ARM chips, pick this if you're unsure. *Don't use on Mac or Windows*. |
| [Prismatic-12b-IQ4_XS.gguf](https://huggingface.co/bartowski/Prismatic-12b-GGUF/blob/main/Prismatic-12b-IQ4_XS.gguf) | IQ4_XS | 6.74GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Prismatic-12b-Q3_K_L.gguf](https://huggingface.co/bartowski/Prismatic-12b-GGUF/blob/main/Prismatic-12b-Q3_K_L.gguf) | Q3_K_L | 6.56GB | false | Lower quality but usable, good for low RAM availability. |
| [Prismatic-12b-Q3_K_M.gguf](https://huggingface.co/bartowski/Prismatic-12b-GGUF/blob/main/Prismatic-12b-Q3_K_M.gguf) | Q3_K_M | 6.08GB | false | Low quality. |
| [Prismatic-12b-IQ3_M.gguf](https://huggingface.co/bartowski/Prismatic-12b-GGUF/blob/main/Prismatic-12b-IQ3_M.gguf) | IQ3_M | 5.72GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Prismatic-12b-Q3_K_S.gguf](https://huggingface.co/bartowski/Prismatic-12b-GGUF/blob/main/Prismatic-12b-Q3_K_S.gguf) | Q3_K_S | 5.53GB | false | Low quality, not recommended. |
| [Prismatic-12b-Q2_K_L.gguf](https://huggingface.co/bartowski/Prismatic-12b-GGUF/blob/main/Prismatic-12b-Q2_K_L.gguf) | Q2_K_L | 5.45GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Prismatic-12b-IQ3_XS.gguf](https://huggingface.co/bartowski/Prismatic-12b-GGUF/blob/main/Prismatic-12b-IQ3_XS.gguf) | IQ3_XS | 5.31GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Prismatic-12b-Q2_K.gguf](https://huggingface.co/bartowski/Prismatic-12b-GGUF/blob/main/Prismatic-12b-Q2_K.gguf) | Q2_K | 4.79GB | false | Very low quality but surprisingly usable. |
| [Prismatic-12b-IQ2_M.gguf](https://huggingface.co/bartowski/Prismatic-12b-GGUF/blob/main/Prismatic-12b-IQ2_M.gguf) | IQ2_M | 4.44GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
| [Prismatic-12b-IQ2_S.gguf](https://huggingface.co/bartowski/Prismatic-12b-GGUF/blob/main/Prismatic-12b-IQ2_S.gguf) | IQ2_S | 4.14GB | false | Low quality, uses SOTA techniques to be usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Prismatic-12b-GGUF --include "Prismatic-12b-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Prismatic-12b-GGUF --include "Prismatic-12b-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Prismatic-12b-Q8_0) or download them all in place (./)
## Q4_0_X_X
These are *NOT* for Metal (Apple) offloading, only ARM chips.
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}