初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Mistral7B-PairRM-SPPO-Iter3-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-16 19:37:13 +08:00
commit 8ff340ed67
24 changed files with 217 additions and 0 deletions

56
.gitattributes vendored Normal file
View File

@@ -0,0 +1,56 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Mistral7B-PairRM-SPPO-Iter3-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Mistral7B-PairRM-SPPO-Iter3-IQ2_S.gguf filter=lfs diff=lfs merge=lfs -text
Mistral7B-PairRM-SPPO-Iter3-IQ2_XS.gguf filter=lfs diff=lfs merge=lfs -text
Mistral7B-PairRM-SPPO-Iter3-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Mistral7B-PairRM-SPPO-Iter3-IQ3_S.gguf filter=lfs diff=lfs merge=lfs -text
Mistral7B-PairRM-SPPO-Iter3-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Mistral7B-PairRM-SPPO-Iter3-IQ3_XXS.gguf filter=lfs diff=lfs merge=lfs -text
Mistral7B-PairRM-SPPO-Iter3-IQ4_NL.gguf filter=lfs diff=lfs merge=lfs -text
Mistral7B-PairRM-SPPO-Iter3-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Mistral7B-PairRM-SPPO-Iter3-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Mistral7B-PairRM-SPPO-Iter3-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Mistral7B-PairRM-SPPO-Iter3-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Mistral7B-PairRM-SPPO-Iter3-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Mistral7B-PairRM-SPPO-Iter3-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Mistral7B-PairRM-SPPO-Iter3-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Mistral7B-PairRM-SPPO-Iter3-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Mistral7B-PairRM-SPPO-Iter3-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Mistral7B-PairRM-SPPO-Iter3-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Mistral7B-PairRM-SPPO-Iter3-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Mistral7B-PairRM-SPPO-Iter3-fp16.gguf filter=lfs diff=lfs merge=lfs -text
Mistral7B-PairRM-SPPO-Iter3.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:087428dc5844f03ebe3d96b993bb0e71d056d37e984fb5eb8162d6bb6da461fe
size 2500713088

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a1a1c205d5221d3d3372d16d911d91caea4fc8b4c2e1504109739498eefea8bd
size 2310920832

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7a1cebb88e2ff14a56b1cf381079d9871583cb0efa200a2b84106fd584b01096
size 2198256256

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:73d5951fdce872b0539e1dc0af16e3b83c87c11aecc607d4c863c5e7287a580a
size 3284892288

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b4a18f826c0bbfdfe3014134142684e833cd084c083231f7b37324c077180aa0
size 3182393984

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4ceae69545c8dbf9282f2069f0167e67b1ebea7e634e89063331f6786eb1e187
size 3018816128

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3d0b3aeb21f6a31f173a1398456e85f365eb36970e7c17c24fcc6efc5e746e9f
size 2827344512

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d60650b5431f57130c0732f95a1bd10af7baedf8bb1637facad8b36b376b08da
size 4125694592

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5fc429bc60e3f4112e803818d00777d3b1516c0f6bfccdb9d48a408243edfe12
size 3907689088

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cb60c5527a87e2ed29dd64ff11ddaa0ac7a84370aca3161f402c3b8bb1a29f87
size 2719242880

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ea10064036a4f802fb53e293f65445d70f716db2c5f4de6636b438fd53f5340c
size 3822025344

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e0f1af44bbaeb5e78a43c26e4b84c6875443a5e23986d9820af1cdc3dcb843d7
size 3518986880

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8e284c0f992a0832c90321246ef4bd2b5e37d646f117a1294a9350be0e7c4831
size 3164568192

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0c2cfb2f98778f53c57a590a18443396b10371732c6366e3cba8d74acd33b620
size 4368439936

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:33e7b754fb2960513fa912a3484ffcce49f2796e030d596846b45701976ddaf2
size 4140374656

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7b2483761dd3aa6699f980043f3e5de7fa78b924aeb85e87b65e9c048cb3f52a
size 5131410048

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0cb17d628edc97cd89a43cb7d1082e1a46f09bf35de251ef4ac1fd0355f7a91b
size 4997716608

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:eae2e55c7cc702cc9ceb252ea2d29c45cedb3c0172ea551cdb4e8efc9b891405
size 5942065792

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:568edd441f3b7a53335a9dd8e2af1e1377a9bdb04d76d71ebd7e691763883659
size 7695858304

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c30b6256ca712f02d8e79c2f7fba5a9123011f95a4e0a4f739b940b9842a4915
size 14484732192

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:849a401647d59d6815167ddb73bc845376f40f0f2205fe591eaa48cee4c58cb0
size 4988166

97
README.md Normal file
View File

@@ -0,0 +1,97 @@
---
license: apache-2.0
datasets:
- openbmb/UltraFeedback
language:
- en
pipeline_tag: text-generation
quantized_by: bartowski
---
## Llamacpp imatrix Quantizations of Mistral7B-PairRM-SPPO-Iter3
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b2777">b2777</a> for quantization.
Original model: https://huggingface.co/UCLA-AGI/Mistral7B-PairRM-SPPO-Iter3
All quants made using imatrix option with dataset provided by Kalomaze [here](https://github.com/ggerganov/llama.cpp/discussions/5263#discussioncomment-8395384)
## Prompt format
```
<s> [INST] {prompt} [/INST]</s>
```
Note that this model does not support a System prompt.
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Description |
| -------- | ---------- | --------- | ----------- |
| [Mistral7B-PairRM-SPPO-Iter3-Q8_0.gguf](https://huggingface.co/bartowski/Mistral7B-PairRM-SPPO-Iter3-GGUF/blob/main/Mistral7B-PairRM-SPPO-Iter3-Q8_0.gguf) | Q8_0 | 7.69GB | Extremely high quality, generally unneeded but max available quant. |
| [Mistral7B-PairRM-SPPO-Iter3-Q6_K.gguf](https://huggingface.co/bartowski/Mistral7B-PairRM-SPPO-Iter3-GGUF/blob/main/Mistral7B-PairRM-SPPO-Iter3-Q6_K.gguf) | Q6_K | 5.94GB | Very high quality, near perfect, *recommended*. |
| [Mistral7B-PairRM-SPPO-Iter3-Q5_K_M.gguf](https://huggingface.co/bartowski/Mistral7B-PairRM-SPPO-Iter3-GGUF/blob/main/Mistral7B-PairRM-SPPO-Iter3-Q5_K_M.gguf) | Q5_K_M | 5.13GB | High quality, *recommended*. |
| [Mistral7B-PairRM-SPPO-Iter3-Q5_K_S.gguf](https://huggingface.co/bartowski/Mistral7B-PairRM-SPPO-Iter3-GGUF/blob/main/Mistral7B-PairRM-SPPO-Iter3-Q5_K_S.gguf) | Q5_K_S | 4.99GB | High quality, *recommended*. |
| [Mistral7B-PairRM-SPPO-Iter3-Q4_K_M.gguf](https://huggingface.co/bartowski/Mistral7B-PairRM-SPPO-Iter3-GGUF/blob/main/Mistral7B-PairRM-SPPO-Iter3-Q4_K_M.gguf) | Q4_K_M | 4.36GB | Good quality, uses about 4.83 bits per weight, *recommended*. |
| [Mistral7B-PairRM-SPPO-Iter3-Q4_K_S.gguf](https://huggingface.co/bartowski/Mistral7B-PairRM-SPPO-Iter3-GGUF/blob/main/Mistral7B-PairRM-SPPO-Iter3-Q4_K_S.gguf) | Q4_K_S | 4.14GB | Slightly lower quality with more space savings, *recommended*. |
| [Mistral7B-PairRM-SPPO-Iter3-IQ4_NL.gguf](https://huggingface.co/bartowski/Mistral7B-PairRM-SPPO-Iter3-GGUF/blob/main/Mistral7B-PairRM-SPPO-Iter3-IQ4_NL.gguf) | IQ4_NL | 4.12GB | Decent quality, slightly smaller than Q4_K_S with similar performance *recommended*. |
| [Mistral7B-PairRM-SPPO-Iter3-IQ4_XS.gguf](https://huggingface.co/bartowski/Mistral7B-PairRM-SPPO-Iter3-GGUF/blob/main/Mistral7B-PairRM-SPPO-Iter3-IQ4_XS.gguf) | IQ4_XS | 3.90GB | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Mistral7B-PairRM-SPPO-Iter3-Q3_K_L.gguf](https://huggingface.co/bartowski/Mistral7B-PairRM-SPPO-Iter3-GGUF/blob/main/Mistral7B-PairRM-SPPO-Iter3-Q3_K_L.gguf) | Q3_K_L | 3.82GB | Lower quality but usable, good for low RAM availability. |
| [Mistral7B-PairRM-SPPO-Iter3-Q3_K_M.gguf](https://huggingface.co/bartowski/Mistral7B-PairRM-SPPO-Iter3-GGUF/blob/main/Mistral7B-PairRM-SPPO-Iter3-Q3_K_M.gguf) | Q3_K_M | 3.51GB | Even lower quality. |
| [Mistral7B-PairRM-SPPO-Iter3-IQ3_M.gguf](https://huggingface.co/bartowski/Mistral7B-PairRM-SPPO-Iter3-GGUF/blob/main/Mistral7B-PairRM-SPPO-Iter3-IQ3_M.gguf) | IQ3_M | 3.28GB | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Mistral7B-PairRM-SPPO-Iter3-IQ3_S.gguf](https://huggingface.co/bartowski/Mistral7B-PairRM-SPPO-Iter3-GGUF/blob/main/Mistral7B-PairRM-SPPO-Iter3-IQ3_S.gguf) | IQ3_S | 3.18GB | Lower quality, new method with decent performance, recommended over Q3_K_S quant, same size with better performance. |
| [Mistral7B-PairRM-SPPO-Iter3-Q3_K_S.gguf](https://huggingface.co/bartowski/Mistral7B-PairRM-SPPO-Iter3-GGUF/blob/main/Mistral7B-PairRM-SPPO-Iter3-Q3_K_S.gguf) | Q3_K_S | 3.16GB | Low quality, not recommended. |
| [Mistral7B-PairRM-SPPO-Iter3-IQ3_XS.gguf](https://huggingface.co/bartowski/Mistral7B-PairRM-SPPO-Iter3-GGUF/blob/main/Mistral7B-PairRM-SPPO-Iter3-IQ3_XS.gguf) | IQ3_XS | 3.01GB | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Mistral7B-PairRM-SPPO-Iter3-IQ3_XXS.gguf](https://huggingface.co/bartowski/Mistral7B-PairRM-SPPO-Iter3-GGUF/blob/main/Mistral7B-PairRM-SPPO-Iter3-IQ3_XXS.gguf) | IQ3_XXS | 2.82GB | Lower quality, new method with decent performance, comparable to Q3 quants. |
| [Mistral7B-PairRM-SPPO-Iter3-Q2_K.gguf](https://huggingface.co/bartowski/Mistral7B-PairRM-SPPO-Iter3-GGUF/blob/main/Mistral7B-PairRM-SPPO-Iter3-Q2_K.gguf) | Q2_K | 2.71GB | Very low quality but surprisingly usable. |
| [Mistral7B-PairRM-SPPO-Iter3-IQ2_M.gguf](https://huggingface.co/bartowski/Mistral7B-PairRM-SPPO-Iter3-GGUF/blob/main/Mistral7B-PairRM-SPPO-Iter3-IQ2_M.gguf) | IQ2_M | 2.50GB | Very low quality, uses SOTA techniques to also be surprisingly usable. |
| [Mistral7B-PairRM-SPPO-Iter3-IQ2_S.gguf](https://huggingface.co/bartowski/Mistral7B-PairRM-SPPO-Iter3-GGUF/blob/main/Mistral7B-PairRM-SPPO-Iter3-IQ2_S.gguf) | IQ2_S | 2.31GB | Very low quality, uses SOTA techniques to be usable. |
| [Mistral7B-PairRM-SPPO-Iter3-IQ2_XS.gguf](https://huggingface.co/bartowski/Mistral7B-PairRM-SPPO-Iter3-GGUF/blob/main/Mistral7B-PairRM-SPPO-Iter3-IQ2_XS.gguf) | IQ2_XS | 2.19GB | Very low quality, uses SOTA techniques to be usable. |
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Mistral7B-PairRM-SPPO-Iter3-GGUF --include "Mistral7B-PairRM-SPPO-Iter3-Q4_K_M.gguf" --local-dir ./ --local-dir-use-symlinks False
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Mistral7B-PairRM-SPPO-Iter3-GGUF --include "Mistral7B-PairRM-SPPO-Iter3-Q8_0.gguf/*" --local-dir Mistral7B-PairRM-SPPO-Iter3-Q8_0 --local-dir-use-symlinks False
```
You can either specify a new local-dir (Mistral7B-PairRM-SPPO-Iter3-Q8_0) or download them all in place (./)
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}