初始化项目,由ModelHub XC社区提供模型

Model: bartowski/MathCoder2-Mistral-7B-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-09-14 05:18:12 +08:00
commit 9c6c7f2efd
28 changed files with 263 additions and 0 deletions

60
.gitattributes vendored Normal file
View File

@@ -0,0 +1,60 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
MathCoder2-Mistral-7B-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-Mistral-7B-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-Mistral-7B-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-Mistral-7B-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-Mistral-7B-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-Mistral-7B-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-Mistral-7B-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-Mistral-7B-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-Mistral-7B-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-Mistral-7B-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-Mistral-7B-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-Mistral-7B-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-Mistral-7B-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-Mistral-7B-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-Mistral-7B-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-Mistral-7B-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-Mistral-7B-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-Mistral-7B-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-Mistral-7B-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-Mistral-7B-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-Mistral-7B-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-Mistral-7B-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-Mistral-7B-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-Mistral-7B-f16.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-Mistral-7B.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e92b176b5e1915c0b53662df2588df5189d6d35167ef46ade47010fe5989e4d4
size 2500713376

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:64c84a997d5b15caea89686b8b0f0b4d27b7b180df128c6c410e9858d1462edb
size 3284892576

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0d07cbdac8165faac9e2f6fce5364eb9d4f7b9147c6b925336312d7e4973f85f
size 3018816416

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:305d842a8530c42ad02d8ca2feb162e9550e844e1c6ec4e59eb2cc9051c204a9
size 3907689376

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c3d420fbdc4f77bda3ee41500d218fa9aa96dee4bb82b663321a01c41bb6e035
size 2719243168

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ffe2c3019ca62c5a0a10d83e5b89b435eb0ddfa7da99c51e1f1d6feb49a55b36
size 2847243168

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7217cbaffdee46ab8a4aac73d3bd36930b0d4fdfd810a1a656b8069f180127ad
size 3822025632

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:16f55dabc509138a38a07d4797c973667a7b666f707882f21e6feb3c66b6d869
size 3518987168

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9ea6efd16aca3393fa21e331937b534862ea9fdcc13890a8b2cc6f8074619759
size 3164568480

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:bf6247e4c618b0c756b8c876c5f9ca04c12c4299fef28c1833630e9af14fd977
size 3936713632

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:99123f494d4d5f46d4a88792f6d161323034dcdde06068f9f737d86e1754c6c7
size 4123597728

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8a498ca837171b2509035f4c1bc5744ff138729c1132b085d5f997e2a4dbfeba
size 4108917664

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a9c89a5f162206271c3255ec42210a96cd43dc9e1be26edba4799be307f0777f
size 4108917664

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c08fdfa07e81c26447270f7bc69b2d1912344d0034c45387a8145368f8c01d5d
size 4108917664

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a462406194a84eb58920e00fc032f18b82d3bf6c75668644bfc43a687a0a8f4e
size 4465720224

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:faca5ac795c27294c40f7ea70fbd4bbbc25ef3c77145314e3e9d6cab8daab196
size 4368440224

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3c428f4640ed5e5a850101cf97cf8f76b8caf3359bd56c764faa63dd6843336e
size 4140374944

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f21dd882f7c89a01c95f2a309095e0b17b919d88b675ff774ae415a9a91b8fba
size 5212306336

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:edff0a2431093a93d7ac83e55280d0d61ddd7ae5848615ae20a9aa1cdd32d3b6
size 5131410336

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9bc3afcfc40cd579240bff7b5af04da533b2fbe0b81866971047565a34718a86
size 4997716896

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:42de8ff214e2c44d09daf5363bb825f4e56f31dfe9987f5fcee6f033cde14b68
size 5942066080

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2a8d9567c49153620a80eed2dcda5a8f61f18e0ca0e283207abe79f0b60eb708
size 6005554080

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7719edc250c730cb82a7db095adc01559b87d02b61741ad6038cf0a0ae4e1d2c
size 7695858592

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:54b037a33c059d42c13dc63991bdd698c6749b7623e7dd276d8784b672099017
size 14484732544

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:69a6ee66270994c5e7d1a7cc6061338031af01b16defe6076e5efff6a683fc0d
size 4988170

127
README.md Normal file
View File

@@ -0,0 +1,127 @@
---
base_model: MathGenie/MathCoder2-Mistral-7B
datasets:
- MathGenie/MathCode-Pile
language:
- en
license: apache-2.0
metrics:
- accuracy
pipeline_tag: text-generation
tags:
- math
quantized_by: bartowski
---
## Llamacpp imatrix Quantizations of MathCoder2-Mistral-7B
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b3901">b3901</a> for quantization.
Original model: https://huggingface.co/MathGenie/MathCoder2-Mistral-7B
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
No prompt format found, check original model page
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [MathCoder2-Mistral-7B-f16.gguf](https://huggingface.co/bartowski/MathCoder2-Mistral-7B-GGUF/blob/main/MathCoder2-Mistral-7B-f16.gguf) | f16 | 14.48GB | false | Full F16 weights. |
| [MathCoder2-Mistral-7B-Q8_0.gguf](https://huggingface.co/bartowski/MathCoder2-Mistral-7B-GGUF/blob/main/MathCoder2-Mistral-7B-Q8_0.gguf) | Q8_0 | 7.70GB | false | Extremely high quality, generally unneeded but max available quant. |
| [MathCoder2-Mistral-7B-Q6_K_L.gguf](https://huggingface.co/bartowski/MathCoder2-Mistral-7B-GGUF/blob/main/MathCoder2-Mistral-7B-Q6_K_L.gguf) | Q6_K_L | 6.01GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [MathCoder2-Mistral-7B-Q6_K.gguf](https://huggingface.co/bartowski/MathCoder2-Mistral-7B-GGUF/blob/main/MathCoder2-Mistral-7B-Q6_K.gguf) | Q6_K | 5.94GB | false | Very high quality, near perfect, *recommended*. |
| [MathCoder2-Mistral-7B-Q5_K_L.gguf](https://huggingface.co/bartowski/MathCoder2-Mistral-7B-GGUF/blob/main/MathCoder2-Mistral-7B-Q5_K_L.gguf) | Q5_K_L | 5.21GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [MathCoder2-Mistral-7B-Q5_K_M.gguf](https://huggingface.co/bartowski/MathCoder2-Mistral-7B-GGUF/blob/main/MathCoder2-Mistral-7B-Q5_K_M.gguf) | Q5_K_M | 5.13GB | false | High quality, *recommended*. |
| [MathCoder2-Mistral-7B-Q5_K_S.gguf](https://huggingface.co/bartowski/MathCoder2-Mistral-7B-GGUF/blob/main/MathCoder2-Mistral-7B-Q5_K_S.gguf) | Q5_K_S | 5.00GB | false | High quality, *recommended*. |
| [MathCoder2-Mistral-7B-Q4_K_L.gguf](https://huggingface.co/bartowski/MathCoder2-Mistral-7B-GGUF/blob/main/MathCoder2-Mistral-7B-Q4_K_L.gguf) | Q4_K_L | 4.47GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [MathCoder2-Mistral-7B-Q4_K_M.gguf](https://huggingface.co/bartowski/MathCoder2-Mistral-7B-GGUF/blob/main/MathCoder2-Mistral-7B-Q4_K_M.gguf) | Q4_K_M | 4.37GB | false | Good quality, default size for must use cases, *recommended*. |
| [MathCoder2-Mistral-7B-Q4_K_S.gguf](https://huggingface.co/bartowski/MathCoder2-Mistral-7B-GGUF/blob/main/MathCoder2-Mistral-7B-Q4_K_S.gguf) | Q4_K_S | 4.14GB | false | Slightly lower quality with more space savings, *recommended*. |
| [MathCoder2-Mistral-7B-Q4_0.gguf](https://huggingface.co/bartowski/MathCoder2-Mistral-7B-GGUF/blob/main/MathCoder2-Mistral-7B-Q4_0.gguf) | Q4_0 | 4.12GB | false | Legacy format, generally not worth using over similarly sized formats |
| [MathCoder2-Mistral-7B-Q4_0_8_8.gguf](https://huggingface.co/bartowski/MathCoder2-Mistral-7B-GGUF/blob/main/MathCoder2-Mistral-7B-Q4_0_8_8.gguf) | Q4_0_8_8 | 4.11GB | false | Optimized for ARM inference. Requires 'sve' support (see link below). *Don't use on Mac or Windows*. |
| [MathCoder2-Mistral-7B-Q4_0_4_8.gguf](https://huggingface.co/bartowski/MathCoder2-Mistral-7B-GGUF/blob/main/MathCoder2-Mistral-7B-Q4_0_4_8.gguf) | Q4_0_4_8 | 4.11GB | false | Optimized for ARM inference. Requires 'i8mm' support (see link below). *Don't use on Mac or Windows*. |
| [MathCoder2-Mistral-7B-Q4_0_4_4.gguf](https://huggingface.co/bartowski/MathCoder2-Mistral-7B-GGUF/blob/main/MathCoder2-Mistral-7B-Q4_0_4_4.gguf) | Q4_0_4_4 | 4.11GB | false | Optimized for ARM inference. Should work well on all ARM chips, pick this if you're unsure. *Don't use on Mac or Windows*. |
| [MathCoder2-Mistral-7B-Q3_K_XL.gguf](https://huggingface.co/bartowski/MathCoder2-Mistral-7B-GGUF/blob/main/MathCoder2-Mistral-7B-Q3_K_XL.gguf) | Q3_K_XL | 3.94GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [MathCoder2-Mistral-7B-IQ4_XS.gguf](https://huggingface.co/bartowski/MathCoder2-Mistral-7B-GGUF/blob/main/MathCoder2-Mistral-7B-IQ4_XS.gguf) | IQ4_XS | 3.91GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [MathCoder2-Mistral-7B-Q3_K_L.gguf](https://huggingface.co/bartowski/MathCoder2-Mistral-7B-GGUF/blob/main/MathCoder2-Mistral-7B-Q3_K_L.gguf) | Q3_K_L | 3.82GB | false | Lower quality but usable, good for low RAM availability. |
| [MathCoder2-Mistral-7B-Q3_K_M.gguf](https://huggingface.co/bartowski/MathCoder2-Mistral-7B-GGUF/blob/main/MathCoder2-Mistral-7B-Q3_K_M.gguf) | Q3_K_M | 3.52GB | false | Low quality. |
| [MathCoder2-Mistral-7B-IQ3_M.gguf](https://huggingface.co/bartowski/MathCoder2-Mistral-7B-GGUF/blob/main/MathCoder2-Mistral-7B-IQ3_M.gguf) | IQ3_M | 3.28GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [MathCoder2-Mistral-7B-Q3_K_S.gguf](https://huggingface.co/bartowski/MathCoder2-Mistral-7B-GGUF/blob/main/MathCoder2-Mistral-7B-Q3_K_S.gguf) | Q3_K_S | 3.16GB | false | Low quality, not recommended. |
| [MathCoder2-Mistral-7B-IQ3_XS.gguf](https://huggingface.co/bartowski/MathCoder2-Mistral-7B-GGUF/blob/main/MathCoder2-Mistral-7B-IQ3_XS.gguf) | IQ3_XS | 3.02GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [MathCoder2-Mistral-7B-Q2_K_L.gguf](https://huggingface.co/bartowski/MathCoder2-Mistral-7B-GGUF/blob/main/MathCoder2-Mistral-7B-Q2_K_L.gguf) | Q2_K_L | 2.85GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [MathCoder2-Mistral-7B-Q2_K.gguf](https://huggingface.co/bartowski/MathCoder2-Mistral-7B-GGUF/blob/main/MathCoder2-Mistral-7B-Q2_K.gguf) | Q2_K | 2.72GB | false | Very low quality but surprisingly usable. |
| [MathCoder2-Mistral-7B-IQ2_M.gguf](https://huggingface.co/bartowski/MathCoder2-Mistral-7B-GGUF/blob/main/MathCoder2-Mistral-7B-IQ2_M.gguf) | IQ2_M | 2.50GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/MathCoder2-Mistral-7B-GGUF --include "MathCoder2-Mistral-7B-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/MathCoder2-Mistral-7B-GGUF --include "MathCoder2-Mistral-7B-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (MathCoder2-Mistral-7B-Q8_0) or download them all in place (./)
## Q4_0_X_X
These are *NOT* for Metal (Apple) offloading, only ARM chips.
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset
Thank you ZeroWw for the inspiration to experiment with embed/output
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}