初始化项目,由ModelHub XC社区提供模型

Model: bartowski/MathCoder2-CodeLlama-7B-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-06 05:45:12 +08:00
commit 6028ee5faa
28 changed files with 263 additions and 0 deletions

60
.gitattributes vendored Normal file
View File

@@ -0,0 +1,60 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
MathCoder2-CodeLlama-7B-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-CodeLlama-7B-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-CodeLlama-7B-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-CodeLlama-7B-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-CodeLlama-7B-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-CodeLlama-7B-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-CodeLlama-7B-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-CodeLlama-7B-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-CodeLlama-7B-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-CodeLlama-7B-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-CodeLlama-7B-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-CodeLlama-7B-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-CodeLlama-7B-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-CodeLlama-7B-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-CodeLlama-7B-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-CodeLlama-7B-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-CodeLlama-7B-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-CodeLlama-7B-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-CodeLlama-7B-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-CodeLlama-7B-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-CodeLlama-7B-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-CodeLlama-7B-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-CodeLlama-7B-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-CodeLlama-7B-f16.gguf filter=lfs diff=lfs merge=lfs -text
MathCoder2-CodeLlama-7B.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d5cf0bce370c726e69c91f0fda0dc10bd7636d95e4b671884d88132ee4b300d8
size 2359825344

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3a143ebd657f1499a156c0f9f68e7c4da28d6f4da59241afcf836992ac54b300
size 3114948032

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cff1006da1cf8cb921b0916dd8f3979f3e1599dc835cfb0799f0df28e9184f94
size 2796606912

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b2c76a72cdb5d25026be1b991b0a3c6e631b230b7009a5f19e9d00fc075b6047
size 3619426240

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:60f00d7c54dabc774c7ad3388f62a7e18df8123edf0358f8662fedf5462727b4
size 2532940736

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5ddb7347216b1f88516db0b02d86686c370c6d43aad8e55394604849802f6352
size 2661004736

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a8bb8a11157fcd8495d153badb317b20a166d835b773c6acb82d5bbaaa1e839d
size 3597194688

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:671f909072ca67d2ac0849599087c06aba006edb4d7a58e2b130b2e3bb1a477f
size 3298088384

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:78427b9d2b6909009d7ce4d970d6e75b0f935024164d131d7c86b5512aad0292
size 2948388288

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:74a808560dc1bea41f775f1e45c1be011f65a48761b639e16bf362f5364a84b9
size 3711940032

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:914c89f23c86b71676b67a51da289d2e05e3629274eca9a763dcacb7a4ff6a6d
size 3837171648

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c22814c7b7d07dfa3480f4bbf3b8d02f0176b91679a6b6030f4e509c6ae36561
size 3825899456

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2af60724e6fc5e7c8401fa5695598225a22282ed137c0765bb2def6ae99f2bce
size 3825899456

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:45ec53bdd08ccf05da2942ba6f249b54b6bf2912ce37b72c22ca9a9cd38677ce
size 3825899456

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:977171ca19e460302842022e701c2602c18610178d52a67698d22b442474532d
size 4178425280

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5acc5f644953a91c11025a5bb347c81b42f65da0953b31c9dd70fe8573e4e367
size 4081096640

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a9b4fb9aecc2194056da77f0e3855e0a4aede60585ce34254d218e8436094ea1
size 3856832448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:17a6ae9da37e67037b0ee866ce9278321ffca871026dc1a7357c913883598bd5
size 4864193984

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:87a9d7304534791fd2c87e4561b0c0a77d6c009cee6079c34c2a48b8c97eb69c
size 4783257536

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b6735f74cdaafa79de6fdeed6d129d97530575e4945cab386e74b4a2015473dd
size 4651792320

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3fe5191c1395f2bd5482073d0e997d45192cc32c95dc04b45f71a842b227bd37
size 5529303488

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c967c03900572f2fe876e847651b184f11b302234e79dfc9a2c7bebc302becff
size 5592823232

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0e8e1ee4838c94d2698f03fbc89facdd20b094b640cc0a0d5e5f6ab8c217895f
size 7161230784

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:14e793ae8c165d7da1704e39a6e160d8dfb82ed201ab03192088d156b128630f
size 13478368416

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:204d0de910b382dc83d1faea64c2bc0f9e635cb20704ec254d9a1209933bac63
size 4562186

127
README.md Normal file
View File

@@ -0,0 +1,127 @@
---
base_model: MathGenie/MathCoder2-CodeLlama-7B
datasets:
- MathGenie/MathCode-Pile
language:
- en
license: apache-2.0
metrics:
- accuracy
pipeline_tag: text-generation
tags:
- math
quantized_by: bartowski
---
## Llamacpp imatrix Quantizations of MathCoder2-CodeLlama-7B
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b3901">b3901</a> for quantization.
Original model: https://huggingface.co/MathGenie/MathCoder2-CodeLlama-7B
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
No prompt format found, check original model page
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [MathCoder2-CodeLlama-7B-f16.gguf](https://huggingface.co/bartowski/MathCoder2-CodeLlama-7B-GGUF/blob/main/MathCoder2-CodeLlama-7B-f16.gguf) | f16 | 13.48GB | false | Full F16 weights. |
| [MathCoder2-CodeLlama-7B-Q8_0.gguf](https://huggingface.co/bartowski/MathCoder2-CodeLlama-7B-GGUF/blob/main/MathCoder2-CodeLlama-7B-Q8_0.gguf) | Q8_0 | 7.16GB | false | Extremely high quality, generally unneeded but max available quant. |
| [MathCoder2-CodeLlama-7B-Q6_K_L.gguf](https://huggingface.co/bartowski/MathCoder2-CodeLlama-7B-GGUF/blob/main/MathCoder2-CodeLlama-7B-Q6_K_L.gguf) | Q6_K_L | 5.59GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [MathCoder2-CodeLlama-7B-Q6_K.gguf](https://huggingface.co/bartowski/MathCoder2-CodeLlama-7B-GGUF/blob/main/MathCoder2-CodeLlama-7B-Q6_K.gguf) | Q6_K | 5.53GB | false | Very high quality, near perfect, *recommended*. |
| [MathCoder2-CodeLlama-7B-Q5_K_L.gguf](https://huggingface.co/bartowski/MathCoder2-CodeLlama-7B-GGUF/blob/main/MathCoder2-CodeLlama-7B-Q5_K_L.gguf) | Q5_K_L | 4.86GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [MathCoder2-CodeLlama-7B-Q5_K_M.gguf](https://huggingface.co/bartowski/MathCoder2-CodeLlama-7B-GGUF/blob/main/MathCoder2-CodeLlama-7B-Q5_K_M.gguf) | Q5_K_M | 4.78GB | false | High quality, *recommended*. |
| [MathCoder2-CodeLlama-7B-Q5_K_S.gguf](https://huggingface.co/bartowski/MathCoder2-CodeLlama-7B-GGUF/blob/main/MathCoder2-CodeLlama-7B-Q5_K_S.gguf) | Q5_K_S | 4.65GB | false | High quality, *recommended*. |
| [MathCoder2-CodeLlama-7B-Q4_K_L.gguf](https://huggingface.co/bartowski/MathCoder2-CodeLlama-7B-GGUF/blob/main/MathCoder2-CodeLlama-7B-Q4_K_L.gguf) | Q4_K_L | 4.18GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [MathCoder2-CodeLlama-7B-Q4_K_M.gguf](https://huggingface.co/bartowski/MathCoder2-CodeLlama-7B-GGUF/blob/main/MathCoder2-CodeLlama-7B-Q4_K_M.gguf) | Q4_K_M | 4.08GB | false | Good quality, default size for must use cases, *recommended*. |
| [MathCoder2-CodeLlama-7B-Q4_K_S.gguf](https://huggingface.co/bartowski/MathCoder2-CodeLlama-7B-GGUF/blob/main/MathCoder2-CodeLlama-7B-Q4_K_S.gguf) | Q4_K_S | 3.86GB | false | Slightly lower quality with more space savings, *recommended*. |
| [MathCoder2-CodeLlama-7B-Q4_0.gguf](https://huggingface.co/bartowski/MathCoder2-CodeLlama-7B-GGUF/blob/main/MathCoder2-CodeLlama-7B-Q4_0.gguf) | Q4_0 | 3.84GB | false | Legacy format, generally not worth using over similarly sized formats |
| [MathCoder2-CodeLlama-7B-Q4_0_8_8.gguf](https://huggingface.co/bartowski/MathCoder2-CodeLlama-7B-GGUF/blob/main/MathCoder2-CodeLlama-7B-Q4_0_8_8.gguf) | Q4_0_8_8 | 3.83GB | false | Optimized for ARM inference. Requires 'sve' support (see link below). *Don't use on Mac or Windows*. |
| [MathCoder2-CodeLlama-7B-Q4_0_4_8.gguf](https://huggingface.co/bartowski/MathCoder2-CodeLlama-7B-GGUF/blob/main/MathCoder2-CodeLlama-7B-Q4_0_4_8.gguf) | Q4_0_4_8 | 3.83GB | false | Optimized for ARM inference. Requires 'i8mm' support (see link below). *Don't use on Mac or Windows*. |
| [MathCoder2-CodeLlama-7B-Q4_0_4_4.gguf](https://huggingface.co/bartowski/MathCoder2-CodeLlama-7B-GGUF/blob/main/MathCoder2-CodeLlama-7B-Q4_0_4_4.gguf) | Q4_0_4_4 | 3.83GB | false | Optimized for ARM inference. Should work well on all ARM chips, pick this if you're unsure. *Don't use on Mac or Windows*. |
| [MathCoder2-CodeLlama-7B-Q3_K_XL.gguf](https://huggingface.co/bartowski/MathCoder2-CodeLlama-7B-GGUF/blob/main/MathCoder2-CodeLlama-7B-Q3_K_XL.gguf) | Q3_K_XL | 3.71GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [MathCoder2-CodeLlama-7B-IQ4_XS.gguf](https://huggingface.co/bartowski/MathCoder2-CodeLlama-7B-GGUF/blob/main/MathCoder2-CodeLlama-7B-IQ4_XS.gguf) | IQ4_XS | 3.62GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [MathCoder2-CodeLlama-7B-Q3_K_L.gguf](https://huggingface.co/bartowski/MathCoder2-CodeLlama-7B-GGUF/blob/main/MathCoder2-CodeLlama-7B-Q3_K_L.gguf) | Q3_K_L | 3.60GB | false | Lower quality but usable, good for low RAM availability. |
| [MathCoder2-CodeLlama-7B-Q3_K_M.gguf](https://huggingface.co/bartowski/MathCoder2-CodeLlama-7B-GGUF/blob/main/MathCoder2-CodeLlama-7B-Q3_K_M.gguf) | Q3_K_M | 3.30GB | false | Low quality. |
| [MathCoder2-CodeLlama-7B-IQ3_M.gguf](https://huggingface.co/bartowski/MathCoder2-CodeLlama-7B-GGUF/blob/main/MathCoder2-CodeLlama-7B-IQ3_M.gguf) | IQ3_M | 3.11GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [MathCoder2-CodeLlama-7B-Q3_K_S.gguf](https://huggingface.co/bartowski/MathCoder2-CodeLlama-7B-GGUF/blob/main/MathCoder2-CodeLlama-7B-Q3_K_S.gguf) | Q3_K_S | 2.95GB | false | Low quality, not recommended. |
| [MathCoder2-CodeLlama-7B-IQ3_XS.gguf](https://huggingface.co/bartowski/MathCoder2-CodeLlama-7B-GGUF/blob/main/MathCoder2-CodeLlama-7B-IQ3_XS.gguf) | IQ3_XS | 2.80GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [MathCoder2-CodeLlama-7B-Q2_K_L.gguf](https://huggingface.co/bartowski/MathCoder2-CodeLlama-7B-GGUF/blob/main/MathCoder2-CodeLlama-7B-Q2_K_L.gguf) | Q2_K_L | 2.66GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [MathCoder2-CodeLlama-7B-Q2_K.gguf](https://huggingface.co/bartowski/MathCoder2-CodeLlama-7B-GGUF/blob/main/MathCoder2-CodeLlama-7B-Q2_K.gguf) | Q2_K | 2.53GB | false | Very low quality but surprisingly usable. |
| [MathCoder2-CodeLlama-7B-IQ2_M.gguf](https://huggingface.co/bartowski/MathCoder2-CodeLlama-7B-GGUF/blob/main/MathCoder2-CodeLlama-7B-IQ2_M.gguf) | IQ2_M | 2.36GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/MathCoder2-CodeLlama-7B-GGUF --include "MathCoder2-CodeLlama-7B-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/MathCoder2-CodeLlama-7B-GGUF --include "MathCoder2-CodeLlama-7B-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (MathCoder2-CodeLlama-7B-Q8_0) or download them all in place (./)
## Q4_0_X_X
These are *NOT* for Metal (Apple) offloading, only ARM chips.
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset
Thank you ZeroWw for the inspiration to experiment with embed/output
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}