初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Replete-Coder-V2-Llama-3.1-8b-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-19 21:01:13 +08:00
commit 135c7fd931
24 changed files with 244 additions and 0 deletions

56
.gitattributes vendored Normal file
View File

@@ -0,0 +1,56 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Replete-Coder-V2-Llama-3.1-8b-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Replete-Coder-V2-Llama-3.1-8b-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Replete-Coder-V2-Llama-3.1-8b-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Replete-Coder-V2-Llama-3.1-8b-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Replete-Coder-V2-Llama-3.1-8b-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Replete-Coder-V2-Llama-3.1-8b-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Replete-Coder-V2-Llama-3.1-8b-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Replete-Coder-V2-Llama-3.1-8b-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Replete-Coder-V2-Llama-3.1-8b-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Replete-Coder-V2-Llama-3.1-8b-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Replete-Coder-V2-Llama-3.1-8b-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Replete-Coder-V2-Llama-3.1-8b-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Replete-Coder-V2-Llama-3.1-8b-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Replete-Coder-V2-Llama-3.1-8b-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Replete-Coder-V2-Llama-3.1-8b-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Replete-Coder-V2-Llama-3.1-8b-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Replete-Coder-V2-Llama-3.1-8b-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Replete-Coder-V2-Llama-3.1-8b-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Replete-Coder-V2-Llama-3.1-8b-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Replete-Coder-V2-Llama-3.1-8b-f16.gguf filter=lfs diff=lfs merge=lfs -text
Replete-Coder-V2-Llama-3.1-8b.imatrix filter=lfs diff=lfs merge=lfs -text

124
README.md Normal file
View File

@@ -0,0 +1,124 @@
---
base_model: Replete-AI/Replete-Coder-V2-Llama-3.1-8b
language:
- en
license: apache-2.0
pipeline_tag: text-generation
tags:
- unsloth
quantized_by: bartowski
---
## Llamacpp imatrix Quantizations of Replete-Coder-V2-Llama-3.1-8b
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b3600">b3600</a> for quantization.
Original model: https://huggingface.co/Replete-AI/Replete-Coder-V2-Llama-3.1-8b
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Torrent files
https://aitorrent.zerroug.de/bartowski-replete-coder-v2-llama-3-1-8b-gguf-torrent/
## Prompt format
```
<|begin_of_text|><|start_header_id|>system<|end_header_id|>
Cutting Knowledge Date: December 2023
Today Date: 26 Jul 2024
{system_prompt}<|eot_id|><|start_header_id|>user<|end_header_id|>
{prompt}<|eot_id|><|start_header_id|>assistant<|end_header_id|>
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Replete-Coder-V2-Llama-3.1-8b-f16.gguf](https://huggingface.co/bartowski/Replete-Coder-V2-Llama-3.1-8b-GGUF/blob/main/Replete-Coder-V2-Llama-3.1-8b-f16.gguf) | f16 | 16.07GB | false | Full F16 weights. |
| [Replete-Coder-V2-Llama-3.1-8b-Q8_0.gguf](https://huggingface.co/bartowski/Replete-Coder-V2-Llama-3.1-8b-GGUF/blob/main/Replete-Coder-V2-Llama-3.1-8b-Q8_0.gguf) | Q8_0 | 8.54GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Replete-Coder-V2-Llama-3.1-8b-Q6_K_L.gguf](https://huggingface.co/bartowski/Replete-Coder-V2-Llama-3.1-8b-GGUF/blob/main/Replete-Coder-V2-Llama-3.1-8b-Q6_K_L.gguf) | Q6_K_L | 6.85GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Replete-Coder-V2-Llama-3.1-8b-Q6_K.gguf](https://huggingface.co/bartowski/Replete-Coder-V2-Llama-3.1-8b-GGUF/blob/main/Replete-Coder-V2-Llama-3.1-8b-Q6_K.gguf) | Q6_K | 6.60GB | false | Very high quality, near perfect, *recommended*. |
| [Replete-Coder-V2-Llama-3.1-8b-Q5_K_L.gguf](https://huggingface.co/bartowski/Replete-Coder-V2-Llama-3.1-8b-GGUF/blob/main/Replete-Coder-V2-Llama-3.1-8b-Q5_K_L.gguf) | Q5_K_L | 6.06GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Replete-Coder-V2-Llama-3.1-8b-Q5_K_M.gguf](https://huggingface.co/bartowski/Replete-Coder-V2-Llama-3.1-8b-GGUF/blob/main/Replete-Coder-V2-Llama-3.1-8b-Q5_K_M.gguf) | Q5_K_M | 5.73GB | false | High quality, *recommended*. |
| [Replete-Coder-V2-Llama-3.1-8b-Q5_K_S.gguf](https://huggingface.co/bartowski/Replete-Coder-V2-Llama-3.1-8b-GGUF/blob/main/Replete-Coder-V2-Llama-3.1-8b-Q5_K_S.gguf) | Q5_K_S | 5.60GB | false | High quality, *recommended*. |
| [Replete-Coder-V2-Llama-3.1-8b-Q4_K_L.gguf](https://huggingface.co/bartowski/Replete-Coder-V2-Llama-3.1-8b-GGUF/blob/main/Replete-Coder-V2-Llama-3.1-8b-Q4_K_L.gguf) | Q4_K_L | 5.31GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Replete-Coder-V2-Llama-3.1-8b-Q4_K_M.gguf](https://huggingface.co/bartowski/Replete-Coder-V2-Llama-3.1-8b-GGUF/blob/main/Replete-Coder-V2-Llama-3.1-8b-Q4_K_M.gguf) | Q4_K_M | 4.92GB | false | Good quality, default size for must use cases, *recommended*. |
| [Replete-Coder-V2-Llama-3.1-8b-Q3_K_XL.gguf](https://huggingface.co/bartowski/Replete-Coder-V2-Llama-3.1-8b-GGUF/blob/main/Replete-Coder-V2-Llama-3.1-8b-Q3_K_XL.gguf) | Q3_K_XL | 4.78GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Replete-Coder-V2-Llama-3.1-8b-Q4_K_S.gguf](https://huggingface.co/bartowski/Replete-Coder-V2-Llama-3.1-8b-GGUF/blob/main/Replete-Coder-V2-Llama-3.1-8b-Q4_K_S.gguf) | Q4_K_S | 4.69GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Replete-Coder-V2-Llama-3.1-8b-IQ4_XS.gguf](https://huggingface.co/bartowski/Replete-Coder-V2-Llama-3.1-8b-GGUF/blob/main/Replete-Coder-V2-Llama-3.1-8b-IQ4_XS.gguf) | IQ4_XS | 4.45GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Replete-Coder-V2-Llama-3.1-8b-Q3_K_L.gguf](https://huggingface.co/bartowski/Replete-Coder-V2-Llama-3.1-8b-GGUF/blob/main/Replete-Coder-V2-Llama-3.1-8b-Q3_K_L.gguf) | Q3_K_L | 4.32GB | false | Lower quality but usable, good for low RAM availability. |
| [Replete-Coder-V2-Llama-3.1-8b-Q3_K_M.gguf](https://huggingface.co/bartowski/Replete-Coder-V2-Llama-3.1-8b-GGUF/blob/main/Replete-Coder-V2-Llama-3.1-8b-Q3_K_M.gguf) | Q3_K_M | 4.02GB | false | Low quality. |
| [Replete-Coder-V2-Llama-3.1-8b-IQ3_M.gguf](https://huggingface.co/bartowski/Replete-Coder-V2-Llama-3.1-8b-GGUF/blob/main/Replete-Coder-V2-Llama-3.1-8b-IQ3_M.gguf) | IQ3_M | 3.78GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Replete-Coder-V2-Llama-3.1-8b-Q2_K_L.gguf](https://huggingface.co/bartowski/Replete-Coder-V2-Llama-3.1-8b-GGUF/blob/main/Replete-Coder-V2-Llama-3.1-8b-Q2_K_L.gguf) | Q2_K_L | 3.69GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Replete-Coder-V2-Llama-3.1-8b-Q3_K_S.gguf](https://huggingface.co/bartowski/Replete-Coder-V2-Llama-3.1-8b-GGUF/blob/main/Replete-Coder-V2-Llama-3.1-8b-Q3_K_S.gguf) | Q3_K_S | 3.66GB | false | Low quality, not recommended. |
| [Replete-Coder-V2-Llama-3.1-8b-IQ3_XS.gguf](https://huggingface.co/bartowski/Replete-Coder-V2-Llama-3.1-8b-GGUF/blob/main/Replete-Coder-V2-Llama-3.1-8b-IQ3_XS.gguf) | IQ3_XS | 3.52GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Replete-Coder-V2-Llama-3.1-8b-Q2_K.gguf](https://huggingface.co/bartowski/Replete-Coder-V2-Llama-3.1-8b-GGUF/blob/main/Replete-Coder-V2-Llama-3.1-8b-Q2_K.gguf) | Q2_K | 3.18GB | false | Very low quality but surprisingly usable. |
| [Replete-Coder-V2-Llama-3.1-8b-IQ2_M.gguf](https://huggingface.co/bartowski/Replete-Coder-V2-Llama-3.1-8b-GGUF/blob/main/Replete-Coder-V2-Llama-3.1-8b-IQ2_M.gguf) | IQ2_M | 2.95GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset
Thank you ZeroWw for the inspiration to experiment with embed/output
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Replete-Coder-V2-Llama-3.1-8b-GGUF --include "Replete-Coder-V2-Llama-3.1-8b-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Replete-Coder-V2-Llama-3.1-8b-GGUF --include "Replete-Coder-V2-Llama-3.1-8b-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Replete-Coder-V2-Llama-3.1-8b-Q8_0) or download them all in place (./)
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7db0de95cb479c08222de46d97bbd3386722c2a9bc45b08d5c76e1a1a31f65ca
size 2948286240

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8800593bbb1c1f8ee7e1bdc05c58a3d0b9572bff80a4c14dffa11cb2b4e403f8
size 3784828704

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8aba1846594b917bb2d0c044ec8d94ba7c4991cd793e6c9611ab5c81ddcf7fe3
size 3518752544

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c30e9c3939443df05bda72bcd768c8dfa2e97a4b7480a1dd8f5b83443a970f82
size 4447668000

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8a8c9c8e9c1b497300c903515678be654818381e761208abbab798764a008e9c
size 3179136800

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d09f4ebbb63a9ed2c470751d3994127cc792d41060d5bfc10eac2a69b2cffe67
size 3692160800

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:85398a3fb6979410506e5977c838b8aa6155fd6f865a76ec14fcab0d35feb9f2
size 4321961760

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:eeca1f1f3ba919e02bdc981d2a3a2f3086cba957191cd2dfbac6ec3324321492
size 4018923296

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:140fbbae5ace75d4aedb9b922f3bbfba09d7d37970877c4532c8f84054e5dd93
size 3664504608

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c1b459f6d5b14d5017fa2a01873ee866e0dfe96f9fa8df5fa0bb44891974afe2
size 4781631264

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:96aa8396d57ad9f81050be4d8eb8a581c1702d47b62881b205ff71c91eea229f
size 5310637856

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9a4c09520b1ea6e033d0864b88ecca56b1358c09a4d0b155d89715e2d7d88168
size 4920739616

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:61a8e51b4739d64c51bdf697d1d8be30328d0d3024d7608f562b7093fe1f80a4
size 4692674336

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6e0b3863d2651d6321ef84e3260313192d8b98914a7f9dc6f5f7792938efc00c
size 6057223968

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7685e4118f4230ba54a95c83ffec57f079b4d3ff77e9814b81b6272cb7acd7bd
size 5732992800

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4c093cedcf6e4a5de91571582463a3f475f76ca9c86293df1c255a0a4c054b96
size 5599299360

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:32747669962c203c91ac39f8a8c7dc6caea748127768c02dc4abd8a36efa662b
size 6596011808

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0cfe9a4e69f4bf9598fe772afdb80834a1c5b6d359cd2884cc39be7a7e016827
size 6850471712

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:13c62e9877ad5f8872161bdd4d6436749d5d62abc53d8c9e9ba9419aaed27258
size 8540776224

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d42b7e88f57a0f14338a84c92798a074fa53c0a8ec77a0d9373280cfd8cb57b3
size 16068896256

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c214d1177ae57fed3d36c03c276d4fe1e6dbdf2cad774c85e0f0b8f444df16f9
size 4988170

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}