初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Llama-3-8B-Instruct-Coder-v2-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-02 21:37:13 +08:00
commit a65a1c08ee
27 changed files with 242 additions and 0 deletions

59
.gitattributes vendored Normal file
View File

@@ -0,0 +1,59 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Instruct-Coder-v2-IQ1_M.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Instruct-Coder-v2-IQ1_S.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Instruct-Coder-v2-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Instruct-Coder-v2-IQ2_S.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Instruct-Coder-v2-IQ2_XS.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Instruct-Coder-v2-IQ2_XXS.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Instruct-Coder-v2-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Instruct-Coder-v2-IQ3_S.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Instruct-Coder-v2-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Instruct-Coder-v2-IQ3_XXS.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Instruct-Coder-v2-IQ4_NL.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Instruct-Coder-v2-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Instruct-Coder-v2-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Instruct-Coder-v2-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Instruct-Coder-v2-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Instruct-Coder-v2-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Instruct-Coder-v2-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Instruct-Coder-v2-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Instruct-Coder-v2-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Instruct-Coder-v2-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Instruct-Coder-v2-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Instruct-Coder-v2-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Instruct-Coder-v2.imatrix filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Instruct-Coder-v2-fp16.gguf filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cd1bd0b64a8c731d1f12c2088e27ef4ceb1667f91dbc61d6d742d77158299de8
size 2161971808

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:77af5c62245e5f23b1308ef2bf6fd70b3e9959bfac90edee2ed085947bf26d16
size 2019627616

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f9676cf8086e0a577609f103055ab7705a8f61de5d3a963f2771c7ca24ff70a8
size 2948280928

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f9275f652939fc682d81647943c92883a7f23b0f230c79aa46598d284840d48b
size 2758488672

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:56e6cf1906d0833240f178f466e81ebd44267f7c8203ed9959ef8f8292033578
size 2605781600

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d4cd08ba9752ac6c355c1dad69a3e1e92a801f9beb20240b7ee4a3e2740bb033
size 2399212128

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8aae5f1b29477793ea74931d676a7e7bbb7b18c27213794d86b6933bf936625e
size 3784823392

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:92d6707ecc73767ea4b43f74d64ad6ebe293e70877d460d878e29f9b041e523d
size 3682325088

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8a5194da681ce7a23ffe8fbae2e9bd3a0fa990e4bc92e739e1958398ee8d12c9
size 3518747232

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8c6e8632d16a9114ccd74b100a13c94efb01f314f619327ca039624ebae314c5
size 3274912352

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c223a80e37aaf083c5891a139f58e5ec93d1a01fcea10b30a37dcea5432622e7
size 4677988960

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:30c57c0bf22acbca2eef91b785d3f1adaf293be529590404a7b72d127a4fb7f4
size 4447662688

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f4cfbca4cc3ce1f270148148ef53cd28ec7283a7cbf4692e3e4261f65011b032
size 3179131488

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:49aaa1f8ea4e16a6a86e6a214967dd3be25ceb29835983541797c85dbc87d1d6
size 4321956448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2d31b34a2d1ccfb4e17023047d6767bd195b10e697922a51a5d94957aea46f15
size 4018917984

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5e24093137a2372f3c97d856b30a53f2247cdda3f4b6bc7dc6517ea39def5bcc
size 3664499296

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8610d8033a2ed28006aa68cc9816b4338b9d6edbf539d1a547781e70ed0559f1
size 4920734304

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:dcd8b5b51b3cc6347d8382358d405d8c9cbb4eb4cacf5a0d743d1797ec3c5ee0
size 4692669024

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1200c91abf342a7cb2038b3393ddf16b2fcfb8ad62a3a713cc3111080f707b69
size 5732987488

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3f3fa07a2d857386fcaf71de2446ce0debf9fa9da06eb9124b40c708f7014edb
size 5599294048

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7d95053dd0d44d5baacda62d5923b559092f6f6e50b1f13536a46f516d238a9a
size 6596006496

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:73a6ce3778782683a4d1a6fcee0fe1f7b325c9a4b1461c588cf0ab16203ccb07
size 8540770912

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:be15f49cc623a4fbfdf868852230bcdc07821b8eb954687575783279178d2b53
size 16068891712

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:97791a0277561d35d21cf759ccada2c69efafa8f4340482f045919c1e8441510
size 4988166

110
README.md Normal file
View File

@@ -0,0 +1,110 @@
---
language:
- en
license: apache-2.0
tags:
- text-generation-inference
- transformers
- unsloth
- llama
- trl
- sft
base_model: NousResearch/Meta-Llama-3-8B-Instruct
quantized_by: bartowski
pipeline_tag: text-generation
---
## Llamacpp imatrix Quantizations of Llama-3-8B-Instruct-Coder-v2
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b2794">b2794</a> for quantization.
Original model: https://huggingface.co/rombodawg/Llama-3-8B-Instruct-Coder-v2
All quants made using imatrix option with dataset provided by Kalomaze [here](https://github.com/ggerganov/llama.cpp/discussions/5263#discussioncomment-8395384)
## Prompt format
```
<|begin_of_text|><|start_header_id|>system<|end_header_id|>
{system_prompt}<|eot_id|><|start_header_id|>user<|end_header_id|>
{prompt}<|eot_id|><|start_header_id|>assistant<|end_header_id|>
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Description |
| -------- | ---------- | --------- | ----------- |
| [Llama-3-8B-Instruct-Coder-v2-Q8_0.gguf](https://huggingface.co/bartowski/Llama-3-8B-Instruct-Coder-v2-GGUF/blob/main/Llama-3-8B-Instruct-Coder-v2-Q8_0.gguf) | Q8_0 | 8.54GB | Extremely high quality, generally unneeded but max available quant. |
| [Llama-3-8B-Instruct-Coder-v2-Q6_K.gguf](https://huggingface.co/bartowski/Llama-3-8B-Instruct-Coder-v2-GGUF/blob/main/Llama-3-8B-Instruct-Coder-v2-Q6_K.gguf) | Q6_K | 6.59GB | Very high quality, near perfect, *recommended*. |
| [Llama-3-8B-Instruct-Coder-v2-Q5_K_M.gguf](https://huggingface.co/bartowski/Llama-3-8B-Instruct-Coder-v2-GGUF/blob/main/Llama-3-8B-Instruct-Coder-v2-Q5_K_M.gguf) | Q5_K_M | 5.73GB | High quality, *recommended*. |
| [Llama-3-8B-Instruct-Coder-v2-Q5_K_S.gguf](https://huggingface.co/bartowski/Llama-3-8B-Instruct-Coder-v2-GGUF/blob/main/Llama-3-8B-Instruct-Coder-v2-Q5_K_S.gguf) | Q5_K_S | 5.59GB | High quality, *recommended*. |
| [Llama-3-8B-Instruct-Coder-v2-Q4_K_M.gguf](https://huggingface.co/bartowski/Llama-3-8B-Instruct-Coder-v2-GGUF/blob/main/Llama-3-8B-Instruct-Coder-v2-Q4_K_M.gguf) | Q4_K_M | 4.92GB | Good quality, uses about 4.83 bits per weight, *recommended*. |
| [Llama-3-8B-Instruct-Coder-v2-Q4_K_S.gguf](https://huggingface.co/bartowski/Llama-3-8B-Instruct-Coder-v2-GGUF/blob/main/Llama-3-8B-Instruct-Coder-v2-Q4_K_S.gguf) | Q4_K_S | 4.69GB | Slightly lower quality with more space savings, *recommended*. |
| [Llama-3-8B-Instruct-Coder-v2-IQ4_NL.gguf](https://huggingface.co/bartowski/Llama-3-8B-Instruct-Coder-v2-GGUF/blob/main/Llama-3-8B-Instruct-Coder-v2-IQ4_NL.gguf) | IQ4_NL | 4.67GB | Decent quality, slightly smaller than Q4_K_S with similar performance *recommended*. |
| [Llama-3-8B-Instruct-Coder-v2-IQ4_XS.gguf](https://huggingface.co/bartowski/Llama-3-8B-Instruct-Coder-v2-GGUF/blob/main/Llama-3-8B-Instruct-Coder-v2-IQ4_XS.gguf) | IQ4_XS | 4.44GB | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Llama-3-8B-Instruct-Coder-v2-Q3_K_L.gguf](https://huggingface.co/bartowski/Llama-3-8B-Instruct-Coder-v2-GGUF/blob/main/Llama-3-8B-Instruct-Coder-v2-Q3_K_L.gguf) | Q3_K_L | 4.32GB | Lower quality but usable, good for low RAM availability. |
| [Llama-3-8B-Instruct-Coder-v2-Q3_K_M.gguf](https://huggingface.co/bartowski/Llama-3-8B-Instruct-Coder-v2-GGUF/blob/main/Llama-3-8B-Instruct-Coder-v2-Q3_K_M.gguf) | Q3_K_M | 4.01GB | Even lower quality. |
| [Llama-3-8B-Instruct-Coder-v2-IQ3_M.gguf](https://huggingface.co/bartowski/Llama-3-8B-Instruct-Coder-v2-GGUF/blob/main/Llama-3-8B-Instruct-Coder-v2-IQ3_M.gguf) | IQ3_M | 3.78GB | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Llama-3-8B-Instruct-Coder-v2-IQ3_S.gguf](https://huggingface.co/bartowski/Llama-3-8B-Instruct-Coder-v2-GGUF/blob/main/Llama-3-8B-Instruct-Coder-v2-IQ3_S.gguf) | IQ3_S | 3.68GB | Lower quality, new method with decent performance, recommended over Q3_K_S quant, same size with better performance. |
| [Llama-3-8B-Instruct-Coder-v2-Q3_K_S.gguf](https://huggingface.co/bartowski/Llama-3-8B-Instruct-Coder-v2-GGUF/blob/main/Llama-3-8B-Instruct-Coder-v2-Q3_K_S.gguf) | Q3_K_S | 3.66GB | Low quality, not recommended. |
| [Llama-3-8B-Instruct-Coder-v2-IQ3_XS.gguf](https://huggingface.co/bartowski/Llama-3-8B-Instruct-Coder-v2-GGUF/blob/main/Llama-3-8B-Instruct-Coder-v2-IQ3_XS.gguf) | IQ3_XS | 3.51GB | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Llama-3-8B-Instruct-Coder-v2-IQ3_XXS.gguf](https://huggingface.co/bartowski/Llama-3-8B-Instruct-Coder-v2-GGUF/blob/main/Llama-3-8B-Instruct-Coder-v2-IQ3_XXS.gguf) | IQ3_XXS | 3.27GB | Lower quality, new method with decent performance, comparable to Q3 quants. |
| [Llama-3-8B-Instruct-Coder-v2-Q2_K.gguf](https://huggingface.co/bartowski/Llama-3-8B-Instruct-Coder-v2-GGUF/blob/main/Llama-3-8B-Instruct-Coder-v2-Q2_K.gguf) | Q2_K | 3.17GB | Very low quality but surprisingly usable. |
| [Llama-3-8B-Instruct-Coder-v2-IQ2_M.gguf](https://huggingface.co/bartowski/Llama-3-8B-Instruct-Coder-v2-GGUF/blob/main/Llama-3-8B-Instruct-Coder-v2-IQ2_M.gguf) | IQ2_M | 2.94GB | Very low quality, uses SOTA techniques to also be surprisingly usable. |
| [Llama-3-8B-Instruct-Coder-v2-IQ2_S.gguf](https://huggingface.co/bartowski/Llama-3-8B-Instruct-Coder-v2-GGUF/blob/main/Llama-3-8B-Instruct-Coder-v2-IQ2_S.gguf) | IQ2_S | 2.75GB | Very low quality, uses SOTA techniques to be usable. |
| [Llama-3-8B-Instruct-Coder-v2-IQ2_XS.gguf](https://huggingface.co/bartowski/Llama-3-8B-Instruct-Coder-v2-GGUF/blob/main/Llama-3-8B-Instruct-Coder-v2-IQ2_XS.gguf) | IQ2_XS | 2.60GB | Very low quality, uses SOTA techniques to be usable. |
| [Llama-3-8B-Instruct-Coder-v2-IQ2_XXS.gguf](https://huggingface.co/bartowski/Llama-3-8B-Instruct-Coder-v2-GGUF/blob/main/Llama-3-8B-Instruct-Coder-v2-IQ2_XXS.gguf) | IQ2_XXS | 2.39GB | Lower quality, uses SOTA techniques to be usable. |
| [Llama-3-8B-Instruct-Coder-v2-IQ1_M.gguf](https://huggingface.co/bartowski/Llama-3-8B-Instruct-Coder-v2-GGUF/blob/main/Llama-3-8B-Instruct-Coder-v2-IQ1_M.gguf) | IQ1_M | 2.16GB | Extremely low quality, *not* recommended. |
| [Llama-3-8B-Instruct-Coder-v2-IQ1_S.gguf](https://huggingface.co/bartowski/Llama-3-8B-Instruct-Coder-v2-GGUF/blob/main/Llama-3-8B-Instruct-Coder-v2-IQ1_S.gguf) | IQ1_S | 2.01GB | Extremely low quality, *not* recommended. |
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Llama-3-8B-Instruct-Coder-v2-GGUF --include "Llama-3-8B-Instruct-Coder-v2-Q4_K_M.gguf" --local-dir ./ --local-dir-use-symlinks False
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Llama-3-8B-Instruct-Coder-v2-GGUF --include "Llama-3-8B-Instruct-Coder-v2-Q8_0.gguf/*" --local-dir Llama-3-8B-Instruct-Coder-v2-Q8_0 --local-dir-use-symlinks False
```
You can either specify a new local-dir (Llama-3-8B-Instruct-Coder-v2-Q8_0) or download them all in place (./)
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}