初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Llama-3-8B-Synthia-v3.5-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-09-08 06:47:12 +08:00
commit caba517093
27 changed files with 232 additions and 0 deletions

59
.gitattributes vendored Normal file
View File

@@ -0,0 +1,59 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Synthia-v3.5-IQ1_M.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Synthia-v3.5-IQ1_S.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Synthia-v3.5-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Synthia-v3.5-IQ2_S.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Synthia-v3.5-IQ2_XS.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Synthia-v3.5-IQ2_XXS.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Synthia-v3.5-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Synthia-v3.5-IQ3_S.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Synthia-v3.5-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Synthia-v3.5-IQ3_XXS.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Synthia-v3.5-IQ4_NL.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Synthia-v3.5-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Synthia-v3.5-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Synthia-v3.5-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Synthia-v3.5-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Synthia-v3.5-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Synthia-v3.5-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Synthia-v3.5-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Synthia-v3.5-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Synthia-v3.5-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Synthia-v3.5-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Synthia-v3.5-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Synthia-v3.5-f16.gguf filter=lfs diff=lfs merge=lfs -text
Llama-3-8B-Synthia-v3.5.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8bb64c90460c7683e335ca2a1f89a7db0c6480d2ffed43a8be40271937ae9aed
size 2161971936

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ed70d0f5225d976f3d9dbe5de16401986b3fd7f82453c567a8af6e693cc518ee
size 2019627744

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c6d310e417f120b713666ed4470f126acfbab32f17c5a9bbdd4943a6a02123d1
size 2948281056

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b83aab4f85f49b2330f88b6d793955633a50d713bd0e39bbe9281ce8a8b3715f
size 2758488800

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d5b6d35a1534e6218e1fd05b5801c387e6b46932e332121ea0322f13a5632c8a
size 2605781728

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6c058c8e92409d109f70b3c8902b9100ec2317d1b974365b7b5135c9023b6071
size 2399212256

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:65d7c3cd7a6c7487fc3645844317ad61253d7db03875cc159e698d8462413ccb
size 3784823520

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:48abc36d9aab6a97c58c5fc94060bec2b09512697d1334444faf3e20f1936a32
size 3682325216

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d93ed5d885367e18df8abb4bde35e09148769aaebbaab134bab52c176098789f
size 3518747360

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:787e4b81c9cff29bb5100fed7fa112669843b62c4b6c5a8dc8c6f1161d961304
size 3274912480

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:98f03a62a405ba2b7a0d0566739a976bf875959ced7fb8681604e81505e04e9d
size 4677989088

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7ab6a49bc37064a7f67fdd5538989a676033a93596b7adcb51a99cd9061c9f3b
size 4447662816

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e5778adf95833f06981e065f33f69a50ab01e40194b5a08f69a5684d376b7886
size 3179131616

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b4293f0fb1027ae0f8fff702516a00b7ac5a9978ecc4ca6df784c0b2996922d2
size 4321956576

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1497d72835307d468e8ceff821c2f9a099ffd607c0b69b7f49d0da8451cc7a10
size 4018918112

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c2c61ab7b49504283baf27ff1dc270ba522dca2daacc624d0d5b0227c2b223ea
size 3664499424

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9f3cc88b94d1b91fa28fb892efcc4a9579509fa64c64abdc299b20e45dd37388
size 4920734432

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5ff2884b830cf586ad9c438fa4b037c4011fef4d16604e8d833eefe9d08f500c
size 4692669152

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8a7287510f6ad1b418975315576aa1a543863645d80f73794983dbef3f13477b
size 5732987616

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:bc0e2a7ebb9c9f014718a3ac689ef7402d1c828af4641276c32787cc7432e2ea
size 5599294176

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:57854ee2f5715ac6a5e99568ade8431d3ad498599196cd9b76fc5d8c6849f6a5
size 6596006624

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:597a8001c565613d16461fa3483dfea2516a1376078a16349e24555ce7360541
size 8540771040

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:eb9ba70447d0e9e958141c89268724e32bfab83b5322234077de2b7fef32eba8
size 16068891072

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:71859bfac7abb8e6f101ff5111692dfbd2d996da1b44011494871e4c0c8cdfdd
size 4988169

100
README.md Normal file
View File

@@ -0,0 +1,100 @@
---
license: llama3
quantized_by: bartowski
pipeline_tag: text-generation
---
## Llamacpp imatrix Quantizations of Llama-3-8B-Synthia-v3.5
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b2901">b2901</a> for quantization.
Original model: https://huggingface.co/migtissera/Llama-3-8B-Synthia-v3.5
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/b6ac44691e994344625687afe3263b3a)
## Prompt format
```
<|begin_of_text|><|start_header_id|>system<|end_header_id|>
{system_prompt}<|eot_id|><|start_header_id|>user<|end_header_id|>
{prompt}<|eot_id|><|start_header_id|>assistant<|end_header_id|>
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Description |
| -------- | ---------- | --------- | ----------- |
| [Llama-3-8B-Synthia-v3.5-Q8_0.gguf](https://huggingface.co/bartowski/Llama-3-8B-Synthia-v3.5-GGUF/blob/main/Llama-3-8B-Synthia-v3.5-Q8_0.gguf) | Q8_0 | 8.54GB | Extremely high quality, generally unneeded but max available quant. |
| [Llama-3-8B-Synthia-v3.5-Q6_K.gguf](https://huggingface.co/bartowski/Llama-3-8B-Synthia-v3.5-GGUF/blob/main/Llama-3-8B-Synthia-v3.5-Q6_K.gguf) | Q6_K | 6.59GB | Very high quality, near perfect, *recommended*. |
| [Llama-3-8B-Synthia-v3.5-Q5_K_M.gguf](https://huggingface.co/bartowski/Llama-3-8B-Synthia-v3.5-GGUF/blob/main/Llama-3-8B-Synthia-v3.5-Q5_K_M.gguf) | Q5_K_M | 5.73GB | High quality, *recommended*. |
| [Llama-3-8B-Synthia-v3.5-Q5_K_S.gguf](https://huggingface.co/bartowski/Llama-3-8B-Synthia-v3.5-GGUF/blob/main/Llama-3-8B-Synthia-v3.5-Q5_K_S.gguf) | Q5_K_S | 5.59GB | High quality, *recommended*. |
| [Llama-3-8B-Synthia-v3.5-Q4_K_M.gguf](https://huggingface.co/bartowski/Llama-3-8B-Synthia-v3.5-GGUF/blob/main/Llama-3-8B-Synthia-v3.5-Q4_K_M.gguf) | Q4_K_M | 4.92GB | Good quality, uses about 4.83 bits per weight, *recommended*. |
| [Llama-3-8B-Synthia-v3.5-Q4_K_S.gguf](https://huggingface.co/bartowski/Llama-3-8B-Synthia-v3.5-GGUF/blob/main/Llama-3-8B-Synthia-v3.5-Q4_K_S.gguf) | Q4_K_S | 4.69GB | Slightly lower quality with more space savings, *recommended*. |
| [Llama-3-8B-Synthia-v3.5-IQ4_NL.gguf](https://huggingface.co/bartowski/Llama-3-8B-Synthia-v3.5-GGUF/blob/main/Llama-3-8B-Synthia-v3.5-IQ4_NL.gguf) | IQ4_NL | 4.67GB | Decent quality, slightly smaller than Q4_K_S with similar performance *recommended*. |
| [Llama-3-8B-Synthia-v3.5-IQ4_XS.gguf](https://huggingface.co/bartowski/Llama-3-8B-Synthia-v3.5-GGUF/blob/main/Llama-3-8B-Synthia-v3.5-IQ4_XS.gguf) | IQ4_XS | 4.44GB | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Llama-3-8B-Synthia-v3.5-Q3_K_L.gguf](https://huggingface.co/bartowski/Llama-3-8B-Synthia-v3.5-GGUF/blob/main/Llama-3-8B-Synthia-v3.5-Q3_K_L.gguf) | Q3_K_L | 4.32GB | Lower quality but usable, good for low RAM availability. |
| [Llama-3-8B-Synthia-v3.5-Q3_K_M.gguf](https://huggingface.co/bartowski/Llama-3-8B-Synthia-v3.5-GGUF/blob/main/Llama-3-8B-Synthia-v3.5-Q3_K_M.gguf) | Q3_K_M | 4.01GB | Even lower quality. |
| [Llama-3-8B-Synthia-v3.5-IQ3_M.gguf](https://huggingface.co/bartowski/Llama-3-8B-Synthia-v3.5-GGUF/blob/main/Llama-3-8B-Synthia-v3.5-IQ3_M.gguf) | IQ3_M | 3.78GB | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Llama-3-8B-Synthia-v3.5-IQ3_S.gguf](https://huggingface.co/bartowski/Llama-3-8B-Synthia-v3.5-GGUF/blob/main/Llama-3-8B-Synthia-v3.5-IQ3_S.gguf) | IQ3_S | 3.68GB | Lower quality, new method with decent performance, recommended over Q3_K_S quant, same size with better performance. |
| [Llama-3-8B-Synthia-v3.5-Q3_K_S.gguf](https://huggingface.co/bartowski/Llama-3-8B-Synthia-v3.5-GGUF/blob/main/Llama-3-8B-Synthia-v3.5-Q3_K_S.gguf) | Q3_K_S | 3.66GB | Low quality, not recommended. |
| [Llama-3-8B-Synthia-v3.5-IQ3_XS.gguf](https://huggingface.co/bartowski/Llama-3-8B-Synthia-v3.5-GGUF/blob/main/Llama-3-8B-Synthia-v3.5-IQ3_XS.gguf) | IQ3_XS | 3.51GB | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Llama-3-8B-Synthia-v3.5-IQ3_XXS.gguf](https://huggingface.co/bartowski/Llama-3-8B-Synthia-v3.5-GGUF/blob/main/Llama-3-8B-Synthia-v3.5-IQ3_XXS.gguf) | IQ3_XXS | 3.27GB | Lower quality, new method with decent performance, comparable to Q3 quants. |
| [Llama-3-8B-Synthia-v3.5-Q2_K.gguf](https://huggingface.co/bartowski/Llama-3-8B-Synthia-v3.5-GGUF/blob/main/Llama-3-8B-Synthia-v3.5-Q2_K.gguf) | Q2_K | 3.17GB | Very low quality but surprisingly usable. |
| [Llama-3-8B-Synthia-v3.5-IQ2_M.gguf](https://huggingface.co/bartowski/Llama-3-8B-Synthia-v3.5-GGUF/blob/main/Llama-3-8B-Synthia-v3.5-IQ2_M.gguf) | IQ2_M | 2.94GB | Very low quality, uses SOTA techniques to also be surprisingly usable. |
| [Llama-3-8B-Synthia-v3.5-IQ2_S.gguf](https://huggingface.co/bartowski/Llama-3-8B-Synthia-v3.5-GGUF/blob/main/Llama-3-8B-Synthia-v3.5-IQ2_S.gguf) | IQ2_S | 2.75GB | Very low quality, uses SOTA techniques to be usable. |
| [Llama-3-8B-Synthia-v3.5-IQ2_XS.gguf](https://huggingface.co/bartowski/Llama-3-8B-Synthia-v3.5-GGUF/blob/main/Llama-3-8B-Synthia-v3.5-IQ2_XS.gguf) | IQ2_XS | 2.60GB | Very low quality, uses SOTA techniques to be usable. |
| [Llama-3-8B-Synthia-v3.5-IQ2_XXS.gguf](https://huggingface.co/bartowski/Llama-3-8B-Synthia-v3.5-GGUF/blob/main/Llama-3-8B-Synthia-v3.5-IQ2_XXS.gguf) | IQ2_XXS | 2.39GB | Lower quality, uses SOTA techniques to be usable. |
| [Llama-3-8B-Synthia-v3.5-IQ1_M.gguf](https://huggingface.co/bartowski/Llama-3-8B-Synthia-v3.5-GGUF/blob/main/Llama-3-8B-Synthia-v3.5-IQ1_M.gguf) | IQ1_M | 2.16GB | Extremely low quality, *not* recommended. |
| [Llama-3-8B-Synthia-v3.5-IQ1_S.gguf](https://huggingface.co/bartowski/Llama-3-8B-Synthia-v3.5-GGUF/blob/main/Llama-3-8B-Synthia-v3.5-IQ1_S.gguf) | IQ1_S | 2.01GB | Extremely low quality, *not* recommended. |
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Llama-3-8B-Synthia-v3.5-GGUF --include "Llama-3-8B-Synthia-v3.5-Q4_K_M.gguf" --local-dir ./ --local-dir-use-symlinks False
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Llama-3-8B-Synthia-v3.5-GGUF --include "Llama-3-8B-Synthia-v3.5-Q8_0.gguf/*" --local-dir Llama-3-8B-Synthia-v3.5-Q8_0 --local-dir-use-symlinks False
```
You can either specify a new local-dir (Llama-3-8B-Synthia-v3.5-Q8_0) or download them all in place (./)
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}