初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Stheno-Hercules-3.1-8B-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-04 06:37:12 +08:00
commit 1394374784
28 changed files with 267 additions and 0 deletions

60
.gitattributes vendored Normal file
View File

@@ -0,0 +1,60 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Stheno-Hercules-3.1-8B-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Stheno-Hercules-3.1-8B-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Stheno-Hercules-3.1-8B-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Stheno-Hercules-3.1-8B-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Stheno-Hercules-3.1-8B-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Stheno-Hercules-3.1-8B-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Stheno-Hercules-3.1-8B-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Stheno-Hercules-3.1-8B-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Stheno-Hercules-3.1-8B-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Stheno-Hercules-3.1-8B-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
Stheno-Hercules-3.1-8B-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
Stheno-Hercules-3.1-8B-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
Stheno-Hercules-3.1-8B-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Stheno-Hercules-3.1-8B-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Stheno-Hercules-3.1-8B-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Stheno-Hercules-3.1-8B-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Stheno-Hercules-3.1-8B-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Stheno-Hercules-3.1-8B-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Stheno-Hercules-3.1-8B-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Stheno-Hercules-3.1-8B-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Stheno-Hercules-3.1-8B-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Stheno-Hercules-3.1-8B-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Stheno-Hercules-3.1-8B-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Stheno-Hercules-3.1-8B-f16.gguf filter=lfs diff=lfs merge=lfs -text
Stheno-Hercules-3.1-8B.imatrix filter=lfs diff=lfs merge=lfs -text

131
README.md Normal file
View File

@@ -0,0 +1,131 @@
---
base_model: ZeroXClem/Stheno-Hercules-3.1-8B
license: apache-2.0
pipeline_tag: text-generation
tags:
- merge
- mergekit
- lazymergekit
- Locutusque/Hercules-6.1-Llama-3.1-8B
- Sao10K/Llama-3.1-8B-Stheno-v3.4
quantized_by: bartowski
---
## Llamacpp imatrix Quantizations of Stheno-Hercules-3.1-8B
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b3878">b3878</a> for quantization.
Original model: https://huggingface.co/ZeroXClem/Stheno-Hercules-3.1-8B
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Stheno-Hercules-3.1-8B-f16.gguf](https://huggingface.co/bartowski/Stheno-Hercules-3.1-8B-GGUF/blob/main/Stheno-Hercules-3.1-8B-f16.gguf) | f16 | 16.07GB | false | Full F16 weights. |
| [Stheno-Hercules-3.1-8B-Q8_0.gguf](https://huggingface.co/bartowski/Stheno-Hercules-3.1-8B-GGUF/blob/main/Stheno-Hercules-3.1-8B-Q8_0.gguf) | Q8_0 | 8.54GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Stheno-Hercules-3.1-8B-Q6_K_L.gguf](https://huggingface.co/bartowski/Stheno-Hercules-3.1-8B-GGUF/blob/main/Stheno-Hercules-3.1-8B-Q6_K_L.gguf) | Q6_K_L | 6.85GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Stheno-Hercules-3.1-8B-Q6_K.gguf](https://huggingface.co/bartowski/Stheno-Hercules-3.1-8B-GGUF/blob/main/Stheno-Hercules-3.1-8B-Q6_K.gguf) | Q6_K | 6.60GB | false | Very high quality, near perfect, *recommended*. |
| [Stheno-Hercules-3.1-8B-Q5_K_L.gguf](https://huggingface.co/bartowski/Stheno-Hercules-3.1-8B-GGUF/blob/main/Stheno-Hercules-3.1-8B-Q5_K_L.gguf) | Q5_K_L | 6.06GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Stheno-Hercules-3.1-8B-Q5_K_M.gguf](https://huggingface.co/bartowski/Stheno-Hercules-3.1-8B-GGUF/blob/main/Stheno-Hercules-3.1-8B-Q5_K_M.gguf) | Q5_K_M | 5.73GB | false | High quality, *recommended*. |
| [Stheno-Hercules-3.1-8B-Q5_K_S.gguf](https://huggingface.co/bartowski/Stheno-Hercules-3.1-8B-GGUF/blob/main/Stheno-Hercules-3.1-8B-Q5_K_S.gguf) | Q5_K_S | 5.60GB | false | High quality, *recommended*. |
| [Stheno-Hercules-3.1-8B-Q4_K_L.gguf](https://huggingface.co/bartowski/Stheno-Hercules-3.1-8B-GGUF/blob/main/Stheno-Hercules-3.1-8B-Q4_K_L.gguf) | Q4_K_L | 5.31GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Stheno-Hercules-3.1-8B-Q4_K_M.gguf](https://huggingface.co/bartowski/Stheno-Hercules-3.1-8B-GGUF/blob/main/Stheno-Hercules-3.1-8B-Q4_K_M.gguf) | Q4_K_M | 4.92GB | false | Good quality, default size for must use cases, *recommended*. |
| [Stheno-Hercules-3.1-8B-Q3_K_XL.gguf](https://huggingface.co/bartowski/Stheno-Hercules-3.1-8B-GGUF/blob/main/Stheno-Hercules-3.1-8B-Q3_K_XL.gguf) | Q3_K_XL | 4.78GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Stheno-Hercules-3.1-8B-Q4_K_S.gguf](https://huggingface.co/bartowski/Stheno-Hercules-3.1-8B-GGUF/blob/main/Stheno-Hercules-3.1-8B-Q4_K_S.gguf) | Q4_K_S | 4.69GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Stheno-Hercules-3.1-8B-Q4_0.gguf](https://huggingface.co/bartowski/Stheno-Hercules-3.1-8B-GGUF/blob/main/Stheno-Hercules-3.1-8B-Q4_0.gguf) | Q4_0 | 4.68GB | false | Legacy format, generally not worth using over similarly sized formats |
| [Stheno-Hercules-3.1-8B-Q4_0_8_8.gguf](https://huggingface.co/bartowski/Stheno-Hercules-3.1-8B-GGUF/blob/main/Stheno-Hercules-3.1-8B-Q4_0_8_8.gguf) | Q4_0_8_8 | 4.66GB | false | Optimized for ARM inference. Requires 'sve' support (see link below). *Don't use on Mac or Windows*. |
| [Stheno-Hercules-3.1-8B-Q4_0_4_8.gguf](https://huggingface.co/bartowski/Stheno-Hercules-3.1-8B-GGUF/blob/main/Stheno-Hercules-3.1-8B-Q4_0_4_8.gguf) | Q4_0_4_8 | 4.66GB | false | Optimized for ARM inference. Requires 'i8mm' support (see link below). *Don't use on Mac or Windows*. |
| [Stheno-Hercules-3.1-8B-Q4_0_4_4.gguf](https://huggingface.co/bartowski/Stheno-Hercules-3.1-8B-GGUF/blob/main/Stheno-Hercules-3.1-8B-Q4_0_4_4.gguf) | Q4_0_4_4 | 4.66GB | false | Optimized for ARM inference. Should work well on all ARM chips, pick this if you're unsure. *Don't use on Mac or Windows*. |
| [Stheno-Hercules-3.1-8B-IQ4_XS.gguf](https://huggingface.co/bartowski/Stheno-Hercules-3.1-8B-GGUF/blob/main/Stheno-Hercules-3.1-8B-IQ4_XS.gguf) | IQ4_XS | 4.45GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Stheno-Hercules-3.1-8B-Q3_K_L.gguf](https://huggingface.co/bartowski/Stheno-Hercules-3.1-8B-GGUF/blob/main/Stheno-Hercules-3.1-8B-Q3_K_L.gguf) | Q3_K_L | 4.32GB | false | Lower quality but usable, good for low RAM availability. |
| [Stheno-Hercules-3.1-8B-Q3_K_M.gguf](https://huggingface.co/bartowski/Stheno-Hercules-3.1-8B-GGUF/blob/main/Stheno-Hercules-3.1-8B-Q3_K_M.gguf) | Q3_K_M | 4.02GB | false | Low quality. |
| [Stheno-Hercules-3.1-8B-IQ3_M.gguf](https://huggingface.co/bartowski/Stheno-Hercules-3.1-8B-GGUF/blob/main/Stheno-Hercules-3.1-8B-IQ3_M.gguf) | IQ3_M | 3.78GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Stheno-Hercules-3.1-8B-Q2_K_L.gguf](https://huggingface.co/bartowski/Stheno-Hercules-3.1-8B-GGUF/blob/main/Stheno-Hercules-3.1-8B-Q2_K_L.gguf) | Q2_K_L | 3.69GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Stheno-Hercules-3.1-8B-Q3_K_S.gguf](https://huggingface.co/bartowski/Stheno-Hercules-3.1-8B-GGUF/blob/main/Stheno-Hercules-3.1-8B-Q3_K_S.gguf) | Q3_K_S | 3.66GB | false | Low quality, not recommended. |
| [Stheno-Hercules-3.1-8B-IQ3_XS.gguf](https://huggingface.co/bartowski/Stheno-Hercules-3.1-8B-GGUF/blob/main/Stheno-Hercules-3.1-8B-IQ3_XS.gguf) | IQ3_XS | 3.52GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Stheno-Hercules-3.1-8B-Q2_K.gguf](https://huggingface.co/bartowski/Stheno-Hercules-3.1-8B-GGUF/blob/main/Stheno-Hercules-3.1-8B-Q2_K.gguf) | Q2_K | 3.18GB | false | Very low quality but surprisingly usable. |
| [Stheno-Hercules-3.1-8B-IQ2_M.gguf](https://huggingface.co/bartowski/Stheno-Hercules-3.1-8B-GGUF/blob/main/Stheno-Hercules-3.1-8B-IQ2_M.gguf) | IQ2_M | 2.95GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Stheno-Hercules-3.1-8B-GGUF --include "Stheno-Hercules-3.1-8B-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Stheno-Hercules-3.1-8B-GGUF --include "Stheno-Hercules-3.1-8B-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Stheno-Hercules-3.1-8B-Q8_0) or download them all in place (./)
## Q4_0_X_X
These are *NOT* for Metal (Apple) offloading, only ARM chips.
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset
Thank you ZeroWw for the inspiration to experiment with embed/output
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2f406873e0e680327f007d3191853c9da64df6b82c915b62ff006b843c144a4c
size 2948282208

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b25deb1599df138c181fbcbfd1d4dba8c3e90674af458f06dddc11ea932c54ab
size 3784824672

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:808ee93825b3b55339a1b46912d1a4e4b53b1a791957a368e6b28db6d81847ad
size 3518748512

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:525c4f6c1b02ac380842c45d9e1d5523488e246b4d4235e02dfdd6f5e2a9f5de
size 4447663968

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d53c724accaadeaf2d50be180257878a1807368a8d5e7a5c2a009de22dacdca9
size 3179132768

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2ff74f5501774b38be470947700ad0f81cff78513faab951c1387a132e2742db
size 3692156768

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3b4bb6882352f4ac808c1893e5d8cd0482308d127605a88845fb3ebeb0bc0a91
size 4321957728

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d3df18eb213737ec2910f56690f26cdd8b364cce9c29696048e49bd7ca04ba8e
size 4018919264

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:11522af7365e79487aee0a9aac1c310b515d6898c68cf0c4ba528c05873b8dcf
size 3664500576

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:606398b940d1d0b3db2ab3650da1902e589b4a064ba76ad98a4041f79c5add07
size 4781627232

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7e8b71d5467037829d8f8aacad117d20a06c84e5cccb889e590dc32455593edf
size 4675893088

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e4f9e0630d5e9a6a849e844f005985b576998811824d658a15ceb70a0fd057af
size 4661213024

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8ed98278c4427e3208f2f9489379ec784e0e0eb3f906413361a56703a6d3c9f8
size 4661213024

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5ecef6fbeccfdec82bd071dab7c57f75a2d4ef6ad12f4880077b45f3f7f553b0
size 4661213024

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d56fd50b3e3608701b7e2b5be57320b41912e47178c0704b50e829db4d3b0f26
size 5310633824

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7eeb322012504169ee96f5cb474bcc6ac442acdfff691b5f6e8fc6eca0a88bcf
size 4920735584

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1ef24126d96824b901baf2fc749b8496772a8f0cbabf745c230c8026e65f882f
size 4692670304

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6f190a86bc185502e810da31e60e091383a1d3904a7592122eb5d248428357c2
size 6057219936

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fdb88529816799d3646902b3e2851ce317917739ab6e87b4408884a969f7853b
size 5732988768

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f33891e99b8fb8b2e42dc5afbceb80aab5896241da1cf58d1fa8ffa664cf936a
size 5599295328

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e08135817239d9d16629b1dbf1a87ad3912f591782e0f16b269ce464675a96ea
size 6596007776

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:377d58e0e2d4b0f0cb658e11046f73c23df4901c452d8665e7cc4f68712fe06f
size 6850467680

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:68e4f60e9ba23c71e2fe834a229288ec888f90f2db9a52d861e437d2ca182c2d
size 8540772192

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8933af62d02fd545c356298473a7c357ed8583a91c8b8be1c3d0306ecb7741ee
size 16068892224

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d3acc016c3373101063288d3e77335dd4c7dd51ee77f25ce4d5d0f589e9fe956
size 4988170

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}