初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Hathor_Gamma-L3-8B-0.6-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-07 14:58:13 +08:00
commit 02587ee346
26 changed files with 230 additions and 0 deletions

58
.gitattributes vendored Normal file
View File

@@ -0,0 +1,58 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Hathor_Gamma-L3-8B-0.6-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Hathor_Gamma-L3-8B-0.6-IQ2_S.gguf filter=lfs diff=lfs merge=lfs -text
Hathor_Gamma-L3-8B-0.6-IQ2_XS.gguf filter=lfs diff=lfs merge=lfs -text
Hathor_Gamma-L3-8B-0.6-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Hathor_Gamma-L3-8B-0.6-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Hathor_Gamma-L3-8B-0.6-IQ3_XXS.gguf filter=lfs diff=lfs merge=lfs -text
Hathor_Gamma-L3-8B-0.6-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Hathor_Gamma-L3-8B-0.6-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Hathor_Gamma-L3-8B-0.6-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Hathor_Gamma-L3-8B-0.6-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Hathor_Gamma-L3-8B-0.6-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Hathor_Gamma-L3-8B-0.6-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Hathor_Gamma-L3-8B-0.6-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Hathor_Gamma-L3-8B-0.6-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Hathor_Gamma-L3-8B-0.6-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Hathor_Gamma-L3-8B-0.6-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Hathor_Gamma-L3-8B-0.6-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Hathor_Gamma-L3-8B-0.6-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Hathor_Gamma-L3-8B-0.6-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Hathor_Gamma-L3-8B-0.6-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Hathor_Gamma-L3-8B-0.6-Q8_0_L.gguf filter=lfs diff=lfs merge=lfs -text
Hathor_Gamma-L3-8B-0.6-f32.gguf filter=lfs diff=lfs merge=lfs -text
Hathor_Gamma-L3-8B-0.6.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:979e0768b6ba7643b11a3d18c5fb59ef578696375b935193731447c0622c35e4
size 2948280896

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:13a3c46612681fcd7411071b5b0904de611022b8edcb0b93fef11de2b3b124c5
size 2758488640

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4337d86edc4d16614d64dd44e085af4dfc409972bb3a5286da6c4b5ae1e1475d
size 2605781568

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e8d011b75b81755d58ec72e2c1070d111bfa3ceb89ae9df290099a4bf091ca53
size 3784823360

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:820c11650d811323f1a41f2d8cb0fd976465b38315dcb0d288f7ec4f7e12050a
size 3518747200

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f1000db3c7e347fe112e67c8f573b365e369a65e6b6f1c1b7865c0d08c9419fa
size 3274912320

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4815f45140a4db60ea535f42b39f1643f535049cbd82b7e663f6f0b491013b48
size 4447662656

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c82f7708770bf1d19c312ab22ae5b8e5d5dd3f670dc496af211794a7bcbf6ad4
size 3179131456

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ac8bfac2e5856706eb0c77ddeff2fcdd356567a73abeb3dbc1e8631a8fe4febe
size 4321956416

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ac2a86fbc1c65d965cc0eb43e999c1decd7296793eccd0ec94f1db83dbf8b468
size 4018917952

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a2987beae6cbc5c662601d808715a74fd798c30dda7b6e07f10c726ae9fe7621
size 3664499264

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6f4b1894f08f0045e03b56756d483e06587d5f8edb1d72b3ae4d9bdcfe1bc268
size 6295638592

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c01ae614fcee69d77af953b422a4a0b6a8418bdd8202a2873ffc3b4342c3aae3
size 4920734272

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:45ff713eb9a80fe76e8765e002968ecf2b291b932a6a83bd1d39707c47e5587f
size 4692668992

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2785b783a49d9d89fe5b5f738ad97edfeb2e1cc58c674742309071f076ae6115
size 7042224704

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:eba85ee172650714bcfbd5d9a6554bca529a9802340ff1f8f2db3e790ece8c63
size 5732987456

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:908741567c65b02c34927fb8724a04e8713cac14f1e72727fefc228d527f8a2a
size 5599294016

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:22512b79f3f01ad40898e04ec11b94beded8c679dd548eba1d72d9d1ab57dc71
size 6596006464

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:325bac5c8ffabcffed6bc3d4bf6bc0dcf7e39f6fd315af52598e446003fddb32
size 7835472448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:14a5daa4f6a0d8291cad5d8d43f43279f72671015ab03ccdc005deb0c614b393
size 8540770880

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ecca8f7e52fa51f2798835a507d9b2db8582ae987517f577f220775840224707
size 9525776960

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3ad4027e8a7a60c50e26f4e3b665dadd6e996f14517b40caafa83a0fc8cf68da
size 32128880928

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9c8eeb232f875b63ebef03d67e5612c0b7f362143fed5a96512298e83e79bdb8
size 4988171

102
README.md Normal file
View File

@@ -0,0 +1,102 @@
---
license: other
language:
- en
quantized_by: bartowski
pipeline_tag: text-generation
---
## Llamacpp imatrix Quantizations of Hathor_Gamma-L3-8B-0.6
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b3197">b3197</a> for quantization.
Original model: https://huggingface.co/Nitral-AI/Hathor_Gamma-L3-8B-0.6
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
## Prompt format
```
<|begin_of_text|><|start_header_id|>system<|end_header_id|>
{system_prompt}<|eot_id|><|start_header_id|>user<|end_header_id|>
{prompt}<|eot_id|><|start_header_id|>assistant<|end_header_id|>
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Description |
| -------- | ---------- | --------- | ----------- |
| [Hathor_Gamma-L3-8B-0.6-Q8_0_L.gguf](https://huggingface.co/bartowski/Hathor_Gamma-L3-8B-0.6-GGUF/blob/main/Hathor_Gamma-L3-8B-0.6-Q8_1.gguf) | Q8_0_L | 9.52GB | *Experimental*, uses f16 for embed and output weights. Please provide any feedback of differences. Extremely high quality, generally unneeded but max available quant. |
| [Hathor_Gamma-L3-8B-0.6-Q8_0.gguf](https://huggingface.co/bartowski/Hathor_Gamma-L3-8B-0.6-GGUF/blob/main/Hathor_Gamma-L3-8B-0.6-Q8_0.gguf) | Q8_0 | 8.54GB | Extremely high quality, generally unneeded but max available quant. |
| [Hathor_Gamma-L3-8B-0.6-Q6_K_L.gguf](https://huggingface.co/bartowski/Hathor_Gamma-L3-8B-0.6-GGUF/blob/main/Hathor_Gamma-L3-8B-0.6-Q6_K_L.gguf) | Q6_K_L | 7.83GB | *Experimental*, uses f16 for embed and output weights. Please provide any feedback of differences. Very high quality, near perfect, *recommended*. |
| [Hathor_Gamma-L3-8B-0.6-Q6_K.gguf](https://huggingface.co/bartowski/Hathor_Gamma-L3-8B-0.6-GGUF/blob/main/Hathor_Gamma-L3-8B-0.6-Q6_K.gguf) | Q6_K | 6.59GB | Very high quality, near perfect, *recommended*. |
| [Hathor_Gamma-L3-8B-0.6-Q5_K_L.gguf](https://huggingface.co/bartowski/Hathor_Gamma-L3-8B-0.6-GGUF/blob/main/Hathor_Gamma-L3-8B-0.6-Q5_K_L.gguf) | Q5_K_L | 7.04GB | *Experimental*, uses f16 for embed and output weights. Please provide any feedback of differences. High quality, *recommended*. |
| [Hathor_Gamma-L3-8B-0.6-Q5_K_M.gguf](https://huggingface.co/bartowski/Hathor_Gamma-L3-8B-0.6-GGUF/blob/main/Hathor_Gamma-L3-8B-0.6-Q5_K_M.gguf) | Q5_K_M | 5.73GB | High quality, *recommended*. |
| [Hathor_Gamma-L3-8B-0.6-Q5_K_S.gguf](https://huggingface.co/bartowski/Hathor_Gamma-L3-8B-0.6-GGUF/blob/main/Hathor_Gamma-L3-8B-0.6-Q5_K_S.gguf) | Q5_K_S | 5.59GB | High quality, *recommended*. |
| [Hathor_Gamma-L3-8B-0.6-Q4_K_L.gguf](https://huggingface.co/bartowski/Hathor_Gamma-L3-8B-0.6-GGUF/blob/main/Hathor_Gamma-L3-8B-0.6-Q4_K_L.gguf) | Q4_K_L | 6.29GB | *Experimental*, uses f16 for embed and output weights. Please provide any feedback of differences. Good quality, uses about 4.83 bits per weight, *recommended*. |
| [Hathor_Gamma-L3-8B-0.6-Q4_K_M.gguf](https://huggingface.co/bartowski/Hathor_Gamma-L3-8B-0.6-GGUF/blob/main/Hathor_Gamma-L3-8B-0.6-Q4_K_M.gguf) | Q4_K_M | 4.92GB | Good quality, uses about 4.83 bits per weight, *recommended*. |
| [Hathor_Gamma-L3-8B-0.6-Q4_K_S.gguf](https://huggingface.co/bartowski/Hathor_Gamma-L3-8B-0.6-GGUF/blob/main/Hathor_Gamma-L3-8B-0.6-Q4_K_S.gguf) | Q4_K_S | 4.69GB | Slightly lower quality with more space savings, *recommended*. |
| [Hathor_Gamma-L3-8B-0.6-IQ4_XS.gguf](https://huggingface.co/bartowski/Hathor_Gamma-L3-8B-0.6-GGUF/blob/main/Hathor_Gamma-L3-8B-0.6-IQ4_XS.gguf) | IQ4_XS | 4.44GB | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Hathor_Gamma-L3-8B-0.6-Q3_K_XL.gguf](https://huggingface.co/bartowski/Hathor_Gamma-L3-8B-0.6-GGUF//main/Hathor_Gamma-L3-8B-0.6-Q3_K_XL.gguf) | Q3_K_XL | | *Experimental*, uses f16 for embed and output weights. Please provide any feedback of differences. Lower quality but usable, good for low RAM availability. |
| [Hathor_Gamma-L3-8B-0.6-Q3_K_L.gguf](https://huggingface.co/bartowski/Hathor_Gamma-L3-8B-0.6-GGUF/blob/main/Hathor_Gamma-L3-8B-0.6-Q3_K_L.gguf) | Q3_K_L | 4.32GB | Lower quality but usable, good for low RAM availability. |
| [Hathor_Gamma-L3-8B-0.6-Q3_K_M.gguf](https://huggingface.co/bartowski/Hathor_Gamma-L3-8B-0.6-GGUF/blob/main/Hathor_Gamma-L3-8B-0.6-Q3_K_M.gguf) | Q3_K_M | 4.01GB | Even lower quality. |
| [Hathor_Gamma-L3-8B-0.6-IQ3_M.gguf](https://huggingface.co/bartowski/Hathor_Gamma-L3-8B-0.6-GGUF/blob/main/Hathor_Gamma-L3-8B-0.6-IQ3_M.gguf) | IQ3_M | 3.78GB | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Hathor_Gamma-L3-8B-0.6-Q3_K_S.gguf](https://huggingface.co/bartowski/Hathor_Gamma-L3-8B-0.6-GGUF/blob/main/Hathor_Gamma-L3-8B-0.6-Q3_K_S.gguf) | Q3_K_S | 3.66GB | Low quality, not recommended. |
| [Hathor_Gamma-L3-8B-0.6-IQ3_XS.gguf](https://huggingface.co/bartowski/Hathor_Gamma-L3-8B-0.6-GGUF/blob/main/Hathor_Gamma-L3-8B-0.6-IQ3_XS.gguf) | IQ3_XS | 3.51GB | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Hathor_Gamma-L3-8B-0.6-IQ3_XXS.gguf](https://huggingface.co/bartowski/Hathor_Gamma-L3-8B-0.6-GGUF/blob/main/Hathor_Gamma-L3-8B-0.6-IQ3_XXS.gguf) | IQ3_XXS | 3.27GB | Lower quality, new method with decent performance, comparable to Q3 quants. |
| [Hathor_Gamma-L3-8B-0.6-Q2_K.gguf](https://huggingface.co/bartowski/Hathor_Gamma-L3-8B-0.6-GGUF/blob/main/Hathor_Gamma-L3-8B-0.6-Q2_K.gguf) | Q2_K | 3.17GB | Very low quality but surprisingly usable. |
| [Hathor_Gamma-L3-8B-0.6-IQ2_M.gguf](https://huggingface.co/bartowski/Hathor_Gamma-L3-8B-0.6-GGUF/blob/main/Hathor_Gamma-L3-8B-0.6-IQ2_M.gguf) | IQ2_M | 2.94GB | Very low quality, uses SOTA techniques to also be surprisingly usable. |
| [Hathor_Gamma-L3-8B-0.6-IQ2_S.gguf](https://huggingface.co/bartowski/Hathor_Gamma-L3-8B-0.6-GGUF/blob/main/Hathor_Gamma-L3-8B-0.6-IQ2_S.gguf) | IQ2_S | 2.75GB | Very low quality, uses SOTA techniques to be usable. |
| [Hathor_Gamma-L3-8B-0.6-IQ2_XS.gguf](https://huggingface.co/bartowski/Hathor_Gamma-L3-8B-0.6-GGUF/blob/main/Hathor_Gamma-L3-8B-0.6-IQ2_XS.gguf) | IQ2_XS | 2.60GB | Very low quality, uses SOTA techniques to be usable. |
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Hathor_Gamma-L3-8B-0.6-GGUF --include "Hathor_Gamma-L3-8B-0.6-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Hathor_Gamma-L3-8B-0.6-GGUF --include "Hathor_Gamma-L3-8B-0.6-Q8_0.gguf/*" --local-dir Hathor_Gamma-L3-8B-0.6-Q8_0
```
You can either specify a new local-dir (Hathor_Gamma-L3-8B-0.6-Q8_0) or download them all in place (./)
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}