初始化项目,由ModelHub XC社区提供模型

Model: bartowski/GeM2-Llamion-14B-LongChat-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-06 11:56:14 +08:00
commit c4a21cd757
23 changed files with 211 additions and 0 deletions

55
.gitattributes vendored Normal file
View File

@@ -0,0 +1,55 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
GeM2-Llamion-14B-LongChat-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
GeM2-Llamion-14B-LongChat-IQ2_S.gguf filter=lfs diff=lfs merge=lfs -text
GeM2-Llamion-14B-LongChat-IQ2_XS.gguf filter=lfs diff=lfs merge=lfs -text
GeM2-Llamion-14B-LongChat-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
GeM2-Llamion-14B-LongChat-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
GeM2-Llamion-14B-LongChat-IQ3_XXS.gguf filter=lfs diff=lfs merge=lfs -text
GeM2-Llamion-14B-LongChat-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
GeM2-Llamion-14B-LongChat-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
GeM2-Llamion-14B-LongChat-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
GeM2-Llamion-14B-LongChat-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
GeM2-Llamion-14B-LongChat-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
GeM2-Llamion-14B-LongChat-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
GeM2-Llamion-14B-LongChat-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
GeM2-Llamion-14B-LongChat-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
GeM2-Llamion-14B-LongChat-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
GeM2-Llamion-14B-LongChat-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
GeM2-Llamion-14B-LongChat-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
GeM2-Llamion-14B-LongChat-f32.gguf/GeM2-Llamion-14B-LongChat-f32-00001-of-00002.gguf filter=lfs diff=lfs merge=lfs -text
GeM2-Llamion-14B-LongChat-f32.gguf/GeM2-Llamion-14B-LongChat-f32-00002-of-00002.gguf filter=lfs diff=lfs merge=lfs -text
GeM2-Llamion-14B-LongChat.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:18e303dde606e5144aba90ef947040db387c3817bb9f3c6f9678357b3d2b12b2
size 5126231840

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8cc585fdae27e5e2ddb43bf1e78b213693f4e1ccf704018f3ac257a556e38dfe
size 4778071840

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:32c724b609c99ed5b3c7b7f0099b2ad342da72eb7d924848fa17676c244e1400
size 4440187680

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:bdddf06f7c9be463504a9e73a3b167622a49d17863cf995cea37fe96858e0001
size 6733077280

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f392e30dd9a4a7cbbc45bbdb83551cea07ebca20b4ee1ffed474f9fa3b65b8ee
size 6082837280

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:379afeff2c902e0b49b2857a0f0ab40a0e0b84b164d0f99aa702739c0b14125b
size 5623895840

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:94a959469d8ed7166b6c31dc94939e52c2967a445c05e355a39fee76d1028c7d
size 7830769440

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d0e74f23f2638625fb4e46f6aa4a85dad91f9e7dd833dcd7a061c5a767e1021c
size 5506361120

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:86811fed9270eb07608521e4ff852519c21c0971d29ea1c28089da6bbc9300c2
size 7754005280

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5d34b4515f8db5bfdc81447a0398f5f0c50d343aa78330438d498cbbf0f68980
size 7124859680

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:83d62bbaa51ced5f17e00bf134b8581f7f4c22ccaf1330b642daeaa59cbbee2e
size 6402325280

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d2c816254c73e366131679a49671e2d7f737fef06fc9f8bee101dd264c4475ba
size 8810962720

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1c5b7b2538c68a89978b4002ca08dcea65836fde8b101d2ba4d9bc366b7715ee
size 8332549920

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8a4362312a75b9373c3d7bd79e1dc9e9be91eb882e0ec3bcd80fd3759c67a136
size 10306903840

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c624874faa9df7d2ff7ff91f996ca7b80d115947feebc5098b65e335a28847e8
size 10028375840

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fd01a0663a22874394419d3ebcc95bb9452afdcbccf720e7711149198cb2c8d4
size 11896341280

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:bb5a52a347d7e585f50c0f04024fb3b3dfcef26f4438830894267fffc700e013
size 15407545120

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8d5e8079757c8aa5af4f6da070677883b1875bc79aa7f5a6ab254b541dd14e10
size 42944947552

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:de29f27ce57f3b165d026c17a510b851f7231850e4bf3b446cdf4099dcdf5d4d
size 15050102112

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c61f24ff601add17d7889f3d139472987bf2f1d2f9d31215c3a0f741b5461298
size 7382099

95
README.md Normal file
View File

@@ -0,0 +1,95 @@
---
license: apache-2.0
quantized_by: bartowski
pipeline_tag: text-generation
---
## Llamacpp imatrix Quantizations of GeM2-Llamion-14B-LongChat
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b3086">b3086</a> for quantization.
Original model: https://huggingface.co/vaiv/GeM2-Llamion-14B-LongChat
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
## Prompt format
```
<|begin_of_text|><|start_header_id|> system<|end_header_id|>
{system_prompt}<|eot_id|><|start_header_id|> user<|end_header_id|>
{prompt}<|eot_id|><|start_header_id|> assistant<|end_header_id|>
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Description |
| -------- | ---------- | --------- | ----------- |
| [GeM2-Llamion-14B-LongChat-Q8_0.gguf](https://huggingface.co/bartowski/GeM2-Llamion-14B-LongChat-GGUF/blob/main/GeM2-Llamion-14B-LongChat-Q8_0.gguf) | Q8_0 | 15.40GB | Extremely high quality, generally unneeded but max available quant. |
| [GeM2-Llamion-14B-LongChat-Q6_K.gguf](https://huggingface.co/bartowski/GeM2-Llamion-14B-LongChat-GGUF/blob/main/GeM2-Llamion-14B-LongChat-Q6_K.gguf) | Q6_K | 11.89GB | Very high quality, near perfect, *recommended*. |
| [GeM2-Llamion-14B-LongChat-Q5_K_M.gguf](https://huggingface.co/bartowski/GeM2-Llamion-14B-LongChat-GGUF/blob/main/GeM2-Llamion-14B-LongChat-Q5_K_M.gguf) | Q5_K_M | 10.30GB | High quality, *recommended*. |
| [GeM2-Llamion-14B-LongChat-Q5_K_S.gguf](https://huggingface.co/bartowski/GeM2-Llamion-14B-LongChat-GGUF/blob/main/GeM2-Llamion-14B-LongChat-Q5_K_S.gguf) | Q5_K_S | 10.02GB | High quality, *recommended*. |
| [GeM2-Llamion-14B-LongChat-Q4_K_M.gguf](https://huggingface.co/bartowski/GeM2-Llamion-14B-LongChat-GGUF/blob/main/GeM2-Llamion-14B-LongChat-Q4_K_M.gguf) | Q4_K_M | 8.81GB | Good quality, uses about 4.83 bits per weight, *recommended*. |
| [GeM2-Llamion-14B-LongChat-Q4_K_S.gguf](https://huggingface.co/bartowski/GeM2-Llamion-14B-LongChat-GGUF/blob/main/GeM2-Llamion-14B-LongChat-Q4_K_S.gguf) | Q4_K_S | 8.33GB | Slightly lower quality with more space savings, *recommended*. |
| [GeM2-Llamion-14B-LongChat-IQ4_XS.gguf](https://huggingface.co/bartowski/GeM2-Llamion-14B-LongChat-GGUF/blob/main/GeM2-Llamion-14B-LongChat-IQ4_XS.gguf) | IQ4_XS | 7.83GB | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [GeM2-Llamion-14B-LongChat-Q3_K_L.gguf](https://huggingface.co/bartowski/GeM2-Llamion-14B-LongChat-GGUF/blob/main/GeM2-Llamion-14B-LongChat-Q3_K_L.gguf) | Q3_K_L | 7.75GB | Lower quality but usable, good for low RAM availability. |
| [GeM2-Llamion-14B-LongChat-Q3_K_M.gguf](https://huggingface.co/bartowski/GeM2-Llamion-14B-LongChat-GGUF/blob/main/GeM2-Llamion-14B-LongChat-Q3_K_M.gguf) | Q3_K_M | 7.12GB | Even lower quality. |
| [GeM2-Llamion-14B-LongChat-IQ3_M.gguf](https://huggingface.co/bartowski/GeM2-Llamion-14B-LongChat-GGUF/blob/main/GeM2-Llamion-14B-LongChat-IQ3_M.gguf) | IQ3_M | 6.73GB | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [GeM2-Llamion-14B-LongChat-Q3_K_S.gguf](https://huggingface.co/bartowski/GeM2-Llamion-14B-LongChat-GGUF/blob/main/GeM2-Llamion-14B-LongChat-Q3_K_S.gguf) | Q3_K_S | 6.40GB | Low quality, not recommended. |
| [GeM2-Llamion-14B-LongChat-IQ3_XS.gguf](https://huggingface.co/bartowski/GeM2-Llamion-14B-LongChat-GGUF/blob/main/GeM2-Llamion-14B-LongChat-IQ3_XS.gguf) | IQ3_XS | 6.08GB | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [GeM2-Llamion-14B-LongChat-IQ3_XXS.gguf](https://huggingface.co/bartowski/GeM2-Llamion-14B-LongChat-GGUF/blob/main/GeM2-Llamion-14B-LongChat-IQ3_XXS.gguf) | IQ3_XXS | 5.62GB | Lower quality, new method with decent performance, comparable to Q3 quants. |
| [GeM2-Llamion-14B-LongChat-Q2_K.gguf](https://huggingface.co/bartowski/GeM2-Llamion-14B-LongChat-GGUF/blob/main/GeM2-Llamion-14B-LongChat-Q2_K.gguf) | Q2_K | 5.50GB | Very low quality but surprisingly usable. |
| [GeM2-Llamion-14B-LongChat-IQ2_M.gguf](https://huggingface.co/bartowski/GeM2-Llamion-14B-LongChat-GGUF/blob/main/GeM2-Llamion-14B-LongChat-IQ2_M.gguf) | IQ2_M | 5.12GB | Very low quality, uses SOTA techniques to also be surprisingly usable. |
| [GeM2-Llamion-14B-LongChat-IQ2_S.gguf](https://huggingface.co/bartowski/GeM2-Llamion-14B-LongChat-GGUF/blob/main/GeM2-Llamion-14B-LongChat-IQ2_S.gguf) | IQ2_S | 4.77GB | Very low quality, uses SOTA techniques to be usable. |
| [GeM2-Llamion-14B-LongChat-IQ2_XS.gguf](https://huggingface.co/bartowski/GeM2-Llamion-14B-LongChat-GGUF/blob/main/GeM2-Llamion-14B-LongChat-IQ2_XS.gguf) | IQ2_XS | 4.44GB | Very low quality, uses SOTA techniques to be usable. |
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/GeM2-Llamion-14B-LongChat-GGUF --include "GeM2-Llamion-14B-LongChat-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/GeM2-Llamion-14B-LongChat-GGUF --include "GeM2-Llamion-14B-LongChat-Q8_0.gguf/*" --local-dir GeM2-Llamion-14B-LongChat-Q8_0
```
You can either specify a new local-dir (GeM2-Llamion-14B-LongChat-Q8_0) or download them all in place (./)
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}