初始化项目,由ModelHub XC社区提供模型

Model: bartowski/MegaBeam-Mistral-7B-512k-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-19 12:48:06 +08:00
commit ffacd64cc7
24 changed files with 238 additions and 0 deletions

57
.gitattributes vendored Normal file
View File

@@ -0,0 +1,57 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
MegaBeam-Mistral-7B-512k-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
MegaBeam-Mistral-7B-512k-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
MegaBeam-Mistral-7B-512k-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
MegaBeam-Mistral-7B-512k-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
MegaBeam-Mistral-7B-512k-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
MegaBeam-Mistral-7B-512k-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
MegaBeam-Mistral-7B-512k-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
MegaBeam-Mistral-7B-512k-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
MegaBeam-Mistral-7B-512k-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
MegaBeam-Mistral-7B-512k-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
MegaBeam-Mistral-7B-512k-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
MegaBeam-Mistral-7B-512k-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
MegaBeam-Mistral-7B-512k-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
MegaBeam-Mistral-7B-512k-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
MegaBeam-Mistral-7B-512k-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
MegaBeam-Mistral-7B-512k-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
MegaBeam-Mistral-7B-512k-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
MegaBeam-Mistral-7B-512k-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
MegaBeam-Mistral-7B-512k-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
MegaBeam-Mistral-7B-512k-f32.gguf filter=lfs diff=lfs merge=lfs -text
MegaBeam-Mistral-7B-512k.imatrix filter=lfs diff=lfs merge=lfs -text
MegaBeam-Mistral-7B-512k-f16.gguf filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:13a070156ca0d9fb155d1868f9cb8ab05a58102da3ba472ab0a093a60ece6afa
size 2500713376

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:32993764913c9a5be2b5451cd427efb0f093f943d5747a81cbfaec67d2eb30b4
size 3284892576

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6da6551ec1c3175b7650bd9fc4e421730383c2566adfc1e4e79b401b98c42e94
size 3018816416

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ce571e5d43dc5f7f92dde14f2f468387e9dcc1fc9bfaf686f75e7deae848efc1
size 3907689376

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6f25509de0d9f1a658ce4686e2f6b784c9bdf1358607159d30cec9d3b4d7f52b
size 2719243168

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0ac4e21acb43ec846a518212901059bd462794fe4980df9534c702f048b27a3b
size 2847243168

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:074fcd41a1f5a4cfcf092de21b19c244950d61744bc075c3184d3846464b701e
size 3822025632

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:450fd6994603c9ead29d17eda50cc25fbe77be84cb0741122279bbda0728fe8d
size 3518987168

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2d091a6ad2cbb8708be977bfe4eb31980005361f584b75d1cf89386ed2577d1e
size 3164568480

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0da322eff85a46660daa1b1aeb22623b9854ec7795bc65253d2183e0d6d4ff5d
size 3936713632

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:829357e8a5a7a721d6c7cdcb4bbf7696ec6e914608bde6e5498905defe19cf82
size 4465720224

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7bba5d61ffda53ae6df7c4001d825fa82bcf9cbb769ae3b4ec7f9fd3c291d753
size 4368440224

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c4b993d4ede942a35118ed6c677e8579a0e533c610ebd6572f71e68186085ea5
size 4140374944

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0d7b75f5ed535a918cd3b93c2f9bea8e0c1ae003234e8cc5a74ec960a41e0711
size 5212306336

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8447ebea5eeb5e8b0d5eed4df43af095adfbc8f66b014f4ea9f50b95b5c90663
size 5131410336

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:84571aa5e7d9e5bf9e1086fd06555efaa0cc180529f26fbe5ba8b0a7dcda0ea5
size 4997716896

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9d139c0fb07db1c239d3367cd49dd17c9886214cb9b27a6e777609ada92ca270
size 5942066080

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1cf982115a7c820918ca681a84bd48116ddbd82fc8ab2a6e085da6acbb7abc79
size 6005554080

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c99568eee9732640dc69b8fb564a3ce3aec528744acbf1d0c23c56d47d9ee81a
size 7695858592

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:83c2a4ce8ecaec28cc92b2c8cef39f38f5e4abdb97dad436247962308a4f8075
size 14484732544

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:406b73017d5bbd0e6b7ad0f94ff71287e0dc79d5bcc8a816603511d8387914ab
size 4988170

117
README.md Normal file
View File

@@ -0,0 +1,117 @@
---
base_model: aws-prototyping/MegaBeam-Mistral-7B-512k
license: apache-2.0
pipeline_tag: text-generation
quantized_by: bartowski
inference: false
---
## Llamacpp imatrix Quantizations of MegaBeam-Mistral-7B-512k
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b3600">b3600</a> for quantization.
Original model: https://huggingface.co/aws-prototyping/MegaBeam-Mistral-7B-512k
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<s> [INST] {prompt} [/INST]
```
Note that this model does not support a System prompt.
## What's new:
Model updated for "improved user experience" and fixing repetition issues
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [MegaBeam-Mistral-7B-512k-f16.gguf](https://huggingface.co/bartowski/MegaBeam-Mistral-7B-512k-GGUF/blob/main/MegaBeam-Mistral-7B-512k-f16.gguf) | f16 | 14.48GB | false | Full F16 weights. |
| [MegaBeam-Mistral-7B-512k-Q8_0.gguf](https://huggingface.co/bartowski/MegaBeam-Mistral-7B-512k-GGUF/blob/main/MegaBeam-Mistral-7B-512k-Q8_0.gguf) | Q8_0 | 7.70GB | false | Extremely high quality, generally unneeded but max available quant. |
| [MegaBeam-Mistral-7B-512k-Q6_K_L.gguf](https://huggingface.co/bartowski/MegaBeam-Mistral-7B-512k-GGUF/blob/main/MegaBeam-Mistral-7B-512k-Q6_K_L.gguf) | Q6_K_L | 6.01GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [MegaBeam-Mistral-7B-512k-Q6_K.gguf](https://huggingface.co/bartowski/MegaBeam-Mistral-7B-512k-GGUF/blob/main/MegaBeam-Mistral-7B-512k-Q6_K.gguf) | Q6_K | 5.94GB | false | Very high quality, near perfect, *recommended*. |
| [MegaBeam-Mistral-7B-512k-Q5_K_L.gguf](https://huggingface.co/bartowski/MegaBeam-Mistral-7B-512k-GGUF/blob/main/MegaBeam-Mistral-7B-512k-Q5_K_L.gguf) | Q5_K_L | 5.21GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [MegaBeam-Mistral-7B-512k-Q5_K_M.gguf](https://huggingface.co/bartowski/MegaBeam-Mistral-7B-512k-GGUF/blob/main/MegaBeam-Mistral-7B-512k-Q5_K_M.gguf) | Q5_K_M | 5.13GB | false | High quality, *recommended*. |
| [MegaBeam-Mistral-7B-512k-Q5_K_S.gguf](https://huggingface.co/bartowski/MegaBeam-Mistral-7B-512k-GGUF/blob/main/MegaBeam-Mistral-7B-512k-Q5_K_S.gguf) | Q5_K_S | 5.00GB | false | High quality, *recommended*. |
| [MegaBeam-Mistral-7B-512k-Q4_K_L.gguf](https://huggingface.co/bartowski/MegaBeam-Mistral-7B-512k-GGUF/blob/main/MegaBeam-Mistral-7B-512k-Q4_K_L.gguf) | Q4_K_L | 4.47GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [MegaBeam-Mistral-7B-512k-Q4_K_M.gguf](https://huggingface.co/bartowski/MegaBeam-Mistral-7B-512k-GGUF/blob/main/MegaBeam-Mistral-7B-512k-Q4_K_M.gguf) | Q4_K_M | 4.37GB | false | Good quality, default size for must use cases, *recommended*. |
| [MegaBeam-Mistral-7B-512k-Q4_K_S.gguf](https://huggingface.co/bartowski/MegaBeam-Mistral-7B-512k-GGUF/blob/main/MegaBeam-Mistral-7B-512k-Q4_K_S.gguf) | Q4_K_S | 4.14GB | false | Slightly lower quality with more space savings, *recommended*. |
| [MegaBeam-Mistral-7B-512k-Q3_K_XL.gguf](https://huggingface.co/bartowski/MegaBeam-Mistral-7B-512k-GGUF/blob/main/MegaBeam-Mistral-7B-512k-Q3_K_XL.gguf) | Q3_K_XL | 3.94GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [MegaBeam-Mistral-7B-512k-IQ4_XS.gguf](https://huggingface.co/bartowski/MegaBeam-Mistral-7B-512k-GGUF/blob/main/MegaBeam-Mistral-7B-512k-IQ4_XS.gguf) | IQ4_XS | 3.91GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [MegaBeam-Mistral-7B-512k-Q3_K_L.gguf](https://huggingface.co/bartowski/MegaBeam-Mistral-7B-512k-GGUF/blob/main/MegaBeam-Mistral-7B-512k-Q3_K_L.gguf) | Q3_K_L | 3.82GB | false | Lower quality but usable, good for low RAM availability. |
| [MegaBeam-Mistral-7B-512k-Q3_K_M.gguf](https://huggingface.co/bartowski/MegaBeam-Mistral-7B-512k-GGUF/blob/main/MegaBeam-Mistral-7B-512k-Q3_K_M.gguf) | Q3_K_M | 3.52GB | false | Low quality. |
| [MegaBeam-Mistral-7B-512k-IQ3_M.gguf](https://huggingface.co/bartowski/MegaBeam-Mistral-7B-512k-GGUF/blob/main/MegaBeam-Mistral-7B-512k-IQ3_M.gguf) | IQ3_M | 3.28GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [MegaBeam-Mistral-7B-512k-Q3_K_S.gguf](https://huggingface.co/bartowski/MegaBeam-Mistral-7B-512k-GGUF/blob/main/MegaBeam-Mistral-7B-512k-Q3_K_S.gguf) | Q3_K_S | 3.16GB | false | Low quality, not recommended. |
| [MegaBeam-Mistral-7B-512k-IQ3_XS.gguf](https://huggingface.co/bartowski/MegaBeam-Mistral-7B-512k-GGUF/blob/main/MegaBeam-Mistral-7B-512k-IQ3_XS.gguf) | IQ3_XS | 3.02GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [MegaBeam-Mistral-7B-512k-Q2_K_L.gguf](https://huggingface.co/bartowski/MegaBeam-Mistral-7B-512k-GGUF/blob/main/MegaBeam-Mistral-7B-512k-Q2_K_L.gguf) | Q2_K_L | 2.85GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [MegaBeam-Mistral-7B-512k-Q2_K.gguf](https://huggingface.co/bartowski/MegaBeam-Mistral-7B-512k-GGUF/blob/main/MegaBeam-Mistral-7B-512k-Q2_K.gguf) | Q2_K | 2.72GB | false | Very low quality but surprisingly usable. |
| [MegaBeam-Mistral-7B-512k-IQ2_M.gguf](https://huggingface.co/bartowski/MegaBeam-Mistral-7B-512k-GGUF/blob/main/MegaBeam-Mistral-7B-512k-IQ2_M.gguf) | IQ2_M | 2.50GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset
Thank you ZeroWw for the inspiration to experiment with embed/output
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/MegaBeam-Mistral-7B-512k-GGUF --include "MegaBeam-Mistral-7B-512k-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/MegaBeam-Mistral-7B-512k-GGUF --include "MegaBeam-Mistral-7B-512k-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (MegaBeam-Mistral-7B-512k-Q8_0) or download them all in place (./)
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}