初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Infinity-Instruct-7M-Gen-Llama3_1-8B-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-27 17:29:09 +08:00
commit 4b6a92dffb
24 changed files with 238 additions and 0 deletions

56
.gitattributes vendored Normal file
View File

@@ -0,0 +1,56 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Infinity-Instruct-7M-Gen-Llama3_1-8B-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Infinity-Instruct-7M-Gen-Llama3_1-8B-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Infinity-Instruct-7M-Gen-Llama3_1-8B-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Infinity-Instruct-7M-Gen-Llama3_1-8B-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Infinity-Instruct-7M-Gen-Llama3_1-8B-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Infinity-Instruct-7M-Gen-Llama3_1-8B-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Infinity-Instruct-7M-Gen-Llama3_1-8B-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Infinity-Instruct-7M-Gen-Llama3_1-8B-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Infinity-Instruct-7M-Gen-Llama3_1-8B-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Infinity-Instruct-7M-Gen-Llama3_1-8B-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Infinity-Instruct-7M-Gen-Llama3_1-8B-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Infinity-Instruct-7M-Gen-Llama3_1-8B-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Infinity-Instruct-7M-Gen-Llama3_1-8B-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Infinity-Instruct-7M-Gen-Llama3_1-8B-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Infinity-Instruct-7M-Gen-Llama3_1-8B-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Infinity-Instruct-7M-Gen-Llama3_1-8B-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Infinity-Instruct-7M-Gen-Llama3_1-8B-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Infinity-Instruct-7M-Gen-Llama3_1-8B-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Infinity-Instruct-7M-Gen-Llama3_1-8B-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Infinity-Instruct-7M-Gen-Llama3_1-8B-f16.gguf filter=lfs diff=lfs merge=lfs -text
Infinity-Instruct-7M-Gen-Llama3_1-8B.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:20a28560e3b7fcdc4211baad0b264544f11fbcf313db84cd3f369e3520560559
size 2948281280

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:10337dff3003d009c2712beca83f1cf914d9023b9b5a867cd2f67c4eff46c960
size 3784823744

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5fa0ea2b5fa568af71c8ae8b8c6222b66b6b72a775223d81ca13086a3f4f7182
size 3518747584

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5fa33fa7807b6fd5bab22dff3df1c4e562d9749a6955c32dc5fa6a45f2a2aed1
size 4447663040

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:687590d59c1591ce6a51f26de9265719c056a4555618e19ac029029c5b89ff2b
size 3179131840

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:44ad5a5590db795572c222c6ca5c29df31716840aef9022d6aab65325fe8bc72
size 3692155840

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:733653c0411e70bfb72dfe65a0cc4fec1afd1f227c047069d8d5244d34510447
size 4321956800

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b53a6e2380833c35d8a54f9437cf16cd7eaa2536205a635cfa092f214f523a80
size 4018918336

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b496de4815fb0706a12118081415118b7d047143019ebea767d22540f8dcc493
size 3664499648

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:caf4e5f8cc736e20d3a5b867909510121997914dde733608d57595ea41e042a0
size 4781626304

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:45783370670e3c0514fdcf6e2c23a6722c17fe0fc960cf7072db50a29cbafe91
size 5310632896

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:706cc39c9942fd5640cfa028e18bdbb5534c4d95e251a55d044e432d9d6b9df0
size 4920734656

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:83e3200561161532dd6c637155a41b0e07eed5f0063e9475fff12e0dc52f1bac
size 4692669376

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2151b45de9424ffc7589b92b9522fbcb95381fd9d84ddb05340b5640b8846f6b
size 6057219008

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3a3c0066ce0d69f54be69a94894cd032460b836091e4497465322a3f859e7440
size 5732987840

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e33d811e5106016471e1702014cfffaa9706982ecb685aaa941f6ff8d02a1519
size 5599294400

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7295c7f2ebbe0669bd66f1eb8a697c9d9580a98e7b7b7882e19e1e9517b35f21
size 6596006848

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:07aee1ac3a37043b73e9ab18c6672626bdf9bc1d510bea74e7f75633f49e34b3
size 6850466752

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3dd34b5592e335c59fa48be45a60476288587ac21b1b42dce57bfef182ce48e9
size 8540771264

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:eabb9faa16dc4eed608dc4e7fe60600d163d7ce9c4deb135bff54027b7265e39
size 16068891264

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:83a5359d5a3eab4bb878782a65bc4d6adb168b1ee057f6994789227274ccae5c
size 4988170

118
README.md Normal file
View File

@@ -0,0 +1,118 @@
---
base_model: BAAI/Infinity-Instruct-7M-Gen-Llama3_1-8B
datasets:
- BAAI/Infinity-Instruct
language:
- en
license: llama3.1
pipeline_tag: text-generation
quantized_by: bartowski
---
## Llamacpp imatrix Quantizations of Infinity-Instruct-7M-Gen-Llama3_1-8B
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b3615">b3615</a> for quantization.
Original model: https://huggingface.co/BAAI/Infinity-Instruct-7M-Gen-Llama3_1-8B
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<|begin_of_text|><|start_header_id|>system<|end_header_id|>
{system_prompt}<|eot_id|><|start_header_id|>user<|end_header_id|>
{prompt}<|eot_id|><|start_header_id|>assistant<|end_header_id|>
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Infinity-Instruct-7M-Gen-Llama3_1-8B-f16.gguf](https://huggingface.co/bartowski/Infinity-Instruct-7M-Gen-Llama3_1-8B-GGUF/blob/main/Infinity-Instruct-7M-Gen-Llama3_1-8B-f16.gguf) | f16 | 16.07GB | false | Full F16 weights. |
| [Infinity-Instruct-7M-Gen-Llama3_1-8B-Q8_0.gguf](https://huggingface.co/bartowski/Infinity-Instruct-7M-Gen-Llama3_1-8B-GGUF/blob/main/Infinity-Instruct-7M-Gen-Llama3_1-8B-Q8_0.gguf) | Q8_0 | 8.54GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Infinity-Instruct-7M-Gen-Llama3_1-8B-Q6_K_L.gguf](https://huggingface.co/bartowski/Infinity-Instruct-7M-Gen-Llama3_1-8B-GGUF/blob/main/Infinity-Instruct-7M-Gen-Llama3_1-8B-Q6_K_L.gguf) | Q6_K_L | 6.85GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Infinity-Instruct-7M-Gen-Llama3_1-8B-Q6_K.gguf](https://huggingface.co/bartowski/Infinity-Instruct-7M-Gen-Llama3_1-8B-GGUF/blob/main/Infinity-Instruct-7M-Gen-Llama3_1-8B-Q6_K.gguf) | Q6_K | 6.60GB | false | Very high quality, near perfect, *recommended*. |
| [Infinity-Instruct-7M-Gen-Llama3_1-8B-Q5_K_L.gguf](https://huggingface.co/bartowski/Infinity-Instruct-7M-Gen-Llama3_1-8B-GGUF/blob/main/Infinity-Instruct-7M-Gen-Llama3_1-8B-Q5_K_L.gguf) | Q5_K_L | 6.06GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Infinity-Instruct-7M-Gen-Llama3_1-8B-Q5_K_M.gguf](https://huggingface.co/bartowski/Infinity-Instruct-7M-Gen-Llama3_1-8B-GGUF/blob/main/Infinity-Instruct-7M-Gen-Llama3_1-8B-Q5_K_M.gguf) | Q5_K_M | 5.73GB | false | High quality, *recommended*. |
| [Infinity-Instruct-7M-Gen-Llama3_1-8B-Q5_K_S.gguf](https://huggingface.co/bartowski/Infinity-Instruct-7M-Gen-Llama3_1-8B-GGUF/blob/main/Infinity-Instruct-7M-Gen-Llama3_1-8B-Q5_K_S.gguf) | Q5_K_S | 5.60GB | false | High quality, *recommended*. |
| [Infinity-Instruct-7M-Gen-Llama3_1-8B-Q4_K_L.gguf](https://huggingface.co/bartowski/Infinity-Instruct-7M-Gen-Llama3_1-8B-GGUF/blob/main/Infinity-Instruct-7M-Gen-Llama3_1-8B-Q4_K_L.gguf) | Q4_K_L | 5.31GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Infinity-Instruct-7M-Gen-Llama3_1-8B-Q4_K_M.gguf](https://huggingface.co/bartowski/Infinity-Instruct-7M-Gen-Llama3_1-8B-GGUF/blob/main/Infinity-Instruct-7M-Gen-Llama3_1-8B-Q4_K_M.gguf) | Q4_K_M | 4.92GB | false | Good quality, default size for must use cases, *recommended*. |
| [Infinity-Instruct-7M-Gen-Llama3_1-8B-Q3_K_XL.gguf](https://huggingface.co/bartowski/Infinity-Instruct-7M-Gen-Llama3_1-8B-GGUF/blob/main/Infinity-Instruct-7M-Gen-Llama3_1-8B-Q3_K_XL.gguf) | Q3_K_XL | 4.78GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Infinity-Instruct-7M-Gen-Llama3_1-8B-Q4_K_S.gguf](https://huggingface.co/bartowski/Infinity-Instruct-7M-Gen-Llama3_1-8B-GGUF/blob/main/Infinity-Instruct-7M-Gen-Llama3_1-8B-Q4_K_S.gguf) | Q4_K_S | 4.69GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Infinity-Instruct-7M-Gen-Llama3_1-8B-IQ4_XS.gguf](https://huggingface.co/bartowski/Infinity-Instruct-7M-Gen-Llama3_1-8B-GGUF/blob/main/Infinity-Instruct-7M-Gen-Llama3_1-8B-IQ4_XS.gguf) | IQ4_XS | 4.45GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Infinity-Instruct-7M-Gen-Llama3_1-8B-Q3_K_L.gguf](https://huggingface.co/bartowski/Infinity-Instruct-7M-Gen-Llama3_1-8B-GGUF/blob/main/Infinity-Instruct-7M-Gen-Llama3_1-8B-Q3_K_L.gguf) | Q3_K_L | 4.32GB | false | Lower quality but usable, good for low RAM availability. |
| [Infinity-Instruct-7M-Gen-Llama3_1-8B-Q3_K_M.gguf](https://huggingface.co/bartowski/Infinity-Instruct-7M-Gen-Llama3_1-8B-GGUF/blob/main/Infinity-Instruct-7M-Gen-Llama3_1-8B-Q3_K_M.gguf) | Q3_K_M | 4.02GB | false | Low quality. |
| [Infinity-Instruct-7M-Gen-Llama3_1-8B-IQ3_M.gguf](https://huggingface.co/bartowski/Infinity-Instruct-7M-Gen-Llama3_1-8B-GGUF/blob/main/Infinity-Instruct-7M-Gen-Llama3_1-8B-IQ3_M.gguf) | IQ3_M | 3.78GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Infinity-Instruct-7M-Gen-Llama3_1-8B-Q2_K_L.gguf](https://huggingface.co/bartowski/Infinity-Instruct-7M-Gen-Llama3_1-8B-GGUF/blob/main/Infinity-Instruct-7M-Gen-Llama3_1-8B-Q2_K_L.gguf) | Q2_K_L | 3.69GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Infinity-Instruct-7M-Gen-Llama3_1-8B-Q3_K_S.gguf](https://huggingface.co/bartowski/Infinity-Instruct-7M-Gen-Llama3_1-8B-GGUF/blob/main/Infinity-Instruct-7M-Gen-Llama3_1-8B-Q3_K_S.gguf) | Q3_K_S | 3.66GB | false | Low quality, not recommended. |
| [Infinity-Instruct-7M-Gen-Llama3_1-8B-IQ3_XS.gguf](https://huggingface.co/bartowski/Infinity-Instruct-7M-Gen-Llama3_1-8B-GGUF/blob/main/Infinity-Instruct-7M-Gen-Llama3_1-8B-IQ3_XS.gguf) | IQ3_XS | 3.52GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Infinity-Instruct-7M-Gen-Llama3_1-8B-Q2_K.gguf](https://huggingface.co/bartowski/Infinity-Instruct-7M-Gen-Llama3_1-8B-GGUF/blob/main/Infinity-Instruct-7M-Gen-Llama3_1-8B-Q2_K.gguf) | Q2_K | 3.18GB | false | Very low quality but surprisingly usable. |
| [Infinity-Instruct-7M-Gen-Llama3_1-8B-IQ2_M.gguf](https://huggingface.co/bartowski/Infinity-Instruct-7M-Gen-Llama3_1-8B-GGUF/blob/main/Infinity-Instruct-7M-Gen-Llama3_1-8B-IQ2_M.gguf) | IQ2_M | 2.95GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset
Thank you ZeroWw for the inspiration to experiment with embed/output
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Infinity-Instruct-7M-Gen-Llama3_1-8B-GGUF --include "Infinity-Instruct-7M-Gen-Llama3_1-8B-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Infinity-Instruct-7M-Gen-Llama3_1-8B-GGUF --include "Infinity-Instruct-7M-Gen-Llama3_1-8B-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Infinity-Instruct-7M-Gen-Llama3_1-8B-Q8_0) or download them all in place (./)
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}