初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Qwen2-7B-Instruct-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-17 17:49:06 +08:00
commit 5e77401ff3
23 changed files with 217 additions and 0 deletions

55
.gitattributes vendored Normal file
View File

@@ -0,0 +1,55 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Qwen2-7B-Instruct-f32.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2-7B-Instruct.imatrix filter=lfs diff=lfs merge=lfs -text
Qwen2-7B-Instruct-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2-7B-Instruct-IQ2_S.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2-7B-Instruct-IQ2_XS.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2-7B-Instruct-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2-7B-Instruct-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2-7B-Instruct-IQ3_XXS.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2-7B-Instruct-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2-7B-Instruct-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2-7B-Instruct-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2-7B-Instruct-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2-7B-Instruct-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2-7B-Instruct-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2-7B-Instruct-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2-7B-Instruct-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2-7B-Instruct-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2-7B-Instruct-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2-7B-Instruct-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2-7B-Instruct-bf16.gguf filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c6b65122950bc41f4e4893b689fb0faff01b47a368157398c77b8697c07e55d6
size 2780340000

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3a2177e22202f0cca7ef0e948ed68705c5ec94ef4743ed8db707b8056a086879
size 2595634976

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:247883eca048b402e4829ae8d72359cb968581e71cbcefe49fb08b6d66673737
size 2469019424

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8dc59415c3941cf07c2a3fe5e5581dc45c8951f82bab7f5614d0bba2e1a34a82
size 3574009632

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:dcd392f72431881477b20b772bcbc7b97c651ee8af6a8e95bbfa5a0c8d8dba39
size 3346253600

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b1ab957f52e5dd77e991c02156f80c3f94e9f4ded2c93e38a8fb99e9407a3e6c
size 3114512160

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d51e881f308be12283e63eefae625840e1d44af4a11ccbd80f8f177d5f7b01f2
size 4218470176

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:71d6be6a0faa192a938b1bd86bbc8d527d0f56b6672164180bf1a0464a80af0c
size 3015937824

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0577a71834e73ae59015880d34dcd5ef549176c50b02dcef3b0acdc975f7479a
size 4088456992

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:27f06e0b1910771eb7bf04fa4a88832d04c009d7b8506ffaa045b9a503b31d16
size 3808388896

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:34aa03f291b5d7b71c6d71062eeed1d2fa4d313b5048cebcb1923656d9806b8c
size 3492366112

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8d0d33f0d9110a04aad1711b1ca02dafc0fa658cd83028bdfa5eff89c294fe76
size 4683071264

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:db87d60e881c760995cac29a8c0ed5d08c8b65cce8b052aaca9d2e3d6f833246
size 4457766688

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9f17976829b0e6d91852da2f3cd3f78d0f110778c858a29494a53d4acee8e69a
size 5444828960

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8696ac6e6c180f3f9528630b9b3cab0d779a4e4dfb12dc9f30c9030075708db4
size 5315174176

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6a5c1a2f31ba0ddb10a40720a6c592a0f7898c7c7b1c752679b0d18bad6b12df
size 6254196512

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3e3382bda013daf109f2d6fb7a990ca149fcaf4c221ec38597d0204c46d15d28
size 8098522912

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:bcfd06d4f9911b033ef0574937de75878e4d7719bd89966203162a31b1a39f6f
size 15237850656

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d8a23f3f11e7143b81b45b0466c728e3b3d322a18bcdf6ebfbe1116182fc140a
size 30468417056

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:23573f736284e14dd701cb2af68ca900ff640c6564ff0e7ce7e79d346db9dade
size 4536679

101
README.md Normal file
View File

@@ -0,0 +1,101 @@
---
license: apache-2.0
language:
- en
pipeline_tag: text-generation
tags:
- chat
quantized_by: bartowski
base_model: Qwen/Qwen2-7B-Instruct
---
# <b>Heads up:</b> currently CUDA offloading is broken unless you enable flash attention
## Llamacpp imatrix Quantizations of Qwen2-7B-Instruct
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> commit <a href="https://github.com/ggerganov/llama.cpp/commit/ee459f40f65810a810151b24eba5b8bd174ceffe">ee459f40f65810a810151b24eba5b8bd174ceffe</a> for quantization.
Original model: https://huggingface.co/Qwen/Qwen2-7B-Instruct
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
## Prompt format
```
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Description |
| -------- | ---------- | --------- | ----------- |
| [Qwen2-7B-Instruct-Q8_0.gguf](https://huggingface.co/bartowski/Qwen2-7B-Instruct-GGUF/blob/main/Qwen2-7B-Instruct-Q8_0.gguf) | Q8_0 | 8.09GB | Extremely high quality, generally unneeded but max available quant. |
| [Qwen2-7B-Instruct-Q6_K.gguf](https://huggingface.co/bartowski/Qwen2-7B-Instruct-GGUF/blob/main/Qwen2-7B-Instruct-Q6_K.gguf) | Q6_K | 6.25GB | Very high quality, near perfect, *recommended*. |
| [Qwen2-7B-Instruct-Q5_K_M.gguf](https://huggingface.co/bartowski/Qwen2-7B-Instruct-GGUF/blob/main/Qwen2-7B-Instruct-Q5_K_M.gguf) | Q5_K_M | 5.44GB | High quality, *recommended*. |
| [Qwen2-7B-Instruct-Q5_K_S.gguf](https://huggingface.co/bartowski/Qwen2-7B-Instruct-GGUF/blob/main/Qwen2-7B-Instruct-Q5_K_S.gguf) | Q5_K_S | 5.31GB | High quality, *recommended*. |
| [Qwen2-7B-Instruct-Q4_K_M.gguf](https://huggingface.co/bartowski/Qwen2-7B-Instruct-GGUF/blob/main/Qwen2-7B-Instruct-Q4_K_M.gguf) | Q4_K_M | 4.68GB | Good quality, uses about 4.83 bits per weight, *recommended*. |
| [Qwen2-7B-Instruct-Q4_K_S.gguf](https://huggingface.co/bartowski/Qwen2-7B-Instruct-GGUF/blob/main/Qwen2-7B-Instruct-Q4_K_S.gguf) | Q4_K_S | 4.45GB | Slightly lower quality with more space savings, *recommended*. |
| [Qwen2-7B-Instruct-IQ4_XS.gguf](https://huggingface.co/bartowski/Qwen2-7B-Instruct-GGUF/blob/main/Qwen2-7B-Instruct-IQ4_XS.gguf) | IQ4_XS | 4.21GB | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Qwen2-7B-Instruct-Q3_K_L.gguf](https://huggingface.co/bartowski/Qwen2-7B-Instruct-GGUF/blob/main/Qwen2-7B-Instruct-Q3_K_L.gguf) | Q3_K_L | 4.08GB | Lower quality but usable, good for low RAM availability. |
| [Qwen2-7B-Instruct-Q3_K_M.gguf](https://huggingface.co/bartowski/Qwen2-7B-Instruct-GGUF/blob/main/Qwen2-7B-Instruct-Q3_K_M.gguf) | Q3_K_M | 3.80GB | Even lower quality. |
| [Qwen2-7B-Instruct-IQ3_M.gguf](https://huggingface.co/bartowski/Qwen2-7B-Instruct-GGUF/blob/main/Qwen2-7B-Instruct-IQ3_M.gguf) | IQ3_M | 3.57GB | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Qwen2-7B-Instruct-Q3_K_S.gguf](https://huggingface.co/bartowski/Qwen2-7B-Instruct-GGUF/blob/main/Qwen2-7B-Instruct-Q3_K_S.gguf) | Q3_K_S | 3.49GB | Low quality, not recommended. |
| [Qwen2-7B-Instruct-IQ3_XS.gguf](https://huggingface.co/bartowski/Qwen2-7B-Instruct-GGUF/blob/main/Qwen2-7B-Instruct-IQ3_XS.gguf) | IQ3_XS | 3.34GB | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Qwen2-7B-Instruct-IQ3_XXS.gguf](https://huggingface.co/bartowski/Qwen2-7B-Instruct-GGUF/blob/main/Qwen2-7B-Instruct-IQ3_XXS.gguf) | IQ3_XXS | 3.11GB | Lower quality, new method with decent performance, comparable to Q3 quants. |
| [Qwen2-7B-Instruct-Q2_K.gguf](https://huggingface.co/bartowski/Qwen2-7B-Instruct-GGUF/blob/main/Qwen2-7B-Instruct-Q2_K.gguf) | Q2_K | 3.01GB | Very low quality but surprisingly usable. |
| [Qwen2-7B-Instruct-IQ2_M.gguf](https://huggingface.co/bartowski/Qwen2-7B-Instruct-GGUF/blob/main/Qwen2-7B-Instruct-IQ2_M.gguf) | IQ2_M | 2.78GB | Very low quality, uses SOTA techniques to also be surprisingly usable. |
| [Qwen2-7B-Instruct-IQ2_S.gguf](https://huggingface.co/bartowski/Qwen2-7B-Instruct-GGUF/blob/main/Qwen2-7B-Instruct-IQ2_S.gguf) | IQ2_S | 2.59GB | Very low quality, uses SOTA techniques to be usable. |
| [Qwen2-7B-Instruct-IQ2_XS.gguf](https://huggingface.co/bartowski/Qwen2-7B-Instruct-GGUF/blob/main/Qwen2-7B-Instruct-IQ2_XS.gguf) | IQ2_XS | 2.46GB | Very low quality, uses SOTA techniques to be usable. |
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Qwen2-7B-Instruct-GGUF --include "Qwen2-7B-Instruct-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Qwen2-7B-Instruct-GGUF --include "Qwen2-7B-Instruct-Q8_0.gguf/*" --local-dir Qwen2-7B-Instruct-Q8_0
```
You can either specify a new local-dir (Qwen2-7B-Instruct-Q8_0) or download them all in place (./)
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}