初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Qwen2.5-1.5B-Instruct-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-18 16:15:06 +08:00
commit 5c800a6b79
28 changed files with 264 additions and 0 deletions

60
.gitattributes vendored Normal file
View File

@@ -0,0 +1,60 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Qwen2.5-1.5B-Instruct-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-1.5B-Instruct-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-1.5B-Instruct-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-1.5B-Instruct-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-1.5B-Instruct-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-1.5B-Instruct-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-1.5B-Instruct-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-1.5B-Instruct-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-1.5B-Instruct-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-1.5B-Instruct-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-1.5B-Instruct-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-1.5B-Instruct-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-1.5B-Instruct-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-1.5B-Instruct-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-1.5B-Instruct-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-1.5B-Instruct-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-1.5B-Instruct-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-1.5B-Instruct-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-1.5B-Instruct-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-1.5B-Instruct-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-1.5B-Instruct-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-1.5B-Instruct-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-1.5B-Instruct-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-1.5B-Instruct-f16.gguf filter=lfs diff=lfs merge=lfs -text
Qwen2.5-1.5B-Instruct.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cee720c998e71ff3f02fbb3d392c7598bc0f845ae08bbb585ef2ff5fcbd45b81
size 601054816

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:aebd579aa34bde75426b7e3b786b089bc366f16da19e8aa60945d27f77e780f0
size 776664320

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9d5bd9c77d11161da7c8c4160913e1592f55c1f75604ede9b856ac7cfee896a6
size 731699296

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6a5911b95eacbfe2d5ac45b529c7de822530e4cb7a58ac162f16ae64e166dae4
size 895731968

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a8880f0de2348db67d00519ef7f4b40326ef67012bf5f2e90bd1d47474e2355c
size 676304992

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fc9639a4e919ba9c57ba81ee89e1afbd3a1daca9fba764345bc1f85f75725e28
size 732825184

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e2909fc76c785568edc57b5a9ff417850d6be87671a8b8a69167fdb2691fae36
size 880163072

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7437ad04011a14fb890074cc783df4c1d537197942337353a824ff7e6115ef9b
size 824178784

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e452220a86586d9116ea5eec444ba82dfe29862d0a5971698d27ae39926ec028
size 760944736

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:eb25b30dbedfb4e95aaabe4f1dcff406d658e2dca7296d833d3f8e359ba523d0
size 936683264

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a5ab77ebbd1d6d70250ff415339935bebac619b7acc508f71545f08231369dc1
size 937535744

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5e1eaf8413dd70a36bce7582d23cace56b6965882f6b127e975c9f0d9aff04bb
size 934955264

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5b3aa1bd68c34167dec0baf7d06188a752dcb2a3799d4c62dfd2b93d5b2aec31
size 934955264

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:33d30fb1b5a3eec99217ed4c1389783ab40b4805886941fe1e76a8d4e0cce2ae
size 934955264

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a76a9bf6606a61d20438449236aabb1174e32d59ba6b208c04bd3fa2d722e4b6
size 1042568960

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1adf0b11065d8ad2e8123ea110d1ec956dab4ab038eab665614adba04b6c3370
size 986048768

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cd382e90a417faa648ea012ebae8b97f94b3836be02947b0522886339ad1e1d7
size 940312832

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b2a97113d8ab7812f9e457d79098095c176beb6a96fe01d72e4ed297cf4983b6
size 1181570816

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cf240adc57e126e86102335f6565fb23e523b28d287c75bdb0759f064e8bb572
size 1125050624

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b682ac80bebdb354a3c74b2dcc13e2c852776f9920d35841298df97c2327ee51
size 1098729728

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1b01b4ea4ccdd5aa6a5972790002e120a5d500a5175be571114d642a8db4d14e
size 1272740096

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:aa556f887a8f6e17c2453c48c77a7adb4bce91d1d481dc81fc6972fec4055e1b
size 1329260288

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7185d306cf45956c8c017cd0d3b05ecc6bc18b3ea8eb5c240dce40e87563db7f
size 1646573312

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b6eaec3509f1d0373d1f4802654c4a7bfcde0768645d2d45e8885f6922b428ee
size 3093669376

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2f85dbb69836a017b9a4dfd0b86ee5b4dae0a0ac9217c57688de13f2dbb16a99
size 2042214

128
README.md Normal file
View File

@@ -0,0 +1,128 @@
---
base_model: Qwen/Qwen2.5-1.5B-Instruct
language:
- en
license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct/blob/main/LICENSE
pipeline_tag: text-generation
tags:
- chat
quantized_by: bartowski
---
## Llamacpp imatrix Quantizations of Qwen2.5-1.5B-Instruct
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b3772">b3772</a> for quantization.
Original model: https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
## What's new:
Update tokenizer
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Qwen2.5-1.5B-Instruct-f16.gguf](https://huggingface.co/bartowski/Qwen2.5-1.5B-Instruct-GGUF/blob/main/Qwen2.5-1.5B-Instruct-f16.gguf) | f16 | 3.09GB | false | Full F16 weights. |
| [Qwen2.5-1.5B-Instruct-Q8_0.gguf](https://huggingface.co/bartowski/Qwen2.5-1.5B-Instruct-GGUF/blob/main/Qwen2.5-1.5B-Instruct-Q8_0.gguf) | Q8_0 | 1.65GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Qwen2.5-1.5B-Instruct-Q6_K_L.gguf](https://huggingface.co/bartowski/Qwen2.5-1.5B-Instruct-GGUF/blob/main/Qwen2.5-1.5B-Instruct-Q6_K_L.gguf) | Q6_K_L | 1.33GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Qwen2.5-1.5B-Instruct-Q6_K.gguf](https://huggingface.co/bartowski/Qwen2.5-1.5B-Instruct-GGUF/blob/main/Qwen2.5-1.5B-Instruct-Q6_K.gguf) | Q6_K | 1.27GB | false | Very high quality, near perfect, *recommended*. |
| [Qwen2.5-1.5B-Instruct-Q5_K_L.gguf](https://huggingface.co/bartowski/Qwen2.5-1.5B-Instruct-GGUF/blob/main/Qwen2.5-1.5B-Instruct-Q5_K_L.gguf) | Q5_K_L | 1.18GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Qwen2.5-1.5B-Instruct-Q5_K_M.gguf](https://huggingface.co/bartowski/Qwen2.5-1.5B-Instruct-GGUF/blob/main/Qwen2.5-1.5B-Instruct-Q5_K_M.gguf) | Q5_K_M | 1.13GB | false | High quality, *recommended*. |
| [Qwen2.5-1.5B-Instruct-Q5_K_S.gguf](https://huggingface.co/bartowski/Qwen2.5-1.5B-Instruct-GGUF/blob/main/Qwen2.5-1.5B-Instruct-Q5_K_S.gguf) | Q5_K_S | 1.10GB | false | High quality, *recommended*. |
| [Qwen2.5-1.5B-Instruct-Q4_K_L.gguf](https://huggingface.co/bartowski/Qwen2.5-1.5B-Instruct-GGUF/blob/main/Qwen2.5-1.5B-Instruct-Q4_K_L.gguf) | Q4_K_L | 1.04GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Qwen2.5-1.5B-Instruct-Q4_K_M.gguf](https://huggingface.co/bartowski/Qwen2.5-1.5B-Instruct-GGUF/blob/main/Qwen2.5-1.5B-Instruct-Q4_K_M.gguf) | Q4_K_M | 0.99GB | false | Good quality, default size for must use cases, *recommended*. |
| [Qwen2.5-1.5B-Instruct-Q4_K_S.gguf](https://huggingface.co/bartowski/Qwen2.5-1.5B-Instruct-GGUF/blob/main/Qwen2.5-1.5B-Instruct-Q4_K_S.gguf) | Q4_K_S | 0.94GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Qwen2.5-1.5B-Instruct-Q4_0.gguf](https://huggingface.co/bartowski/Qwen2.5-1.5B-Instruct-GGUF/blob/main/Qwen2.5-1.5B-Instruct-Q4_0.gguf) | Q4_0 | 0.94GB | false | Legacy format, generally not worth using over similarly sized formats |
| [Qwen2.5-1.5B-Instruct-Q3_K_XL.gguf](https://huggingface.co/bartowski/Qwen2.5-1.5B-Instruct-GGUF/blob/main/Qwen2.5-1.5B-Instruct-Q3_K_XL.gguf) | Q3_K_XL | 0.94GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Qwen2.5-1.5B-Instruct-Q4_0_8_8.gguf](https://huggingface.co/bartowski/Qwen2.5-1.5B-Instruct-GGUF/blob/main/Qwen2.5-1.5B-Instruct-Q4_0_8_8.gguf) | Q4_0_8_8 | 0.93GB | false | Optimized for ARM inference. Requires 'sve' support (see link below). |
| [Qwen2.5-1.5B-Instruct-Q4_0_4_8.gguf](https://huggingface.co/bartowski/Qwen2.5-1.5B-Instruct-GGUF/blob/main/Qwen2.5-1.5B-Instruct-Q4_0_4_8.gguf) | Q4_0_4_8 | 0.93GB | false | Optimized for ARM inference. Requires 'i8mm' support (see link below). |
| [Qwen2.5-1.5B-Instruct-Q4_0_4_4.gguf](https://huggingface.co/bartowski/Qwen2.5-1.5B-Instruct-GGUF/blob/main/Qwen2.5-1.5B-Instruct-Q4_0_4_4.gguf) | Q4_0_4_4 | 0.93GB | false | Optimized for ARM inference. Should work well on all ARM chips, pick this if you're unsure. |
| [Qwen2.5-1.5B-Instruct-IQ4_XS.gguf](https://huggingface.co/bartowski/Qwen2.5-1.5B-Instruct-GGUF/blob/main/Qwen2.5-1.5B-Instruct-IQ4_XS.gguf) | IQ4_XS | 0.90GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Qwen2.5-1.5B-Instruct-Q3_K_L.gguf](https://huggingface.co/bartowski/Qwen2.5-1.5B-Instruct-GGUF/blob/main/Qwen2.5-1.5B-Instruct-Q3_K_L.gguf) | Q3_K_L | 0.88GB | false | Lower quality but usable, good for low RAM availability. |
| [Qwen2.5-1.5B-Instruct-IQ3_M.gguf](https://huggingface.co/bartowski/Qwen2.5-1.5B-Instruct-GGUF/blob/main/Qwen2.5-1.5B-Instruct-IQ3_M.gguf) | IQ3_M | 0.78GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Qwen2.5-1.5B-Instruct-GGUF --include "Qwen2.5-1.5B-Instruct-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Qwen2.5-1.5B-Instruct-GGUF --include "Qwen2.5-1.5B-Instruct-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Qwen2.5-1.5B-Instruct-Q8_0) or download them all in place (./)
## Q4_0_X_X
These are *NOT* for Metal (Apple) offloading, only ARM chips.
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset
Thank you ZeroWw for the inspiration to experiment with embed/output
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}