初始化项目,由ModelHub XC社区提供模型

Model: bartowski/GRMR-2B-Instruct-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-06 07:41:14 +08:00
commit 60201559e5
30 changed files with 317 additions and 0 deletions

62
.gitattributes vendored Normal file
View File

@@ -0,0 +1,62 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
GRMR-2B-Instruct-f16.gguf filter=lfs diff=lfs merge=lfs -text
GRMR-2B-Instruct-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
GRMR-2B-Instruct-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
GRMR-2B-Instruct-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
GRMR-2B-Instruct-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
GRMR-2B-Instruct-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
GRMR-2B-Instruct-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
GRMR-2B-Instruct-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
GRMR-2B-Instruct-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
GRMR-2B-Instruct-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
GRMR-2B-Instruct-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
GRMR-2B-Instruct-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
GRMR-2B-Instruct-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
GRMR-2B-Instruct-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
GRMR-2B-Instruct-IQ4_NL.gguf filter=lfs diff=lfs merge=lfs -text
GRMR-2B-Instruct-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
GRMR-2B-Instruct-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
GRMR-2B-Instruct-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
GRMR-2B-Instruct-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
GRMR-2B-Instruct-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
GRMR-2B-Instruct-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
GRMR-2B-Instruct-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
GRMR-2B-Instruct-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
GRMR-2B-Instruct-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
GRMR-2B-Instruct-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
GRMR-2B-Instruct-f32.gguf filter=lfs diff=lfs merge=lfs -text
GRMR-2B-Instruct.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:90159095d79dd03c42d9cf1d5ff67792acc6a4d2532a17b4d81f691d639acfc7
size 1088013984

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ff7941195ac5e196a47754bed373f9c72f767732309b6bb7a581005738de8a41
size 1393561248

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e4f7775a8e6bc5b9f62b560f93a9c20a531d8d5d9bf09672082f9b165fce2940
size 1314211488

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:bbad3625e3694e429d47f46caf96941a82df7d5feb29df79cd38ae47cc4cdc58
size 1629509280

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:976c3c15c30fcaae4b1bc554b671921da506ec8704e5cea44e5015eb24af5d08
size 1566250656

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:222e63a4a9b73f3b34baa1d0384445a1d2b4284125b92c09a163a6ccf4f8b632
size 1229829792

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a682496c1eddcd1e91410d190e42fc7919d7ddbce79da0ef29dc7218541ec2e5
size 1372677792

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:17516b38f2c42396f5504fdb7708913338d25f5ba5ba378a7f186c3635b89dfa
size 1550436000

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2aac86918ae6d2039597cef211e01d30366fd7d705dba842cc7c4c6345190852
size 1461667488

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8c255f9460848f808e386372cf6ad8091581a7849ffac042deff3092331eab0e
size 1360660128

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:68ece83daa5d7dd28a900e3a473beaa2e97a880ff01d76e1a9af00f52a4ae7b0
size 1693284000

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b584149fa609b742b24cdc3c9fa18802d85892d97e4968be5facbf2f7d7104a2
size 1633490592

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:bd1f65a9e31fe062e00e165432a1d67eced29bcc221ffe0d8488239f39a6631b
size 1629509280

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c5a63f574df88ca0096a79ef7543067ed0df03e8c86f0ea2cea03e71520a41ae
size 1629509280

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:daeb9a66ccca5871596a6c66dcfcf2227a1dc753d54d611b2d3e7c2f2ad6b728
size 1629509280

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:30b4d5d61a457800ffc7a6ea6c4c21eeb880aead1486698c957f3e9602793a3c
size 1851430560

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6695dd2808852b6de40ddec7ee97732c1aac14e832935d16d33e8b72c77e0397
size 1708582560

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4ba7b1c9c8755d5a78f1136b5565711e90f3f86a012b276519a5e7f60f39429a
size 1638651552

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cb499aca3ac6825995c142fbc2587697a77cbd2216bed4c9a5c1a7856dc396dd
size 2066126496

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:16d452464b20cdfb43e2e37181d8ef92a3570cdd9c0847f454ab3010241b9345
size 1923278496

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:82c3b385cf384da5aea7fe075896df4505eabdbf37b683aa2eb601703c8cda16
size 1882543776

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c6cd4c795236f7d345b0c0b5a7d12d13a7c62d5731aada1d29619ea2b8637fe8
size 2151392928

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:66bc308652bd035a0be2b4259a7384b5db7ae3ef56e5b9fdb26eab395a0a5e71
size 2294240928

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5c92e82474ab2f06e77697634212a2c222edb50ed846f495528e3a627230bec1
size 2784495264

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6f964881c56f412ccc6cb84cb03c790ab7a7997fbd484231b9fb37cff9c09e45
size 5235213984

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fdfc71fb67d1f57174d3e9778276c59e16567f84c38d1cefe6cf737aca8e3b30
size 10463413632

3
GRMR-2B-Instruct.imatrix Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:762e19b2826742dcb8e2bd802279463b7781d1e9230638090ce2d81d3ea0c568
size 2375572

173
README.md Normal file
View File

@@ -0,0 +1,173 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
language:
- en
base_model: qingy2024/GRMR-2B-Instruct
tags:
- text-generation-inference
- transformers
- unsloth
- gemma2
- trl
license: apache-2.0
---
## Llamacpp imatrix Quantizations of GRMR-2B-Instruct
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4273">b4273</a> for quantization.
Original model: https://huggingface.co/qingy2024/GRMR-2B-Instruct
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
No prompt format found, check original model page
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [GRMR-2B-Instruct-f32.gguf](https://huggingface.co/bartowski/GRMR-2B-Instruct-GGUF/blob/main/GRMR-2B-Instruct-f32.gguf) | f32 | 10.46GB | false | Full F32 weights. |
| [GRMR-2B-Instruct-f16.gguf](https://huggingface.co/bartowski/GRMR-2B-Instruct-GGUF/blob/main/GRMR-2B-Instruct-f16.gguf) | f16 | 5.24GB | false | Full F16 weights. |
| [GRMR-2B-Instruct-Q8_0.gguf](https://huggingface.co/bartowski/GRMR-2B-Instruct-GGUF/blob/main/GRMR-2B-Instruct-Q8_0.gguf) | Q8_0 | 2.78GB | false | Extremely high quality, generally unneeded but max available quant. |
| [GRMR-2B-Instruct-Q6_K_L.gguf](https://huggingface.co/bartowski/GRMR-2B-Instruct-GGUF/blob/main/GRMR-2B-Instruct-Q6_K_L.gguf) | Q6_K_L | 2.29GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [GRMR-2B-Instruct-Q6_K.gguf](https://huggingface.co/bartowski/GRMR-2B-Instruct-GGUF/blob/main/GRMR-2B-Instruct-Q6_K.gguf) | Q6_K | 2.15GB | false | Very high quality, near perfect, *recommended*. |
| [GRMR-2B-Instruct-Q5_K_L.gguf](https://huggingface.co/bartowski/GRMR-2B-Instruct-GGUF/blob/main/GRMR-2B-Instruct-Q5_K_L.gguf) | Q5_K_L | 2.07GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [GRMR-2B-Instruct-Q5_K_M.gguf](https://huggingface.co/bartowski/GRMR-2B-Instruct-GGUF/blob/main/GRMR-2B-Instruct-Q5_K_M.gguf) | Q5_K_M | 1.92GB | false | High quality, *recommended*. |
| [GRMR-2B-Instruct-Q5_K_S.gguf](https://huggingface.co/bartowski/GRMR-2B-Instruct-GGUF/blob/main/GRMR-2B-Instruct-Q5_K_S.gguf) | Q5_K_S | 1.88GB | false | High quality, *recommended*. |
| [GRMR-2B-Instruct-Q4_K_L.gguf](https://huggingface.co/bartowski/GRMR-2B-Instruct-GGUF/blob/main/GRMR-2B-Instruct-Q4_K_L.gguf) | Q4_K_L | 1.85GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [GRMR-2B-Instruct-Q4_K_M.gguf](https://huggingface.co/bartowski/GRMR-2B-Instruct-GGUF/blob/main/GRMR-2B-Instruct-Q4_K_M.gguf) | Q4_K_M | 1.71GB | false | Good quality, default size for most use cases, *recommended*. |
| [GRMR-2B-Instruct-Q3_K_XL.gguf](https://huggingface.co/bartowski/GRMR-2B-Instruct-GGUF/blob/main/GRMR-2B-Instruct-Q3_K_XL.gguf) | Q3_K_XL | 1.69GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [GRMR-2B-Instruct-Q4_K_S.gguf](https://huggingface.co/bartowski/GRMR-2B-Instruct-GGUF/blob/main/GRMR-2B-Instruct-Q4_K_S.gguf) | Q4_K_S | 1.64GB | false | Slightly lower quality with more space savings, *recommended*. |
| [GRMR-2B-Instruct-Q4_0_8_8.gguf](https://huggingface.co/bartowski/GRMR-2B-Instruct-GGUF/blob/main/GRMR-2B-Instruct-Q4_0_8_8.gguf) | Q4_0_8_8 | 1.63GB | false | Optimized for ARM and AVX inference. Requires 'sve' support for ARM (see details below). *Don't use on Mac*. |
| [GRMR-2B-Instruct-Q4_0_4_8.gguf](https://huggingface.co/bartowski/GRMR-2B-Instruct-GGUF/blob/main/GRMR-2B-Instruct-Q4_0_4_8.gguf) | Q4_0_4_8 | 1.63GB | false | Optimized for ARM inference. Requires 'i8mm' support (see details below). *Don't use on Mac*. |
| [GRMR-2B-Instruct-Q4_0_4_4.gguf](https://huggingface.co/bartowski/GRMR-2B-Instruct-GGUF/blob/main/GRMR-2B-Instruct-Q4_0_4_4.gguf) | Q4_0_4_4 | 1.63GB | false | Optimized for ARM inference. Should work well on all ARM chips, not for use with GPUs. *Don't use on Mac*. |
| [GRMR-2B-Instruct-Q4_0.gguf](https://huggingface.co/bartowski/GRMR-2B-Instruct-GGUF/blob/main/GRMR-2B-Instruct-Q4_0.gguf) | Q4_0 | 1.63GB | false | Legacy format, offers online repacking for ARM CPU inference. |
| [GRMR-2B-Instruct-IQ4_NL.gguf](https://huggingface.co/bartowski/GRMR-2B-Instruct-GGUF/blob/main/GRMR-2B-Instruct-IQ4_NL.gguf) | IQ4_NL | 1.63GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [GRMR-2B-Instruct-IQ4_XS.gguf](https://huggingface.co/bartowski/GRMR-2B-Instruct-GGUF/blob/main/GRMR-2B-Instruct-IQ4_XS.gguf) | IQ4_XS | 1.57GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [GRMR-2B-Instruct-Q3_K_L.gguf](https://huggingface.co/bartowski/GRMR-2B-Instruct-GGUF/blob/main/GRMR-2B-Instruct-Q3_K_L.gguf) | Q3_K_L | 1.55GB | false | Lower quality but usable, good for low RAM availability. |
| [GRMR-2B-Instruct-Q3_K_M.gguf](https://huggingface.co/bartowski/GRMR-2B-Instruct-GGUF/blob/main/GRMR-2B-Instruct-Q3_K_M.gguf) | Q3_K_M | 1.46GB | false | Low quality. |
| [GRMR-2B-Instruct-IQ3_M.gguf](https://huggingface.co/bartowski/GRMR-2B-Instruct-GGUF/blob/main/GRMR-2B-Instruct-IQ3_M.gguf) | IQ3_M | 1.39GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [GRMR-2B-Instruct-Q2_K_L.gguf](https://huggingface.co/bartowski/GRMR-2B-Instruct-GGUF/blob/main/GRMR-2B-Instruct-Q2_K_L.gguf) | Q2_K_L | 1.37GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [GRMR-2B-Instruct-Q3_K_S.gguf](https://huggingface.co/bartowski/GRMR-2B-Instruct-GGUF/blob/main/GRMR-2B-Instruct-Q3_K_S.gguf) | Q3_K_S | 1.36GB | false | Low quality, not recommended. |
| [GRMR-2B-Instruct-IQ3_XS.gguf](https://huggingface.co/bartowski/GRMR-2B-Instruct-GGUF/blob/main/GRMR-2B-Instruct-IQ3_XS.gguf) | IQ3_XS | 1.31GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [GRMR-2B-Instruct-Q2_K.gguf](https://huggingface.co/bartowski/GRMR-2B-Instruct-GGUF/blob/main/GRMR-2B-Instruct-Q2_K.gguf) | Q2_K | 1.23GB | false | Very low quality but surprisingly usable. |
| [GRMR-2B-Instruct-IQ2_M.gguf](https://huggingface.co/bartowski/GRMR-2B-Instruct-GGUF/blob/main/GRMR-2B-Instruct-IQ2_M.gguf) | IQ2_M | 1.09GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/GRMR-2B-Instruct-GGUF --include "GRMR-2B-Instruct-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/GRMR-2B-Instruct-GGUF --include "GRMR-2B-Instruct-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (GRMR-2B-Instruct-Q8_0) or download them all in place (./)
</details>
## Q4_0_X_X information
New: Thanks to efforts made to have online repacking of weights in [this PR](https://github.com/ggerganov/llama.cpp/pull/9921), you can now just use Q4_0 if your llama.cpp has been compiled for your ARM device.
Similarly, if you want to get slightly better performance, you can use IQ4_NL thanks to [this PR](https://github.com/ggerganov/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information</summary>
These are *NOT* for Metal (Apple) or GPU (nvidia/AMD/intel) offloading, only ARM chips (and certain AVX2/AVX512 CPUs).
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
If you're using a CPU that supports AVX2 or AVX512 (typically server CPUs and AMD's latest Zen5 CPUs) and are not offloading to a GPU, the Q4_0_8_8 may offer a nice speed as well:
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}