初始化项目,由ModelHub XC社区提供模型

Model: bartowski/MiniThinky-v2-1B-Llama-3.2-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-25 08:07:05 +08:00
commit d7255fdf13
28 changed files with 308 additions and 0 deletions

60
.gitattributes vendored Normal file
View File

@@ -0,0 +1,60 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
MiniThinky-v2-1B-Llama-3.2-f16.gguf filter=lfs diff=lfs merge=lfs -text
MiniThinky-v2-1B-Llama-3.2-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
MiniThinky-v2-1B-Llama-3.2-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
MiniThinky-v2-1B-Llama-3.2-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
MiniThinky-v2-1B-Llama-3.2-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
MiniThinky-v2-1B-Llama-3.2-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
MiniThinky-v2-1B-Llama-3.2-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
MiniThinky-v2-1B-Llama-3.2-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
MiniThinky-v2-1B-Llama-3.2-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
MiniThinky-v2-1B-Llama-3.2-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
MiniThinky-v2-1B-Llama-3.2-Q4_1.gguf filter=lfs diff=lfs merge=lfs -text
MiniThinky-v2-1B-Llama-3.2-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
MiniThinky-v2-1B-Llama-3.2-IQ4_NL.gguf filter=lfs diff=lfs merge=lfs -text
MiniThinky-v2-1B-Llama-3.2-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
MiniThinky-v2-1B-Llama-3.2-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
MiniThinky-v2-1B-Llama-3.2-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
MiniThinky-v2-1B-Llama-3.2-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
MiniThinky-v2-1B-Llama-3.2-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
MiniThinky-v2-1B-Llama-3.2-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
MiniThinky-v2-1B-Llama-3.2-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
MiniThinky-v2-1B-Llama-3.2-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
MiniThinky-v2-1B-Llama-3.2-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
MiniThinky-v2-1B-Llama-3.2-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
MiniThinky-v2-1B-Llama-3.2-f32.gguf filter=lfs diff=lfs merge=lfs -text
MiniThinky-v2-1B-Llama-3.2.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9e1ab2723a2e9497ea967eff8f50170befff34f080b788f90f17d57196da1b29
size 515446048

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:877cc0f2078c8d58d97c0d6d9d05577b7d725e763e08775c5e4d8f34d151bae2
size 657286432

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0f4449df9128777129b426e78e316f892c0d5e5eeaaf7619ebf78d4e0227369a
size 621110560

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3b220eab48d6b27a767bcbad785d59c032eb6bdd3dfd4f580f171252ad91569a
size 773023008

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a89466a698002834f16c6c7c2793ebf3a4bf71ffa13642faf02f06bba740e1f8
size 743138592

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:af3ac9c94c5e0a528ec6a5a9c3be7fd331d216400447b86a22a5bf5d344260ff
size 580871456

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3e44bd98f8ebc2caba7f969368b738a0375b2fbaad674b3e6b1a55b38af8d00a
size 644486432

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6d58ae621d91ee950e804131469f19cac99d8721262e565350c6853ad7acf422
size 732521760

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:293348a4fc9456ff29852db9dbef576943455d732201cdd234bc3daaeefe1389
size 690840864

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:11a94ae92d657691154d4dd383865ca322903dc91fca768a8f543476cfd49a83
size 641688864

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e771249dda0e4e477f27def70e2feeec589c0eb8678f4d37b4359d2a2456440c
size 796136736

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:75e24fec3ffcf2566fb7d1ba8099a6e0bd38afdd5a9aad668536f260203ecf1c
size 773023008

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:25331e9d2ad602f7752f6e9bf324a0dd6bd2449a6fc6d9b80768d8b65e6662f4
size 831743264

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2e151326aae8895abfd86875f6ff768f5bc1590f741341a1e8ef9144c12b14cf
size 871306528

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:086857b6364afd757a123eea0474bede09f25608783e7a6fcf2f88d8cb322ca1
size 807691552

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d603e73f6a957ee3dc74e7822780c5806c06b949af97aba949c1303e880153f8
size 775644448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6fcfaf9abf8f3640c9566bc1ef74d4cbca84e26ec2a6bafe4fb7f104d5bab6e2
size 975115552

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d624dcb7a0b56262276037d9ec3724f9feec692f4ccb9088483ee4f6001d16b0
size 911500576

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:35999e7860e1890620339b7c499449ffcbf1931e6dcec3fd78570e601c69f9c4
size 892560672

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a21887dd6ce747d5f2e933fc3b1f40789119ae38e1f753650218db0ba6957e11
size 1021797664

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:011deeaf8eac22ca92f948f744419a2ce52d826f75ea82f74317aa36d122fb1a
size 1085412640

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4e0e924149dfdaf57dedab472a7c77abd262b4bc58b684475a002b162e6ef020
size 1321080096

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cd3da08e8c177ce3a245dcdfc15d7d669eba70b96d44053ea87c922393b2831f
size 2479592736

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:926781822c8b22179e86c736929e91bd4af56b2bbaaee10036eb5cc36b86c7fb
size 4951086080

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:dadef0afe16569826322fba112fd78fe380a929351282e82850b9b0aa70141f7
size 1314426

172
README.md Normal file
View File

@@ -0,0 +1,172 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
tags:
- trl
- sft
base_model: ngxson/MiniThinky-v2-1B-Llama-3.2
datasets:
- ngxson/MiniThinky-dataset
---
## Llamacpp imatrix Quantizations of MiniThinky-v2-1B-Llama-3.2
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4418">b4418</a> for quantization.
Original model: https://huggingface.co/ngxson/MiniThinky-v2-1B-Llama-3.2
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<|begin_of_text|><|start_header_id|>system<|end_header_id|>
{system_prompt}<|eot_id|><|start_header_id|>user<|end_header_id|>
{prompt}<|eot_id|><|start_header_id|>assistant<|end_header_id|>
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [MiniThinky-v2-1B-Llama-3.2-f32.gguf](https://huggingface.co/bartowski/MiniThinky-v2-1B-Llama-3.2-GGUF/blob/main/MiniThinky-v2-1B-Llama-3.2-f32.gguf) | f32 | 4.95GB | false | Full F32 weights. |
| [MiniThinky-v2-1B-Llama-3.2-f16.gguf](https://huggingface.co/bartowski/MiniThinky-v2-1B-Llama-3.2-GGUF/blob/main/MiniThinky-v2-1B-Llama-3.2-f16.gguf) | f16 | 2.48GB | false | Full F16 weights. |
| [MiniThinky-v2-1B-Llama-3.2-Q8_0.gguf](https://huggingface.co/bartowski/MiniThinky-v2-1B-Llama-3.2-GGUF/blob/main/MiniThinky-v2-1B-Llama-3.2-Q8_0.gguf) | Q8_0 | 1.32GB | false | Extremely high quality, generally unneeded but max available quant. |
| [MiniThinky-v2-1B-Llama-3.2-Q6_K_L.gguf](https://huggingface.co/bartowski/MiniThinky-v2-1B-Llama-3.2-GGUF/blob/main/MiniThinky-v2-1B-Llama-3.2-Q6_K_L.gguf) | Q6_K_L | 1.09GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [MiniThinky-v2-1B-Llama-3.2-Q6_K.gguf](https://huggingface.co/bartowski/MiniThinky-v2-1B-Llama-3.2-GGUF/blob/main/MiniThinky-v2-1B-Llama-3.2-Q6_K.gguf) | Q6_K | 1.02GB | false | Very high quality, near perfect, *recommended*. |
| [MiniThinky-v2-1B-Llama-3.2-Q5_K_L.gguf](https://huggingface.co/bartowski/MiniThinky-v2-1B-Llama-3.2-GGUF/blob/main/MiniThinky-v2-1B-Llama-3.2-Q5_K_L.gguf) | Q5_K_L | 0.98GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [MiniThinky-v2-1B-Llama-3.2-Q5_K_M.gguf](https://huggingface.co/bartowski/MiniThinky-v2-1B-Llama-3.2-GGUF/blob/main/MiniThinky-v2-1B-Llama-3.2-Q5_K_M.gguf) | Q5_K_M | 0.91GB | false | High quality, *recommended*. |
| [MiniThinky-v2-1B-Llama-3.2-Q5_K_S.gguf](https://huggingface.co/bartowski/MiniThinky-v2-1B-Llama-3.2-GGUF/blob/main/MiniThinky-v2-1B-Llama-3.2-Q5_K_S.gguf) | Q5_K_S | 0.89GB | false | High quality, *recommended*. |
| [MiniThinky-v2-1B-Llama-3.2-Q4_K_L.gguf](https://huggingface.co/bartowski/MiniThinky-v2-1B-Llama-3.2-GGUF/blob/main/MiniThinky-v2-1B-Llama-3.2-Q4_K_L.gguf) | Q4_K_L | 0.87GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [MiniThinky-v2-1B-Llama-3.2-Q4_1.gguf](https://huggingface.co/bartowski/MiniThinky-v2-1B-Llama-3.2-GGUF/blob/main/MiniThinky-v2-1B-Llama-3.2-Q4_1.gguf) | Q4_1 | 0.83GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| [MiniThinky-v2-1B-Llama-3.2-Q4_K_M.gguf](https://huggingface.co/bartowski/MiniThinky-v2-1B-Llama-3.2-GGUF/blob/main/MiniThinky-v2-1B-Llama-3.2-Q4_K_M.gguf) | Q4_K_M | 0.81GB | false | Good quality, default size for most use cases, *recommended*. |
| [MiniThinky-v2-1B-Llama-3.2-Q3_K_XL.gguf](https://huggingface.co/bartowski/MiniThinky-v2-1B-Llama-3.2-GGUF/blob/main/MiniThinky-v2-1B-Llama-3.2-Q3_K_XL.gguf) | Q3_K_XL | 0.80GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [MiniThinky-v2-1B-Llama-3.2-Q4_K_S.gguf](https://huggingface.co/bartowski/MiniThinky-v2-1B-Llama-3.2-GGUF/blob/main/MiniThinky-v2-1B-Llama-3.2-Q4_K_S.gguf) | Q4_K_S | 0.78GB | false | Slightly lower quality with more space savings, *recommended*. |
| [MiniThinky-v2-1B-Llama-3.2-Q4_0.gguf](https://huggingface.co/bartowski/MiniThinky-v2-1B-Llama-3.2-GGUF/blob/main/MiniThinky-v2-1B-Llama-3.2-Q4_0.gguf) | Q4_0 | 0.77GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| [MiniThinky-v2-1B-Llama-3.2-IQ4_NL.gguf](https://huggingface.co/bartowski/MiniThinky-v2-1B-Llama-3.2-GGUF/blob/main/MiniThinky-v2-1B-Llama-3.2-IQ4_NL.gguf) | IQ4_NL | 0.77GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [MiniThinky-v2-1B-Llama-3.2-IQ4_XS.gguf](https://huggingface.co/bartowski/MiniThinky-v2-1B-Llama-3.2-GGUF/blob/main/MiniThinky-v2-1B-Llama-3.2-IQ4_XS.gguf) | IQ4_XS | 0.74GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [MiniThinky-v2-1B-Llama-3.2-Q3_K_L.gguf](https://huggingface.co/bartowski/MiniThinky-v2-1B-Llama-3.2-GGUF/blob/main/MiniThinky-v2-1B-Llama-3.2-Q3_K_L.gguf) | Q3_K_L | 0.73GB | false | Lower quality but usable, good for low RAM availability. |
| [MiniThinky-v2-1B-Llama-3.2-Q3_K_M.gguf](https://huggingface.co/bartowski/MiniThinky-v2-1B-Llama-3.2-GGUF/blob/main/MiniThinky-v2-1B-Llama-3.2-Q3_K_M.gguf) | Q3_K_M | 0.69GB | false | Low quality. |
| [MiniThinky-v2-1B-Llama-3.2-IQ3_M.gguf](https://huggingface.co/bartowski/MiniThinky-v2-1B-Llama-3.2-GGUF/blob/main/MiniThinky-v2-1B-Llama-3.2-IQ3_M.gguf) | IQ3_M | 0.66GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [MiniThinky-v2-1B-Llama-3.2-Q3_K_S.gguf](https://huggingface.co/bartowski/MiniThinky-v2-1B-Llama-3.2-GGUF/blob/main/MiniThinky-v2-1B-Llama-3.2-Q3_K_S.gguf) | Q3_K_S | 0.64GB | false | Low quality, not recommended. |
| [MiniThinky-v2-1B-Llama-3.2-Q2_K_L.gguf](https://huggingface.co/bartowski/MiniThinky-v2-1B-Llama-3.2-GGUF/blob/main/MiniThinky-v2-1B-Llama-3.2-Q2_K_L.gguf) | Q2_K_L | 0.64GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [MiniThinky-v2-1B-Llama-3.2-IQ3_XS.gguf](https://huggingface.co/bartowski/MiniThinky-v2-1B-Llama-3.2-GGUF/blob/main/MiniThinky-v2-1B-Llama-3.2-IQ3_XS.gguf) | IQ3_XS | 0.62GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [MiniThinky-v2-1B-Llama-3.2-Q2_K.gguf](https://huggingface.co/bartowski/MiniThinky-v2-1B-Llama-3.2-GGUF/blob/main/MiniThinky-v2-1B-Llama-3.2-Q2_K.gguf) | Q2_K | 0.58GB | false | Very low quality but surprisingly usable. |
| [MiniThinky-v2-1B-Llama-3.2-IQ2_M.gguf](https://huggingface.co/bartowski/MiniThinky-v2-1B-Llama-3.2-GGUF/blob/main/MiniThinky-v2-1B-Llama-3.2-IQ2_M.gguf) | IQ2_M | 0.52GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/MiniThinky-v2-1B-Llama-3.2-GGUF --include "MiniThinky-v2-1B-Llama-3.2-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/MiniThinky-v2-1B-Llama-3.2-GGUF --include "MiniThinky-v2-1B-Llama-3.2-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (MiniThinky-v2-1B-Llama-3.2-Q8_0) or download them all in place (./)
</details>
## ARM/AVX information
Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.
Now, however, there is something called "online repacking" for weights. details in [this PR](https://github.com/ggerganov/llama.cpp/pull/9921). If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.
As of llama.cpp build [b4282](https://github.com/ggerganov/llama.cpp/releases/tag/b4282) you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.
Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to [this PR](https://github.com/ggerganov/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information (deprecated</summary>
I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}