初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Llama-SmolTalk-3.2-1B-Instruct-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-09-01 12:42:14 +08:00
commit f178ecd692
24 changed files with 295 additions and 0 deletions

56
.gitattributes vendored Normal file
View File

@@ -0,0 +1,56 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Llama-SmolTalk-3.2-1B-Instruct-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Llama-SmolTalk-3.2-1B-Instruct-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Llama-SmolTalk-3.2-1B-Instruct-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Llama-SmolTalk-3.2-1B-Instruct-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Llama-SmolTalk-3.2-1B-Instruct-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Llama-SmolTalk-3.2-1B-Instruct-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Llama-SmolTalk-3.2-1B-Instruct-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Llama-SmolTalk-3.2-1B-Instruct-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Llama-SmolTalk-3.2-1B-Instruct-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Llama-SmolTalk-3.2-1B-Instruct-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
Llama-SmolTalk-3.2-1B-Instruct-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
Llama-SmolTalk-3.2-1B-Instruct-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
Llama-SmolTalk-3.2-1B-Instruct-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Llama-SmolTalk-3.2-1B-Instruct-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Llama-SmolTalk-3.2-1B-Instruct-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Llama-SmolTalk-3.2-1B-Instruct-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Llama-SmolTalk-3.2-1B-Instruct-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Llama-SmolTalk-3.2-1B-Instruct-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Llama-SmolTalk-3.2-1B-Instruct-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Llama-SmolTalk-3.2-1B-Instruct-f16.gguf filter=lfs diff=lfs merge=lfs -text
Llama-SmolTalk-3.2-1B-Instruct.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:239803e59e38cc0ade05da63f6e4c566519ba0a9159daa12117dc2c1727c3b4c
size 657290656

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:73a5edf3f71a0336474bba894045358df923e870723c95528dd18735476602c5
size 743142816

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ef3605bfb81a1e8a8aff40723f1a3f719e6056092eac8d27541fb2540020aa76
size 580875680

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2c47afbce36a5838483124084c3cf640dbc265a6ed75d1d06a2e5872024991cd
size 644490656

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e04c5e996b4c83462c57c3b0d71890318804b640cff6589831d58e03ccd7c84a
size 732525984

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:826fef4f663dcb7f741ccc873fd08159df000c96de0f109c5b228cdf0b9af6ab
size 796140960

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0e6be0a67cdf084eec2a3e080ac183c4af8df902ee0b5dd7fa92f19e01065ea9
size 773027232

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f52c073ca1426b7f73a39a7f1bd7897ffe027c4b25137e726f8d013a5abfdecd
size 770930080

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:09576b98af302cbd597f516e34d3d8656393ebb2a0476da66fd4735dd5512f8e
size 770930080

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:419b516a6563b2db97d754f75e20c187919f6414074ade8558fc505c1e90f504
size 770930080

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f7f0b1b5f534be7c34946585eb8ad6a1f357b13f2eb0c31821c7c6d105f4b5a1
size 871310752

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8ad49fcc72ec4bc0701b79a437659329fd441c6b5564f08b1d3236384b4f7eba
size 807695776

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e497fba5f1e89b9296b1dcbcb0dfed2ab83c9ab4d076cd40faa611a8d1be0f4c
size 775648672

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:643b20e88d9406a6ddb9a85d57ea5ecdff09b002f054fb320376c0823041add1
size 975119776

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:66ca559a23bd46e33e9e247488ef645d6a9dfc3d4f3d5b7b14f18fce9f4f72b4
size 911504800

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8181129a3dd4c87901ff026453a8e060de8b03b8acc4053889b1846a33131d55
size 892564896

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:57d6be7d6789cf714ee970294ebb4093e4005820826c967d70b20c170b3a1d63
size 1021801888

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:012b5a63968d6fc8c6ed803918456721a7c373218e2a884a94c8d0a71ebad04c
size 1085416864

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3fb8faac49ab09edaa055ff438df3f5c513fddae51ef5768b3da0f6031ff5216
size 1321084320

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:24e22faa83803228b21085a8114ad61d4400688a14f781c0b50c5bc1b32fa62e
size 2479596672

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:deaa0471f97ad47b9d6b7e01a9f6bbade2d88c50c2da2ee34755733ede115303
size 1314426

175
README.md Normal file
View File

@@ -0,0 +1,175 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
datasets:
- HuggingFaceTB/smoltalk
base_model: prithivMLmods/Llama-SmolTalk-3.2-1B-Instruct
tags:
- Llama
- Llama-CPP
- SmolTalk
- ollama
- bin
license: creativeml-openrail-m
language:
- en
---
## Llamacpp imatrix Quantizations of Llama-SmolTalk-3.2-1B-Instruct
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4132">b4132</a> for quantization.
Original model: https://huggingface.co/prithivMLmods/Llama-SmolTalk-3.2-1B-Instruct
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<|begin_of_text|><|start_header_id|>system<|end_header_id|>
Cutting Knowledge Date: December 2023
Today Date: 26 July 2024
{system_prompt}<|eot_id|><|start_header_id|>user<|end_header_id|>
{prompt}<|eot_id|><|start_header_id|>assistant<|end_header_id|>
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Llama-SmolTalk-3.2-1B-Instruct-f16.gguf](https://huggingface.co/bartowski/Llama-SmolTalk-3.2-1B-Instruct-GGUF/blob/main/Llama-SmolTalk-3.2-1B-Instruct-f16.gguf) | f16 | 2.48GB | false | Full F16 weights. |
| [Llama-SmolTalk-3.2-1B-Instruct-Q8_0.gguf](https://huggingface.co/bartowski/Llama-SmolTalk-3.2-1B-Instruct-GGUF/blob/main/Llama-SmolTalk-3.2-1B-Instruct-Q8_0.gguf) | Q8_0 | 1.32GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Llama-SmolTalk-3.2-1B-Instruct-Q6_K_L.gguf](https://huggingface.co/bartowski/Llama-SmolTalk-3.2-1B-Instruct-GGUF/blob/main/Llama-SmolTalk-3.2-1B-Instruct-Q6_K_L.gguf) | Q6_K_L | 1.09GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Llama-SmolTalk-3.2-1B-Instruct-Q6_K.gguf](https://huggingface.co/bartowski/Llama-SmolTalk-3.2-1B-Instruct-GGUF/blob/main/Llama-SmolTalk-3.2-1B-Instruct-Q6_K.gguf) | Q6_K | 1.02GB | false | Very high quality, near perfect, *recommended*. |
| [Llama-SmolTalk-3.2-1B-Instruct-Q5_K_L.gguf](https://huggingface.co/bartowski/Llama-SmolTalk-3.2-1B-Instruct-GGUF/blob/main/Llama-SmolTalk-3.2-1B-Instruct-Q5_K_L.gguf) | Q5_K_L | 0.98GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Llama-SmolTalk-3.2-1B-Instruct-Q5_K_M.gguf](https://huggingface.co/bartowski/Llama-SmolTalk-3.2-1B-Instruct-GGUF/blob/main/Llama-SmolTalk-3.2-1B-Instruct-Q5_K_M.gguf) | Q5_K_M | 0.91GB | false | High quality, *recommended*. |
| [Llama-SmolTalk-3.2-1B-Instruct-Q5_K_S.gguf](https://huggingface.co/bartowski/Llama-SmolTalk-3.2-1B-Instruct-GGUF/blob/main/Llama-SmolTalk-3.2-1B-Instruct-Q5_K_S.gguf) | Q5_K_S | 0.89GB | false | High quality, *recommended*. |
| [Llama-SmolTalk-3.2-1B-Instruct-Q4_K_L.gguf](https://huggingface.co/bartowski/Llama-SmolTalk-3.2-1B-Instruct-GGUF/blob/main/Llama-SmolTalk-3.2-1B-Instruct-Q4_K_L.gguf) | Q4_K_L | 0.87GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Llama-SmolTalk-3.2-1B-Instruct-Q4_K_M.gguf](https://huggingface.co/bartowski/Llama-SmolTalk-3.2-1B-Instruct-GGUF/blob/main/Llama-SmolTalk-3.2-1B-Instruct-Q4_K_M.gguf) | Q4_K_M | 0.81GB | false | Good quality, default size for most use cases, *recommended*. |
| [Llama-SmolTalk-3.2-1B-Instruct-Q3_K_XL.gguf](https://huggingface.co/bartowski/Llama-SmolTalk-3.2-1B-Instruct-GGUF/blob/main/Llama-SmolTalk-3.2-1B-Instruct-Q3_K_XL.gguf) | Q3_K_XL | 0.80GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Llama-SmolTalk-3.2-1B-Instruct-Q4_K_S.gguf](https://huggingface.co/bartowski/Llama-SmolTalk-3.2-1B-Instruct-GGUF/blob/main/Llama-SmolTalk-3.2-1B-Instruct-Q4_K_S.gguf) | Q4_K_S | 0.78GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Llama-SmolTalk-3.2-1B-Instruct-Q4_0_8_8.gguf](https://huggingface.co/bartowski/Llama-SmolTalk-3.2-1B-Instruct-GGUF/blob/main/Llama-SmolTalk-3.2-1B-Instruct-Q4_0_8_8.gguf) | Q4_0_8_8 | 0.77GB | false | Optimized for ARM and AVX inference. Requires 'sve' support for ARM (see details below). *Don't use on Mac*. |
| [Llama-SmolTalk-3.2-1B-Instruct-Q4_0_4_8.gguf](https://huggingface.co/bartowski/Llama-SmolTalk-3.2-1B-Instruct-GGUF/blob/main/Llama-SmolTalk-3.2-1B-Instruct-Q4_0_4_8.gguf) | Q4_0_4_8 | 0.77GB | false | Optimized for ARM inference. Requires 'i8mm' support (see details below). *Don't use on Mac*. |
| [Llama-SmolTalk-3.2-1B-Instruct-Q4_0_4_4.gguf](https://huggingface.co/bartowski/Llama-SmolTalk-3.2-1B-Instruct-GGUF/blob/main/Llama-SmolTalk-3.2-1B-Instruct-Q4_0_4_4.gguf) | Q4_0_4_4 | 0.77GB | false | Optimized for ARM inference. Should work well on all ARM chips, not for use with GPUs. *Don't use on Mac*. |
| [Llama-SmolTalk-3.2-1B-Instruct-Q4_0.gguf](https://huggingface.co/bartowski/Llama-SmolTalk-3.2-1B-Instruct-GGUF/blob/main/Llama-SmolTalk-3.2-1B-Instruct-Q4_0.gguf) | Q4_0 | 0.77GB | false | Legacy format, generally not worth using over similarly sized formats |
| [Llama-SmolTalk-3.2-1B-Instruct-IQ4_XS.gguf](https://huggingface.co/bartowski/Llama-SmolTalk-3.2-1B-Instruct-GGUF/blob/main/Llama-SmolTalk-3.2-1B-Instruct-IQ4_XS.gguf) | IQ4_XS | 0.74GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Llama-SmolTalk-3.2-1B-Instruct-Q3_K_L.gguf](https://huggingface.co/bartowski/Llama-SmolTalk-3.2-1B-Instruct-GGUF/blob/main/Llama-SmolTalk-3.2-1B-Instruct-Q3_K_L.gguf) | Q3_K_L | 0.73GB | false | Lower quality but usable, good for low RAM availability. |
| [Llama-SmolTalk-3.2-1B-Instruct-IQ3_M.gguf](https://huggingface.co/bartowski/Llama-SmolTalk-3.2-1B-Instruct-GGUF/blob/main/Llama-SmolTalk-3.2-1B-Instruct-IQ3_M.gguf) | IQ3_M | 0.66GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Llama-SmolTalk-3.2-1B-Instruct-Q2_K_L.gguf](https://huggingface.co/bartowski/Llama-SmolTalk-3.2-1B-Instruct-GGUF/blob/main/Llama-SmolTalk-3.2-1B-Instruct-Q2_K_L.gguf) | Q2_K_L | 0.64GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Llama-SmolTalk-3.2-1B-Instruct-Q2_K.gguf](https://huggingface.co/bartowski/Llama-SmolTalk-3.2-1B-Instruct-GGUF/blob/main/Llama-SmolTalk-3.2-1B-Instruct-Q2_K.gguf) | Q2_K | 0.58GB | false | Very low quality but surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Llama-SmolTalk-3.2-1B-Instruct-GGUF --include "Llama-SmolTalk-3.2-1B-Instruct-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Llama-SmolTalk-3.2-1B-Instruct-GGUF --include "Llama-SmolTalk-3.2-1B-Instruct-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Llama-SmolTalk-3.2-1B-Instruct-Q8_0) or download them all in place (./)
</details>
## Q4_0_X_X information
<details>
<summary>Click to view Q4_0_X_X information</summary>
These are *NOT* for Metal (Apple) or GPU (nvidia/AMD/intel) offloading, only ARM chips (and certain AVX2/AVX512 CPUs).
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
If you're using a CPU that supports AVX2 or AVX512 (typically server CPUs and AMD's latest Zen5 CPUs) and are not offloading to a GPU, the Q4_0_8_8 may offer a nice speed as well:
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}