初始化项目,由ModelHub XC社区提供模型

Model: bartowski/ALLaM-AI_ALLaM-7B-Instruct-preview-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-09-26 21:07:15 +08:00
commit b26332b53f
28 changed files with 290 additions and 0 deletions

47
.gitattributes vendored Normal file
View File

@@ -0,0 +1,47 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bin.* filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zstandard filter=lfs diff=lfs merge=lfs -text
*.tfevents* filter=lfs diff=lfs merge=lfs -text
*.db* filter=lfs diff=lfs merge=lfs -text
*.ark* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*data* filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.meta filter=lfs diff=lfs merge=lfs -text
**/*ckpt*.index filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.gguf* filter=lfs diff=lfs merge=lfs -text
*.ggml filter=lfs diff=lfs merge=lfs -text
*.llamafile* filter=lfs diff=lfs merge=lfs -text
*.pt2 filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:deeabf7e87b4e8cc6b8d63a4d800c865eee4ab86ea510f2f61931d18cf23761c
size 2507124512

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cd8d1aa8f28c7047be90f3b8d65adc8169e151d65af387a9c84dc1382ebc6b42
size 3279646496

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1779c14c78d37d8cfb75d2625a8e2878330e49b28a56b7cab32a566dd8656801
size 2961305376

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fd79836d16e2667ce138d366bc3f1a09361b7186346a8ec6cab35c41bb827141
size 2732764960

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9fa9af4522bccf1826d02c077339e0c43bd55bddbb48651f4e361ae4866fe3d5
size 4007997216

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:db5ccec5b058004c5ac466fbee429f26ece0e805297d79b284aac1e79220c26c
size 3797430048

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1f4310d8165404559f1e1b21ae038fa92376dfb45efe3921a7e7ecb74a74e593
size 2684333856

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:72c5d50e6c384d9b9d491efe211dfcf4b687678cc420c5968f7b6339eea3970b
size 2940333856

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d1b21ddcbfec3807f0b5543ababddc48deb203ab0086c6691052b6be7d122cc3
size 3761893152

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:583b5034a169d71b26dedd7bf5436c6b8d5694797207449b6e5e80fd40aadb5f
size 3462786848

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:bca1a97b6bf951c71025eea65e16451a27d958836d1576a5fb83a79936fa7eee
size 3113086752

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3caa2199b61fba776f830afa10ecafb15fdb718c2ee81b36e11b90bbfdb86429
size 3991269152

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8f56ec27c5453b20a28872cdc3c25e7859bd5858aed5cb0fc45e2bf2dacca498
size 4019269408

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3a243b7f0b9a3092fb5458bf037b8f7e1e844104fa79e541ff3537a57a16949a
size 4429131552

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:67c7cc8dcfbc978426cf7c5a9bfc5e857b97f70d191e65c330dcaaefd142be83
size 4457754400

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:46b1186de822abaa6c054125a209cc80bc5b97d873f68f9046d961341d995873
size 4263194400

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:57d131e54ff174c2ccd52dc47b01e66d578220932b9f0a6ee09f64ba02a38cb5
size 4038930208

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f05d6d8d76ea8256074174f04eaa70f2c8eb160087dc45fbfb1c1d045550e500
size 5143523104

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:71b3ccd5d0a82231d10b1a09a9a7eba70e7afc9dad8663f9b1e544ae5adc0d6a
size 4981731104

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2844d97b9095afb0f533838938b2b943cdeeac0558f069fe53f427226606c418
size 4850265888

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7a004b796407ad0df4feb6181c76ad85e3361006fa29b7fd6d70c3ae3b4b7f79
size 5745176352

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:aa06c0748e05334827310a339459d93647777d8ecb3f12159d1a905b04f23578
size 5872152352

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5a2c4d24eac044110a544c5bf7a6997a33c11123d38036cf98bbbce7dd569d5d
size 7440559904

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b3dccb5836a5e05f522dd16bcc7541fddb2d29124ba5db75fed85b646aceffed
size 14003334624

Binary file not shown.

170
README.md Normal file
View File

@@ -0,0 +1,170 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
language:
- ar
- en
tags: []
license: apache-2.0
base_model: ALLaM-AI/ALLaM-7B-Instruct-preview
---
## Llamacpp imatrix Quantizations of ALLaM-7B-Instruct-preview by ALLaM-AI
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4867">b4867</a> for quantization.
Original model: https://huggingface.co/ALLaM-AI/ALLaM-7B-Instruct-preview
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
Run them directly with [llama.cpp](https://github.com/ggerganov/llama.cpp), or any other llama.cpp based project
## Prompt format
No prompt format found, check original model page
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [ALLaM-7B-Instruct-preview-bf16.gguf](https://huggingface.co/bartowski/ALLaM-AI_ALLaM-7B-Instruct-preview-GGUF/blob/main/ALLaM-AI_ALLaM-7B-Instruct-preview-bf16.gguf) | bf16 | 14.00GB | false | Full BF16 weights. |
| [ALLaM-7B-Instruct-preview-Q8_0.gguf](https://huggingface.co/bartowski/ALLaM-AI_ALLaM-7B-Instruct-preview-GGUF/blob/main/ALLaM-AI_ALLaM-7B-Instruct-preview-Q8_0.gguf) | Q8_0 | 7.44GB | false | Extremely high quality, generally unneeded but max available quant. |
| [ALLaM-7B-Instruct-preview-Q6_K_L.gguf](https://huggingface.co/bartowski/ALLaM-AI_ALLaM-7B-Instruct-preview-GGUF/blob/main/ALLaM-AI_ALLaM-7B-Instruct-preview-Q6_K_L.gguf) | Q6_K_L | 5.87GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [ALLaM-7B-Instruct-preview-Q6_K.gguf](https://huggingface.co/bartowski/ALLaM-AI_ALLaM-7B-Instruct-preview-GGUF/blob/main/ALLaM-AI_ALLaM-7B-Instruct-preview-Q6_K.gguf) | Q6_K | 5.75GB | false | Very high quality, near perfect, *recommended*. |
| [ALLaM-7B-Instruct-preview-Q5_K_L.gguf](https://huggingface.co/bartowski/ALLaM-AI_ALLaM-7B-Instruct-preview-GGUF/blob/main/ALLaM-AI_ALLaM-7B-Instruct-preview-Q5_K_L.gguf) | Q5_K_L | 5.14GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [ALLaM-7B-Instruct-preview-Q5_K_M.gguf](https://huggingface.co/bartowski/ALLaM-AI_ALLaM-7B-Instruct-preview-GGUF/blob/main/ALLaM-AI_ALLaM-7B-Instruct-preview-Q5_K_M.gguf) | Q5_K_M | 4.98GB | false | High quality, *recommended*. |
| [ALLaM-7B-Instruct-preview-Q5_K_S.gguf](https://huggingface.co/bartowski/ALLaM-AI_ALLaM-7B-Instruct-preview-GGUF/blob/main/ALLaM-AI_ALLaM-7B-Instruct-preview-Q5_K_S.gguf) | Q5_K_S | 4.85GB | false | High quality, *recommended*. |
| [ALLaM-7B-Instruct-preview-Q4_K_L.gguf](https://huggingface.co/bartowski/ALLaM-AI_ALLaM-7B-Instruct-preview-GGUF/blob/main/ALLaM-AI_ALLaM-7B-Instruct-preview-Q4_K_L.gguf) | Q4_K_L | 4.46GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [ALLaM-7B-Instruct-preview-Q4_1.gguf](https://huggingface.co/bartowski/ALLaM-AI_ALLaM-7B-Instruct-preview-GGUF/blob/main/ALLaM-AI_ALLaM-7B-Instruct-preview-Q4_1.gguf) | Q4_1 | 4.43GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| [ALLaM-7B-Instruct-preview-Q4_K_M.gguf](https://huggingface.co/bartowski/ALLaM-AI_ALLaM-7B-Instruct-preview-GGUF/blob/main/ALLaM-AI_ALLaM-7B-Instruct-preview-Q4_K_M.gguf) | Q4_K_M | 4.26GB | false | Good quality, default size for most use cases, *recommended*. |
| [ALLaM-7B-Instruct-preview-Q4_K_S.gguf](https://huggingface.co/bartowski/ALLaM-AI_ALLaM-7B-Instruct-preview-GGUF/blob/main/ALLaM-AI_ALLaM-7B-Instruct-preview-Q4_K_S.gguf) | Q4_K_S | 4.04GB | false | Slightly lower quality with more space savings, *recommended*. |
| [ALLaM-7B-Instruct-preview-Q4_0.gguf](https://huggingface.co/bartowski/ALLaM-AI_ALLaM-7B-Instruct-preview-GGUF/blob/main/ALLaM-AI_ALLaM-7B-Instruct-preview-Q4_0.gguf) | Q4_0 | 4.02GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| [ALLaM-7B-Instruct-preview-IQ4_NL.gguf](https://huggingface.co/bartowski/ALLaM-AI_ALLaM-7B-Instruct-preview-GGUF/blob/main/ALLaM-AI_ALLaM-7B-Instruct-preview-IQ4_NL.gguf) | IQ4_NL | 4.01GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [ALLaM-7B-Instruct-preview-Q3_K_XL.gguf](https://huggingface.co/bartowski/ALLaM-AI_ALLaM-7B-Instruct-preview-GGUF/blob/main/ALLaM-AI_ALLaM-7B-Instruct-preview-Q3_K_XL.gguf) | Q3_K_XL | 3.99GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [ALLaM-7B-Instruct-preview-IQ4_XS.gguf](https://huggingface.co/bartowski/ALLaM-AI_ALLaM-7B-Instruct-preview-GGUF/blob/main/ALLaM-AI_ALLaM-7B-Instruct-preview-IQ4_XS.gguf) | IQ4_XS | 3.80GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [ALLaM-7B-Instruct-preview-Q3_K_L.gguf](https://huggingface.co/bartowski/ALLaM-AI_ALLaM-7B-Instruct-preview-GGUF/blob/main/ALLaM-AI_ALLaM-7B-Instruct-preview-Q3_K_L.gguf) | Q3_K_L | 3.76GB | false | Lower quality but usable, good for low RAM availability. |
| [ALLaM-7B-Instruct-preview-Q3_K_M.gguf](https://huggingface.co/bartowski/ALLaM-AI_ALLaM-7B-Instruct-preview-GGUF/blob/main/ALLaM-AI_ALLaM-7B-Instruct-preview-Q3_K_M.gguf) | Q3_K_M | 3.46GB | false | Low quality. |
| [ALLaM-7B-Instruct-preview-IQ3_M.gguf](https://huggingface.co/bartowski/ALLaM-AI_ALLaM-7B-Instruct-preview-GGUF/blob/main/ALLaM-AI_ALLaM-7B-Instruct-preview-IQ3_M.gguf) | IQ3_M | 3.28GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [ALLaM-7B-Instruct-preview-Q3_K_S.gguf](https://huggingface.co/bartowski/ALLaM-AI_ALLaM-7B-Instruct-preview-GGUF/blob/main/ALLaM-AI_ALLaM-7B-Instruct-preview-Q3_K_S.gguf) | Q3_K_S | 3.11GB | false | Low quality, not recommended. |
| [ALLaM-7B-Instruct-preview-IQ3_XS.gguf](https://huggingface.co/bartowski/ALLaM-AI_ALLaM-7B-Instruct-preview-GGUF/blob/main/ALLaM-AI_ALLaM-7B-Instruct-preview-IQ3_XS.gguf) | IQ3_XS | 2.96GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [ALLaM-7B-Instruct-preview-Q2_K_L.gguf](https://huggingface.co/bartowski/ALLaM-AI_ALLaM-7B-Instruct-preview-GGUF/blob/main/ALLaM-AI_ALLaM-7B-Instruct-preview-Q2_K_L.gguf) | Q2_K_L | 2.94GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [ALLaM-7B-Instruct-preview-IQ3_XXS.gguf](https://huggingface.co/bartowski/ALLaM-AI_ALLaM-7B-Instruct-preview-GGUF/blob/main/ALLaM-AI_ALLaM-7B-Instruct-preview-IQ3_XXS.gguf) | IQ3_XXS | 2.73GB | false | Lower quality, new method with decent performance, comparable to Q3 quants. |
| [ALLaM-7B-Instruct-preview-Q2_K.gguf](https://huggingface.co/bartowski/ALLaM-AI_ALLaM-7B-Instruct-preview-GGUF/blob/main/ALLaM-AI_ALLaM-7B-Instruct-preview-Q2_K.gguf) | Q2_K | 2.68GB | false | Very low quality but surprisingly usable. |
| [ALLaM-7B-Instruct-preview-IQ2_M.gguf](https://huggingface.co/bartowski/ALLaM-AI_ALLaM-7B-Instruct-preview-GGUF/blob/main/ALLaM-AI_ALLaM-7B-Instruct-preview-IQ2_M.gguf) | IQ2_M | 2.51GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/ALLaM-AI_ALLaM-7B-Instruct-preview-GGUF --include "ALLaM-AI_ALLaM-7B-Instruct-preview-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/ALLaM-AI_ALLaM-7B-Instruct-preview-GGUF --include "ALLaM-AI_ALLaM-7B-Instruct-preview-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (ALLaM-AI_ALLaM-7B-Instruct-preview-Q8_0) or download them all in place (./)
</details>
## ARM/AVX information
Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.
Now, however, there is something called "online repacking" for weights. details in [this PR](https://github.com/ggerganov/llama.cpp/pull/9921). If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.
As of llama.cpp build [b4282](https://github.com/ggerganov/llama.cpp/releases/tag/b4282) you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.
Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to [this PR](https://github.com/ggerganov/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information (deprecated</summary>
I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Thank you to LM Studio for sponsoring my work.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}