初始化项目,由ModelHub XC社区提供模型

Model: bartowski/RA_Reasoner-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-24 15:31:13 +08:00
commit 3c484cd265
27 changed files with 307 additions and 0 deletions

59
.gitattributes vendored Normal file
View File

@@ -0,0 +1,59 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
RA_Reasoner-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
RA_Reasoner-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
RA_Reasoner-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
RA_Reasoner-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
RA_Reasoner-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
RA_Reasoner-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
RA_Reasoner-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
RA_Reasoner-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
RA_Reasoner-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
RA_Reasoner-Q4_1.gguf filter=lfs diff=lfs merge=lfs -text
RA_Reasoner-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
RA_Reasoner-IQ4_NL.gguf filter=lfs diff=lfs merge=lfs -text
RA_Reasoner-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
RA_Reasoner-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
RA_Reasoner-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
RA_Reasoner-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
RA_Reasoner-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
RA_Reasoner-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
RA_Reasoner-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
RA_Reasoner-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
RA_Reasoner-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
RA_Reasoner-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
RA_Reasoner-f16.gguf filter=lfs diff=lfs merge=lfs -text
RA_Reasoner.imatrix filter=lfs diff=lfs merge=lfs -text

3
RA_Reasoner-IQ2_M.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0d415a3f580ab23a41979f4de706827fba5ce8d8e8120fda5e750767324e86d1
size 3592345312

3
RA_Reasoner-IQ3_M.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:65ea55da7ef69e61228ee848bf847b4393aba39443c3a48978f9e89fe3864d0a
size 4704986848

3
RA_Reasoner-IQ3_XS.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f398237107422a8a0b6388aad8d2a6d7c050876e53b95740250b196f6739eb76
size 4368479968

3
RA_Reasoner-IQ4_NL.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1be7a1468f3d1db89db1ffe2e0403fa76d7fc955e7e58f8e1f83e311a751f635
size 5906347744

3
RA_Reasoner-IQ4_XS.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:76e8b4321757ed4b4bcd1f49c96f520df8dcf55696bc768522213b2cfeca013c
size 5596886752

3
RA_Reasoner-Q2_K.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a6c2bd9f7007aa4e50b1f1bbf6a3e9f5eb2d4feb682171d8aa5211e477c5f246
size 3924047584

3
RA_Reasoner-Q2_K_L.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5423b2ae8d75d5a040b4112b1b7928159ae75ed5e5143f91f0d981377727695f
size 4317263584

3
RA_Reasoner-Q3_K_L.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a6c8ef27c9cff9c18b46420391e74c261cf6072d50cba46f0115a028955c7d85
size 5450807008

3
RA_Reasoner-Q3_K_M.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:980fe27a958d1e5be30a58f8a27f178b5e111c876c82201fdc8412e489f47359
size 5052479200

3
RA_Reasoner-Q3_K_S.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5e65b9e7f678916d223ee939c69d5c1cdbfb2a6912d0e8a8d59df94665b17271
size 4591138528

3
RA_Reasoner-Q3_K_XL.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e1a8d37129f538d01e6e492c099da309727cca14a86883721ca7df8770d7b367
size 5803128544

3
RA_Reasoner-Q4_0.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3fb6e6afa7ed9ebd16b2fb73d5a361abb1331f065b11ef7c555f5374ab26e111
size 5928466144

3
RA_Reasoner-Q4_1.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f7d691dcdf58139dc61808e17dc59b03be0dea289d3a24921c16a83cb5fc1676
size 6525269728

3
RA_Reasoner-Q4_K_L.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6b4506eed174b2c54e25beee44dc2f94666077fef3d60dbff254cc86a4f5a90f
size 6586365664

3
RA_Reasoner-Q4_K_M.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:80235e05a27ea36f711f78a260b143c1caea2db56ed906dfc6c829fd49087310
size 6287521504

3
RA_Reasoner-Q4_K_S.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9ec3b5bd7a0f7d1842b790ab4f254552102e1f009a9bc436c46cd94032f81f2b
size 5952157408

3
RA_Reasoner-Q5_K_L.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4334eff7182c75b5327a770e424ae65723e66fd18a46e7a83d7cba6296f03705
size 7589066464

3
RA_Reasoner-Q5_K_M.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1c76b0cfcc4d681b0f5257b6566f4178bda1dc03b45c878952ccb1fa62a617d7
size 7340553952

3
RA_Reasoner-Q5_K_S.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:b6ea8ffe7675f2d5ade87c4fa8765c2e71c524b8b7246bba7763147a99303ab3
size 7144191712

3
RA_Reasoner-Q6_K.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fed85cf1935bbba6da5ea7b7f3f09c43bb2c2ec4a8bd5f5c07316073ba734191
size 8459400928

3
RA_Reasoner-Q6_K_L.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:425aea03294f07bbb1b26ef2a5f035280ab47c9add6345e46c69e532abcde044
size 8654436064

3
RA_Reasoner-Q8_0.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7bc433368251245a06b4d0add7f13685b33a8457a952979f8fc6b7aa8e40ae9f
size 10955241184

3
RA_Reasoner-f16.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3029cebc457390cfd8783f1be3ce2314051c367938fc7269586751c20d454b50
size 20616558048

3
RA_Reasoner.imatrix Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8796fb7cb04063e7636775d12a746b020c461d820dde3a3bb08190556e90090c
size 6644818

175
README.md Normal file
View File

@@ -0,0 +1,175 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
language:
- en
base_model: Daemontatox/RA_Reasoner
license: apache-2.0
tags:
- text-generation-inference
- transformers
- unsloth
- llama
- trl
---
## Llamacpp imatrix Quantizations of RA_Reasoner
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4381">b4381</a> for quantization.
Original model: https://huggingface.co/Daemontatox/RA_Reasoner
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<|system|>
{system_prompt}
<|user|>
{prompt}
<|assistant|>
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [RA_Reasoner-f16.gguf](https://huggingface.co/bartowski/RA_Reasoner-GGUF/blob/main/RA_Reasoner-f16.gguf) | f16 | 20.62GB | false | Full F16 weights. |
| [RA_Reasoner-Q8_0.gguf](https://huggingface.co/bartowski/RA_Reasoner-GGUF/blob/main/RA_Reasoner-Q8_0.gguf) | Q8_0 | 10.96GB | false | Extremely high quality, generally unneeded but max available quant. |
| [RA_Reasoner-Q6_K_L.gguf](https://huggingface.co/bartowski/RA_Reasoner-GGUF/blob/main/RA_Reasoner-Q6_K_L.gguf) | Q6_K_L | 8.65GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [RA_Reasoner-Q6_K.gguf](https://huggingface.co/bartowski/RA_Reasoner-GGUF/blob/main/RA_Reasoner-Q6_K.gguf) | Q6_K | 8.46GB | false | Very high quality, near perfect, *recommended*. |
| [RA_Reasoner-Q5_K_L.gguf](https://huggingface.co/bartowski/RA_Reasoner-GGUF/blob/main/RA_Reasoner-Q5_K_L.gguf) | Q5_K_L | 7.59GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [RA_Reasoner-Q5_K_M.gguf](https://huggingface.co/bartowski/RA_Reasoner-GGUF/blob/main/RA_Reasoner-Q5_K_M.gguf) | Q5_K_M | 7.34GB | false | High quality, *recommended*. |
| [RA_Reasoner-Q5_K_S.gguf](https://huggingface.co/bartowski/RA_Reasoner-GGUF/blob/main/RA_Reasoner-Q5_K_S.gguf) | Q5_K_S | 7.14GB | false | High quality, *recommended*. |
| [RA_Reasoner-Q4_K_L.gguf](https://huggingface.co/bartowski/RA_Reasoner-GGUF/blob/main/RA_Reasoner-Q4_K_L.gguf) | Q4_K_L | 6.59GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [RA_Reasoner-Q4_1.gguf](https://huggingface.co/bartowski/RA_Reasoner-GGUF/blob/main/RA_Reasoner-Q4_1.gguf) | Q4_1 | 6.53GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| [RA_Reasoner-Q4_K_M.gguf](https://huggingface.co/bartowski/RA_Reasoner-GGUF/blob/main/RA_Reasoner-Q4_K_M.gguf) | Q4_K_M | 6.29GB | false | Good quality, default size for most use cases, *recommended*. |
| [RA_Reasoner-Q4_K_S.gguf](https://huggingface.co/bartowski/RA_Reasoner-GGUF/blob/main/RA_Reasoner-Q4_K_S.gguf) | Q4_K_S | 5.95GB | false | Slightly lower quality with more space savings, *recommended*. |
| [RA_Reasoner-Q4_0.gguf](https://huggingface.co/bartowski/RA_Reasoner-GGUF/blob/main/RA_Reasoner-Q4_0.gguf) | Q4_0 | 5.93GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| [RA_Reasoner-IQ4_NL.gguf](https://huggingface.co/bartowski/RA_Reasoner-GGUF/blob/main/RA_Reasoner-IQ4_NL.gguf) | IQ4_NL | 5.91GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [RA_Reasoner-Q3_K_XL.gguf](https://huggingface.co/bartowski/RA_Reasoner-GGUF/blob/main/RA_Reasoner-Q3_K_XL.gguf) | Q3_K_XL | 5.80GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [RA_Reasoner-IQ4_XS.gguf](https://huggingface.co/bartowski/RA_Reasoner-GGUF/blob/main/RA_Reasoner-IQ4_XS.gguf) | IQ4_XS | 5.60GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [RA_Reasoner-Q3_K_L.gguf](https://huggingface.co/bartowski/RA_Reasoner-GGUF/blob/main/RA_Reasoner-Q3_K_L.gguf) | Q3_K_L | 5.45GB | false | Lower quality but usable, good for low RAM availability. |
| [RA_Reasoner-Q3_K_M.gguf](https://huggingface.co/bartowski/RA_Reasoner-GGUF/blob/main/RA_Reasoner-Q3_K_M.gguf) | Q3_K_M | 5.05GB | false | Low quality. |
| [RA_Reasoner-IQ3_M.gguf](https://huggingface.co/bartowski/RA_Reasoner-GGUF/blob/main/RA_Reasoner-IQ3_M.gguf) | IQ3_M | 4.70GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [RA_Reasoner-Q3_K_S.gguf](https://huggingface.co/bartowski/RA_Reasoner-GGUF/blob/main/RA_Reasoner-Q3_K_S.gguf) | Q3_K_S | 4.59GB | false | Low quality, not recommended. |
| [RA_Reasoner-IQ3_XS.gguf](https://huggingface.co/bartowski/RA_Reasoner-GGUF/blob/main/RA_Reasoner-IQ3_XS.gguf) | IQ3_XS | 4.37GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [RA_Reasoner-Q2_K_L.gguf](https://huggingface.co/bartowski/RA_Reasoner-GGUF/blob/main/RA_Reasoner-Q2_K_L.gguf) | Q2_K_L | 4.32GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [RA_Reasoner-Q2_K.gguf](https://huggingface.co/bartowski/RA_Reasoner-GGUF/blob/main/RA_Reasoner-Q2_K.gguf) | Q2_K | 3.92GB | false | Very low quality but surprisingly usable. |
| [RA_Reasoner-IQ2_M.gguf](https://huggingface.co/bartowski/RA_Reasoner-GGUF/blob/main/RA_Reasoner-IQ2_M.gguf) | IQ2_M | 3.59GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/RA_Reasoner-GGUF --include "RA_Reasoner-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/RA_Reasoner-GGUF --include "RA_Reasoner-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (RA_Reasoner-Q8_0) or download them all in place (./)
</details>
## ARM/AVX information
Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.
Now, however, there is something called "online repacking" for weights. details in [this PR](https://github.com/ggerganov/llama.cpp/pull/9921). If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.
As of llama.cpp build [b4282](https://github.com/ggerganov/llama.cpp/releases/tag/b4282) you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.
Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to [this PR](https://github.com/ggerganov/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information (deprecated</summary>
I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}