初始化项目,由ModelHub XC社区提供模型

Model: bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-06 02:59:12 +08:00
commit 3e877dbbc2
30 changed files with 318 additions and 0 deletions

62
.gitattributes vendored Normal file
View File

@@ -0,0 +1,62 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
FuseChat-Llama-3.1-8B-Instruct-f16.gguf filter=lfs diff=lfs merge=lfs -text
FuseChat-Llama-3.1-8B-Instruct-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
FuseChat-Llama-3.1-8B-Instruct-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
FuseChat-Llama-3.1-8B-Instruct-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
FuseChat-Llama-3.1-8B-Instruct-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
FuseChat-Llama-3.1-8B-Instruct-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
FuseChat-Llama-3.1-8B-Instruct-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
FuseChat-Llama-3.1-8B-Instruct-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
FuseChat-Llama-3.1-8B-Instruct-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
FuseChat-Llama-3.1-8B-Instruct-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
FuseChat-Llama-3.1-8B-Instruct-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
FuseChat-Llama-3.1-8B-Instruct-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
FuseChat-Llama-3.1-8B-Instruct-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
FuseChat-Llama-3.1-8B-Instruct-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
FuseChat-Llama-3.1-8B-Instruct-IQ4_NL.gguf filter=lfs diff=lfs merge=lfs -text
FuseChat-Llama-3.1-8B-Instruct-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
FuseChat-Llama-3.1-8B-Instruct-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
FuseChat-Llama-3.1-8B-Instruct-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
FuseChat-Llama-3.1-8B-Instruct-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
FuseChat-Llama-3.1-8B-Instruct-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
FuseChat-Llama-3.1-8B-Instruct-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
FuseChat-Llama-3.1-8B-Instruct-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
FuseChat-Llama-3.1-8B-Instruct-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
FuseChat-Llama-3.1-8B-Instruct-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
FuseChat-Llama-3.1-8B-Instruct-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
FuseChat-Llama-3.1-8B-Instruct-f32.gguf filter=lfs diff=lfs merge=lfs -text
FuseChat-Llama-3.1-8B-Instruct.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0b1d4a8d176d2ce2f5f4bec6fc89e84cbb366c605c319c6357c0f365454e875a
size 2948281408

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:764dbd468c7141f9eb955a56314c0504b888b7e64f3a3e263478a3ad482a5418
size 3784823872

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6e3df2e28b5564266c0ed9bed554a2aa9ca7c2879f35260cf9bd06e4f4e7026c
size 3518747712

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2921b6a5b2b07d06949c60722cb530056bb41bd406a72615e5f4c30be544c3a5
size 4677989440

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7c88abe2624bc4074b6c04eacb16bc64d03bab91e1b3fafd754133b73caa3843
size 4447663168

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:74499a4c46716fda522744183ae8fcc4ba21ec8e68057c97b461b7459dffee8e
size 3179131968

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8718f417de680bfaedd321ef15540e021d48f075d25b7cc657a9df4d2815bc2a
size 3692155968

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0ceeac21e2d1b3ea757517f0af52e2b61c692012225479d7e9eb6fe5edd6a8dc
size 4321956928

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:86c7cc3e56c245b8f27b8b5561e148d16f1e9900fd1d0597e00ccab631cedc6f
size 4018918464

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:30d61765a386a9af3fbb317a8a043082139b032dfa107b7cda07b7037347390a
size 3664499776

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d2d7cbe726dbce4e97ce61c96939e187379fa7ff8eb350e1bce744070eb75ea9
size 4781626432

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fe2b751492ff5fb549c0a3240b80c97cbc202774ca10dfb386bb6b57eb88baf5
size 4675892288

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8b63f782ad216152a675c5641f24f1a7e844de177ac5e23dd9214275a7691113
size 4661212224

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c120dc93cda06ceb38b5b82bbf77646f18e7e461e264680864d9232a0982df18
size 4661212224

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c1439f05037871198d48d02b289b7cc8a9e42a9040da4aed630dc580e11ad259
size 4661212224

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:317910cdabc7d4b0b2db7088a0fed8771a96763d78c040cbabb4555a8519b993
size 5310633024

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fe58c8c9b695e36e6b0ee5e4d81ff71ea0a4f1a11fa7bb16e8d6f1b35a58dff6
size 4920734784

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7dd6ea62e9ec9c52a4e5781a3eef83302b33a2a72ca856c1e95441cfbd480803
size 4692669504

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:54ca31c9f0acd477a88673232ede6b692a11f52d0fd2c88645aea63f2c95bfca
size 6057219136

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6064b4df51bdffad8dfab002d5539ac210d70a7608dbcd04b7e33e726c91b17e
size 5732987968

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ffd09ae8fac7ce39e7ab6ef40f062c453ca0685685e68f31a520b4bba729c6be
size 5599294528

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9db796877aa735df6726ed4bc872eb43f21dad4b06d3e10f15d2c07794a459fc
size 6596006976

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9f65bb0cbe07c699ea84b19d4a506331874b709384bf52ac6365053f0da882d7
size 6850466880

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:92bf386f532338452d8ed4b62913a4564eba49ff295c4369ff0d925dad8985fe
size 8540771392

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:dbb5538f6950c755376c911451f0c30a1b447693e400fc1e3b3df76480c8dd4c
size 16068891712

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:091804ac3cab4d05a271551a398ec8dca604ba39b013179c1e92b72d64146394
size 32128881408

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a205a1362d7d247886a617e1abc5fb02dd2989a6894cae2130e4ab8cda68661d
size 4988170

174
README.md Normal file
View File

@@ -0,0 +1,174 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
base_model: FuseAI/FuseChat-Llama-3.1-8B-Instruct
---
## Llamacpp imatrix Quantizations of FuseChat-Llama-3.1-8B-Instruct
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4273">b4273</a> for quantization.
Original model: https://huggingface.co/FuseAI/FuseChat-Llama-3.1-8B-Instruct
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<|begin_of_text|><|start_header_id|>system<|end_header_id|>
{system_prompt}<|eot_id|><|start_header_id|>user<|end_header_id|>
{prompt}<|eot_id|><|start_header_id|>assistant<|end_header_id|>
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [FuseChat-Llama-3.1-8B-Instruct-f32.gguf](https://huggingface.co/bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF/blob/main/FuseChat-Llama-3.1-8B-Instruct-f32.gguf) | f32 | 32.13GB | false | Full F32 weights. |
| [FuseChat-Llama-3.1-8B-Instruct-f16.gguf](https://huggingface.co/bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF/blob/main/FuseChat-Llama-3.1-8B-Instruct-f16.gguf) | f16 | 16.07GB | false | Full F16 weights. |
| [FuseChat-Llama-3.1-8B-Instruct-Q8_0.gguf](https://huggingface.co/bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF/blob/main/FuseChat-Llama-3.1-8B-Instruct-Q8_0.gguf) | Q8_0 | 8.54GB | false | Extremely high quality, generally unneeded but max available quant. |
| [FuseChat-Llama-3.1-8B-Instruct-Q6_K_L.gguf](https://huggingface.co/bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF/blob/main/FuseChat-Llama-3.1-8B-Instruct-Q6_K_L.gguf) | Q6_K_L | 6.85GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [FuseChat-Llama-3.1-8B-Instruct-Q6_K.gguf](https://huggingface.co/bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF/blob/main/FuseChat-Llama-3.1-8B-Instruct-Q6_K.gguf) | Q6_K | 6.60GB | false | Very high quality, near perfect, *recommended*. |
| [FuseChat-Llama-3.1-8B-Instruct-Q5_K_L.gguf](https://huggingface.co/bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF/blob/main/FuseChat-Llama-3.1-8B-Instruct-Q5_K_L.gguf) | Q5_K_L | 6.06GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [FuseChat-Llama-3.1-8B-Instruct-Q5_K_M.gguf](https://huggingface.co/bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF/blob/main/FuseChat-Llama-3.1-8B-Instruct-Q5_K_M.gguf) | Q5_K_M | 5.73GB | false | High quality, *recommended*. |
| [FuseChat-Llama-3.1-8B-Instruct-Q5_K_S.gguf](https://huggingface.co/bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF/blob/main/FuseChat-Llama-3.1-8B-Instruct-Q5_K_S.gguf) | Q5_K_S | 5.60GB | false | High quality, *recommended*. |
| [FuseChat-Llama-3.1-8B-Instruct-Q4_K_L.gguf](https://huggingface.co/bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF/blob/main/FuseChat-Llama-3.1-8B-Instruct-Q4_K_L.gguf) | Q4_K_L | 5.31GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [FuseChat-Llama-3.1-8B-Instruct-Q4_K_M.gguf](https://huggingface.co/bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF/blob/main/FuseChat-Llama-3.1-8B-Instruct-Q4_K_M.gguf) | Q4_K_M | 4.92GB | false | Good quality, default size for most use cases, *recommended*. |
| [FuseChat-Llama-3.1-8B-Instruct-Q3_K_XL.gguf](https://huggingface.co/bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF/blob/main/FuseChat-Llama-3.1-8B-Instruct-Q3_K_XL.gguf) | Q3_K_XL | 4.78GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [FuseChat-Llama-3.1-8B-Instruct-Q4_K_S.gguf](https://huggingface.co/bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF/blob/main/FuseChat-Llama-3.1-8B-Instruct-Q4_K_S.gguf) | Q4_K_S | 4.69GB | false | Slightly lower quality with more space savings, *recommended*. |
| [FuseChat-Llama-3.1-8B-Instruct-Q4_0.gguf](https://huggingface.co/bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF/blob/main/FuseChat-Llama-3.1-8B-Instruct-Q4_0.gguf) | Q4_0 | 4.68GB | false | Legacy format, offers online repacking for ARM CPU inference. |
| [FuseChat-Llama-3.1-8B-Instruct-IQ4_NL.gguf](https://huggingface.co/bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF/blob/main/FuseChat-Llama-3.1-8B-Instruct-IQ4_NL.gguf) | IQ4_NL | 4.68GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [FuseChat-Llama-3.1-8B-Instruct-Q4_0_8_8.gguf](https://huggingface.co/bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF/blob/main/FuseChat-Llama-3.1-8B-Instruct-Q4_0_8_8.gguf) | Q4_0_8_8 | 4.66GB | false | Optimized for ARM and AVX inference. Requires 'sve' support for ARM (see details below). *Don't use on Mac*. |
| [FuseChat-Llama-3.1-8B-Instruct-Q4_0_4_8.gguf](https://huggingface.co/bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF/blob/main/FuseChat-Llama-3.1-8B-Instruct-Q4_0_4_8.gguf) | Q4_0_4_8 | 4.66GB | false | Optimized for ARM inference. Requires 'i8mm' support (see details below). *Don't use on Mac*. |
| [FuseChat-Llama-3.1-8B-Instruct-Q4_0_4_4.gguf](https://huggingface.co/bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF/blob/main/FuseChat-Llama-3.1-8B-Instruct-Q4_0_4_4.gguf) | Q4_0_4_4 | 4.66GB | false | Optimized for ARM inference. Should work well on all ARM chips, not for use with GPUs. *Don't use on Mac*. |
| [FuseChat-Llama-3.1-8B-Instruct-IQ4_XS.gguf](https://huggingface.co/bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF/blob/main/FuseChat-Llama-3.1-8B-Instruct-IQ4_XS.gguf) | IQ4_XS | 4.45GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [FuseChat-Llama-3.1-8B-Instruct-Q3_K_L.gguf](https://huggingface.co/bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF/blob/main/FuseChat-Llama-3.1-8B-Instruct-Q3_K_L.gguf) | Q3_K_L | 4.32GB | false | Lower quality but usable, good for low RAM availability. |
| [FuseChat-Llama-3.1-8B-Instruct-Q3_K_M.gguf](https://huggingface.co/bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF/blob/main/FuseChat-Llama-3.1-8B-Instruct-Q3_K_M.gguf) | Q3_K_M | 4.02GB | false | Low quality. |
| [FuseChat-Llama-3.1-8B-Instruct-IQ3_M.gguf](https://huggingface.co/bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF/blob/main/FuseChat-Llama-3.1-8B-Instruct-IQ3_M.gguf) | IQ3_M | 3.78GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [FuseChat-Llama-3.1-8B-Instruct-Q2_K_L.gguf](https://huggingface.co/bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF/blob/main/FuseChat-Llama-3.1-8B-Instruct-Q2_K_L.gguf) | Q2_K_L | 3.69GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [FuseChat-Llama-3.1-8B-Instruct-Q3_K_S.gguf](https://huggingface.co/bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF/blob/main/FuseChat-Llama-3.1-8B-Instruct-Q3_K_S.gguf) | Q3_K_S | 3.66GB | false | Low quality, not recommended. |
| [FuseChat-Llama-3.1-8B-Instruct-IQ3_XS.gguf](https://huggingface.co/bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF/blob/main/FuseChat-Llama-3.1-8B-Instruct-IQ3_XS.gguf) | IQ3_XS | 3.52GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [FuseChat-Llama-3.1-8B-Instruct-Q2_K.gguf](https://huggingface.co/bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF/blob/main/FuseChat-Llama-3.1-8B-Instruct-Q2_K.gguf) | Q2_K | 3.18GB | false | Very low quality but surprisingly usable. |
| [FuseChat-Llama-3.1-8B-Instruct-IQ2_M.gguf](https://huggingface.co/bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF/blob/main/FuseChat-Llama-3.1-8B-Instruct-IQ2_M.gguf) | IQ2_M | 2.95GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF --include "FuseChat-Llama-3.1-8B-Instruct-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/FuseChat-Llama-3.1-8B-Instruct-GGUF --include "FuseChat-Llama-3.1-8B-Instruct-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (FuseChat-Llama-3.1-8B-Instruct-Q8_0) or download them all in place (./)
</details>
## ARM/AVX information
Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.
Now, however, there is something called "online repacking" for weights. details in [this PR](https://github.com/ggerganov/llama.cpp/pull/9921). If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.
As of llama.cpp build [b4282](https://github.com/ggerganov/llama.cpp/releases/tag/b4282) you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.
Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to [this PR](https://github.com/ggerganov/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information</summary>
These are *NOT* for Metal (Apple) or GPU (nvidia/AMD/intel) offloading, only ARM chips (and certain AVX2/AVX512 CPUs).
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
If you're using a CPU that supports AVX2 or AVX512 (typically server CPUs and AMD's latest Zen5 CPUs) and are not offloading to a GPU, the Q4_0_8_8 may offer a nice speed as well:
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}