初始化项目,由ModelHub XC社区提供模型

Model: bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-09-28 06:19:13 +08:00
commit 8815bb2aad
31 changed files with 316 additions and 0 deletions

63
.gitattributes vendored Normal file
View File

@@ -0,0 +1,63 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
KoboldAI_LLaMA2-13B-Tiefighter-f16.gguf filter=lfs diff=lfs merge=lfs -text
KoboldAI_LLaMA2-13B-Tiefighter-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
KoboldAI_LLaMA2-13B-Tiefighter-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
KoboldAI_LLaMA2-13B-Tiefighter-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
KoboldAI_LLaMA2-13B-Tiefighter-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
KoboldAI_LLaMA2-13B-Tiefighter-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
KoboldAI_LLaMA2-13B-Tiefighter-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
KoboldAI_LLaMA2-13B-Tiefighter-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
KoboldAI_LLaMA2-13B-Tiefighter-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
KoboldAI_LLaMA2-13B-Tiefighter-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
KoboldAI_LLaMA2-13B-Tiefighter-Q4_1.gguf filter=lfs diff=lfs merge=lfs -text
KoboldAI_LLaMA2-13B-Tiefighter-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
KoboldAI_LLaMA2-13B-Tiefighter-IQ4_NL.gguf filter=lfs diff=lfs merge=lfs -text
KoboldAI_LLaMA2-13B-Tiefighter-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
KoboldAI_LLaMA2-13B-Tiefighter-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
KoboldAI_LLaMA2-13B-Tiefighter-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
KoboldAI_LLaMA2-13B-Tiefighter-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
KoboldAI_LLaMA2-13B-Tiefighter-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
KoboldAI_LLaMA2-13B-Tiefighter-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
KoboldAI_LLaMA2-13B-Tiefighter-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
KoboldAI_LLaMA2-13B-Tiefighter-IQ3_XXS.gguf filter=lfs diff=lfs merge=lfs -text
KoboldAI_LLaMA2-13B-Tiefighter-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
KoboldAI_LLaMA2-13B-Tiefighter-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
KoboldAI_LLaMA2-13B-Tiefighter-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
KoboldAI_LLaMA2-13B-Tiefighter-IQ2_S.gguf filter=lfs diff=lfs merge=lfs -text
KoboldAI_LLaMA2-13B-Tiefighter-f32/KoboldAI_LLaMA2-13B-Tiefighter-f32-00001-of-00002.gguf filter=lfs diff=lfs merge=lfs -text
KoboldAI_LLaMA2-13B-Tiefighter-f32/KoboldAI_LLaMA2-13B-Tiefighter-f32-00002-of-00002.gguf filter=lfs diff=lfs merge=lfs -text
KoboldAI_LLaMA2-13B-Tiefighter.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:908074af5ce2c56151475d1dc31128864e83e39daa9f32cd44ee0d04623a93db
size 4517579328

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:063fd6dff27be41bbbc086dd790af4ac05b5d74dea415f57e5cf34c28a6866d8
size 4197681728

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:bb8cf66593739c6d7ccfa12fb0e69df67b5c4b027fb7301dcbccef627d8d8d53
size 5984510528

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:92e803d3c254d925c023606217eb003c26e19574d8189b271402b666e80d70a2
size 5361611328

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:68d5ffd3f7b708c23b1992ba20a39a29363a04fa6b6e9e70e89cbf83c72245b4
size 4960561728

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:60043bba47fc01acbfac6a51347d076ce2200112dd99966d63d5436593829ddb
size 7365835328

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3f0c998a23e5857b9abdd07738ff2affcdb13e503163bf54ae646f1db0b740a6
size 6964222528

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:55adea4ccc1dbc50d26f68825b23c6b8518f2d30c0189f139db88372a43455fe
size 4854270528

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9e5c75a3d74323c73298cad60dd55d61e526136e4626520aa57d1ea18e07e818
size 5014270528

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:593b25d0c34511e801265abdbb187d8ebb81325626dfed588be2716c2e63e25e
size 6929560128

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7a39e778a8e0347f1bbb038bbdf4cd6b9526985755bc3dcd46dde9557f3068cc
size 6337770048

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d1421ca0db62605c845dc6fa559c8db032d5d6d63680ecf9349dea6d77aca32e
size 5658980928

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:76e4db778749051e17df28996c4251274591daec9d2d72f128c12c216108351a
size 7072920128

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cab6bf21df8f09c749d733b569cf766c992692e5e90c11fe259b5bd569b0813b
size 7387953728

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:54415eed4dff8964233fea42c61802efdc1039f71e68c30ff5bbbb5fdf80949d
size 8169060928

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0c433d88520828df710cf1b9d3bbd371f1ff7250a0c26175eec150e895d55b45
size 7987556928

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:57bdb6a01c5ee92e8483ebc47bd94bd0da133f67b904ec41f3ceb8a9d9412780
size 7865956928

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:6c89a0bec66e9b3dbcb3ec385933db417c46c51480d175e85be34f5538ac9b7d
size 7423179328

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:bbb6ba333f8eae4405bf0640a52b0e66b54b5fe5f761ef97f3a4148a9253b617
size 9331044928

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:009b39bab557b12e11dbaffb0fe510b163469a279acac695f059c1ecc9d576a3
size 9229924928

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d1e79ceb6662ba606fc6a044efee0467796b1eb01674d6805f3762696bb56c41
size 8972286528

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2dcee683735f0216291b9d305963319f5d45ccd275a88e5f93aaa8bc9e2a3045
size 10679140928

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c65e47cceaa94de674059b64b75b9c48659c7b8d8949ed6c46bbd4298b767e3d
size 10758500928

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7d598647674d465b99f9d0a4dc407c109697d0719ff64190af0f993afab492ed
size 13831320128

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fae85eabeeec19f8e4af9b647dc021e2b231cafe83b25f80cb192e8a7e805b04
size 26033304128

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8a0fb409d648ffe27728a7ba7e78558bf1ad928d6f6f77eb0ba61d989eaed04b
size 39989456416

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:60ecf37af792422eb1b84bf964e108b1409d589bd74bf6bc69cc5a2632f44a4c
size 12074746816

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:102c6b1ff7f5599a0d78329ec80af1d24a866dcfa0bbbe9b2453321d1e6616a9
size 7136338

168
README.md Normal file
View File

@@ -0,0 +1,168 @@
---
quantized_by: bartowski
pipeline_tag: text-generation
license: llama2
base_model: KoboldAI/LLaMA2-13B-Tiefighter
---
## Llamacpp imatrix Quantizations of LLaMA2-13B-Tiefighter by KoboldAI
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b4585">b4585</a> for quantization.
Original model: https://huggingface.co/KoboldAI/LLaMA2-13B-Tiefighter
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
Run them directly with [llama.cpp](https://github.com/ggerganov/llama.cpp), or any other llama.cpp based project
## Prompt format
No prompt format found, check original model page
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [LLaMA2-13B-Tiefighter-f32.gguf](https://huggingface.co/bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF/tree/main/KoboldAI_LLaMA2-13B-Tiefighter-f32) | f32 | 52.06GB | true | Full F32 weights. |
| [LLaMA2-13B-Tiefighter-f16.gguf](https://huggingface.co/bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF/blob/main/KoboldAI_LLaMA2-13B-Tiefighter-f16.gguf) | f16 | 26.03GB | false | Full F16 weights. |
| [LLaMA2-13B-Tiefighter-Q8_0.gguf](https://huggingface.co/bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF/blob/main/KoboldAI_LLaMA2-13B-Tiefighter-Q8_0.gguf) | Q8_0 | 13.83GB | false | Extremely high quality, generally unneeded but max available quant. |
| [LLaMA2-13B-Tiefighter-Q6_K_L.gguf](https://huggingface.co/bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF/blob/main/KoboldAI_LLaMA2-13B-Tiefighter-Q6_K_L.gguf) | Q6_K_L | 10.76GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [LLaMA2-13B-Tiefighter-Q6_K.gguf](https://huggingface.co/bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF/blob/main/KoboldAI_LLaMA2-13B-Tiefighter-Q6_K.gguf) | Q6_K | 10.68GB | false | Very high quality, near perfect, *recommended*. |
| [LLaMA2-13B-Tiefighter-Q5_K_L.gguf](https://huggingface.co/bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF/blob/main/KoboldAI_LLaMA2-13B-Tiefighter-Q5_K_L.gguf) | Q5_K_L | 9.33GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [LLaMA2-13B-Tiefighter-Q5_K_M.gguf](https://huggingface.co/bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF/blob/main/KoboldAI_LLaMA2-13B-Tiefighter-Q5_K_M.gguf) | Q5_K_M | 9.23GB | false | High quality, *recommended*. |
| [LLaMA2-13B-Tiefighter-Q5_K_S.gguf](https://huggingface.co/bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF/blob/main/KoboldAI_LLaMA2-13B-Tiefighter-Q5_K_S.gguf) | Q5_K_S | 8.97GB | false | High quality, *recommended*. |
| [LLaMA2-13B-Tiefighter-Q4_1.gguf](https://huggingface.co/bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF/blob/main/KoboldAI_LLaMA2-13B-Tiefighter-Q4_1.gguf) | Q4_1 | 8.17GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| [LLaMA2-13B-Tiefighter-Q4_K_L.gguf](https://huggingface.co/bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF/blob/main/KoboldAI_LLaMA2-13B-Tiefighter-Q4_K_L.gguf) | Q4_K_L | 7.99GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [LLaMA2-13B-Tiefighter-Q4_K_M.gguf](https://huggingface.co/bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF/blob/main/KoboldAI_LLaMA2-13B-Tiefighter-Q4_K_M.gguf) | Q4_K_M | 7.87GB | false | Good quality, default size for most use cases, *recommended*. |
| [LLaMA2-13B-Tiefighter-Q4_K_S.gguf](https://huggingface.co/bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF/blob/main/KoboldAI_LLaMA2-13B-Tiefighter-Q4_K_S.gguf) | Q4_K_S | 7.42GB | false | Slightly lower quality with more space savings, *recommended*. |
| [LLaMA2-13B-Tiefighter-Q4_0.gguf](https://huggingface.co/bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF/blob/main/KoboldAI_LLaMA2-13B-Tiefighter-Q4_0.gguf) | Q4_0 | 7.39GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| [LLaMA2-13B-Tiefighter-IQ4_NL.gguf](https://huggingface.co/bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF/blob/main/KoboldAI_LLaMA2-13B-Tiefighter-IQ4_NL.gguf) | IQ4_NL | 7.37GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| [LLaMA2-13B-Tiefighter-Q3_K_XL.gguf](https://huggingface.co/bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF/blob/main/KoboldAI_LLaMA2-13B-Tiefighter-Q3_K_XL.gguf) | Q3_K_XL | 7.07GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [LLaMA2-13B-Tiefighter-IQ4_XS.gguf](https://huggingface.co/bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF/blob/main/KoboldAI_LLaMA2-13B-Tiefighter-IQ4_XS.gguf) | IQ4_XS | 6.96GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [LLaMA2-13B-Tiefighter-Q3_K_L.gguf](https://huggingface.co/bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF/blob/main/KoboldAI_LLaMA2-13B-Tiefighter-Q3_K_L.gguf) | Q3_K_L | 6.93GB | false | Lower quality but usable, good for low RAM availability. |
| [LLaMA2-13B-Tiefighter-Q3_K_M.gguf](https://huggingface.co/bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF/blob/main/KoboldAI_LLaMA2-13B-Tiefighter-Q3_K_M.gguf) | Q3_K_M | 6.34GB | false | Low quality. |
| [LLaMA2-13B-Tiefighter-IQ3_M.gguf](https://huggingface.co/bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF/blob/main/KoboldAI_LLaMA2-13B-Tiefighter-IQ3_M.gguf) | IQ3_M | 5.98GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [LLaMA2-13B-Tiefighter-Q3_K_S.gguf](https://huggingface.co/bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF/blob/main/KoboldAI_LLaMA2-13B-Tiefighter-Q3_K_S.gguf) | Q3_K_S | 5.66GB | false | Low quality, not recommended. |
| [LLaMA2-13B-Tiefighter-IQ3_XS.gguf](https://huggingface.co/bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF/blob/main/KoboldAI_LLaMA2-13B-Tiefighter-IQ3_XS.gguf) | IQ3_XS | 5.36GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [LLaMA2-13B-Tiefighter-Q2_K_L.gguf](https://huggingface.co/bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF/blob/main/KoboldAI_LLaMA2-13B-Tiefighter-Q2_K_L.gguf) | Q2_K_L | 5.01GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [LLaMA2-13B-Tiefighter-IQ3_XXS.gguf](https://huggingface.co/bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF/blob/main/KoboldAI_LLaMA2-13B-Tiefighter-IQ3_XXS.gguf) | IQ3_XXS | 4.96GB | false | Lower quality, new method with decent performance, comparable to Q3 quants. |
| [LLaMA2-13B-Tiefighter-Q2_K.gguf](https://huggingface.co/bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF/blob/main/KoboldAI_LLaMA2-13B-Tiefighter-Q2_K.gguf) | Q2_K | 4.85GB | false | Very low quality but surprisingly usable. |
| [LLaMA2-13B-Tiefighter-IQ2_M.gguf](https://huggingface.co/bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF/blob/main/KoboldAI_LLaMA2-13B-Tiefighter-IQ2_M.gguf) | IQ2_M | 4.52GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
| [LLaMA2-13B-Tiefighter-IQ2_S.gguf](https://huggingface.co/bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF/blob/main/KoboldAI_LLaMA2-13B-Tiefighter-IQ2_S.gguf) | IQ2_S | 4.20GB | false | Low quality, uses SOTA techniques to be usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
## Downloading using huggingface-cli
<details>
<summary>Click to view download instructions</summary>
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF --include "KoboldAI_LLaMA2-13B-Tiefighter-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/KoboldAI_LLaMA2-13B-Tiefighter-GGUF --include "KoboldAI_LLaMA2-13B-Tiefighter-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (KoboldAI_LLaMA2-13B-Tiefighter-Q8_0) or download them all in place (./)
</details>
## ARM/AVX information
Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.
Now, however, there is something called "online repacking" for weights. details in [this PR](https://github.com/ggerganov/llama.cpp/pull/9921). If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.
As of llama.cpp build [b4282](https://github.com/ggerganov/llama.cpp/releases/tag/b4282) you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.
Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to [this PR](https://github.com/ggerganov/llama.cpp/pull/10541) which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.
<details>
<summary>Click to view Q4_0_X_X information (deprecated</summary>
I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.
<details>
<summary>Click to view benchmarks on an AVX2 system (EPYC7702)</summary>
| model | size | params | backend | threads | test | t/s | % (vs Q4_0) |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% |
| qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% |
| qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% |
| qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% |
Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation
</details>
</details>
## Which file should I choose?
<details>
<summary>Click here for details</summary>
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
</details>
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Thank you to LM Studio for sponsoring my work.
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}