初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Human-Like-Qwen2.5-7B-Instruct-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-09-01 14:42:13 +08:00
commit 7efa45324a
28 changed files with 355 additions and 0 deletions

60
.gitattributes vendored Normal file
View File

@@ -0,0 +1,60 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Humanish-Qwen2.5-7B-Instruct-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Humanish-Qwen2.5-7B-Instruct-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Humanish-Qwen2.5-7B-Instruct-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Humanish-Qwen2.5-7B-Instruct-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Humanish-Qwen2.5-7B-Instruct-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Humanish-Qwen2.5-7B-Instruct-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Humanish-Qwen2.5-7B-Instruct-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Humanish-Qwen2.5-7B-Instruct-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Humanish-Qwen2.5-7B-Instruct-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Humanish-Qwen2.5-7B-Instruct-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
Humanish-Qwen2.5-7B-Instruct-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
Humanish-Qwen2.5-7B-Instruct-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
Humanish-Qwen2.5-7B-Instruct-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Humanish-Qwen2.5-7B-Instruct-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Humanish-Qwen2.5-7B-Instruct-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Humanish-Qwen2.5-7B-Instruct-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Humanish-Qwen2.5-7B-Instruct-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Humanish-Qwen2.5-7B-Instruct-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Humanish-Qwen2.5-7B-Instruct-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Humanish-Qwen2.5-7B-Instruct-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Humanish-Qwen2.5-7B-Instruct-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Humanish-Qwen2.5-7B-Instruct-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Humanish-Qwen2.5-7B-Instruct-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Humanish-Qwen2.5-7B-Instruct-f16.gguf filter=lfs diff=lfs merge=lfs -text
Humanish-Qwen2.5-7B-Instruct.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:24290482b8caa5da6993a4ea40bd6cc889f6dea4fba2bacbdb532760bcc7c2e4
size 2780340672

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8a0738d83172ae2e14e65ba59212a393b4be39644d3b574c6141ff202b269729
size 3574010304

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:e207abc57230f122146199002d5c3c5ac59660447aa16302cea23f121bc92762
size 3346254272

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:eac4cbeb870144315a0312a068c592c9048844f9bd6eeacebc6c2a02fe08a085
size 4218470848

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f44389447bcd3f82c86c2f088d25e59ada175930a8886c65d83480060aef95d3
size 3015938496

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0779951e3390b48ca34a114b72ca94510acf3a485ff7a3495c3589bb589e4ee2
size 3548162496

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5b3481b51be82d8af4a7d7a0be5a145908ab30f2473d7f9035c446ecffcc5f6d
size 4088457664

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a4b6310f051c62597748ca58518811d5ba752a2f3153206b221e7471f6e5b25e
size 3808389568

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:307b9558c759208c60d19d0677f2537cc8729cb0aa612d173597a65f5b48ec71
size 3492366784

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:153557cd1b644cd89db47cb1686ac67a7060b59bb6be7d0a7b13fd79b961a90c
size 4565330368

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7775a5f1d5ff5e2c44b1e907a31aba8f09156eca9ccc6310a4bbd7e52f10326a
size 4444119488

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:64cd3dcf41a653f5c3012a6a86032ca9643e9ec9d69da47558d069b3d961b7b0
size 4431389120

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4f5f6dbb9d569d7cacb0530c15b265c6553c0968269d81dd606a756bec83fb38
size 4431389120

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ce06a76dd43550b92bd667840bff7f472850a555eb9955b367881eb30d1a9af8
size 4431389120

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0f64f01683ecb70478fe454a2361c36fe78a9c92c2864d69cb0b4041e6ef5f1c
size 5087562176

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fcbf5949d2e27e4de14db68bdf74845386dce66a5022b1c2a0c8c857bb61ff85
size 4683071936

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:98ba40dc153050d64f373ba894490baa52263396f610f2a095dbd212dd62bfbf
size 4457767360

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:65faf4b63c4f55b33f87dd89483a5d6e7fc90f3fd9468d39ee0627f821718164
size 5781195200

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:758b6126c88f4f8ad4c3b71162232f9e568c6c4e01234994d776abf1039b3791
size 5444829632

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4799962aa3832e880bc92dcbc2b3b3093a5619b698b174ee83d69e52a8503387
size 5315174848

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0176d9252a6e8474a19e8444c871c20974b24291953ddfdf618ca671b4915555
size 6254197184

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:7dee935a9ad06d3b676b84c5d04ed1ed2d9a9e76d6ea585d568e11e437a5b92c
size 6518180288

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d88664500a8e7794edf8418f10b0be0105a97cbb096acbfbc2b1a6d9d6de2b52
size 8098523584

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4304e3f5878b2d96516404d059e134e719f1a5df86d9ae39cc5a74b0724fcd63
size 15237851264

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a4e36cf877e45af90951a84775359c09d9a78db4f38e58fa26088ad1bc3545cf
size 4536678

219
README.md Normal file
View File

@@ -0,0 +1,219 @@
---
base_model: HumanLLMs/Humanish-Qwen2.5-7B-Instruct
license: apache-2.0
pipeline_tag: text-generation
tags:
- axolotl
- dpo
- trl
- generated_from_trainer
quantized_by: bartowski
model-index:
- name: Humanish-Qwen2.5-7B-Instruct
results:
- task:
type: text-generation
name: Text Generation
dataset:
name: IFEval (0-Shot)
type: HuggingFaceH4/ifeval
args:
num_few_shot: 0
metrics:
- type: inst_level_strict_acc and prompt_level_strict_acc
value: 72.84
name: strict accuracy
source:
url: https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard?query=HumanLLMs/Humanish-Qwen2.5-7B-Instruct
name: Open LLM Leaderboard
- task:
type: text-generation
name: Text Generation
dataset:
name: BBH (3-Shot)
type: BBH
args:
num_few_shot: 3
metrics:
- type: acc_norm
value: 34.48
name: normalized accuracy
source:
url: https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard?query=HumanLLMs/Humanish-Qwen2.5-7B-Instruct
name: Open LLM Leaderboard
- task:
type: text-generation
name: Text Generation
dataset:
name: MATH Lvl 5 (4-Shot)
type: hendrycks/competition_math
args:
num_few_shot: 4
metrics:
- type: exact_match
value: 0.0
name: exact match
source:
url: https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard?query=HumanLLMs/Humanish-Qwen2.5-7B-Instruct
name: Open LLM Leaderboard
- task:
type: text-generation
name: Text Generation
dataset:
name: GPQA (0-shot)
type: Idavidrein/gpqa
args:
num_few_shot: 0
metrics:
- type: acc_norm
value: 6.49
name: acc_norm
source:
url: https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard?query=HumanLLMs/Humanish-Qwen2.5-7B-Instruct
name: Open LLM Leaderboard
- task:
type: text-generation
name: Text Generation
dataset:
name: MuSR (0-shot)
type: TAUR-Lab/MuSR
args:
num_few_shot: 0
metrics:
- type: acc_norm
value: 8.42
name: acc_norm
source:
url: https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard?query=HumanLLMs/Humanish-Qwen2.5-7B-Instruct
name: Open LLM Leaderboard
- task:
type: text-generation
name: Text Generation
dataset:
name: MMLU-PRO (5-shot)
type: TIGER-Lab/MMLU-Pro
config: main
split: test
args:
num_few_shot: 5
metrics:
- type: acc
value: 37.76
name: accuracy
source:
url: https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard?query=HumanLLMs/Humanish-Qwen2.5-7B-Instruct
name: Open LLM Leaderboard
---
## Llamacpp imatrix Quantizations of Humanish-Qwen2.5-7B-Instruct
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b3878">b3878</a> for quantization.
Original model: https://huggingface.co/HumanLLMs/Humanish-Qwen2.5-7B-Instruct
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
No prompt format found, check original model page
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Humanish-Qwen2.5-7B-Instruct-f16.gguf](https://huggingface.co/bartowski/Humanish-Qwen2.5-7B-Instruct-GGUF/blob/main/Humanish-Qwen2.5-7B-Instruct-f16.gguf) | f16 | 15.24GB | false | Full F16 weights. |
| [Humanish-Qwen2.5-7B-Instruct-Q8_0.gguf](https://huggingface.co/bartowski/Humanish-Qwen2.5-7B-Instruct-GGUF/blob/main/Humanish-Qwen2.5-7B-Instruct-Q8_0.gguf) | Q8_0 | 8.10GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Humanish-Qwen2.5-7B-Instruct-Q6_K_L.gguf](https://huggingface.co/bartowski/Humanish-Qwen2.5-7B-Instruct-GGUF/blob/main/Humanish-Qwen2.5-7B-Instruct-Q6_K_L.gguf) | Q6_K_L | 6.52GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Humanish-Qwen2.5-7B-Instruct-Q6_K.gguf](https://huggingface.co/bartowski/Humanish-Qwen2.5-7B-Instruct-GGUF/blob/main/Humanish-Qwen2.5-7B-Instruct-Q6_K.gguf) | Q6_K | 6.25GB | false | Very high quality, near perfect, *recommended*. |
| [Humanish-Qwen2.5-7B-Instruct-Q5_K_L.gguf](https://huggingface.co/bartowski/Humanish-Qwen2.5-7B-Instruct-GGUF/blob/main/Humanish-Qwen2.5-7B-Instruct-Q5_K_L.gguf) | Q5_K_L | 5.78GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Humanish-Qwen2.5-7B-Instruct-Q5_K_M.gguf](https://huggingface.co/bartowski/Humanish-Qwen2.5-7B-Instruct-GGUF/blob/main/Humanish-Qwen2.5-7B-Instruct-Q5_K_M.gguf) | Q5_K_M | 5.44GB | false | High quality, *recommended*. |
| [Humanish-Qwen2.5-7B-Instruct-Q5_K_S.gguf](https://huggingface.co/bartowski/Humanish-Qwen2.5-7B-Instruct-GGUF/blob/main/Humanish-Qwen2.5-7B-Instruct-Q5_K_S.gguf) | Q5_K_S | 5.32GB | false | High quality, *recommended*. |
| [Humanish-Qwen2.5-7B-Instruct-Q4_K_L.gguf](https://huggingface.co/bartowski/Humanish-Qwen2.5-7B-Instruct-GGUF/blob/main/Humanish-Qwen2.5-7B-Instruct-Q4_K_L.gguf) | Q4_K_L | 5.09GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Humanish-Qwen2.5-7B-Instruct-Q4_K_M.gguf](https://huggingface.co/bartowski/Humanish-Qwen2.5-7B-Instruct-GGUF/blob/main/Humanish-Qwen2.5-7B-Instruct-Q4_K_M.gguf) | Q4_K_M | 4.68GB | false | Good quality, default size for must use cases, *recommended*. |
| [Humanish-Qwen2.5-7B-Instruct-Q3_K_XL.gguf](https://huggingface.co/bartowski/Humanish-Qwen2.5-7B-Instruct-GGUF/blob/main/Humanish-Qwen2.5-7B-Instruct-Q3_K_XL.gguf) | Q3_K_XL | 4.57GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Humanish-Qwen2.5-7B-Instruct-Q4_K_S.gguf](https://huggingface.co/bartowski/Humanish-Qwen2.5-7B-Instruct-GGUF/blob/main/Humanish-Qwen2.5-7B-Instruct-Q4_K_S.gguf) | Q4_K_S | 4.46GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Humanish-Qwen2.5-7B-Instruct-Q4_0.gguf](https://huggingface.co/bartowski/Humanish-Qwen2.5-7B-Instruct-GGUF/blob/main/Humanish-Qwen2.5-7B-Instruct-Q4_0.gguf) | Q4_0 | 4.44GB | false | Legacy format, generally not worth using over similarly sized formats |
| [Humanish-Qwen2.5-7B-Instruct-Q4_0_8_8.gguf](https://huggingface.co/bartowski/Humanish-Qwen2.5-7B-Instruct-GGUF/blob/main/Humanish-Qwen2.5-7B-Instruct-Q4_0_8_8.gguf) | Q4_0_8_8 | 4.43GB | false | Optimized for ARM inference. Requires 'sve' support (see link below). |
| [Humanish-Qwen2.5-7B-Instruct-Q4_0_4_8.gguf](https://huggingface.co/bartowski/Humanish-Qwen2.5-7B-Instruct-GGUF/blob/main/Humanish-Qwen2.5-7B-Instruct-Q4_0_4_8.gguf) | Q4_0_4_8 | 4.43GB | false | Optimized for ARM inference. Requires 'i8mm' support (see link below). |
| [Humanish-Qwen2.5-7B-Instruct-Q4_0_4_4.gguf](https://huggingface.co/bartowski/Humanish-Qwen2.5-7B-Instruct-GGUF/blob/main/Humanish-Qwen2.5-7B-Instruct-Q4_0_4_4.gguf) | Q4_0_4_4 | 4.43GB | false | Optimized for ARM inference. Should work well on all ARM chips, pick this if you're unsure. |
| [Humanish-Qwen2.5-7B-Instruct-IQ4_XS.gguf](https://huggingface.co/bartowski/Humanish-Qwen2.5-7B-Instruct-GGUF/blob/main/Humanish-Qwen2.5-7B-Instruct-IQ4_XS.gguf) | IQ4_XS | 4.22GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Humanish-Qwen2.5-7B-Instruct-Q3_K_L.gguf](https://huggingface.co/bartowski/Humanish-Qwen2.5-7B-Instruct-GGUF/blob/main/Humanish-Qwen2.5-7B-Instruct-Q3_K_L.gguf) | Q3_K_L | 4.09GB | false | Lower quality but usable, good for low RAM availability. |
| [Humanish-Qwen2.5-7B-Instruct-Q3_K_M.gguf](https://huggingface.co/bartowski/Humanish-Qwen2.5-7B-Instruct-GGUF/blob/main/Humanish-Qwen2.5-7B-Instruct-Q3_K_M.gguf) | Q3_K_M | 3.81GB | false | Low quality. |
| [Humanish-Qwen2.5-7B-Instruct-IQ3_M.gguf](https://huggingface.co/bartowski/Humanish-Qwen2.5-7B-Instruct-GGUF/blob/main/Humanish-Qwen2.5-7B-Instruct-IQ3_M.gguf) | IQ3_M | 3.57GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Humanish-Qwen2.5-7B-Instruct-Q2_K_L.gguf](https://huggingface.co/bartowski/Humanish-Qwen2.5-7B-Instruct-GGUF/blob/main/Humanish-Qwen2.5-7B-Instruct-Q2_K_L.gguf) | Q2_K_L | 3.55GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Humanish-Qwen2.5-7B-Instruct-Q3_K_S.gguf](https://huggingface.co/bartowski/Humanish-Qwen2.5-7B-Instruct-GGUF/blob/main/Humanish-Qwen2.5-7B-Instruct-Q3_K_S.gguf) | Q3_K_S | 3.49GB | false | Low quality, not recommended. |
| [Humanish-Qwen2.5-7B-Instruct-IQ3_XS.gguf](https://huggingface.co/bartowski/Humanish-Qwen2.5-7B-Instruct-GGUF/blob/main/Humanish-Qwen2.5-7B-Instruct-IQ3_XS.gguf) | IQ3_XS | 3.35GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Humanish-Qwen2.5-7B-Instruct-Q2_K.gguf](https://huggingface.co/bartowski/Humanish-Qwen2.5-7B-Instruct-GGUF/blob/main/Humanish-Qwen2.5-7B-Instruct-Q2_K.gguf) | Q2_K | 3.02GB | false | Very low quality but surprisingly usable. |
| [Humanish-Qwen2.5-7B-Instruct-IQ2_M.gguf](https://huggingface.co/bartowski/Humanish-Qwen2.5-7B-Instruct-GGUF/blob/main/Humanish-Qwen2.5-7B-Instruct-IQ2_M.gguf) | IQ2_M | 2.78GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Humanish-Qwen2.5-7B-Instruct-GGUF --include "Humanish-Qwen2.5-7B-Instruct-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Humanish-Qwen2.5-7B-Instruct-GGUF --include "Humanish-Qwen2.5-7B-Instruct-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Humanish-Qwen2.5-7B-Instruct-Q8_0) or download them all in place (./)
## Q4_0_X_X
These are *NOT* for Metal (Apple) offloading, only ARM chips.
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset
Thank you ZeroWw for the inspiration to experiment with embed/output
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}