初始化项目,由ModelHub XC社区提供模型

Model: bartowski/Dans-PersonalityEngine-v1.0.0-8b-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-09-14 05:24:12 +08:00
commit 1973fcbf71
28 changed files with 304 additions and 0 deletions

60
.gitattributes vendored Normal file
View File

@@ -0,0 +1,60 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
Dans-PersonalityEngine-v1.0.0-8b-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
Dans-PersonalityEngine-v1.0.0-8b-Q6_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Dans-PersonalityEngine-v1.0.0-8b-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
Dans-PersonalityEngine-v1.0.0-8b-Q5_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Dans-PersonalityEngine-v1.0.0-8b-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Dans-PersonalityEngine-v1.0.0-8b-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Dans-PersonalityEngine-v1.0.0-8b-Q4_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Dans-PersonalityEngine-v1.0.0-8b-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Dans-PersonalityEngine-v1.0.0-8b-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Dans-PersonalityEngine-v1.0.0-8b-Q4_0_8_8.gguf filter=lfs diff=lfs merge=lfs -text
Dans-PersonalityEngine-v1.0.0-8b-Q4_0_4_8.gguf filter=lfs diff=lfs merge=lfs -text
Dans-PersonalityEngine-v1.0.0-8b-Q4_0_4_4.gguf filter=lfs diff=lfs merge=lfs -text
Dans-PersonalityEngine-v1.0.0-8b-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
Dans-PersonalityEngine-v1.0.0-8b-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
Dans-PersonalityEngine-v1.0.0-8b-Q3_K_XL.gguf filter=lfs diff=lfs merge=lfs -text
Dans-PersonalityEngine-v1.0.0-8b-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Dans-PersonalityEngine-v1.0.0-8b-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
Dans-PersonalityEngine-v1.0.0-8b-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
Dans-PersonalityEngine-v1.0.0-8b-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
Dans-PersonalityEngine-v1.0.0-8b-IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
Dans-PersonalityEngine-v1.0.0-8b-Q2_K_L.gguf filter=lfs diff=lfs merge=lfs -text
Dans-PersonalityEngine-v1.0.0-8b-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
Dans-PersonalityEngine-v1.0.0-8b-IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
Dans-PersonalityEngine-v1.0.0-8b-f16.gguf filter=lfs diff=lfs merge=lfs -text
Dans-PersonalityEngine-v1.0.0-8b.imatrix filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f5612ca856540ebfac43d5d76266182a57bd4417608be58a99484b87298c72e1
size 2948283712

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1fe8894f5b3687521fc6ee44a90ae7b90be82849567dc8f2908cf44043b6e1c8
size 3784826176

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:446a236089f7efcad4a4fa9a5431837817ce50ebff589b8ad52cb7186c301f23
size 3518750016

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:1abcdb3329e159d9d85718d8b4a39b29cdeb1ddb2c0fe3422f2335b87808c341
size 4447665472

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5c64e0ce7314eb38e99ff38de55ab344a1496ebf0c407b62aa3c5e873db05e38
size 3179134272

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f1017bcf429af2f41d62f22503c6f39d368e7b68b1cc29ab95416e45389d4dcc
size 3692158272

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0b040d76fdff634c17ab0d05a4dfce41f209b929f7d3105a586e0e941f7c3a66
size 4321959232

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:dc123c5815a8bf84ce17acd5fedd8cf4f28a26a37a23dbce08c0a5c2be7d148e
size 4018920768

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:397ee800492691e2d14204798bb10c19e9b82a9bcb477ccc296f5001a3070a36
size 3664502080

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:9906843cd1393127b6391998ba10c412d93a1696828ffa59cc7507f4fec8a6d0
size 4781628736

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ee4296386499868418e5e0d856dea21b87ee1f4360bf8754a05f3b8d621520d4
size 4675894592

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c20cd4d606ff9e438026dbac364a9e5cefc4ffa35d2781dc746a4152407d78e1
size 4661214528

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:32b88163da07fd6f562073e43841f5071e4f43d8871336d792d9c0c66aa20859
size 4661214528

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:0bcac64345b85aeb707013af2b9cc77f42003c10c9a313c0ffd715e6d06bcb0c
size 4661214528

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f1b63d73f4032a2851e36372b121b5a33ffdf04431d507d9460deb3e2a34fec7
size 5310635328

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:193b66434c9962e278bb171a21e652f0d3f299f04e86c95f9f75ec5aa8ff006e
size 4920737088

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c4eb37bf7163ad14195f019ba268dd86e462d7d01d19ff6b8f8073c999cf21c9
size 4692671808

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ade502522b3e3f55436acb0dab47763c2abc14725c3e483f8a32819348627cfc
size 6057221440

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:758d27254d9544140c82f321c0df86be87444b48a0a0a72a80076c760cdc071a
size 5732990272

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a600d6c93d9f97a963a63dab432b44d9f92c37d2c71874234759cab125bbef52
size 5599296832

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d8213aa6834a0a4eeda817cf557d96e67002e877e68ac1a73c0316a808fb2e3f
size 6596009280

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8e0da6d5daa590a930d44f0416452186746e89082d0f95e164e0a723c240c161
size 6850469184

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c75a46734658d81c6ff6dbd5f7f60b68cd869f981be67a85f764e72f73e20ec3
size 8540773696

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:85912be3411156f7c88ac776fe2649853add3106f4bd234cc2451e125c76d0dd
size 16068893728

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:13e8498340904c387be19354d7ff8a1a5cda65be06022c3299ea3fecc1a0a091
size 4988170

168
README.md Normal file
View File

@@ -0,0 +1,168 @@
---
base_model: PocketDoc/Dans-PersonalityEngine-v1.0.0-8b
datasets:
- PocketDoc/Dans-MemoryCore-CoreCurriculum-Small
- PocketDoc/Dans-Prosemaxx-Gutenberg
- PocketDoc/Dans-Prosemaxx-Cowriter-S
- PocketDoc/Dans-Prosemaxx-Adventure
- PocketDoc/Dans-Prosemaxx-Opus-Writing
- PocketDoc/Dans-Assistantmaxx-Sharegpt
- PocketDoc/Dans-Assistantmaxx-OpenAssistant2
- PocketDoc/Dans-Assistantmaxx-Opus-instruct-1
- PocketDoc/Dans-Assistantmaxx-Opus-instruct-2
- PocketDoc/Dans-Assistantmaxx-Opus-instruct-3
- PocketDoc/Dans-Assistantmaxx-Opus-Multi-Instruct
- PocketDoc/Dans-Assistantmaxx-sonnetorca-subset
- PocketDoc/Dans-Assistantmaxx-NoRobots
- AquaV/Energetic-Materials-Sharegpt
- AquaV/Chemical-Biological-Safety-Applications-Sharegpt
- AquaV/US-Army-Survival-Sharegpt
- AquaV/Resistance-Sharegpt
- AquaV/Interrogation-Sharegpt
- AquaV/Multi-Environment-Operations-Sharegpt
- PocketDoc/Dans-Mathmaxx
- PJMixers/Math-Multiturn-1K-ShareGPT
- PocketDoc/Dans-Benchmaxx
- PocketDoc/Dans-Codemaxx-LeetCode
- PocketDoc/Dans-Codemaxx-CodeFeedback-Conversations
- PocketDoc/Dans-Codemaxx-CodeFeedback-SingleTurn
- PocketDoc/Dans-Taskmaxx
- PocketDoc/Dans-Taskmaxx-DataPrepper
- PocketDoc/Dans-Taskmaxx-ConcurrentQA-Reworked
- PocketDoc/Dans-Systemmaxx
- PocketDoc/Dans-Toolmaxx-Agent
- PocketDoc/Dans-Toolmaxx-ShellCommands
- PocketDoc/Dans-ASCIIMaxx-Wordart
- PocketDoc/Dans-Personamaxx
- PocketDoc/DansTestYard
language:
- en
license: apache-2.0
pipeline_tag: text-generation
tags:
- chemistry
- biology
- code
- climate
- text-generation-inference
quantized_by: bartowski
---
## Llamacpp imatrix Quantizations of Dans-PersonalityEngine-v1.0.0-8b
Using <a href="https://github.com/ggerganov/llama.cpp/">llama.cpp</a> release <a href="https://github.com/ggerganov/llama.cpp/releases/tag/b3901">b3901</a> for quantization.
Original model: https://huggingface.co/PocketDoc/Dans-PersonalityEngine-v1.0.0-8b
All quants made using imatrix option with dataset from [here](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8)
Run them in [LM Studio](https://lmstudio.ai/)
## Prompt format
```
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
```
## Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
| -------- | ---------- | --------- | ----- | ----------- |
| [Dans-PersonalityEngine-v1.0.0-8b-f16.gguf](https://huggingface.co/bartowski/Dans-PersonalityEngine-v1.0.0-8b-GGUF/blob/main/Dans-PersonalityEngine-v1.0.0-8b-f16.gguf) | f16 | 16.07GB | false | Full F16 weights. |
| [Dans-PersonalityEngine-v1.0.0-8b-Q8_0.gguf](https://huggingface.co/bartowski/Dans-PersonalityEngine-v1.0.0-8b-GGUF/blob/main/Dans-PersonalityEngine-v1.0.0-8b-Q8_0.gguf) | Q8_0 | 8.54GB | false | Extremely high quality, generally unneeded but max available quant. |
| [Dans-PersonalityEngine-v1.0.0-8b-Q6_K_L.gguf](https://huggingface.co/bartowski/Dans-PersonalityEngine-v1.0.0-8b-GGUF/blob/main/Dans-PersonalityEngine-v1.0.0-8b-Q6_K_L.gguf) | Q6_K_L | 6.85GB | false | Uses Q8_0 for embed and output weights. Very high quality, near perfect, *recommended*. |
| [Dans-PersonalityEngine-v1.0.0-8b-Q6_K.gguf](https://huggingface.co/bartowski/Dans-PersonalityEngine-v1.0.0-8b-GGUF/blob/main/Dans-PersonalityEngine-v1.0.0-8b-Q6_K.gguf) | Q6_K | 6.60GB | false | Very high quality, near perfect, *recommended*. |
| [Dans-PersonalityEngine-v1.0.0-8b-Q5_K_L.gguf](https://huggingface.co/bartowski/Dans-PersonalityEngine-v1.0.0-8b-GGUF/blob/main/Dans-PersonalityEngine-v1.0.0-8b-Q5_K_L.gguf) | Q5_K_L | 6.06GB | false | Uses Q8_0 for embed and output weights. High quality, *recommended*. |
| [Dans-PersonalityEngine-v1.0.0-8b-Q5_K_M.gguf](https://huggingface.co/bartowski/Dans-PersonalityEngine-v1.0.0-8b-GGUF/blob/main/Dans-PersonalityEngine-v1.0.0-8b-Q5_K_M.gguf) | Q5_K_M | 5.73GB | false | High quality, *recommended*. |
| [Dans-PersonalityEngine-v1.0.0-8b-Q5_K_S.gguf](https://huggingface.co/bartowski/Dans-PersonalityEngine-v1.0.0-8b-GGUF/blob/main/Dans-PersonalityEngine-v1.0.0-8b-Q5_K_S.gguf) | Q5_K_S | 5.60GB | false | High quality, *recommended*. |
| [Dans-PersonalityEngine-v1.0.0-8b-Q4_K_L.gguf](https://huggingface.co/bartowski/Dans-PersonalityEngine-v1.0.0-8b-GGUF/blob/main/Dans-PersonalityEngine-v1.0.0-8b-Q4_K_L.gguf) | Q4_K_L | 5.31GB | false | Uses Q8_0 for embed and output weights. Good quality, *recommended*. |
| [Dans-PersonalityEngine-v1.0.0-8b-Q4_K_M.gguf](https://huggingface.co/bartowski/Dans-PersonalityEngine-v1.0.0-8b-GGUF/blob/main/Dans-PersonalityEngine-v1.0.0-8b-Q4_K_M.gguf) | Q4_K_M | 4.92GB | false | Good quality, default size for must use cases, *recommended*. |
| [Dans-PersonalityEngine-v1.0.0-8b-Q3_K_XL.gguf](https://huggingface.co/bartowski/Dans-PersonalityEngine-v1.0.0-8b-GGUF/blob/main/Dans-PersonalityEngine-v1.0.0-8b-Q3_K_XL.gguf) | Q3_K_XL | 4.78GB | false | Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability. |
| [Dans-PersonalityEngine-v1.0.0-8b-Q4_K_S.gguf](https://huggingface.co/bartowski/Dans-PersonalityEngine-v1.0.0-8b-GGUF/blob/main/Dans-PersonalityEngine-v1.0.0-8b-Q4_K_S.gguf) | Q4_K_S | 4.69GB | false | Slightly lower quality with more space savings, *recommended*. |
| [Dans-PersonalityEngine-v1.0.0-8b-Q4_0.gguf](https://huggingface.co/bartowski/Dans-PersonalityEngine-v1.0.0-8b-GGUF/blob/main/Dans-PersonalityEngine-v1.0.0-8b-Q4_0.gguf) | Q4_0 | 4.68GB | false | Legacy format, generally not worth using over similarly sized formats |
| [Dans-PersonalityEngine-v1.0.0-8b-Q4_0_8_8.gguf](https://huggingface.co/bartowski/Dans-PersonalityEngine-v1.0.0-8b-GGUF/blob/main/Dans-PersonalityEngine-v1.0.0-8b-Q4_0_8_8.gguf) | Q4_0_8_8 | 4.66GB | false | Optimized for ARM inference. Requires 'sve' support (see link below). *Don't use on Mac or Windows*. |
| [Dans-PersonalityEngine-v1.0.0-8b-Q4_0_4_8.gguf](https://huggingface.co/bartowski/Dans-PersonalityEngine-v1.0.0-8b-GGUF/blob/main/Dans-PersonalityEngine-v1.0.0-8b-Q4_0_4_8.gguf) | Q4_0_4_8 | 4.66GB | false | Optimized for ARM inference. Requires 'i8mm' support (see link below). *Don't use on Mac or Windows*. |
| [Dans-PersonalityEngine-v1.0.0-8b-Q4_0_4_4.gguf](https://huggingface.co/bartowski/Dans-PersonalityEngine-v1.0.0-8b-GGUF/blob/main/Dans-PersonalityEngine-v1.0.0-8b-Q4_0_4_4.gguf) | Q4_0_4_4 | 4.66GB | false | Optimized for ARM inference. Should work well on all ARM chips, pick this if you're unsure. *Don't use on Mac or Windows*. |
| [Dans-PersonalityEngine-v1.0.0-8b-IQ4_XS.gguf](https://huggingface.co/bartowski/Dans-PersonalityEngine-v1.0.0-8b-GGUF/blob/main/Dans-PersonalityEngine-v1.0.0-8b-IQ4_XS.gguf) | IQ4_XS | 4.45GB | false | Decent quality, smaller than Q4_K_S with similar performance, *recommended*. |
| [Dans-PersonalityEngine-v1.0.0-8b-Q3_K_L.gguf](https://huggingface.co/bartowski/Dans-PersonalityEngine-v1.0.0-8b-GGUF/blob/main/Dans-PersonalityEngine-v1.0.0-8b-Q3_K_L.gguf) | Q3_K_L | 4.32GB | false | Lower quality but usable, good for low RAM availability. |
| [Dans-PersonalityEngine-v1.0.0-8b-Q3_K_M.gguf](https://huggingface.co/bartowski/Dans-PersonalityEngine-v1.0.0-8b-GGUF/blob/main/Dans-PersonalityEngine-v1.0.0-8b-Q3_K_M.gguf) | Q3_K_M | 4.02GB | false | Low quality. |
| [Dans-PersonalityEngine-v1.0.0-8b-IQ3_M.gguf](https://huggingface.co/bartowski/Dans-PersonalityEngine-v1.0.0-8b-GGUF/blob/main/Dans-PersonalityEngine-v1.0.0-8b-IQ3_M.gguf) | IQ3_M | 3.78GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| [Dans-PersonalityEngine-v1.0.0-8b-Q2_K_L.gguf](https://huggingface.co/bartowski/Dans-PersonalityEngine-v1.0.0-8b-GGUF/blob/main/Dans-PersonalityEngine-v1.0.0-8b-Q2_K_L.gguf) | Q2_K_L | 3.69GB | false | Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable. |
| [Dans-PersonalityEngine-v1.0.0-8b-Q3_K_S.gguf](https://huggingface.co/bartowski/Dans-PersonalityEngine-v1.0.0-8b-GGUF/blob/main/Dans-PersonalityEngine-v1.0.0-8b-Q3_K_S.gguf) | Q3_K_S | 3.66GB | false | Low quality, not recommended. |
| [Dans-PersonalityEngine-v1.0.0-8b-IQ3_XS.gguf](https://huggingface.co/bartowski/Dans-PersonalityEngine-v1.0.0-8b-GGUF/blob/main/Dans-PersonalityEngine-v1.0.0-8b-IQ3_XS.gguf) | IQ3_XS | 3.52GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| [Dans-PersonalityEngine-v1.0.0-8b-Q2_K.gguf](https://huggingface.co/bartowski/Dans-PersonalityEngine-v1.0.0-8b-GGUF/blob/main/Dans-PersonalityEngine-v1.0.0-8b-Q2_K.gguf) | Q2_K | 3.18GB | false | Very low quality but surprisingly usable. |
| [Dans-PersonalityEngine-v1.0.0-8b-IQ2_M.gguf](https://huggingface.co/bartowski/Dans-PersonalityEngine-v1.0.0-8b-GGUF/blob/main/Dans-PersonalityEngine-v1.0.0-8b-IQ2_M.gguf) | IQ2_M | 2.95GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
## Embed/output weights
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.
Thanks!
## Downloading using huggingface-cli
First, make sure you have hugginface-cli installed:
```
pip install -U "huggingface_hub[cli]"
```
Then, you can target the specific file you want:
```
huggingface-cli download bartowski/Dans-PersonalityEngine-v1.0.0-8b-GGUF --include "Dans-PersonalityEngine-v1.0.0-8b-Q4_K_M.gguf" --local-dir ./
```
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
```
huggingface-cli download bartowski/Dans-PersonalityEngine-v1.0.0-8b-GGUF --include "Dans-PersonalityEngine-v1.0.0-8b-Q8_0/*" --local-dir ./
```
You can either specify a new local-dir (Dans-PersonalityEngine-v1.0.0-8b-Q8_0) or download them all in place (./)
## Q4_0_X_X
These are *NOT* for Metal (Apple) offloading, only ARM chips.
If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons [on the original pull request](https://github.com/ggerganov/llama.cpp/pull/5780#pullrequestreview-21657544660)
To check which one would work best for your ARM chip, you can check [AArch64 SoC features](https://gpages.juszkiewicz.com.pl/arm-socs-table/arm-socs.html) (thanks EloyOn!).
## Which file should I choose?
A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9)
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
[llama.cpp feature matrix](https://github.com/ggerganov/llama.cpp/wiki/Feature-matrix)
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
The I-quants are *not* compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.
## Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset
Thank you ZeroWw for the inspiration to experiment with embed/output
Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

1
configuration.json Normal file
View File

@@ -0,0 +1 @@
{"framework": "pytorch", "task": "text-generation", "allow_remote": true}