初始化项目,由ModelHub XC社区提供模型
Model: Lewdiculous/Nyanade_Stunna-Maid-7B-v0.2-GGUF-IQ-Imatrix Source: Original Platform
This commit is contained in:
48
.gitattributes
vendored
Normal file
48
.gitattributes
vendored
Normal file
@@ -0,0 +1,48 @@
|
|||||||
|
*.7z filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.arrow filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.bin filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.ftz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.gz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.h5 filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.joblib filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.model filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.npy filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.npz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.onnx filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.ot filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.parquet filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pb filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pickle filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pkl filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pt filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pth filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.rar filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
||||||
|
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tar filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tflite filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tgz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.wasm filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.xz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.zip filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.zst filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
||||||
|
imatrix.dat filter=lfs diff=lfs merge=lfs -text
|
||||||
|
mmproj/mmproj-model-f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
Nyanade_Stunna-Maid-7B-v0.2-F16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
Nyanade_Stunna-Maid-7B-v0.2-IQ3_M-imat.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
Nyanade_Stunna-Maid-7B-v0.2-IQ3_XXS-imat.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
Nyanade_Stunna-Maid-7B-v0.2-IQ4_NL-imat.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
Nyanade_Stunna-Maid-7B-v0.2-IQ4_XS-imat.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
Nyanade_Stunna-Maid-7B-v0.2-Q4_K_M-imat.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
Nyanade_Stunna-Maid-7B-v0.2-Q4_K_S-imat.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
Nyanade_Stunna-Maid-7B-v0.2-Q5_K_M-imat.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
Nyanade_Stunna-Maid-7B-v0.2-Q5_K_S-imat.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
Nyanade_Stunna-Maid-7B-v0.2-Q6_K-imat.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
Nyanade_Stunna-Maid-7B-v0.2-Q8_0-imat.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
3
Nyanade_Stunna-Maid-7B-v0.2-F16.gguf
Normal file
3
Nyanade_Stunna-Maid-7B-v0.2-F16.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:3e7e73be4b09d578f2ac34459a068344fc327a4421fd7e422042b816fa374507
|
||||||
|
size 14484731648
|
||||||
3
Nyanade_Stunna-Maid-7B-v0.2-IQ3_M-imat.gguf
Normal file
3
Nyanade_Stunna-Maid-7B-v0.2-IQ3_M-imat.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:c21aa90de932590507daafbe6bc210aa213150d0e432e6d587a3d2609e4f87cd
|
||||||
|
size 3284891456
|
||||||
3
Nyanade_Stunna-Maid-7B-v0.2-IQ3_XXS-imat.gguf
Normal file
3
Nyanade_Stunna-Maid-7B-v0.2-IQ3_XXS-imat.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:ad2cc6608fe2027c614ac30633cdeca12b0206b9c946973389eecd8712ef4a0a
|
||||||
|
size 2827343680
|
||||||
3
Nyanade_Stunna-Maid-7B-v0.2-IQ4_NL-imat.gguf
Normal file
3
Nyanade_Stunna-Maid-7B-v0.2-IQ4_NL-imat.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:f3723b1b9fef737aa7c810cf9b174b9f95a7beb3cdd86234d2d444851efb0936
|
||||||
|
size 4125693760
|
||||||
3
Nyanade_Stunna-Maid-7B-v0.2-IQ4_XS-imat.gguf
Normal file
3
Nyanade_Stunna-Maid-7B-v0.2-IQ4_XS-imat.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:90e4a16fa5a46a620e5bc7402f80a2c50a0f00b0f574c6a0dde1c087de6cd6d4
|
||||||
|
size 3907688256
|
||||||
3
Nyanade_Stunna-Maid-7B-v0.2-Q4_K_M-imat.gguf
Normal file
3
Nyanade_Stunna-Maid-7B-v0.2-Q4_K_M-imat.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:149c6d4c2f795b9dddcf9829293392ce9d7af875f056fca3884c541fd64938a1
|
||||||
|
size 4368439104
|
||||||
3
Nyanade_Stunna-Maid-7B-v0.2-Q4_K_S-imat.gguf
Normal file
3
Nyanade_Stunna-Maid-7B-v0.2-Q4_K_S-imat.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:64d605180d6b509713678c6389ccad8043a8272c5717ecb0153401da0db424ce
|
||||||
|
size 4140373824
|
||||||
3
Nyanade_Stunna-Maid-7B-v0.2-Q5_K_M-imat.gguf
Normal file
3
Nyanade_Stunna-Maid-7B-v0.2-Q5_K_M-imat.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:bab8c2ac7aaa198004a1682eeef65413dc86beb9cff25870229bd8d1d71b5f35
|
||||||
|
size 5131409216
|
||||||
3
Nyanade_Stunna-Maid-7B-v0.2-Q5_K_S-imat.gguf
Normal file
3
Nyanade_Stunna-Maid-7B-v0.2-Q5_K_S-imat.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:8aa9531165b2c9aa3e1889ce34cd9185fb8d49527511ecb6253b3a77111af94f
|
||||||
|
size 4997715776
|
||||||
3
Nyanade_Stunna-Maid-7B-v0.2-Q6_K-imat.gguf
Normal file
3
Nyanade_Stunna-Maid-7B-v0.2-Q6_K-imat.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:1d0a4bf5c0f3aeb225703a3f6ca8973d22eda143f93f4013516b884173053edb
|
||||||
|
size 5942064960
|
||||||
3
Nyanade_Stunna-Maid-7B-v0.2-Q8_0-imat.gguf
Normal file
3
Nyanade_Stunna-Maid-7B-v0.2-Q8_0-imat.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:6d41f50e0df258bc8329ab457d1d1368601456c1594b2c45f9f8259706124966
|
||||||
|
size 7695857472
|
||||||
130
README.md
Normal file
130
README.md
Normal file
@@ -0,0 +1,130 @@
|
|||||||
|
---
|
||||||
|
inference: false
|
||||||
|
tags:
|
||||||
|
- gguf
|
||||||
|
- quantized
|
||||||
|
- roleplay
|
||||||
|
- multimodal
|
||||||
|
- vision
|
||||||
|
- llava
|
||||||
|
- sillytavern
|
||||||
|
- merge
|
||||||
|
- mistral
|
||||||
|
- conversational
|
||||||
|
license: other
|
||||||
|
---
|
||||||
|
|
||||||
|
> [!TIP]
|
||||||
|
> **Support:** <br>
|
||||||
|
> My upload speeds have been cooked and unstable lately. <br>
|
||||||
|
> Realistically I'd need to move to get a better provider. <br>
|
||||||
|
> If you **want** and you are able to... <br>
|
||||||
|
> [**You can support my various endeavors here (Ko-fi).**](https://ko-fi.com/Lewdiculous) <br>
|
||||||
|
> I apologize for disrupting your experience.
|
||||||
|
|
||||||
|
|
||||||
|
# #Roleplay #Multimodal #Vision #Based #Unhinged #Unaligned
|
||||||
|
|
||||||
|
In this repository you can find **GGUF-IQ-Imatrix** quants for [ChaoticNeutrals/Nyanade_Stunna-Maid-7B-v0.2](https://huggingface.co/ChaoticNeutrals/Nyanade_Stunna-Maid-7B-v0.2) and if needed you can get some basic SillyTavern presets [here](https://huggingface.co/Lewdiculous/Model-Requests/tree/main/data/presets/lewdicu-3.0.2-mistral-0.2), if you have issues with repetitiveness or lack or variety in responses I recommend changing the **Temperature** to 1.15, **MinP** to 0.075, **RepPen** to 1.15 and **RepPenRange** to 1024.
|
||||||
|
|
||||||
|
> [!TIP]
|
||||||
|
> **Vision:** <br>
|
||||||
|
> This is a **#multimodal** model that also has optional **#vision** capabilities. <br> Expand the relevant sections bellow and read the full card information if you also want to make use that functionality.
|
||||||
|
>
|
||||||
|
> **Quant options:** <br>
|
||||||
|
> Reading bellow you can also find quant option recommendations for some common GPU VRAM capacities.
|
||||||
|
|
||||||
|
**"Unhinged RP with the spice of the previous 0.420 remixes, 32k context and vision capabilities."**
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
|
||||||
|
# General recommendations for quant options:
|
||||||
|
|
||||||
|
<details><summary>
|
||||||
|
⇲ Click here to expand/hide general common recommendations.
|
||||||
|
</summary>
|
||||||
|
|
||||||
|
*Assuming a context size of 8192 for simplicity and 1GB of Operating System VRAM overhead with some safety margin to avoid overflowing buffers...* <br> <br>
|
||||||
|
**For 11-12GB VRAM:** <br> A GPU with **11-12GB** of VRAM capacity can comfortably use the **Q6_K-imat** quant option and run it at good speeds. <br> This is the same with or without using #vision capabilities. <br> <br>
|
||||||
|
**For 8GB VRAM:** <br> If not using #vision, for GPUs with **8GB** of VRAM capacity the **Q5_K_M-imat** quant option will fit comfortably and should run at good speeds. <br> If **you are** also using #vision from this model opt for the **Q4_K_M-imat** quant option to avoid filling the buffers and potential slowdowns. <br><br>
|
||||||
|
**For 6GB VRAM:** <br> If not using #vision, for GPUs with **6GB** of VRAM capacity the **IQ3_M-imat** quant option should fit comfortably to run at good speeds. <br> If **you are** also using #vision from this model opt for the **IQ3_XXS-imat** quant option. <br><br>
|
||||||
|
|
||||||
|
</details><br>
|
||||||
|
|
||||||
|
|
||||||
|
# Quantization process information:
|
||||||
|
|
||||||
|
<details><summary>
|
||||||
|
⇲ Click here to expand/hide more information about this topic.
|
||||||
|
</summary>
|
||||||
|
|
||||||
|
```python
|
||||||
|
quantization_options = [
|
||||||
|
"IQ3_M", "IQ3_XXS",
|
||||||
|
"Q4_K_M", "Q4_K_S", "IQ4_XS", "IQ4_NL",
|
||||||
|
"Q5_K_M", "Q5_K_S",
|
||||||
|
"Q6_K",
|
||||||
|
"Q8_0"
|
||||||
|
]
|
||||||
|
```
|
||||||
|
|
||||||
|
**Steps performed:**
|
||||||
|
|
||||||
|
```
|
||||||
|
Base⇢ GGUF(F16)⇢ Imatrix-Data(F16)⇢ GGUF(Imatrix-Quants)
|
||||||
|
```
|
||||||
|
The latest of **llama.cpp** available at the time was used, with [imatrix-with-rp-ex.txt](https://huggingface.co/Lewdiculous/Model-Requests/blob/main/data/imatrix/imatrix-with-rp-ex.txt) as calibration data.
|
||||||
|
|
||||||
|
</details><br>
|
||||||
|
|
||||||
|
# What does "Imatrix" mean?
|
||||||
|
|
||||||
|
<details><summary>
|
||||||
|
⇲ Click here to expand/hide more information about this topic.
|
||||||
|
</summary>
|
||||||
|
|
||||||
|
It stands for **Importance Matrix**, a technique used to improve the quality of quantized models.
|
||||||
|
The **Imatrix** is calculated based on calibration data, and it helps determine the importance of different model activations during the quantization process.
|
||||||
|
The idea is to preserve the most important information during quantization, which can help reduce the loss of model performance, especially when the calibration data is diverse.
|
||||||
|
[[1]](https://github.com/ggerganov/llama.cpp/discussions/5006) [[2]](https://github.com/ggerganov/llama.cpp/discussions/5263#discussioncomment-8395384)
|
||||||
|
|
||||||
|
> [!NOTE]
|
||||||
|
> For imatrix data generation, kalomaze's `groups_merged.txt` with additional roleplay chats was used, you can find it [here](https://huggingface.co/Lewdiculous/Model-Requests/blob/main/data/imatrix/imatrix-with-rp-ex.txt) for reference. This was just to add a bit more diversity to the data with the intended use case in mind.
|
||||||
|
|
||||||
|
</details><br>
|
||||||
|
|
||||||
|
# Vision/multimodal capabilities:
|
||||||
|
|
||||||
|
<details><summary>
|
||||||
|
⇲ Click here to expand/hide how this would work in practice in a roleplay chat.
|
||||||
|
</summary>
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
</details><br>
|
||||||
|
|
||||||
|
<details><summary>
|
||||||
|
⇲ Click here to expand/hide how your SillyTavern Image Captions extension settings should look.
|
||||||
|
</summary>
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
</details><br>
|
||||||
|
|
||||||
|
# Required for vision functionality:
|
||||||
|
|
||||||
|
> [!WARNING]
|
||||||
|
> To use the multimodal capabilities of this model, such as **vision**, you also need to load the specified **mmproj** file, you can get it [here](https://huggingface.co/cjpais/llava-1.6-mistral-7b-gguf/blob/main/mmproj-model-f16.gguf) or as uploaded in the **mmproj** folder in the repository.
|
||||||
|
|
||||||
|
1: Make sure you are using the latest version of [KoboldCpp](https://github.com/LostRuins/koboldcpp).
|
||||||
|
|
||||||
|
2: Load the **mmproj file** by using the corresponding section in the interface:
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
2.1: For **CLI** users, you can load the **mmproj file** by adding the respective flag to your usual command:
|
||||||
|
|
||||||
|
```
|
||||||
|
--mmproj your-mmproj-file.gguf
|
||||||
|
```
|
||||||
2422
imatrix-with-rp-ex.txt
Normal file
2422
imatrix-with-rp-ex.txt
Normal file
File diff suppressed because it is too large
Load Diff
3
imatrix.dat
Normal file
3
imatrix.dat
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:3bbde349ce73bfddce1bd5c4326511105ad996c7bc815bc100e6fcbb5f53f00a
|
||||||
|
size 4988126
|
||||||
3
mmproj/mmproj-model-f16.gguf
Normal file
3
mmproj/mmproj-model-f16.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:00205ee8a0d7a381900cd031e43105f86aa0d8c07bf329851e85c71a26632d16
|
||||||
|
size 624451168
|
||||||
Reference in New Issue
Block a user