57 lines
3.0 KiB
Markdown
57 lines
3.0 KiB
Markdown
|
|
---
|
||
|
|
license: apache-2.0
|
||
|
|
language:
|
||
|
|
- en
|
||
|
|
base_model:
|
||
|
|
- DarwinAnim8or/Trouper-12B
|
||
|
|
pipeline_tag: text-generation
|
||
|
|
tags:
|
||
|
|
- character_roleplay
|
||
|
|
- creative_writing
|
||
|
|
- roleplay
|
||
|
|
- llama-cpp
|
||
|
|
---
|
||
|
|
|
||
|
|
# <a href="https://huggingface.co/DarwinAnim8or/Trouper-12B">Trouper-12B</a> GGUF by <a href="https://huggingface.co/DarwinAnim8or">DarwinAnim8or</a>
|
||
|
|
|
||
|
|
From the original model card:
|
||
|
|
|
||
|
|
>A character roleplay model trained on the custom "Actors" dataset, fine-tuned from Mistral-Nemo-Base-12B. This model was made to expand on the things I learned from TinyRP, and to overcome certain limitations I found from it; also on an entirely new dataset made just for this model.
|
||
|
|
>
|
||
|
|
>This model writes more naturally, less like "AI"; even more so than the 24B model I'm also releasing. I suppose this is because the 12B model saw less synthethic data, and is thus less likely to use phrases typical in AI writing & prose.
|
||
|
|
|
||
|
|
Using an importance matrix actually increases KLD (lowers quality) for `Q8_0` and `Q6_K`, so I didn't use one for those quants. For anything smaller the importance matrix provides better quality with no inference performance penalty.
|
||
|
|
|
||
|
|
The importance matrix was computed using the [Bartowski](https://huggingface.co/bartowski)'s [`calibration_datav3.txt`](https://gist.github.com/bartowski1182/eb213dccb3571f863da82e99418f81e8) dataset, which has been proven to produce good results for all kinds of tasks.
|
||
|
|
|
||
|
|
The KLD metric was measured against the wikitext-2 validation split with 4096 tokens of context.
|
||
|
|
|
||
|
|
| Quant | BPW | Size (GiB) | KLD99 | iMatrix |
|
||
|
|
|---------|------:|-----------:|:--------:|:-------:|
|
||
|
|
| BF16 | 16.00 | 22.81 | N/A | N/A |
|
||
|
|
| Q8_0 | 8.50 | 12.12 | 0.005840 | |
|
||
|
|
| Q6_K | 6.56 | 9.36 | 0.023252 | |
|
||
|
|
| Q5_K_M | 5.70 | 8.12 | 0.044965 | X |
|
||
|
|
| Q5_K_S | 5.56 | 7.93 | 0.053320 | X |
|
||
|
|
| Q4_K_M | 4.88 | 6.96 | 0.130179 | X |
|
||
|
|
| Q4_K_S | 4.65 | 6.62 | 0.156863 | X |
|
||
|
|
| Q3_K_L | 4.28 | 6.10 | 0.391981 | X |
|
||
|
|
| Q3_K_M | 3.97 | 5.66 | 0.453535 | X |
|
||
|
|
| Q3_K_S | 3.61 | 5.15 | 1.682329 | X |
|
||
|
|
| Q2_K | 3.12 | 4.45 | 2.297512 | X |
|
||
|
|
| Q2_K_S | 2.93 | 4.18 | 2.207151 | X |
|
||
|
|
| IQ4_NL | 4.63 | 6.60 | 0.165374 | X |
|
||
|
|
| IQ4_XS | 4.40 | 6.27 | 0.171350 | X |
|
||
|
|
| IQ3_M | 3.73 | 5.32 | 0.513056 | X |
|
||
|
|
| IQ3_S | 3.63 | 5.17 | 0.544215 | X |
|
||
|
|
| IQ3_XS | 3.46 | 4.93 | 0.733339 | X |
|
||
|
|
| IQ3_XXS | 3.23 | 4.60 | 1.081350 | X |
|
||
|
|
| IQ2_M | 2.89 | 4.12 | 1.818141 | X |
|
||
|
|
| IQ2_S | 2.70 | 3.85 | 2.526351 | X |
|
||
|
|
| IQ2_XS | 2.55 | 3.64 | 2.840284 | X |
|
||
|
|
| IQ2_XXS | 2.34 | 3.34 | 3.772680 | X |
|
||
|
|
| IQ1_M | 2.10 | 2.99 | 5.374491 | X |
|
||
|
|
| IQ1_S | 1.95 | 2.79 | 6.467200 | X |
|
||
|
|
|
||
|
|

|