Files
ModelHub XC e247992041 初始化项目,由ModelHub XC社区提供模型
Model: AnkitAI/Mistral-Heretica-12B-GGUF
Source: Original Platform
2026-08-06 00:43:17 +08:00

107 lines
4.2 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
base_model: mrcuddle/Mistral-Heretica-12B
base_model_relation: quantized
quantized_by: AnkitAI
library_name: gguf
pipeline_tag: text-generation
license: other
language:
- en
tags:
- gguf
- llama.cpp
- quantized
- mistral
- merge
- roleplay
- creative-writing
---
# Mistral-Heretica-12B-GGUF
![Mistral Heretica 12B GGUF](banner.svg)
GGUF quantizations of [mrcuddle/Mistral-Heretica-12B](https://huggingface.co/mrcuddle/Mistral-Heretica-12B).
Original model: a [Task Arithmetic](https://arxiv.org/abs/2212.04089) merge of Mistral-Nemo-Instruct-2407, [Lumimaid-v0.2-12B](https://huggingface.co/NeverSleep/Lumimaid-v0.2-12B), and [absolute-heresy](https://huggingface.co/MuXodious/Mistral-Nemo-Instruct-2407-absolute-heresy) — tuned for uncensored roleplay and creative writing. 12B params, Mistral architecture.
Quantized with [llama.cpp](https://github.com/ggml-org/llama.cpp) (release b9890). Run these in [LM Studio](https://lmstudio.ai/), [Ollama](https://ollama.com/), [KoboldCpp](https://github.com/LostRuins/koboldcpp), [SillyTavern](https://github.com/SillyTavern/SillyTavern), or `llama.cpp` directly.
## Download
| File | Quant | Size | Description |
|---|---|---|---|
| [Q8_0](./Mistral-Heretica-12B-GGUF-Q8_0.gguf) | Q8_0 | 12 GB | Maximum quality, near-lossless. Overkill for most. |
| [Q6_K](./Mistral-Heretica-12B-GGUF-Q6_K.gguf) | Q6_K | 9.4 GB | Very high quality, effectively lossless. |
| [Q5_K_M](./Mistral-Heretica-12B-GGUF-Q5_K_M.gguf) | Q5_K_M | 8.1 GB | High quality. Balanced pick. |
| [Q4_K_M](./Mistral-Heretica-12B-GGUF-Q4_K_M.gguf) | Q4_K_M | 7.0 GB | Good quality, best size/speed tradeoff. **Recommended default.** |
F16 (23 GB) is the unquantized conversion — only needed if you want to re-quantize yourself.
## Which file should I choose?
Pick the largest quant that fits in your RAM/VRAM with room to spare (leave ~12 GB for context).
- 8 GB RAM/VRAM → **Q4_K_M**
- 12 GB → **Q5_K_M** or **Q6_K**
- 16 GB+ → **Q6_K** or **Q8_0**
For roleplay/creative use, Q4_K_M is plenty — quality loss over Q8 is negligible in practice.
## Prompt format
Mistral instruct format (primary):
```
<s>[INST] {prompt} [/INST]
```
Most front-ends (SillyTavern, LM Studio) apply this automatically when you select a **Mistral / Mistral-Nemo** template. The Lumimaid component was also trained with ChatML, so ChatML works too — but Mistral format is the safe default.
## Recommended settings
Mistral-Nemo (this model's base) is **temperature-sensitive** — high temperature makes it incoherent. Keep it low:
| Setting | Value | Note |
|---|---|---|
| Temperature | **0.5 0.7** | Do not exceed ~1.0. Lower = more coherent, higher = more creative. |
| Min-P | **0.05** | Primary truncation. Leave Top-P/Top-K off if using Min-P. |
| Repetition penalty | 1.05 1.1 | Or use DRY (0.8 / 1.75 / 2) if your front-end supports it. |
| Context | up to 16k reliable | Nemo's trained context is 128k but quality degrades well before that. |
Start at temp 0.6 / min-p 0.05 and adjust from there.
## Run it
**llama.cpp:**
```bash
llama-cli -m Mistral-Heretica-12B-GGUF-Q4_K_M.gguf --jinja -p "Write a short noir opening."
```
**Download one file with the HF CLI (skip the rest):**
```bash
hf download AnkitAI/Mistral-Heretica-12B-GGUF \
Mistral-Heretica-12B-GGUF-Q4_K_M.gguf --local-dir .
```
**Ollama:**
```bash
ollama run hf.co/AnkitAI/Mistral-Heretica-12B-GGUF:Q4_K_M
```
## Verified
Q4_K_M smoke-tested: loads correctly, follows instructions (valid JSON output), and produces coherent creative prose. ~56 tok/s generation on Apple Silicon (M-series).
## License
The base model declares no explicit license. This merge inherits terms from its source models — notably **Lumimaid-v0.2-12B, which is CC-BY-NC-4.0 (non-commercial)**. Treat this model as **non-commercial** unless you verify otherwise with the original authors. These are quantizations only; all model rights belong to the original creators.
## Credits
- Original model: [mrcuddle](https://huggingface.co/mrcuddle)
- Merge components: [NeverSleep](https://huggingface.co/NeverSleep), [MuXodious](https://huggingface.co/MuXodious), [Mistral AI](https://huggingface.co/mistralai)
- Quantization tooling: [llama.cpp](https://github.com/ggml-org/llama.cpp)
More models: [ankitaglawe.com](https://ankitaglawe.com)