--- base_model: mrcuddle/Mistral-Heretica-12B base_model_relation: quantized quantized_by: AnkitAI library_name: gguf pipeline_tag: text-generation license: other language: - en tags: - gguf - llama.cpp - quantized - mistral - merge - roleplay - creative-writing --- # Mistral-Heretica-12B-GGUF ![Mistral Heretica 12B GGUF](banner.svg) GGUF quantizations of [mrcuddle/Mistral-Heretica-12B](https://huggingface.co/mrcuddle/Mistral-Heretica-12B). Original model: a [Task Arithmetic](https://arxiv.org/abs/2212.04089) merge of Mistral-Nemo-Instruct-2407, [Lumimaid-v0.2-12B](https://huggingface.co/NeverSleep/Lumimaid-v0.2-12B), and [absolute-heresy](https://huggingface.co/MuXodious/Mistral-Nemo-Instruct-2407-absolute-heresy) — tuned for uncensored roleplay and creative writing. 12B params, Mistral architecture. Quantized with [llama.cpp](https://github.com/ggml-org/llama.cpp) (release b9890). Run these in [LM Studio](https://lmstudio.ai/), [Ollama](https://ollama.com/), [KoboldCpp](https://github.com/LostRuins/koboldcpp), [SillyTavern](https://github.com/SillyTavern/SillyTavern), or `llama.cpp` directly. ## Download | File | Quant | Size | Description | |---|---|---|---| | [Q8_0](./Mistral-Heretica-12B-GGUF-Q8_0.gguf) | Q8_0 | 12 GB | Maximum quality, near-lossless. Overkill for most. | | [Q6_K](./Mistral-Heretica-12B-GGUF-Q6_K.gguf) | Q6_K | 9.4 GB | Very high quality, effectively lossless. | | [Q5_K_M](./Mistral-Heretica-12B-GGUF-Q5_K_M.gguf) | Q5_K_M | 8.1 GB | High quality. Balanced pick. | | [Q4_K_M](./Mistral-Heretica-12B-GGUF-Q4_K_M.gguf) | Q4_K_M | 7.0 GB | Good quality, best size/speed tradeoff. **Recommended default.** | F16 (23 GB) is the unquantized conversion — only needed if you want to re-quantize yourself. ## Which file should I choose? Pick the largest quant that fits in your RAM/VRAM with room to spare (leave ~1–2 GB for context). - 8 GB RAM/VRAM → **Q4_K_M** - 12 GB → **Q5_K_M** or **Q6_K** - 16 GB+ → **Q6_K** or **Q8_0** For roleplay/creative use, Q4_K_M is plenty — quality loss over Q8 is negligible in practice. ## Prompt format Mistral instruct format (primary): ``` [INST] {prompt} [/INST] ``` Most front-ends (SillyTavern, LM Studio) apply this automatically when you select a **Mistral / Mistral-Nemo** template. The Lumimaid component was also trained with ChatML, so ChatML works too — but Mistral format is the safe default. ## Recommended settings Mistral-Nemo (this model's base) is **temperature-sensitive** — high temperature makes it incoherent. Keep it low: | Setting | Value | Note | |---|---|---| | Temperature | **0.5 – 0.7** | Do not exceed ~1.0. Lower = more coherent, higher = more creative. | | Min-P | **0.05** | Primary truncation. Leave Top-P/Top-K off if using Min-P. | | Repetition penalty | 1.05 – 1.1 | Or use DRY (0.8 / 1.75 / 2) if your front-end supports it. | | Context | up to 16k reliable | Nemo's trained context is 128k but quality degrades well before that. | Start at temp 0.6 / min-p 0.05 and adjust from there. ## Run it **llama.cpp:** ```bash llama-cli -m Mistral-Heretica-12B-GGUF-Q4_K_M.gguf --jinja -p "Write a short noir opening." ``` **Download one file with the HF CLI (skip the rest):** ```bash hf download AnkitAI/Mistral-Heretica-12B-GGUF \ Mistral-Heretica-12B-GGUF-Q4_K_M.gguf --local-dir . ``` **Ollama:** ```bash ollama run hf.co/AnkitAI/Mistral-Heretica-12B-GGUF:Q4_K_M ``` ## Verified Q4_K_M smoke-tested: loads correctly, follows instructions (valid JSON output), and produces coherent creative prose. ~5–6 tok/s generation on Apple Silicon (M-series). ## License The base model declares no explicit license. This merge inherits terms from its source models — notably **Lumimaid-v0.2-12B, which is CC-BY-NC-4.0 (non-commercial)**. Treat this model as **non-commercial** unless you verify otherwise with the original authors. These are quantizations only; all model rights belong to the original creators. ## Credits - Original model: [mrcuddle](https://huggingface.co/mrcuddle) - Merge components: [NeverSleep](https://huggingface.co/NeverSleep), [MuXodious](https://huggingface.co/MuXodious), [Mistral AI](https://huggingface.co/mistralai) - Quantization tooling: [llama.cpp](https://github.com/ggml-org/llama.cpp) More models: [ankitaglawe.com](https://ankitaglawe.com)