From e24799204140e65294cb92abae401d7a47d135c4 Mon Sep 17 00:00:00 2001 From: ModelHub XC Date: Thu, 6 Aug 2026 00:43:17 +0800 Subject: [PATCH] =?UTF-8?q?=E5=88=9D=E5=A7=8B=E5=8C=96=E9=A1=B9=E7=9B=AE?= =?UTF-8?q?=EF=BC=8C=E7=94=B1ModelHub=20XC=E7=A4=BE=E5=8C=BA=E6=8F=90?= =?UTF-8?q?=E4=BE=9B=E6=A8=A1=E5=9E=8B?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Model: AnkitAI/Mistral-Heretica-12B-GGUF Source: Original Platform --- .gitattributes | 39 +++++++++ Mistral-Heretica-12B-GGUF-Q4_K_M.gguf | 3 + Mistral-Heretica-12B-GGUF-Q5_K_M.gguf | 3 + Mistral-Heretica-12B-GGUF-Q6_K.gguf | 3 + Mistral-Heretica-12B-GGUF-Q8_0.gguf | 3 + README.md | 106 ++++++++++++++++++++++++ banner.svg | 113 ++++++++++++++++++++++++++ 7 files changed, 270 insertions(+) create mode 100644 .gitattributes create mode 100644 Mistral-Heretica-12B-GGUF-Q4_K_M.gguf create mode 100644 Mistral-Heretica-12B-GGUF-Q5_K_M.gguf create mode 100644 Mistral-Heretica-12B-GGUF-Q6_K.gguf create mode 100644 Mistral-Heretica-12B-GGUF-Q8_0.gguf create mode 100644 README.md create mode 100644 banner.svg diff --git a/.gitattributes b/.gitattributes new file mode 100644 index 0000000..ed753ad --- /dev/null +++ b/.gitattributes @@ -0,0 +1,39 @@ +*.7z filter=lfs diff=lfs merge=lfs -text +*.arrow filter=lfs diff=lfs merge=lfs -text +*.bin filter=lfs diff=lfs merge=lfs -text +*.bz2 filter=lfs diff=lfs merge=lfs -text +*.ckpt filter=lfs diff=lfs merge=lfs -text +*.ftz filter=lfs diff=lfs merge=lfs -text +*.gz filter=lfs diff=lfs merge=lfs -text +*.h5 filter=lfs diff=lfs merge=lfs -text +*.joblib filter=lfs diff=lfs merge=lfs -text +*.lfs.* filter=lfs diff=lfs merge=lfs -text +*.mlmodel filter=lfs diff=lfs merge=lfs -text +*.model filter=lfs diff=lfs merge=lfs -text +*.msgpack filter=lfs diff=lfs merge=lfs -text +*.npy filter=lfs diff=lfs merge=lfs -text +*.npz filter=lfs diff=lfs merge=lfs -text +*.onnx filter=lfs diff=lfs merge=lfs -text +*.ot filter=lfs diff=lfs merge=lfs -text +*.parquet filter=lfs diff=lfs merge=lfs -text +*.pb filter=lfs diff=lfs merge=lfs -text +*.pickle filter=lfs diff=lfs merge=lfs -text +*.pkl filter=lfs diff=lfs merge=lfs -text +*.pt filter=lfs diff=lfs merge=lfs -text +*.pth filter=lfs diff=lfs merge=lfs -text +*.rar filter=lfs diff=lfs merge=lfs -text +*.safetensors filter=lfs diff=lfs merge=lfs -text +saved_model/**/* filter=lfs diff=lfs merge=lfs -text +*.tar.* filter=lfs diff=lfs merge=lfs -text +*.tar filter=lfs diff=lfs merge=lfs -text +*.tflite filter=lfs diff=lfs merge=lfs -text +*.tgz filter=lfs diff=lfs merge=lfs -text +*.wasm filter=lfs diff=lfs merge=lfs -text +*.xz filter=lfs diff=lfs merge=lfs -text +*.zip filter=lfs diff=lfs merge=lfs -text +*.zst filter=lfs diff=lfs merge=lfs -text +*tfevents* filter=lfs diff=lfs merge=lfs -text +Mistral-Heretica-12B-GGUF-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-Heretica-12B-GGUF-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-Heretica-12B-GGUF-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-Heretica-12B-GGUF-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text diff --git a/Mistral-Heretica-12B-GGUF-Q4_K_M.gguf b/Mistral-Heretica-12B-GGUF-Q4_K_M.gguf new file mode 100644 index 0000000..13cc254 --- /dev/null +++ b/Mistral-Heretica-12B-GGUF-Q4_K_M.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:71ddaefd2890ba0fb8334aabead76161e6242e7702def3b778e13ccd0857b7cf +size 7477208288 diff --git a/Mistral-Heretica-12B-GGUF-Q5_K_M.gguf b/Mistral-Heretica-12B-GGUF-Q5_K_M.gguf new file mode 100644 index 0000000..51c9599 --- /dev/null +++ b/Mistral-Heretica-12B-GGUF-Q5_K_M.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:b924505842f6283f4713633d1851459c9e4afe8d81f003498fa837d65a42b1cc +size 8727635168 diff --git a/Mistral-Heretica-12B-GGUF-Q6_K.gguf b/Mistral-Heretica-12B-GGUF-Q6_K.gguf new file mode 100644 index 0000000..db43108 --- /dev/null +++ b/Mistral-Heretica-12B-GGUF-Q6_K.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:4716218454f2e668714bce8a13eb850e0ffd2ed644e8a5c0337d0a831b80953a +size 10056213728 diff --git a/Mistral-Heretica-12B-GGUF-Q8_0.gguf b/Mistral-Heretica-12B-GGUF-Q8_0.gguf new file mode 100644 index 0000000..2a61557 --- /dev/null +++ b/Mistral-Heretica-12B-GGUF-Q8_0.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:11c6dc6359ebb422595dfae095ddfa8314edf9a23fb2af294d044adda5b6bdf7 +size 13022373088 diff --git a/README.md b/README.md new file mode 100644 index 0000000..2f0b570 --- /dev/null +++ b/README.md @@ -0,0 +1,106 @@ +--- +base_model: mrcuddle/Mistral-Heretica-12B +base_model_relation: quantized +quantized_by: AnkitAI +library_name: gguf +pipeline_tag: text-generation +license: other +language: +- en +tags: +- gguf +- llama.cpp +- quantized +- mistral +- merge +- roleplay +- creative-writing +--- + +# Mistral-Heretica-12B-GGUF + +![Mistral Heretica 12B GGUF](banner.svg) + +GGUF quantizations of [mrcuddle/Mistral-Heretica-12B](https://huggingface.co/mrcuddle/Mistral-Heretica-12B). + +Original model: a [Task Arithmetic](https://arxiv.org/abs/2212.04089) merge of Mistral-Nemo-Instruct-2407, [Lumimaid-v0.2-12B](https://huggingface.co/NeverSleep/Lumimaid-v0.2-12B), and [absolute-heresy](https://huggingface.co/MuXodious/Mistral-Nemo-Instruct-2407-absolute-heresy) — tuned for uncensored roleplay and creative writing. 12B params, Mistral architecture. + +Quantized with [llama.cpp](https://github.com/ggml-org/llama.cpp) (release b9890). Run these in [LM Studio](https://lmstudio.ai/), [Ollama](https://ollama.com/), [KoboldCpp](https://github.com/LostRuins/koboldcpp), [SillyTavern](https://github.com/SillyTavern/SillyTavern), or `llama.cpp` directly. + +## Download + +| File | Quant | Size | Description | +|---|---|---|---| +| [Q8_0](./Mistral-Heretica-12B-GGUF-Q8_0.gguf) | Q8_0 | 12 GB | Maximum quality, near-lossless. Overkill for most. | +| [Q6_K](./Mistral-Heretica-12B-GGUF-Q6_K.gguf) | Q6_K | 9.4 GB | Very high quality, effectively lossless. | +| [Q5_K_M](./Mistral-Heretica-12B-GGUF-Q5_K_M.gguf) | Q5_K_M | 8.1 GB | High quality. Balanced pick. | +| [Q4_K_M](./Mistral-Heretica-12B-GGUF-Q4_K_M.gguf) | Q4_K_M | 7.0 GB | Good quality, best size/speed tradeoff. **Recommended default.** | + +F16 (23 GB) is the unquantized conversion — only needed if you want to re-quantize yourself. + +## Which file should I choose? + +Pick the largest quant that fits in your RAM/VRAM with room to spare (leave ~1–2 GB for context). + +- 8 GB RAM/VRAM → **Q4_K_M** +- 12 GB → **Q5_K_M** or **Q6_K** +- 16 GB+ → **Q6_K** or **Q8_0** + +For roleplay/creative use, Q4_K_M is plenty — quality loss over Q8 is negligible in practice. + +## Prompt format + +Mistral instruct format (primary): + +``` +[INST] {prompt} [/INST] +``` + +Most front-ends (SillyTavern, LM Studio) apply this automatically when you select a **Mistral / Mistral-Nemo** template. The Lumimaid component was also trained with ChatML, so ChatML works too — but Mistral format is the safe default. + +## Recommended settings + +Mistral-Nemo (this model's base) is **temperature-sensitive** — high temperature makes it incoherent. Keep it low: + +| Setting | Value | Note | +|---|---|---| +| Temperature | **0.5 – 0.7** | Do not exceed ~1.0. Lower = more coherent, higher = more creative. | +| Min-P | **0.05** | Primary truncation. Leave Top-P/Top-K off if using Min-P. | +| Repetition penalty | 1.05 – 1.1 | Or use DRY (0.8 / 1.75 / 2) if your front-end supports it. | +| Context | up to 16k reliable | Nemo's trained context is 128k but quality degrades well before that. | + +Start at temp 0.6 / min-p 0.05 and adjust from there. + +## Run it + +**llama.cpp:** +```bash +llama-cli -m Mistral-Heretica-12B-GGUF-Q4_K_M.gguf --jinja -p "Write a short noir opening." +``` + +**Download one file with the HF CLI (skip the rest):** +```bash +hf download AnkitAI/Mistral-Heretica-12B-GGUF \ + Mistral-Heretica-12B-GGUF-Q4_K_M.gguf --local-dir . +``` + +**Ollama:** +```bash +ollama run hf.co/AnkitAI/Mistral-Heretica-12B-GGUF:Q4_K_M +``` + +## Verified + +Q4_K_M smoke-tested: loads correctly, follows instructions (valid JSON output), and produces coherent creative prose. ~5–6 tok/s generation on Apple Silicon (M-series). + +## License + +The base model declares no explicit license. This merge inherits terms from its source models — notably **Lumimaid-v0.2-12B, which is CC-BY-NC-4.0 (non-commercial)**. Treat this model as **non-commercial** unless you verify otherwise with the original authors. These are quantizations only; all model rights belong to the original creators. + +## Credits + +- Original model: [mrcuddle](https://huggingface.co/mrcuddle) +- Merge components: [NeverSleep](https://huggingface.co/NeverSleep), [MuXodious](https://huggingface.co/MuXodious), [Mistral AI](https://huggingface.co/mistralai) +- Quantization tooling: [llama.cpp](https://github.com/ggml-org/llama.cpp) + +More models: [ankitaglawe.com](https://ankitaglawe.com) diff --git a/banner.svg b/banner.svg new file mode 100644 index 0000000..d2c0ca4 --- /dev/null +++ b/banner.svg @@ -0,0 +1,113 @@ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +