Model: Lambent/Arsenic-Shahrazad-12B-v4.4.1-mtp Source: Original Platform
base_model, library_name, tags, license
| base_model | library_name | tags | license | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Lambent/Arsenic-Shahrazad-12B-v4.4 | transformers |
|
cc-by-nc-4.0 |
Fork of llama.cpp to run the MTP head: https://github.com/lambentLogic/llama.cpp/tree/mtp-llama-mistral
Arsenic-Shahrazad-12B-v4.4.1-mtp
A co-adapted variant of Arsenic-Shahrazad-12B-v4.4, trained alongside a Multi-Token-Prediction (MTP) draft head.
The draft head was initially trained with the main model frozen, and then trained alongside the main model.
This does increase speed slightly, especially at lower temperatures, but is mostly an exploration - likely not as effective as Qwen's yet.
Usable as an ordinary chat model on its own; pair it with the MTP head GGUF for the speedup.
What "co-adapted" means here
v4.4.1 was produced by training the base model (QLoRA, very low LR) and an MTP draft head together on her own temperature-1 self-generated samples — spanning creative writing, conversation, explanation, debate, poetry/constrained form, translation, and research, plus persona-conditioned variations for breadth.
Speedup (llama.cpp MTP speculative decoding)
Draft acceptance and throughput vs. the initial v4.4 frozen-trunk MTP head, both Q5_K_M trunk + Q8 (T=0.4-baked) head, n_max=1, on in-distribution prompts:
| temp | v4.4.1 accept | v4.4 head accept | v4.4.1 tok/s | v4.4 tok/s |
|---|---|---|---|---|
| 0.3 | 0.418 | 0.373 | 81.5 | 75.5 |
| 0.5 | 0.393 | 0.360 | 75.6 | 73.9 |
| 0.7 | 0.416 | 0.371 | 76.8 | 74.1 |
≈ +10% relative draft acceptance across the temperature range.
Edit: After upstream improvements to MTP merged in, the tok/s have improved (calculated with the Q5_K_M model and Q8_0 head):
| temp | draft accept | tok/s (MTP) | tok/s (no MTP) | speedup |
|---|---|---|---|---|
| 0.3 | 0.43 | 83.8 | 76.0 | +10% |
| 0.5 | 0.42 | 82.3 | 74.6 | +10% |
| 0.7 | 0.44 | 81.9 | 73.3 | +12% |
Usage
With MTP acceleration (llama.cpp)
Requires the MTP draft-head GGUF and a llama.cpp build with draft-mtp support, not yet implemented:
llama-server -m Arsenic-Shahrazad-12B-v4.4.1-Q5_K_M.gguf \
--spec-type draft-mtp -md mtp-Arsenic-v4.4.1-T04-Q8_0.gguf \
--spec-draft-n-max 1 --spec-draft-p-min 0.0 \
-ngl 99 -c 8192