--- license: apache-2.0 base_model: prism-ml/Ternary-Bonsai-8B tags: - gguf - rdna4 - gfx1201 - ternary - speculative-decoding - the-rock8 - llama.cpp library_name: gguf --- # Bonsai-8B Ternary — RDNA4 (The Rock8) 🦆 **RDNA4 (gfx1201) GGUF builds of [prism-ml/Ternary-Bonsai-8B](https://huggingface.co/prism-ml/Ternary-Bonsai-8B)** — a dense Qwen3-8B trained natively ternary (QAT, 1.58-bit). Part of the **[The Rock8 — RDNA4 fp8](https://huggingface.co/collections/Gorilla4X/the-rock8-rdna4-fp8-6a547070f667cb41db0bc2ed)** collection. This repo is the **async speculative-decoding showcase** for The Rock8 — the config that hits **+66% decode on a dual-R9700 box, byte-identical output.** ## Files | File | Size | Role | |---|---|---| | `Ternary-Bonsai-8B-F16.gguf` | 16.4 GB | **Verify target** — ternary weights in F16 storage, runs on RDNA4 WMMA | | `Ternary-Bonsai-8B-Q2_0.gguf` | 2.18 GB | **The ternary self-draft** — the cheap drafter that makes async win | | `Ternary-Bonsai-8B-F8E4M3.gguf` | 8.6 GB | Native RDNA4 **fp8 (E4M3)** build | ## The async spec-decode win (measured on gfx1201, dual R9700) Using the **ternary Q2 draft** to speculate for the **F16 target**, with The Rock8's async pipeline (`LLAMA_SPEC_ASYNC=2`) — draft-gen on GPU1 **‖** verify on GPU0: | Config | Decode t/s | Accept | Output | |---|---|---|---| | Sequential, 1-GPU | 63.30 ± 0.53 | 100% | baseline | | **Async pipeline, 2-GPU** | **105.08 ± 0.70** | **100%** | **byte-identical ✅** | **+66% decode, zero quality loss** (re-validated 2026-07-13 on the TheRock ROCm 7.13 build; peaks ~111 t/s / +75% on an unloaded box). ### Why it works — "intel per byte" A natively-ternary model is **its own near-lossless, cheap draft**. The Q2 draft runs on the **VALU/mmvq** path while the F16 target verifies on the **WMMA tensor cores** — *different execution units*, so draft-gen and verify genuinely overlap instead of fighting for the same silicon. That disjoint-compute overlap is the whole trick, and it's why async pays off here but not on a same-precision self-draft (draft costs as much as verify → nothing to hide). > Note: the async pipeline is a **2-GPU** lever and needs a **dense** target (plain-attention KV supports the pipeline's partial-rollback). Hybrid SSM/GatedDeltaNet targets need a core-level rollback fix — see the collection notes. ### Not Bonsai-specific — async wins on off-the-shelf models too The pipeline is a general lever: it wins whenever the **verify is heavy and the draft is cheap**. Same trick, vanilla Qwen3-8B (byte-identical output in every case): | Target | Draft | Async Δ | |---|---|---| | Qwen3-8B fp8 (light) | Q4_K_M | −9% (too light — loses) | | Qwen3-8B **BF16** (heavy) | Q4_K_M | **+14.5%** | | Qwen3-8B **BF16** (heavy) | Q2_K | **+21.9%** | | **Bonsai F16** (heavy) | **ternary Q2** (VALU) | **+66%** | Swapping the *identical* Q4 draft from an fp8 target to a BF16 target flips −9% → +14.5% — **the target's verify-weight is the deciding lever**, not the draft. Bonsai tops the table because its ternary draft runs *entirely off* the WMMA units (max overlap); a heavier target × a cheaper draft ⇒ a bigger win. ## Usage (The Rock8 fork) ```bash # async 2-GPU: ternary Q2 draft ‖ F16 verify LLAMA_SPEC_ASYNC=2 ./llama-speculative-simple \ -m Ternary-Bonsai-8B-F16.gguf -dev ROCm0 \ -md Ternary-Bonsai-8B-Q2_0.gguf -devd ROCm1 \ --spec-type draft-simple --spec-draft-n-max 4 \ -c 2048 --temp 0 -n 130 \ -p "What do you call a dried grape? Answer in one word." ``` Build: [github.com/The-Monk/The-Rock8](https://github.com/The-Monk/The-Rock8) (RDNA4 native-fp8 llama.cpp fork + Podman appliance). ## Attribution & license Base model **[prism-ml/Ternary-Bonsai-8B](https://huggingface.co/prism-ml/Ternary-Bonsai-8B)** by PrismML, **Apache-2.0**. These are GGUF conversions/quantizations for RDNA4; all credit for the model and its ternary QAT training to PrismML. Distributed under the same Apache-2.0 license. 🦆 *Got any weights?*