76 lines
4.0 KiB
Markdown
76 lines
4.0 KiB
Markdown
|
|
---
|
|||
|
|
license: apache-2.0
|
|||
|
|
base_model: prism-ml/Ternary-Bonsai-8B
|
|||
|
|
tags:
|
|||
|
|
- gguf
|
|||
|
|
- rdna4
|
|||
|
|
- gfx1201
|
|||
|
|
- ternary
|
|||
|
|
- speculative-decoding
|
|||
|
|
- the-rock8
|
|||
|
|
- llama.cpp
|
|||
|
|
library_name: gguf
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# Bonsai-8B Ternary — RDNA4 (The Rock8) 🦆
|
|||
|
|
|
|||
|
|
**RDNA4 (gfx1201) GGUF builds of [prism-ml/Ternary-Bonsai-8B](https://huggingface.co/prism-ml/Ternary-Bonsai-8B)** — a dense Qwen3-8B trained natively ternary (QAT, 1.58-bit). Part of the **[The Rock8 — RDNA4 fp8](https://huggingface.co/collections/Gorilla4X/the-rock8-rdna4-fp8-6a547070f667cb41db0bc2ed)** collection.
|
|||
|
|
|
|||
|
|
This repo is the **async speculative-decoding showcase** for The Rock8 — the config that hits **+66% decode on a dual-R9700 box, byte-identical output.**
|
|||
|
|
|
|||
|
|
## Files
|
|||
|
|
|
|||
|
|
| File | Size | Role |
|
|||
|
|
|---|---|---|
|
|||
|
|
| `Ternary-Bonsai-8B-F16.gguf` | 16.4 GB | **Verify target** — ternary weights in F16 storage, runs on RDNA4 WMMA |
|
|||
|
|
| `Ternary-Bonsai-8B-Q2_0.gguf` | 2.18 GB | **The ternary self-draft** — the cheap drafter that makes async win |
|
|||
|
|
| `Ternary-Bonsai-8B-F8E4M3.gguf` | 8.6 GB | Native RDNA4 **fp8 (E4M3)** build |
|
|||
|
|
|
|||
|
|
## The async spec-decode win (measured on gfx1201, dual R9700)
|
|||
|
|
|
|||
|
|
Using the **ternary Q2 draft** to speculate for the **F16 target**, with The Rock8's async pipeline (`LLAMA_SPEC_ASYNC=2`) — draft-gen on GPU1 **‖** verify on GPU0:
|
|||
|
|
|
|||
|
|
| Config | Decode t/s | Accept | Output |
|
|||
|
|
|---|---|---|---|
|
|||
|
|
| Sequential, 1-GPU | 63.30 ± 0.53 | 100% | baseline |
|
|||
|
|
| **Async pipeline, 2-GPU** | **105.08 ± 0.70** | **100%** | **byte-identical ✅** |
|
|||
|
|
|
|||
|
|
**+66% decode, zero quality loss** (re-validated 2026-07-13 on the TheRock ROCm 7.13 build; peaks ~111 t/s / +75% on an unloaded box).
|
|||
|
|
|
|||
|
|
### Why it works — "intel per byte"
|
|||
|
|
A natively-ternary model is **its own near-lossless, cheap draft**. The Q2 draft runs on the **VALU/mmvq** path while the F16 target verifies on the **WMMA tensor cores** — *different execution units*, so draft-gen and verify genuinely overlap instead of fighting for the same silicon. That disjoint-compute overlap is the whole trick, and it's why async pays off here but not on a same-precision self-draft (draft costs as much as verify → nothing to hide).
|
|||
|
|
|
|||
|
|
> Note: the async pipeline is a **2-GPU** lever and needs a **dense** target (plain-attention KV supports the pipeline's partial-rollback). Hybrid SSM/GatedDeltaNet targets need a core-level rollback fix — see the collection notes.
|
|||
|
|
|
|||
|
|
### Not Bonsai-specific — async wins on off-the-shelf models too
|
|||
|
|
The pipeline is a general lever: it wins whenever the **verify is heavy and the draft is cheap**. Same trick, vanilla Qwen3-8B (byte-identical output in every case):
|
|||
|
|
|
|||
|
|
| Target | Draft | Async Δ |
|
|||
|
|
|---|---|---|
|
|||
|
|
| Qwen3-8B fp8 (light) | Q4_K_M | −9% (too light — loses) |
|
|||
|
|
| Qwen3-8B **BF16** (heavy) | Q4_K_M | **+14.5%** |
|
|||
|
|
| Qwen3-8B **BF16** (heavy) | Q2_K | **+21.9%** |
|
|||
|
|
| **Bonsai F16** (heavy) | **ternary Q2** (VALU) | **+66%** |
|
|||
|
|
|
|||
|
|
Swapping the *identical* Q4 draft from an fp8 target to a BF16 target flips −9% → +14.5% — **the target's verify-weight is the deciding lever**, not the draft. Bonsai tops the table because its ternary draft runs *entirely off* the WMMA units (max overlap); a heavier target × a cheaper draft ⇒ a bigger win.
|
|||
|
|
|
|||
|
|
## Usage (The Rock8 fork)
|
|||
|
|
|
|||
|
|
```bash
|
|||
|
|
# async 2-GPU: ternary Q2 draft ‖ F16 verify
|
|||
|
|
LLAMA_SPEC_ASYNC=2 ./llama-speculative-simple \
|
|||
|
|
-m Ternary-Bonsai-8B-F16.gguf -dev ROCm0 \
|
|||
|
|
-md Ternary-Bonsai-8B-Q2_0.gguf -devd ROCm1 \
|
|||
|
|
--spec-type draft-simple --spec-draft-n-max 4 \
|
|||
|
|
-c 2048 --temp 0 -n 130 \
|
|||
|
|
-p "What do you call a dried grape? Answer in one word."
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Build: [github.com/The-Monk/The-Rock8](https://github.com/The-Monk/The-Rock8) (RDNA4 native-fp8 llama.cpp fork + Podman appliance).
|
|||
|
|
|
|||
|
|
## Attribution & license
|
|||
|
|
|
|||
|
|
Base model **[prism-ml/Ternary-Bonsai-8B](https://huggingface.co/prism-ml/Ternary-Bonsai-8B)** by PrismML, **Apache-2.0**. These are GGUF conversions/quantizations for RDNA4; all credit for the model and its ternary QAT training to PrismML. Distributed under the same Apache-2.0 license.
|
|||
|
|
|
|||
|
|
🦆 *Got any weights?*
|