Files
Bonsai-8B-Ternary-RDNA4/README.md
ModelHub XC 7afbe47738 初始化项目,由ModelHub XC社区提供模型
Model: Gorilla4X/Bonsai-8B-Ternary-RDNA4
Source: Original Platform
2026-09-18 07:58:16 +08:00

76 lines
4.0 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
license: apache-2.0
base_model: prism-ml/Ternary-Bonsai-8B
tags:
- gguf
- rdna4
- gfx1201
- ternary
- speculative-decoding
- the-rock8
- llama.cpp
library_name: gguf
---
# Bonsai-8B Ternary — RDNA4 (The Rock8) 🦆
**RDNA4 (gfx1201) GGUF builds of [prism-ml/Ternary-Bonsai-8B](https://huggingface.co/prism-ml/Ternary-Bonsai-8B)** — a dense Qwen3-8B trained natively ternary (QAT, 1.58-bit). Part of the **[The Rock8 — RDNA4 fp8](https://huggingface.co/collections/Gorilla4X/the-rock8-rdna4-fp8-6a547070f667cb41db0bc2ed)** collection.
This repo is the **async speculative-decoding showcase** for The Rock8 — the config that hits **+66% decode on a dual-R9700 box, byte-identical output.**
## Files
| File | Size | Role |
|---|---|---|
| `Ternary-Bonsai-8B-F16.gguf` | 16.4 GB | **Verify target** — ternary weights in F16 storage, runs on RDNA4 WMMA |
| `Ternary-Bonsai-8B-Q2_0.gguf` | 2.18 GB | **The ternary self-draft** — the cheap drafter that makes async win |
| `Ternary-Bonsai-8B-F8E4M3.gguf` | 8.6 GB | Native RDNA4 **fp8 (E4M3)** build |
## The async spec-decode win (measured on gfx1201, dual R9700)
Using the **ternary Q2 draft** to speculate for the **F16 target**, with The Rock8's async pipeline (`LLAMA_SPEC_ASYNC=2`) — draft-gen on GPU1 **‖** verify on GPU0:
| Config | Decode t/s | Accept | Output |
|---|---|---|---|
| Sequential, 1-GPU | 63.30 ± 0.53 | 100% | baseline |
| **Async pipeline, 2-GPU** | **105.08 ± 0.70** | **100%** | **byte-identical ✅** |
**+66% decode, zero quality loss** (re-validated 2026-07-13 on the TheRock ROCm 7.13 build; peaks ~111 t/s / +75% on an unloaded box).
### Why it works — "intel per byte"
A natively-ternary model is **its own near-lossless, cheap draft**. The Q2 draft runs on the **VALU/mmvq** path while the F16 target verifies on the **WMMA tensor cores** — *different execution units*, so draft-gen and verify genuinely overlap instead of fighting for the same silicon. That disjoint-compute overlap is the whole trick, and it's why async pays off here but not on a same-precision self-draft (draft costs as much as verify → nothing to hide).
> Note: the async pipeline is a **2-GPU** lever and needs a **dense** target (plain-attention KV supports the pipeline's partial-rollback). Hybrid SSM/GatedDeltaNet targets need a core-level rollback fix — see the collection notes.
### Not Bonsai-specific — async wins on off-the-shelf models too
The pipeline is a general lever: it wins whenever the **verify is heavy and the draft is cheap**. Same trick, vanilla Qwen3-8B (byte-identical output in every case):
| Target | Draft | Async Δ |
|---|---|---|
| Qwen3-8B fp8 (light) | Q4_K_M | −9% (too light — loses) |
| Qwen3-8B **BF16** (heavy) | Q4_K_M | **+14.5%** |
| Qwen3-8B **BF16** (heavy) | Q2_K | **+21.9%** |
| **Bonsai F16** (heavy) | **ternary Q2** (VALU) | **+66%** |
Swapping the *identical* Q4 draft from an fp8 target to a BF16 target flips −9% → +14.5% — **the target's verify-weight is the deciding lever**, not the draft. Bonsai tops the table because its ternary draft runs *entirely off* the WMMA units (max overlap); a heavier target × a cheaper draft ⇒ a bigger win.
## Usage (The Rock8 fork)
```bash
# async 2-GPU: ternary Q2 draft ‖ F16 verify
LLAMA_SPEC_ASYNC=2 ./llama-speculative-simple \
-m Ternary-Bonsai-8B-F16.gguf -dev ROCm0 \
-md Ternary-Bonsai-8B-Q2_0.gguf -devd ROCm1 \
--spec-type draft-simple --spec-draft-n-max 4 \
-c 2048 --temp 0 -n 130 \
-p "What do you call a dried grape? Answer in one word."
```
Build: [github.com/The-Monk/The-Rock8](https://github.com/The-Monk/The-Rock8) (RDNA4 native-fp8 llama.cpp fork + Podman appliance).
## Attribution & license
Base model **[prism-ml/Ternary-Bonsai-8B](https://huggingface.co/prism-ml/Ternary-Bonsai-8B)** by PrismML, **Apache-2.0**. These are GGUF conversions/quantizations for RDNA4; all credit for the model and its ternary QAT training to PrismML. Distributed under the same Apache-2.0 license.
🦆 *Got any weights?*