From 7afbe4773812ea433a7bc4024f81e1cde5c19d56 Mon Sep 17 00:00:00 2001 From: ModelHub XC Date: Fri, 18 Sep 2026 07:58:16 +0800 Subject: [PATCH] =?UTF-8?q?=E5=88=9D=E5=A7=8B=E5=8C=96=E9=A1=B9=E7=9B=AE?= =?UTF-8?q?=EF=BC=8C=E7=94=B1ModelHub=20XC=E7=A4=BE=E5=8C=BA=E6=8F=90?= =?UTF-8?q?=E4=BE=9B=E6=A8=A1=E5=9E=8B?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Model: Gorilla4X/Bonsai-8B-Ternary-RDNA4 Source: Original Platform --- .gitattributes | 38 ++++++++++++++++++ README.md | 75 +++++++++++++++++++++++++++++++++++ Ternary-Bonsai-8B-F16.gguf | 3 ++ Ternary-Bonsai-8B-F8E4M3.gguf | 3 ++ Ternary-Bonsai-8B-Q2_0.gguf | 3 ++ 5 files changed, 122 insertions(+) create mode 100644 .gitattributes create mode 100644 README.md create mode 100644 Ternary-Bonsai-8B-F16.gguf create mode 100644 Ternary-Bonsai-8B-F8E4M3.gguf create mode 100644 Ternary-Bonsai-8B-Q2_0.gguf diff --git a/.gitattributes b/.gitattributes new file mode 100644 index 0000000..aafd32f --- /dev/null +++ b/.gitattributes @@ -0,0 +1,38 @@ +*.7z filter=lfs diff=lfs merge=lfs -text +*.arrow filter=lfs diff=lfs merge=lfs -text +*.bin filter=lfs diff=lfs merge=lfs -text +*.bz2 filter=lfs diff=lfs merge=lfs -text +*.ckpt filter=lfs diff=lfs merge=lfs -text +*.ftz filter=lfs diff=lfs merge=lfs -text +*.gz filter=lfs diff=lfs merge=lfs -text +*.h5 filter=lfs diff=lfs merge=lfs -text +*.joblib filter=lfs diff=lfs merge=lfs -text +*.lfs.* filter=lfs diff=lfs merge=lfs -text +*.mlmodel filter=lfs diff=lfs merge=lfs -text +*.model filter=lfs diff=lfs merge=lfs -text +*.msgpack filter=lfs diff=lfs merge=lfs -text +*.npy filter=lfs diff=lfs merge=lfs -text +*.npz filter=lfs diff=lfs merge=lfs -text +*.onnx filter=lfs diff=lfs merge=lfs -text +*.ot filter=lfs diff=lfs merge=lfs -text +*.parquet filter=lfs diff=lfs merge=lfs -text +*.pb filter=lfs diff=lfs merge=lfs -text +*.pickle filter=lfs diff=lfs merge=lfs -text +*.pkl filter=lfs diff=lfs merge=lfs -text +*.pt filter=lfs diff=lfs merge=lfs -text +*.pth filter=lfs diff=lfs merge=lfs -text +*.rar filter=lfs diff=lfs merge=lfs -text +*.safetensors filter=lfs diff=lfs merge=lfs -text +saved_model/**/* filter=lfs diff=lfs merge=lfs -text +*.tar.* filter=lfs diff=lfs merge=lfs -text +*.tar filter=lfs diff=lfs merge=lfs -text +*.tflite filter=lfs diff=lfs merge=lfs -text +*.tgz filter=lfs diff=lfs merge=lfs -text +*.wasm filter=lfs diff=lfs merge=lfs -text +*.xz filter=lfs diff=lfs merge=lfs -text +*.zip filter=lfs diff=lfs merge=lfs -text +*.zst filter=lfs diff=lfs merge=lfs -text +*tfevents* filter=lfs diff=lfs merge=lfs -text +Ternary-Bonsai-8B-Q2_0.gguf filter=lfs diff=lfs merge=lfs -text +Ternary-Bonsai-8B-F8E4M3.gguf filter=lfs diff=lfs merge=lfs -text +Ternary-Bonsai-8B-F16.gguf filter=lfs diff=lfs merge=lfs -text diff --git a/README.md b/README.md new file mode 100644 index 0000000..e3c92fa --- /dev/null +++ b/README.md @@ -0,0 +1,75 @@ +--- +license: apache-2.0 +base_model: prism-ml/Ternary-Bonsai-8B +tags: +- gguf +- rdna4 +- gfx1201 +- ternary +- speculative-decoding +- the-rock8 +- llama.cpp +library_name: gguf +--- + +# Bonsai-8B Ternary — RDNA4 (The Rock8) 🦆 + +**RDNA4 (gfx1201) GGUF builds of [prism-ml/Ternary-Bonsai-8B](https://huggingface.co/prism-ml/Ternary-Bonsai-8B)** — a dense Qwen3-8B trained natively ternary (QAT, 1.58-bit). Part of the **[The Rock8 — RDNA4 fp8](https://huggingface.co/collections/Gorilla4X/the-rock8-rdna4-fp8-6a547070f667cb41db0bc2ed)** collection. + +This repo is the **async speculative-decoding showcase** for The Rock8 — the config that hits **+66% decode on a dual-R9700 box, byte-identical output.** + +## Files + +| File | Size | Role | +|---|---|---| +| `Ternary-Bonsai-8B-F16.gguf` | 16.4 GB | **Verify target** — ternary weights in F16 storage, runs on RDNA4 WMMA | +| `Ternary-Bonsai-8B-Q2_0.gguf` | 2.18 GB | **The ternary self-draft** — the cheap drafter that makes async win | +| `Ternary-Bonsai-8B-F8E4M3.gguf` | 8.6 GB | Native RDNA4 **fp8 (E4M3)** build | + +## The async spec-decode win (measured on gfx1201, dual R9700) + +Using the **ternary Q2 draft** to speculate for the **F16 target**, with The Rock8's async pipeline (`LLAMA_SPEC_ASYNC=2`) — draft-gen on GPU1 **‖** verify on GPU0: + +| Config | Decode t/s | Accept | Output | +|---|---|---|---| +| Sequential, 1-GPU | 63.30 ± 0.53 | 100% | baseline | +| **Async pipeline, 2-GPU** | **105.08 ± 0.70** | **100%** | **byte-identical ✅** | + +**+66% decode, zero quality loss** (re-validated 2026-07-13 on the TheRock ROCm 7.13 build; peaks ~111 t/s / +75% on an unloaded box). + +### Why it works — "intel per byte" +A natively-ternary model is **its own near-lossless, cheap draft**. The Q2 draft runs on the **VALU/mmvq** path while the F16 target verifies on the **WMMA tensor cores** — *different execution units*, so draft-gen and verify genuinely overlap instead of fighting for the same silicon. That disjoint-compute overlap is the whole trick, and it's why async pays off here but not on a same-precision self-draft (draft costs as much as verify → nothing to hide). + +> Note: the async pipeline is a **2-GPU** lever and needs a **dense** target (plain-attention KV supports the pipeline's partial-rollback). Hybrid SSM/GatedDeltaNet targets need a core-level rollback fix — see the collection notes. + +### Not Bonsai-specific — async wins on off-the-shelf models too +The pipeline is a general lever: it wins whenever the **verify is heavy and the draft is cheap**. Same trick, vanilla Qwen3-8B (byte-identical output in every case): + +| Target | Draft | Async Δ | +|---|---|---| +| Qwen3-8B fp8 (light) | Q4_K_M | −9% (too light — loses) | +| Qwen3-8B **BF16** (heavy) | Q4_K_M | **+14.5%** | +| Qwen3-8B **BF16** (heavy) | Q2_K | **+21.9%** | +| **Bonsai F16** (heavy) | **ternary Q2** (VALU) | **+66%** | + +Swapping the *identical* Q4 draft from an fp8 target to a BF16 target flips −9% → +14.5% — **the target's verify-weight is the deciding lever**, not the draft. Bonsai tops the table because its ternary draft runs *entirely off* the WMMA units (max overlap); a heavier target × a cheaper draft ⇒ a bigger win. + +## Usage (The Rock8 fork) + +```bash +# async 2-GPU: ternary Q2 draft ‖ F16 verify +LLAMA_SPEC_ASYNC=2 ./llama-speculative-simple \ + -m Ternary-Bonsai-8B-F16.gguf -dev ROCm0 \ + -md Ternary-Bonsai-8B-Q2_0.gguf -devd ROCm1 \ + --spec-type draft-simple --spec-draft-n-max 4 \ + -c 2048 --temp 0 -n 130 \ + -p "What do you call a dried grape? Answer in one word." +``` + +Build: [github.com/The-Monk/The-Rock8](https://github.com/The-Monk/The-Rock8) (RDNA4 native-fp8 llama.cpp fork + Podman appliance). + +## Attribution & license + +Base model **[prism-ml/Ternary-Bonsai-8B](https://huggingface.co/prism-ml/Ternary-Bonsai-8B)** by PrismML, **Apache-2.0**. These are GGUF conversions/quantizations for RDNA4; all credit for the model and its ternary QAT training to PrismML. Distributed under the same Apache-2.0 license. + +🦆 *Got any weights?* diff --git a/Ternary-Bonsai-8B-F16.gguf b/Ternary-Bonsai-8B-F16.gguf new file mode 100644 index 0000000..8d4e035 --- /dev/null +++ b/Ternary-Bonsai-8B-F16.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:a6abfaf896c1e36db825112fc0a18e49adea05eeca1c6b2fba4d785ca7e947ff +size 16383663200 diff --git a/Ternary-Bonsai-8B-F8E4M3.gguf b/Ternary-Bonsai-8B-F8E4M3.gguf new file mode 100644 index 0000000..5bc1208 --- /dev/null +++ b/Ternary-Bonsai-8B-F8E4M3.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:d65c37c9a150afa5b8501d6201d324c92285d67bf9b1d194706e868d694e14b6 +size 8556732672 diff --git a/Ternary-Bonsai-8B-Q2_0.gguf b/Ternary-Bonsai-8B-Q2_0.gguf new file mode 100644 index 0000000..434453c --- /dev/null +++ b/Ternary-Bonsai-8B-Q2_0.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:3c8d70470a5d97e5a2b9410ddd899cb740116591462626c60cb2fead6448f60b +size 2182184672