初始化项目,由ModelHub XC社区提供模型
Model: Gorilla4X/Bonsai-8B-Ternary-RDNA4 Source: Original Platform
This commit is contained in:
38
.gitattributes
vendored
Normal file
38
.gitattributes
vendored
Normal file
@@ -0,0 +1,38 @@
|
||||
*.7z filter=lfs diff=lfs merge=lfs -text
|
||||
*.arrow filter=lfs diff=lfs merge=lfs -text
|
||||
*.bin filter=lfs diff=lfs merge=lfs -text
|
||||
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
||||
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
||||
*.ftz filter=lfs diff=lfs merge=lfs -text
|
||||
*.gz filter=lfs diff=lfs merge=lfs -text
|
||||
*.h5 filter=lfs diff=lfs merge=lfs -text
|
||||
*.joblib filter=lfs diff=lfs merge=lfs -text
|
||||
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
||||
*.model filter=lfs diff=lfs merge=lfs -text
|
||||
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
||||
*.npy filter=lfs diff=lfs merge=lfs -text
|
||||
*.npz filter=lfs diff=lfs merge=lfs -text
|
||||
*.onnx filter=lfs diff=lfs merge=lfs -text
|
||||
*.ot filter=lfs diff=lfs merge=lfs -text
|
||||
*.parquet filter=lfs diff=lfs merge=lfs -text
|
||||
*.pb filter=lfs diff=lfs merge=lfs -text
|
||||
*.pickle filter=lfs diff=lfs merge=lfs -text
|
||||
*.pkl filter=lfs diff=lfs merge=lfs -text
|
||||
*.pt filter=lfs diff=lfs merge=lfs -text
|
||||
*.pth filter=lfs diff=lfs merge=lfs -text
|
||||
*.rar filter=lfs diff=lfs merge=lfs -text
|
||||
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
||||
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar filter=lfs diff=lfs merge=lfs -text
|
||||
*.tflite filter=lfs diff=lfs merge=lfs -text
|
||||
*.tgz filter=lfs diff=lfs merge=lfs -text
|
||||
*.wasm filter=lfs diff=lfs merge=lfs -text
|
||||
*.xz filter=lfs diff=lfs merge=lfs -text
|
||||
*.zip filter=lfs diff=lfs merge=lfs -text
|
||||
*.zst filter=lfs diff=lfs merge=lfs -text
|
||||
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
||||
Ternary-Bonsai-8B-Q2_0.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Ternary-Bonsai-8B-F8E4M3.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Ternary-Bonsai-8B-F16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
75
README.md
Normal file
75
README.md
Normal file
@@ -0,0 +1,75 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
base_model: prism-ml/Ternary-Bonsai-8B
|
||||
tags:
|
||||
- gguf
|
||||
- rdna4
|
||||
- gfx1201
|
||||
- ternary
|
||||
- speculative-decoding
|
||||
- the-rock8
|
||||
- llama.cpp
|
||||
library_name: gguf
|
||||
---
|
||||
|
||||
# Bonsai-8B Ternary — RDNA4 (The Rock8) 🦆
|
||||
|
||||
**RDNA4 (gfx1201) GGUF builds of [prism-ml/Ternary-Bonsai-8B](https://huggingface.co/prism-ml/Ternary-Bonsai-8B)** — a dense Qwen3-8B trained natively ternary (QAT, 1.58-bit). Part of the **[The Rock8 — RDNA4 fp8](https://huggingface.co/collections/Gorilla4X/the-rock8-rdna4-fp8-6a547070f667cb41db0bc2ed)** collection.
|
||||
|
||||
This repo is the **async speculative-decoding showcase** for The Rock8 — the config that hits **+66% decode on a dual-R9700 box, byte-identical output.**
|
||||
|
||||
## Files
|
||||
|
||||
| File | Size | Role |
|
||||
|---|---|---|
|
||||
| `Ternary-Bonsai-8B-F16.gguf` | 16.4 GB | **Verify target** — ternary weights in F16 storage, runs on RDNA4 WMMA |
|
||||
| `Ternary-Bonsai-8B-Q2_0.gguf` | 2.18 GB | **The ternary self-draft** — the cheap drafter that makes async win |
|
||||
| `Ternary-Bonsai-8B-F8E4M3.gguf` | 8.6 GB | Native RDNA4 **fp8 (E4M3)** build |
|
||||
|
||||
## The async spec-decode win (measured on gfx1201, dual R9700)
|
||||
|
||||
Using the **ternary Q2 draft** to speculate for the **F16 target**, with The Rock8's async pipeline (`LLAMA_SPEC_ASYNC=2`) — draft-gen on GPU1 **‖** verify on GPU0:
|
||||
|
||||
| Config | Decode t/s | Accept | Output |
|
||||
|---|---|---|---|
|
||||
| Sequential, 1-GPU | 63.30 ± 0.53 | 100% | baseline |
|
||||
| **Async pipeline, 2-GPU** | **105.08 ± 0.70** | **100%** | **byte-identical ✅** |
|
||||
|
||||
**+66% decode, zero quality loss** (re-validated 2026-07-13 on the TheRock ROCm 7.13 build; peaks ~111 t/s / +75% on an unloaded box).
|
||||
|
||||
### Why it works — "intel per byte"
|
||||
A natively-ternary model is **its own near-lossless, cheap draft**. The Q2 draft runs on the **VALU/mmvq** path while the F16 target verifies on the **WMMA tensor cores** — *different execution units*, so draft-gen and verify genuinely overlap instead of fighting for the same silicon. That disjoint-compute overlap is the whole trick, and it's why async pays off here but not on a same-precision self-draft (draft costs as much as verify → nothing to hide).
|
||||
|
||||
> Note: the async pipeline is a **2-GPU** lever and needs a **dense** target (plain-attention KV supports the pipeline's partial-rollback). Hybrid SSM/GatedDeltaNet targets need a core-level rollback fix — see the collection notes.
|
||||
|
||||
### Not Bonsai-specific — async wins on off-the-shelf models too
|
||||
The pipeline is a general lever: it wins whenever the **verify is heavy and the draft is cheap**. Same trick, vanilla Qwen3-8B (byte-identical output in every case):
|
||||
|
||||
| Target | Draft | Async Δ |
|
||||
|---|---|---|
|
||||
| Qwen3-8B fp8 (light) | Q4_K_M | −9% (too light — loses) |
|
||||
| Qwen3-8B **BF16** (heavy) | Q4_K_M | **+14.5%** |
|
||||
| Qwen3-8B **BF16** (heavy) | Q2_K | **+21.9%** |
|
||||
| **Bonsai F16** (heavy) | **ternary Q2** (VALU) | **+66%** |
|
||||
|
||||
Swapping the *identical* Q4 draft from an fp8 target to a BF16 target flips −9% → +14.5% — **the target's verify-weight is the deciding lever**, not the draft. Bonsai tops the table because its ternary draft runs *entirely off* the WMMA units (max overlap); a heavier target × a cheaper draft ⇒ a bigger win.
|
||||
|
||||
## Usage (The Rock8 fork)
|
||||
|
||||
```bash
|
||||
# async 2-GPU: ternary Q2 draft ‖ F16 verify
|
||||
LLAMA_SPEC_ASYNC=2 ./llama-speculative-simple \
|
||||
-m Ternary-Bonsai-8B-F16.gguf -dev ROCm0 \
|
||||
-md Ternary-Bonsai-8B-Q2_0.gguf -devd ROCm1 \
|
||||
--spec-type draft-simple --spec-draft-n-max 4 \
|
||||
-c 2048 --temp 0 -n 130 \
|
||||
-p "What do you call a dried grape? Answer in one word."
|
||||
```
|
||||
|
||||
Build: [github.com/The-Monk/The-Rock8](https://github.com/The-Monk/The-Rock8) (RDNA4 native-fp8 llama.cpp fork + Podman appliance).
|
||||
|
||||
## Attribution & license
|
||||
|
||||
Base model **[prism-ml/Ternary-Bonsai-8B](https://huggingface.co/prism-ml/Ternary-Bonsai-8B)** by PrismML, **Apache-2.0**. These are GGUF conversions/quantizations for RDNA4; all credit for the model and its ternary QAT training to PrismML. Distributed under the same Apache-2.0 license.
|
||||
|
||||
🦆 *Got any weights?*
|
||||
3
Ternary-Bonsai-8B-F16.gguf
Normal file
3
Ternary-Bonsai-8B-F16.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:a6abfaf896c1e36db825112fc0a18e49adea05eeca1c6b2fba4d785ca7e947ff
|
||||
size 16383663200
|
||||
3
Ternary-Bonsai-8B-F8E4M3.gguf
Normal file
3
Ternary-Bonsai-8B-F8E4M3.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:d65c37c9a150afa5b8501d6201d324c92285d67bf9b1d194706e868d694e14b6
|
||||
size 8556732672
|
||||
3
Ternary-Bonsai-8B-Q2_0.gguf
Normal file
3
Ternary-Bonsai-8B-Q2_0.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:3c8d70470a5d97e5a2b9410ddd899cb740116591462626c60cb2fead6448f60b
|
||||
size 2182184672
|
||||
Reference in New Issue
Block a user