Model: donghyunli/Meta-Llama-3-8B-KronQ-W3A16-g128-fake Source: Original Platform
base_model, language, license, pipeline_tag, library_name, tags
| base_model | language | license | pipeline_tag | library_name | tags | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| meta-llama/Meta-Llama-3-8B |
|
llama3 | text-generation | transformers |
|
Meta-Llama-3-8B — KronQ W3A16 g128 (fake-quant fp16)
Paper: arXiv:2607.07964 · Code: GitHub
⚠️ Fake-quant fp16 checkpoint. 3-bit group-128 weights stored in fp16 (KronQ does not pack int3) — same size as bf16, for PPL/accuracy reproduction. For deployable low-bit see the W4A16-g128 / W2A16-g128 (packed) repos.
Meta-Llama-3-8B quantized to 3-bit weights (group 128) with KronQ, exported as a standard fp16 model.
Results (WikiText-2, seqlen 2048)
Perplexity: 6.959
Zero-shot accuracy:
| PIQA | ARC-E | ARC-C | HellaSwag | WinoGrande | BoolQ | OBQA | Average |
|---|---|---|---|---|---|---|---|
| 78.67 | 74.20 | 49.06 | 74.92 | 71.67 | 81.68 | 42.40 | 67.51 |
(lm-evaluation-harness 0-shot.)
Usage
Loads as a standard fp16 model (no KronQ code):
from transformers import AutoModelForCausalLM
m = AutoModelForCausalLM.from_pretrained("donghyunli/Meta-Llama-3-8B-KronQ-W3A16-g128-fake", torch_dtype="float16", device_map="auto")
Recipe
Group-128 asymmetric W3, weight-only, --alpha 0.25, --act_order, BiIP, raw H_G.
License
Derivative of Meta-Llama-3-8B — llama3 license.
Description