Model: donghyunli/Meta-Llama-3-8B-KronQ-W3A16-fake Source: Original Platform
base_model, language, license, pipeline_tag, library_name, tags
| base_model | language | license | pipeline_tag | library_name | tags | |||||
|---|---|---|---|---|---|---|---|---|---|---|
| meta-llama/Meta-Llama-3-8B |
|
llama3 | text-generation | transformers |
|
Meta-Llama-3-8B — KronQ W3A16 (fake-quant fp16)
Paper: arXiv:2607.07964 · Code: GitHub
⚠️ Fake-quant fp16 checkpoint. The 3-bit weights are stored in fp16 (KronQ does not pack int3), so this repo is the same size as bf16 — no compression or speedup, for PPL / accuracy reproduction only. For deployable low-bit, see the W4A16 (packed int4) repo.
Meta-Llama-3-8B quantized to 3-bit weights with KronQ (Kronecker-factored Hessian quantization), exported as a standard fp16 model.
Results (WikiText-2, seqlen 2048)
Perplexity: 7.09
Zero-shot accuracy:
| PIQA | ARC-E | ARC-C | HellaSwag | WinoGrande | BoolQ | OBQA | Average |
|---|---|---|---|---|---|---|---|
| 77.53 | 74.54 | 50.17 | 74.92 | 71.74 | 81.13 | 41.20 | 67.32 |
(lm-evaluation-harness, 0-shot. acc_norm for PIQA/HellaSwag/ARC/OBQA, acc for WinoGrande/BoolQ.)
Usage
Loads as a standard fp16 model — no KronQ code needed:
from transformers import AutoModelForCausalLM, AutoTokenizer
m = AutoModelForCausalLM.from_pretrained("donghyunli/Meta-Llama-3-8B-KronQ-W3A16-fake", torch_dtype="float16").cuda()
tok = AutoTokenizer.from_pretrained("donghyunli/Meta-Llama-3-8B-KronQ-W3A16-fake")
Recipe
Per-channel asymmetric W3, weight-only (a_bits=16), --alpha 0.25, bidirectional incoherence processing (BiIP), act_order, raw H_G. Calibrated on 128 WikiText-2 sequences.
License
Derivative of Meta-Llama-3-8B — subject to the Llama 3 Community License.