base_model, language, license, pipeline_tag, library_name, tags
base_model language license pipeline_tag library_name tags
meta-llama/Meta-Llama-3-8B
en
llama3 text-generation transformers
kronq
quantization
fake-quant
fp16

Meta-Llama-3-8B — KronQ W3A16 (fake-quant fp16)

Paper: arXiv:2607.07964 · Code: GitHub

⚠️ Fake-quant fp16 checkpoint. The 3-bit weights are stored in fp16 (KronQ does not pack int3), so this repo is the same size as bf16 — no compression or speedup, for PPL / accuracy reproduction only. For deployable low-bit, see the W4A16 (packed int4) repo.

Meta-Llama-3-8B quantized to 3-bit weights with KronQ (Kronecker-factored Hessian quantization), exported as a standard fp16 model.

Results (WikiText-2, seqlen 2048)

Perplexity: 7.09

Zero-shot accuracy:

PIQA ARC-E ARC-C HellaSwag WinoGrande BoolQ OBQA Average
77.53 74.54 50.17 74.92 71.74 81.13 41.20 67.32

(lm-evaluation-harness, 0-shot. acc_norm for PIQA/HellaSwag/ARC/OBQA, acc for WinoGrande/BoolQ.)

Usage

Loads as a standard fp16 model — no KronQ code needed:

from transformers import AutoModelForCausalLM, AutoTokenizer
m = AutoModelForCausalLM.from_pretrained("donghyunli/Meta-Llama-3-8B-KronQ-W3A16-fake", torch_dtype="float16").cuda()
tok = AutoTokenizer.from_pretrained("donghyunli/Meta-Llama-3-8B-KronQ-W3A16-fake")

Recipe

Per-channel asymmetric W3, weight-only (a_bits=16), --alpha 0.25, bidirectional incoherence processing (BiIP), act_order, raw H_G. Calibrated on 128 WikiText-2 sequences.

License

Derivative of Meta-Llama-3-8B — subject to the Llama 3 Community License.

Description
Model synced from source: donghyunli/Meta-Llama-3-8B-KronQ-W3A16-fake
Readme 2.6 MiB