--- base_model: meta-llama/Meta-Llama-3-8B language: - en license: llama3 pipeline_tag: text-generation library_name: transformers tags: - kronq - quantization - fake-quant - fp16 --- # Meta-Llama-3-8B — KronQ W3A16 (fake-quant fp16) **Paper:** [arXiv:2607.07964](https://arxiv.org/abs/2607.07964) · **Code:** [GitHub](https://github.com/Intelligent-Computing-Lab-Panda/KronQ) > ⚠️ **Fake-quant fp16 checkpoint.** The 3-bit weights are stored in **fp16** > (KronQ does not pack int3), so this repo is the **same size as bf16** — no > compression or speedup, for **PPL / accuracy reproduction only**. For > deployable low-bit, see the W4A16 (packed int4) repo. [Meta-Llama-3-8B](https://huggingface.co/meta-llama/Meta-Llama-3-8B) quantized to 3-bit weights with **KronQ** (Kronecker-factored Hessian quantization), exported as a standard fp16 model. ## Results (WikiText-2, seqlen 2048) **Perplexity:** **7.09** **Zero-shot accuracy:** | PIQA | ARC-E | ARC-C | HellaSwag | WinoGrande | BoolQ | OBQA | Average | |---|---|---|---|---|---|---|---| | 77.53 | 74.54 | 50.17 | 74.92 | 71.74 | 81.13 | 41.20 | **67.32** | (lm-evaluation-harness, 0-shot. `acc_norm` for PIQA/HellaSwag/ARC/OBQA, `acc` for WinoGrande/BoolQ.) ## Usage Loads as a **standard fp16 model** — no KronQ code needed: ```python from transformers import AutoModelForCausalLM, AutoTokenizer m = AutoModelForCausalLM.from_pretrained("donghyunli/Meta-Llama-3-8B-KronQ-W3A16-fake", torch_dtype="float16").cuda() tok = AutoTokenizer.from_pretrained("donghyunli/Meta-Llama-3-8B-KronQ-W3A16-fake") ``` ## Recipe Per-channel asymmetric W3, weight-only (a_bits=16), `--alpha 0.25`, bidirectional incoherence processing (BiIP), `act_order`, raw H_G. Calibrated on 128 WikiText-2 sequences. ## License Derivative of Meta-Llama-3-8B — subject to the [Llama 3 Community License](https://llama.meta.com/llama3/license/).