Files
ModelHub XC 9685982964 初始化项目,由ModelHub XC社区提供模型
Model: donghyunli/Llama-2-7b-KronQ-W3A16-fake
Source: Original Platform
2026-07-20 11:52:12 +08:00

1.8 KiB

base_model, language, license, pipeline_tag, library_name, tags
base_model language license pipeline_tag library_name tags
meta-llama/Llama-2-7b-hf
en
llama2 text-generation transformers
kronq
quantization
fake-quant
fp16

Llama-2-7b — KronQ W3A16 (fake-quant fp16)

Paper: arXiv:2607.07964 · Code: GitHub

⚠️ Fake-quant fp16 checkpoint. The 3-bit weights are stored in fp16 (KronQ does not pack int3), so this repo is the same size as bf16 — no compression or speedup, for PPL / accuracy reproduction only. For deployable low-bit, see the W4A16 / W2A16 (packed int4/int2) repos.

Llama-2-7b quantized to 3-bit weights with KronQ (Kronecker-factored Hessian quantization), exported as a standard fp16 model.

Results (WikiText-2, seqlen 2048)

Perplexity: 5.84

Zero-shot accuracy:

PIQA ARC-E ARC-C HellaSwag WinoGrande BoolQ OBQA Average
77.09 72.39 42.58 72.04 67.64 75.35 41.60 64.10

(lm-evaluation-harness, 0-shot. acc_norm for PIQA/HellaSwag/ARC/OBQA, acc for WinoGrande/BoolQ.)

Usage

Loads as a standard fp16 model — no KronQ code needed:

from transformers import AutoModelForCausalLM, AutoTokenizer
m = AutoModelForCausalLM.from_pretrained("donghyunli/Llama-2-7b-KronQ-W3A16-fake", torch_dtype="float16").cuda()
tok = AutoTokenizer.from_pretrained("donghyunli/Llama-2-7b-KronQ-W3A16-fake")

Recipe

Per-channel asymmetric W3, weight-only (a_bits=16), --alpha 0.25, bidirectional incoherence processing (BiIP), act_order, raw H_G. Calibrated on 128 WikiText-2 sequences.

License

Derivative of Llama-2-7b — subject to the Llama 2 Community License.