初始化项目,由ModelHub XC社区提供模型

Model: Entrit/Mistral-7B-v0.3-trit-uniform-d4
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-06-16 04:44:19 +08:00
commit 284e2a6c1b
8 changed files with 275884 additions and 0 deletions

35
.gitattributes vendored Normal file
View File

@@ -0,0 +1,35 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text

1
.mirrored Normal file
View File

@@ -0,0 +1 @@
2026-04-16T00:25:13.410997

60
README.md Normal file
View File

@@ -0,0 +1,60 @@
---
license: apache-2.0
base_model: mistralai/Mistral-7B-v0.3
tags:
- quantization
- ternary
- balanced-ternary
- tritllm
library_name: transformers
---
# Mistral-7B-v0.3-trit-uniform-d4
Balanced ternary quantization of [`mistralai/Mistral-7B-v0.3`](https://huggingface.co/mistralai/Mistral-7B-v0.3) at depth **d=4** (81 levels per weight, **6.64 bits per weight**).
Produced with the codec from **"Balanced Ternary Post-Training Quantization for Large Language Models"** (Stentzel, 2026). See [Entrit/tritllm-codec](https://huggingface.co/Entrit/tritllm-codec) for the codec source.
## Quick load
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("Entrit/Mistral-7B-v0.3-trit-uniform-d4")
tokenizer = AutoTokenizer.from_pretrained("Entrit/Mistral-7B-v0.3-trit-uniform-d4")
```
The weights are dequantized to FP16 for stock-`transformers` compatibility. The on-disk size is therefore the same as the FP16 source. The 6.64-bpw figure refers to the *information content* of the quantized matrices and is what matters for inference on hardware that consumes the packed trit format directly (see [Entrit/tritllm-kernel](https://huggingface.co/Entrit/tritllm-kernel)).
## Quantization details
| Field | Value |
|---|---|
| Source model | [`mistralai/Mistral-7B-v0.3`](https://huggingface.co/mistralai/Mistral-7B-v0.3) |
| Depth | d=4 (81 levels) |
| Bits per weight | 6.64 |
| Group size | 16 |
| Scale codebook | 27-entry log-spaced (scale_depth=3) |
| Method | Uniform PTQ |
| Quantized layers | all 2D linear matrices |
| Kept FP16 | `lm_head`, token embeddings, all `*_norm` layers |
| Codec | tritllm v2 |
## Citation
```
@article{stentzel2026ternaryptq,
title = {Balanced Ternary Post-Training Quantization for Large Language Models},
author = {Stentzel, Eric},
year = 2026,
note = {Entrit Systems}
}
```
## Reproducibility
```bash
git clone https://huggingface.co/Entrit/tritllm-codec
cd tritllm-codec
python quantize_model_v2.py --model mistralai/Mistral-7B-v0.3 --configs uniform-d4 --out ./out
```

30
config.json Normal file
View File

@@ -0,0 +1,30 @@
{
"architectures": [
"MistralForCausalLM"
],
"attention_dropout": 0.0,
"bos_token_id": 1,
"dtype": "float16",
"eos_token_id": 2,
"head_dim": 128,
"hidden_act": "silu",
"hidden_size": 4096,
"initializer_range": 0.02,
"intermediate_size": 14336,
"max_position_embeddings": 32768,
"model_type": "mistral",
"num_attention_heads": 32,
"num_hidden_layers": 32,
"num_key_value_heads": 8,
"pad_token_id": null,
"rms_norm_eps": 1e-05,
"rope_parameters": {
"rope_theta": 1000000.0,
"rope_type": "default"
},
"sliding_window": null,
"tie_word_embeddings": false,
"transformers_version": "5.5.3",
"use_cache": true,
"vocab_size": 32768
}

6
generation_config.json Normal file
View File

@@ -0,0 +1,6 @@
{
"_from_model_config": true,
"bos_token_id": 1,
"eos_token_id": 2,
"transformers_version": "5.5.3"
}

3
model.safetensors Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:777971af468a5e908e76c36f0df364e5599582ec6373194f083b98b6e29d42c9
size 14496080848

275733
tokenizer.json Normal file

File diff suppressed because it is too large Load Diff

16
tokenizer_config.json Normal file
View File

@@ -0,0 +1,16 @@
{
"add_prefix_space": true,
"backend": "tokenizers",
"bos_token": "<s>",
"clean_up_tokenization_spaces": false,
"eos_token": "</s>",
"is_local": true,
"legacy": false,
"model_max_length": 1000000000000000019884624838656,
"pad_token": "</s>",
"sp_model_kwargs": {},
"spaces_between_special_tokens": false,
"tokenizer_class": "TokenizersBackend",
"unk_token": "<unk>",
"use_default_system_prompt": false
}