Files
ModelHub XC 25bd811ce1 初始化项目,由ModelHub XC社区提供模型
Model: JongYeop/Qwen2.5-14B-Instruct-FP8-W8A8
Source: Original Platform
2026-07-31 20:23:37 +08:00

145 lines
3.7 KiB
Markdown

---
language:
- en
license: apache-2.0
library_name: transformers
tags:
- quantization
- fp8
- compressed-tensors
- qwen2
- text-generation
- 8bit
base_model: Qwen/Qwen2.5-14B-Instruct
pipeline_tag: text-generation
model-index:
- name: Qwen2.5-14B-Instruct-FP8-W8A8
results: []
quantization:
quant_method: compressed-tensors
bits: 8
type: float
format: float-quantized
strategy: tensor
symmetric: true
---
# Qwen2.5-14B-Instruct-FP8-W8A8
## Model Description
This is an FP8 quantized version of [Qwen/Qwen2.5-14B-Instruct](https://huggingface.co/Qwen/Qwen2.5-14B-Instruct) using the compressed-tensors quantization method.
- **Base Model**: Qwen/Qwen2.5-14B-Instruct
- **Quantization Method**: compressed-tensors
- **Quantization Type**: FP8 W8A8 (8-bit Weight and Activation)
- **Model Size**: ~16.3GB (compared to ~28GB for BF16)
- **Compression Ratio**: ~1.7x
## Quantization Configuration
This model uses **FP8 quantization** with per-tensor quantization for both weights and activations:
### Weights
- **Precision**: FP8 (8-bit floating point)
- **Strategy**: Per-tensor
- **Group Size**: None (per-tensor)
- **Symmetric**: Yes
- **Dynamic**: No (static quantization)
- **Observer**: MinMax
### Activations
- **Precision**: FP8 (8-bit floating point)
- **Strategy**: Per-tensor
- **Group Size**: None (per-tensor)
- **Symmetric**: Yes
- **Dynamic**: No (static quantization)
- **Observer**: MinMax
### Other Details
- **Format**: float-quantized
- **KV Cache**: Not quantized
- **Ignored Layers**: lm_head
- **Target Layers**: Linear layers
- **Quantization Version**: 0.11.0
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "JongYeop/Qwen2.5-14B-Instruct-FP8-W8A8"
# Load tokenizer
tokenizer = AutoTokenizer.from_pretrained(model_id)
# Load quantized model
model = AutoModelForCausalLM.from_pretrained(
model_id,
device_map="auto",
torch_dtype="auto"
)
# Generate text
messages = [
{"role": "user", "content": "What is machine learning?"}
]
input_ids = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
return_tensors="pt"
).to(model.device)
outputs = model.generate(
input_ids,
max_new_tokens=256,
do_sample=True,
temperature=0.7,
top_p=0.9,
)
response = tokenizer.decode(outputs[0][input_ids.shape[-1]:], skip_special_tokens=True)
print(response)
```
## Model Architecture
- **Architecture**: Qwen2ForCausalLM
- **Hidden Size**: 5120
- **Intermediate Size**: 13824
- **Number of Layers**: 48
- **Number of Attention Heads**: 40
- **Number of KV Heads**: 8
- **Vocabulary Size**: 152064
- **Max Position Embeddings**: 32768
## Intended Use
This quantized model is intended for efficient inference with reduced memory footprint while maintaining high accuracy. It is suitable for:
- Production deployment with reduced memory requirements
- High throughput inference scenarios
- GPU inference with FP8 support
- Applications where accuracy is important but memory savings are desired
## Limitations
- Best performance is achieved on hardware with native FP8 support (e.g., NVIDIA H100, Ada Lovelace, Blackwell GPUs)
- Requires compatible inference engines that support FP8 computation
- Static quantization may not adapt to varying input distributions
## Performance Notes
- **Memory Usage**: ~1.7x reduction compared to BF16
- **Speed**: Requires hardware with FP8 tensor core support for optimal performance
- **Accuracy**: Generally retains most of the original model's accuracy, with minimal degradation compared to FP4
## Citation
If you use this model, please cite the original Qwen2.5 paper and the compressed-tensors library.
## License
Same as the base model: [Apache 2.0](https://huggingface.co/Qwen/Qwen2.5-14B-Instruct)