149 lines
3.9 KiB
Markdown
149 lines
3.9 KiB
Markdown
---
|
|
language:
|
|
- en
|
|
license: apache-2.0
|
|
library_name: transformers
|
|
tags:
|
|
- quantization
|
|
- fp4
|
|
- nvfp4
|
|
- compressed-tensors
|
|
- qwen2
|
|
- text-generation
|
|
- 4bit
|
|
base_model: Qwen/Qwen2.5-14B-Instruct
|
|
pipeline_tag: text-generation
|
|
model-index:
|
|
- name: Qwen2.5-14B-Instruct-FP4-W4A4
|
|
results: []
|
|
quantization:
|
|
quant_method: compressed-tensors
|
|
bits: 4
|
|
type: float
|
|
format: nvfp4-pack-quantized
|
|
strategy: tensor_group
|
|
group_size: 16
|
|
symmetric: true
|
|
---
|
|
|
|
# Qwen2.5-14B-Instruct-FP4-W4A4
|
|
|
|
## Model Description
|
|
|
|
This is an NVFP4 (NVIDIA FP4) quantized version of [Qwen/Qwen2.5-14B-Instruct](https://huggingface.co/Qwen/Qwen2.5-14B-Instruct) using the compressed-tensors quantization method.
|
|
|
|
- **Base Model**: Qwen/Qwen2.5-14B-Instruct
|
|
- **Quantization Method**: compressed-tensors
|
|
- **Quantization Type**: NVFP4 W4A4 (4-bit Weight and Activation)
|
|
- **Model Size**: ~10.5GB (compared to ~28GB for BF16)
|
|
- **Compression Ratio**: ~2.7x
|
|
|
|
## Quantization Configuration
|
|
|
|
This model uses **NVFP4 (NVIDIA FP4) quantization** with grouped quantization for both weights and activations:
|
|
|
|
### Weights
|
|
- **Precision**: NVFP4 (4-bit floating point)
|
|
- **Strategy**: Tensor-group (grouped quantization)
|
|
- **Group Size**: 16
|
|
- **Symmetric**: Yes
|
|
- **Dynamic**: No (static quantization)
|
|
- **Observer**: MinMax
|
|
|
|
### Activations
|
|
- **Precision**: NVFP4 (4-bit floating point)
|
|
- **Strategy**: Tensor-group (grouped quantization)
|
|
- **Group Size**: 16
|
|
- **Symmetric**: Yes
|
|
- **Dynamic**: Local (dynamic quantization with local calibration)
|
|
- **Observer**: MinMax
|
|
|
|
### Other Details
|
|
- **Format**: nvfp4-pack-quantized (packed 4-bit format)
|
|
- **KV Cache**: Not quantized
|
|
- **Ignored Layers**: lm_head
|
|
- **Target Layers**: Linear layers
|
|
- **Quantization Version**: 0.11.0
|
|
|
|
## Usage
|
|
|
|
```python
|
|
from transformers import AutoModelForCausalLM, AutoTokenizer
|
|
|
|
model_id = "JongYeop/Qwen2.5-14B-Instruct-FP4-W4A4"
|
|
|
|
# Load tokenizer
|
|
tokenizer = AutoTokenizer.from_pretrained(model_id)
|
|
|
|
# Load quantized model
|
|
model = AutoModelForCausalLM.from_pretrained(
|
|
model_id,
|
|
device_map="auto",
|
|
torch_dtype="auto"
|
|
)
|
|
|
|
# Generate text
|
|
messages = [
|
|
{"role": "user", "content": "What is machine learning?"}
|
|
]
|
|
|
|
input_ids = tokenizer.apply_chat_template(
|
|
messages,
|
|
add_generation_prompt=True,
|
|
return_tensors="pt"
|
|
).to(model.device)
|
|
|
|
outputs = model.generate(
|
|
input_ids,
|
|
max_new_tokens=256,
|
|
do_sample=True,
|
|
temperature=0.7,
|
|
top_p=0.9,
|
|
)
|
|
|
|
response = tokenizer.decode(outputs[0][input_ids.shape[-1]:], skip_special_tokens=True)
|
|
print(response)
|
|
```
|
|
|
|
## Model Architecture
|
|
|
|
- **Architecture**: Qwen2ForCausalLM
|
|
- **Hidden Size**: 5120
|
|
- **Intermediate Size**: 13824
|
|
- **Number of Layers**: 48
|
|
- **Number of Attention Heads**: 40
|
|
- **Number of KV Heads**: 8
|
|
- **Vocabulary Size**: 152064
|
|
- **Max Position Embeddings**: 32768
|
|
|
|
## Intended Use
|
|
|
|
This quantized model is intended for efficient inference with significantly reduced memory footprint while maintaining reasonable performance. It is suitable for:
|
|
|
|
- Resource-constrained environments
|
|
- Edge deployment
|
|
- Applications requiring minimal memory usage
|
|
- High throughput scenarios
|
|
- GPU inference with FP4 support
|
|
|
|
## Limitations
|
|
|
|
- FP4 quantization may result in more accuracy loss compared to FP8 or INT8 quantization
|
|
- Best performance is achieved on hardware with native FP4 support (e.g., NVIDIA H100, Ada Lovelace, Blackwell GPUs)
|
|
- Dynamic activation quantization may introduce additional runtime overhead
|
|
- Grouped quantization requires compatible inference engines
|
|
|
|
## Performance Notes
|
|
|
|
- **Memory Usage**: ~2.7x reduction compared to BF16
|
|
- **Speed**: Requires hardware with FP4 tensor core support for optimal performance
|
|
- **Accuracy**: May experience some degradation compared to higher precision formats
|
|
|
|
## Citation
|
|
|
|
If you use this model, please cite the original Qwen2.5 paper and the compressed-tensors library.
|
|
|
|
## License
|
|
|
|
Same as the base model: [Apache 2.0](https://huggingface.co/Qwen/Qwen2.5-14B-Instruct)
|