23a1c646288c287cf6cabf5f987a5c0a716be319
Model: liodon-ai/Qwen2.5-Coder-7B-Instruct-imatrix-GGUF Source: Original Platform
license, base_model, base_model_relation, pipeline_tag, library_name, tags, quantized_by
| license | base_model | base_model_relation | pipeline_tag | library_name | tags | quantized_by | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| other | Qwen/Qwen2.5-Coder-7B-Instruct | quantized | text-generation | gguf |
|
liodon-ai |
Qwen2.5-Coder-7B-Instruct — iMatrix GGUF
GGUF quantizations of Qwen/Qwen2.5-Coder-7B-Instruct, published by Liodon AI.
Quick Start
llama.cpp
llama-cli -hf liodon-ai/Qwen2.5-Coder-7B-Instruct-imatrix-GGUF:Q4_K_M
Ollama
ollama run hf.co/liodon-ai/Qwen2.5-Coder-7B-Instruct-imatrix-GGUF:Q4_K_M
LM Studio / Jan — search liodon-ai/Qwen2.5-Coder-7B-Instruct-imatrix-GGUF and pick your quant.
Quants
| Quant | Size | VRAM est. | Notes |
|---|---|---|---|
IQ2_M |
2.78 GB | ~3 GB | 2-bit, iMatrix — smallest usable |
IQ3_M |
3.57 GB | ~4 GB | 3-bit, iMatrix — great quality/size tradeoff |
IQ4_XS |
4.22 GB | ~5 GB | 4-bit extra-small, iMatrix |
Q4_K_M |
4.68 GB | ~5 GB | 4-bit, iMatrix-calibrated (recommended) |
Q5_K_M |
5.44 GB | ~6 GB | 5-bit, iMatrix-calibrated |
Q6_K |
6.25 GB | ~7 GB | 6-bit, iMatrix-calibrated, near-lossless |
Q8_0 |
8.10 GB | ~9 GB | 8-bit, essentially lossless |
What is iMatrix?
Standard quantization treats all weights equally. iMatrix runs 128 calibration chunks through the full-precision model to find which weights matter most, then allocates more precision where it counts. At Q2/Q3/Q4 this means noticeably better coherence and instruction-following — same file size, better output.
Calibration: 2M tokens of WikiText-103.
Also see plain (non-iMatrix) quants:
liodon-ai/Qwen2.5-Coder-7B-Instruct-GGUF
Source
- Model: Qwen/Qwen2.5-Coder-7B-Instruct
- License: other
Quantized by Liodon AI
Description