68 lines
1.9 KiB
Markdown
68 lines
1.9 KiB
Markdown
|
|
---
|
||
|
|
license: other
|
||
|
|
base_model: LiquidAI/LFM2.5-8B-A1B
|
||
|
|
base_model_relation: quantized
|
||
|
|
pipeline_tag: text-generation
|
||
|
|
library_name: gguf
|
||
|
|
tags:
|
||
|
|
- gguf
|
||
|
|
- ollama
|
||
|
|
- local-llm
|
||
|
|
- llama.cpp
|
||
|
|
- lm-studio
|
||
|
|
- quantized
|
||
|
|
- imatrix
|
||
|
|
- sub-4-bit
|
||
|
|
quantized_by: liodon-ai
|
||
|
|
---
|
||
|
|
|
||
|
|
# LFM2.5-8B-A1B — iMatrix GGUF
|
||
|
|
|
||
|
|
GGUF quantizations of [LiquidAI/LFM2.5-8B-A1B](https://huggingface.co/LiquidAI/LFM2.5-8B-A1B), published by [Liodon AI](https://huggingface.co/liodon-ai).
|
||
|
|
|
||
|
|
## Quick Start
|
||
|
|
|
||
|
|
**llama.cpp**
|
||
|
|
```bash
|
||
|
|
llama-cli -hf liodon-ai/LFM2.5-8B-A1B-imatrix-GGUF:Q4_K_M
|
||
|
|
```
|
||
|
|
|
||
|
|
**Ollama**
|
||
|
|
```bash
|
||
|
|
ollama run hf.co/liodon-ai/LFM2.5-8B-A1B-imatrix-GGUF:Q4_K_M
|
||
|
|
```
|
||
|
|
|
||
|
|
**LM Studio / Jan** — search `liodon-ai/LFM2.5-8B-A1B-imatrix-GGUF` and pick your quant.
|
||
|
|
|
||
|
|
## Quants
|
||
|
|
|
||
|
|
| Quant | Size | VRAM est. | Notes |
|
||
|
|
|-------|------|-----------|-------|
|
||
|
|
| `IQ2_M` | 2.84 GB | ~3 GB | 2-bit, iMatrix — smallest usable |
|
||
|
|
| `IQ3_M` | 3.78 GB | ~4 GB | 3-bit, iMatrix — great quality/size tradeoff |
|
||
|
|
| `IQ4_XS` | 4.59 GB | ~5 GB | 4-bit extra-small, iMatrix |
|
||
|
|
| `Q4_K_M` | 5.16 GB | ~6 GB | 4-bit, iMatrix-calibrated (recommended) |
|
||
|
|
| `Q5_K_M` | 6.03 GB | ~7 GB | 5-bit, iMatrix-calibrated |
|
||
|
|
| `Q6_K` | 6.96 GB | ~8 GB | 6-bit, iMatrix-calibrated, near-lossless |
|
||
|
|
| `Q8_0` | 9.01 GB | ~10 GB | 8-bit, essentially lossless |
|
||
|
|
|
||
|
|
|
||
|
|
## What is iMatrix?
|
||
|
|
|
||
|
|
Standard quantization treats all weights equally. iMatrix runs 128 calibration chunks through
|
||
|
|
the full-precision model to find which weights matter most, then allocates more precision where
|
||
|
|
it counts. At Q2/Q3/Q4 this means noticeably better coherence and instruction-following —
|
||
|
|
**same file size, better output**.
|
||
|
|
|
||
|
|
Calibration: 2M tokens of [WikiText-103](https://huggingface.co/datasets/wikitext).
|
||
|
|
|
||
|
|
> Also see plain (non-iMatrix) quants: `liodon-ai/LFM2.5-8B-A1B-GGUF`
|
||
|
|
|
||
|
|
## Source
|
||
|
|
|
||
|
|
- **Model**: [LiquidAI/LFM2.5-8B-A1B](https://huggingface.co/LiquidAI/LFM2.5-8B-A1B)
|
||
|
|
- **License**: other
|
||
|
|
|
||
|
|
---
|
||
|
|
*Quantized by [Liodon AI](https://huggingface.co/liodon-ai)*
|