commit bcc765f1fac4a8de9bc08a636a391d44de36b4a0 Author: ModelHub XC Date: Sat Aug 8 07:39:17 2026 +0800 初始化项目,由ModelHub XC社区提供模型 Model: AtomicChat/lfm25-8b-a1b-GGUF Source: Original Platform diff --git a/.gitattributes b/.gitattributes new file mode 100644 index 0000000..b6811a6 --- /dev/null +++ b/.gitattributes @@ -0,0 +1,47 @@ +*.7z filter=lfs diff=lfs merge=lfs -text +*.arrow filter=lfs diff=lfs merge=lfs -text +*.bin filter=lfs diff=lfs merge=lfs -text +*.bz2 filter=lfs diff=lfs merge=lfs -text +*.ckpt filter=lfs diff=lfs merge=lfs -text +*.ftz filter=lfs diff=lfs merge=lfs -text +*.gz filter=lfs diff=lfs merge=lfs -text +*.h5 filter=lfs diff=lfs merge=lfs -text +*.joblib filter=lfs diff=lfs merge=lfs -text +*.lfs.* filter=lfs diff=lfs merge=lfs -text +*.mlmodel filter=lfs diff=lfs merge=lfs -text +*.model filter=lfs diff=lfs merge=lfs -text +*.msgpack filter=lfs diff=lfs merge=lfs -text +*.npy filter=lfs diff=lfs merge=lfs -text +*.npz filter=lfs diff=lfs merge=lfs -text +*.onnx filter=lfs diff=lfs merge=lfs -text +*.ot filter=lfs diff=lfs merge=lfs -text +*.parquet filter=lfs diff=lfs merge=lfs -text +*.pb filter=lfs diff=lfs merge=lfs -text +*.pickle filter=lfs diff=lfs merge=lfs -text +*.pkl filter=lfs diff=lfs merge=lfs -text +*.pt filter=lfs diff=lfs merge=lfs -text +*.pth filter=lfs diff=lfs merge=lfs -text +*.rar filter=lfs diff=lfs merge=lfs -text +*.safetensors filter=lfs diff=lfs merge=lfs -text +saved_model/**/* filter=lfs diff=lfs merge=lfs -text +*.tar.* filter=lfs diff=lfs merge=lfs -text +*.tar filter=lfs diff=lfs merge=lfs -text +*.tflite filter=lfs diff=lfs merge=lfs -text +*.tgz filter=lfs diff=lfs merge=lfs -text +*.wasm filter=lfs diff=lfs merge=lfs -text +*.xz filter=lfs diff=lfs merge=lfs -text +*.zip filter=lfs diff=lfs merge=lfs -text +*.zst filter=lfs diff=lfs merge=lfs -text +*tfevents* filter=lfs diff=lfs merge=lfs -text +lfm25-8b-a1b-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text +lfm25-8b-a1b-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text +lfm25-8b-a1b-Q2_K.gguf filter=lfs diff=lfs merge=lfs -text +lfm25-8b-a1b-Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text +lfm25-8b-a1b-Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text +lfm25-8b-a1b-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text +lfm25-8b-a1b-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text +lfm25-8b-a1b-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text +lfm25-8b-a1b-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text +lfm25-8b-a1b-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text +lfm25-8b-a1b-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text +lfm25-8b-a1b-UD-Q4_K_XL.gguf filter=lfs diff=lfs merge=lfs -text diff --git a/README.md b/README.md new file mode 100644 index 0000000..7112b17 --- /dev/null +++ b/README.md @@ -0,0 +1,136 @@ +--- +license: other +license_link: LICENSE +thumbnail: https://huggingface.co/AtomicChat/lfm25-8b-a1b-GGUF/resolve/main/hero.png +base_model: +- LiquidAI/LFM2.5-8B-A1B +base_model_relation: quantized +quantized_by: AtomicChat +pipeline_tag: text-generation +library_name: gguf +tags: +- atomic-chat +- lfm2.5 +- liquidai +- gguf +- llama.cpp +- quantized +--- + +
+ +
+Atomic Chat +Join Discord +GitHub +
+ +
+ +LFM2.5 8B A1B + +
+Base model: LiquidAI/LFM2.5-8B-A1B +
+
+ +**LFM2.5 8B A1B**, self-quantized to GGUF by [Atomic Chat](https://atomic.chat). Built straight from Liquid AI's original weights with a per-tensor importance matrix, so this is not a repack of somebody else's files. Runs fully offline. + +## Highlights + +- **8.5B parameters**: the weights this repo quantizes. +- **Context length**: 128,000 tokens (125K), as published by Liquid AI. +- **24 layers**: Mixture-of-Experts. +- **Full imatrix ladder**: every quant is calibrated with an importance matrix. +- **On-device personal assistant**: Designed to power real-life applications, chaining tool calls, and following complex instructions on all devices. +- **Compressed performance**: Competitive with much larger dense and MoE models on instruction following and agentic tasks. +- **Unmatched throughput**: Fastest in its size class on both CPU and GPU inference, with day-one support for llama.cpp, MLX, vLLM, and SGLang. + +> [!NOTE] +> These GGUFs are **self-quantized from the original weights**, not a repack. The importance matrix keeps low-bit quants closer to the full-precision model. + +> [!IMPORTANT] +> Always pass `--jinja` so the **LFM2.5 8B A1B chat template** is applied. Without it the model can emit malformed turns. + +## Model Overview + +| Property | Value | +|---|---| +| Base model | `LiquidAI/LFM2.5-8B-A1B` | +| Parameters | 8.5B | +| Layers | 24 | +| Experts | 32 routed (top-4) | +| Context length | 128,000 tokens (125K) | +| Vocabulary | 128,000 | +| Modalities | Text | +| Architecture | Mixture-of-Experts, 32 experts (top-4), 32 attention heads over 8 KV heads, `Lfm2MoeForCausalLM` | +| This repo | GGUF quants (imatrix). Quants: `Q2_K`, `IQ3_M`, `Q3_K_M`, `Q3_K_L`, `IQ4_XS`, `Q4_K_S`, `Q4_K_M`, `UD-Q4_K_XL`, `Q5_K_S`, `Q5_K_M`, `Q6_K`, `Q8_0` | + +LFM2.5 8B A1B benchmark scores + +Scores are Liquid AI's published results for the base `LiquidAI/LFM2.5-8B-A1B`, not our own measurements. Quantization preserves the large majority of this; `Q4_K_M` and up stay close to full precision. + +## Choosing a quant + +| Quant | Size | Notes | +|---|---|---| +| `Q2_K` | 3.2 GB | Smallest K-quant. Minimal RAM, clear quality drop. | +| `IQ3_M` | 3.8 GB | Beats Q3 at a similar size thanks to imatrix. Best low-RAM pick. | +| `Q3_K_M` | 4.1 GB | Low quality but usable. | +| `Q3_K_L` | 4.4 GB | A step above Q3_K_M. | +| `IQ4_XS` | 4.6 GB | Excellent quality for size. Recommended low-bit. | +| `Q4_K_S` | 4.9 GB | Compact 4-bit, fast. | +| **`Q4_K_M`** | 5.2 GB | **Recommended default. Best balance of size, speed and quality.** | +| `UD-Q4_K_XL` | 5.2 GB | Dynamic. Embeddings and output kept at Q8_0 for higher quality at a Q4 footprint. | +| `Q5_K_S` | 5.9 GB | Higher quality, slightly more compact than Q5_K_M. | +| `Q5_K_M` | 6.0 GB | Higher quality, low loss. | +| `Q6_K` | 7.0 GB | Near lossless, noticeably lighter than Q8_0. | +| `Q8_0` | 9.0 GB | Effectively lossless, reference quality. | + +> [!TIP] +> Pick the largest file that fits your (V)RAM with room for context. `Q4_K_M` or `UD-Q4_K_XL` is the sweet spot for most setups; `Q6_K` or `Q8_0` for maximum fidelity. + +## Get started + +Run LFM2.5 8B A1B locally with: + +- **[Atomic Chat](https://atomic.chat):** the easiest path. Open the app, search `AtomicChat/lfm25-8b-a1b-GGUF`, pick a quant, hit **Use this model**. +- **llama.cpp:** `llama-server -hf AtomicChat/lfm25-8b-a1b-GGUF:Q4_K_M --jinja -c 8192` +- **Ollama:** `ollama run hf.co/AtomicChat/lfm25-8b-a1b-GGUF:Q4_K_M` +- **LM Studio / Jan:** search the repo id, download any quant. + +## Best practices + +| Parameter | Value | +|---|---| +| temperature | 0.2 | +| top_k | 80 | +| repetition_penalty | 1.05 | + +Liquid AI's recommended sampling configuration for `LiquidAI/LFM2.5-8B-A1B`. + +## Run in llama.cpp + +```bash +git clone https://github.com/ggml-org/llama.cpp +cmake llama.cpp -B llama.cpp/build -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=ON +cmake --build llama.cpp/build --config Release -j --target llama-cli llama-server +``` + +```bash +./llama.cpp/build/bin/llama-server \ + -hf AtomicChat/lfm25-8b-a1b-GGUF:Q4_K_M \ + --jinja -ngl 99 -c 8192 -fa on +``` + +## How these were made + +1. Download `LiquidAI/LFM2.5-8B-A1B` (original weights). +2. Convert to f16 GGUF with [llama.cpp](https://github.com/ggml-org/llama.cpp). +3. Build an importance matrix over our calibration corpus. +4. Quantize the ladder with `--imatrix`. +5. `UD-Q4_K_XL` additionally pins the token-embedding and output tensors to `Q8_0`. + +## License + +Original model by Liquid AI, released under the other license. Full terms: [other](LICENSE). Quantized by Atomic Chat. diff --git a/bench.json b/bench.json new file mode 100644 index 0000000..c8c06e3 --- /dev/null +++ b/bench.json @@ -0,0 +1,6 @@ +{ + "model": "lfm25-8b-a1b", "chunks": 40, "f16_ppl": "36.0986", "unsloth_repo": "unsloth/LFM2.5-8B-A1B-GGUF", + "rows": [ + {"quant":"Q4_K_M","our_kld":"0.126876","our_agree":"82.896","our_ppl":"2.792977","uns_kld":"0.074829","uns_agree":"87.029","uns_ppl":"0.225532"} + ] +} diff --git a/benchmark.png b/benchmark.png new file mode 100644 index 0000000..0cc4de8 Binary files /dev/null and b/benchmark.png differ diff --git a/hero.png b/hero.png new file mode 100644 index 0000000..39823b4 Binary files /dev/null and b/hero.png differ diff --git a/lfm25-8b-a1b-IQ3_M.gguf b/lfm25-8b-a1b-IQ3_M.gguf new file mode 100644 index 0000000..5001ae8 --- /dev/null +++ b/lfm25-8b-a1b-IQ3_M.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:2b007291d6ffa0b987966f85eeba2505ebd4f6d6717eb64f70d9bbac6a74f16e +size 3778752032 diff --git a/lfm25-8b-a1b-IQ4_XS.gguf b/lfm25-8b-a1b-IQ4_XS.gguf new file mode 100644 index 0000000..29a16de --- /dev/null +++ b/lfm25-8b-a1b-IQ4_XS.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:025bbb71cf0d43e45a718fd92897233d820faa06b52b4ba52b7c22d1a2999a29 +size 4588301856 diff --git a/lfm25-8b-a1b-Q2_K.gguf b/lfm25-8b-a1b-Q2_K.gguf new file mode 100644 index 0000000..33021cd --- /dev/null +++ b/lfm25-8b-a1b-Q2_K.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:4bafc634954578c9e68aef058656b4fa7da639cc4baade0cdce6f6af588eb6f2 +size 3190435360 diff --git a/lfm25-8b-a1b-Q3_K_L.gguf b/lfm25-8b-a1b-Q3_K_L.gguf new file mode 100644 index 0000000..1025b55 --- /dev/null +++ b/lfm25-8b-a1b-Q3_K_L.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:ed612ef1055216c3789144926d688b63dbae4b51bb9ad46010db955334e716c3 +size 4436864544 diff --git a/lfm25-8b-a1b-Q3_K_M.gguf b/lfm25-8b-a1b-Q3_K_M.gguf new file mode 100644 index 0000000..99e6114 --- /dev/null +++ b/lfm25-8b-a1b-Q3_K_M.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:593b211efe50169e13f4ec81a0889956d9024e61bd3a162d7635928665d2bf38 +size 4108398112 diff --git a/lfm25-8b-a1b-Q4_K_M.gguf b/lfm25-8b-a1b-Q4_K_M.gguf new file mode 100644 index 0000000..68e5c93 --- /dev/null +++ b/lfm25-8b-a1b-Q4_K_M.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:72955dd6c06a8833e660af0aa1cc9a4d722cdc92abe94a61715ae7b866d67bf2 +size 5155565088 diff --git a/lfm25-8b-a1b-Q4_K_S.gguf b/lfm25-8b-a1b-Q4_K_S.gguf new file mode 100644 index 0000000..498321a --- /dev/null +++ b/lfm25-8b-a1b-Q4_K_S.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:c6feceb7b07b4252f72c158c2394eae3392304af805900841013934fd3ce5831 +size 4863553056 diff --git a/lfm25-8b-a1b-Q5_K_M.gguf b/lfm25-8b-a1b-Q5_K_M.gguf new file mode 100644 index 0000000..85b75a8 --- /dev/null +++ b/lfm25-8b-a1b-Q5_K_M.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:d547400a4486771d93ca195204965b172ae2c0ca2dbd3344fa7f44b4b2c716c3 +size 6030339616 diff --git a/lfm25-8b-a1b-Q5_K_S.gguf b/lfm25-8b-a1b-Q5_K_S.gguf new file mode 100644 index 0000000..afb3ab9 --- /dev/null +++ b/lfm25-8b-a1b-Q5_K_S.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:992df0d7b33873c3c79033de757b18f46ac7f0184c9c1ae92bbcd1ee1a46a757 +size 5870186016 diff --git a/lfm25-8b-a1b-Q6_K.gguf b/lfm25-8b-a1b-Q6_K.gguf new file mode 100644 index 0000000..b13ed14 --- /dev/null +++ b/lfm25-8b-a1b-Q6_K.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:1e9fbc921e3d7f93ed81db80cce917dbc8d142a1d28c365c95abe30cce9050b9 +size 6959787552 diff --git a/lfm25-8b-a1b-Q8_0.gguf b/lfm25-8b-a1b-Q8_0.gguf new file mode 100644 index 0000000..e192f2b --- /dev/null +++ b/lfm25-8b-a1b-Q8_0.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:67b3b20eaab2d0eb87eb971ceb1e0d7b229c8bc3b90cbb46d33f0ecfa5b17e63 +size 9010196000 diff --git a/lfm25-8b-a1b-UD-Q4_K_XL.gguf b/lfm25-8b-a1b-UD-Q4_K_XL.gguf new file mode 100644 index 0000000..7fc6fca --- /dev/null +++ b/lfm25-8b-a1b-UD-Q4_K_XL.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:283b12943743c5bb54b7f9fc8f7076c8ea32d163a06fb0e6f525178e7232c588 +size 5219053088 diff --git a/pill_atomic_v3.png b/pill_atomic_v3.png new file mode 100644 index 0000000..d8f38e5 Binary files /dev/null and b/pill_atomic_v3.png differ diff --git a/pill_discord_v3.png b/pill_discord_v3.png new file mode 100644 index 0000000..43c5083 Binary files /dev/null and b/pill_discord_v3.png differ diff --git a/pill_github_v3.png b/pill_github_v3.png new file mode 100644 index 0000000..c1d320e Binary files /dev/null and b/pill_github_v3.png differ