47 lines
1.3 KiB
Markdown
47 lines
1.3 KiB
Markdown
---
|
|
license: other
|
|
language:
|
|
- en
|
|
base_model:
|
|
- openbmb/MiniCPM5-1B
|
|
tags:
|
|
- agentic
|
|
- coding
|
|
- agent
|
|
- gguf
|
|
---
|
|
|
|
# MiniCPM5-1B-Agentic-v7
|
|
|
|
Created by GLM-5.2. Model 7 of 8 in the agentic post-training series.
|
|
|
|
**Variant**: 3-way Soup
|
|
|
|
## Evaluation
|
|
|
|
| Metric | Score |
|
|
|---|---|
|
|
| Real-World Tasks | 28.7% (3.44/12) |
|
|
| Unique Tasks Solved | 8/12 |
|
|
|
|
GGUFs available: f16, q8_0, q5_k_m, q4_k_m, q3_k_m, q2_k
|
|
|
|
## Quantization Recommendations
|
|
|
|
This is a 1B model — heavier quantization degrades output quality significantly.
|
|
|
|
| Quant | Quality | Size | Recommendation |
|
|
|---|---|---|---|
|
|
| f16 | Full | ~2.1GB | Best quality |
|
|
| q8_0 | Excellent | ~1.1GB | **Recommended** — near-identical to f16 |
|
|
| q5_k_m | Good | ~0.8GB | Reasoning OK, response may degrade on longer outputs |
|
|
| q4_k_m | Fair | ~0.7GB | Reasoning OK, response degrades into repetition |
|
|
| q3_k_m | Poor | ~0.6GB | Not recommended |
|
|
| q2_k | Poor | ~0.5GB | Not recommended |
|
|
|
|
**For production use, prefer q8_0 or f16.** The model uses reasoning tokens; lower quantizations break the transition from reasoning to response.
|
|
|
|
## Chat Template
|
|
|
|
The GGUF chat template defaults to `enable_thinking=true`, so the model will always produce reasoning followed by response. If your inference engine supports `enable_thinking=false`, you can skip reasoning for faster responses.
|