Files
minicpm5-1B-GLM-5.2-Agentic-v8/README.md
ModelHub XC 181efe2e43 初始化项目,由ModelHub XC社区提供模型
Model: hudsongouge/minicpm5-1B-GLM-5.2-Agentic-v8
Source: Original Platform
2026-07-24 07:36:11 +08:00

49 lines
1.4 KiB
Markdown

---
license: other
language:
- en
base_model:
- openbmb/MiniCPM5-1B
tags:
- agentic
- coding
- agent
- gguf
---
# MiniCPM5-1B-Agentic-v8
Created by GLM-5.2. Model 8 of 8 in the agentic post-training series.
**Variant**: 4-way Soup (best overall)
## Evaluation
| Metric | Score |
|---|---|
| Real-World Tasks | 33.3% (4/12) |
| Unique Tasks Solved | 8/12 |
| Consistent (5/5) | rw_http_server |
| Often (4/5) | rw_venv_setup |
GGUFs available: f16, q8_0, q5_k_m, q4_k_m, q3_k_m, q2_k
## Quantization Recommendations
This is a 1B model — heavier quantization degrades output quality significantly.
| Quant | Quality | Size | Recommendation |
|---|---|---|---|
| f16 | Full | ~2.1GB | Best quality |
| q8_0 | Excellent | ~1.1GB | **Recommended** — near-identical to f16 |
| q5_k_m | Good | ~0.8GB | Reasoning OK, response may degrade on longer outputs |
| q4_k_m | Fair | ~0.7GB | Reasoning OK, response degrades into repetition |
| q3_k_m | Poor | ~0.6GB | Not recommended |
| q2_k | Poor | ~0.5GB | Not recommended |
**For production use, prefer q8_0 or f16.** The model uses reasoning tokens; lower quantizations break the transition from reasoning to response.
## Chat Template
The GGUF chat template defaults to `enable_thinking=true`, so the model will always produce reasoning followed by response. If your inference engine supports `enable_thinking=false`, you can skip reasoning for faster responses.