--- license: other language: - en base_model: - openbmb/MiniCPM5-1B tags: - agentic - coding - agent - gguf --- # MiniCPM5-1B-Agentic-v7 Created by GLM-5.2. Model 7 of 8 in the agentic post-training series. **Variant**: 3-way Soup ## Evaluation | Metric | Score | |---|---| | Real-World Tasks | 28.7% (3.44/12) | | Unique Tasks Solved | 8/12 | GGUFs available: f16, q8_0, q5_k_m, q4_k_m, q3_k_m, q2_k ## Quantization Recommendations This is a 1B model — heavier quantization degrades output quality significantly. | Quant | Quality | Size | Recommendation | |---|---|---|---| | f16 | Full | ~2.1GB | Best quality | | q8_0 | Excellent | ~1.1GB | **Recommended** — near-identical to f16 | | q5_k_m | Good | ~0.8GB | Reasoning OK, response may degrade on longer outputs | | q4_k_m | Fair | ~0.7GB | Reasoning OK, response degrades into repetition | | q3_k_m | Poor | ~0.6GB | Not recommended | | q2_k | Poor | ~0.5GB | Not recommended | **For production use, prefer q8_0 or f16.** The model uses reasoning tokens; lower quantizations break the transition from reasoning to response. ## Chat Template The GGUF chat template defaults to `enable_thinking=true`, so the model will always produce reasoning followed by response. If your inference engine supports `enable_thinking=false`, you can skip reasoning for faster responses.