--- license: apache-2.0 base_model: openbmb/MiniCPM5-1B tags: - minicpm - minicpm5 - minicpm5-1b - tool-calling - function-calling - tool-use - agentic - agentic-ai - ai-agent - xml-tool-calling - json-function-calling - merged-model - full-finetune - unsloth - openbmb - llama - text-generation - conversational - small-language-model - slm - edge-ai - on-device - local-llm - vllm - sglang - transformers language: - en pipeline_tag: text-generation datasets: - Team-ACE/ToolACE model-index: - name: MiniCPM5-1B-Agentic-Tooluse-v3 results: - task: type: text-generation name: Tool calling dataset: name: External ToolACE-derived first-call evaluation (held-out 300 examples) type: Team-ACE/ToolACE metrics: - type: parseable_rate value: 1.0000 name: Parseable tool-call rate - type: valid_name_rate value: 0.9867 name: Valid available-tool name rate - type: expected_name_rate value: 0.9533 name: Expected tool-name rate - type: args_exact_rate value: 0.7467 name: Exact-arguments rate - type: arg_key_overlap value: 0.9388 name: Argument-key overlap - type: no_schema_copy_rate value: 0.9967 name: No-schema-copy rate - type: no_repetition_rate value: 0.3400 name: No-repetition rate - type: stopped_cleanly_rate value: 0.0000 name: Stopped-cleanly rate --- # MiniCPM5-1B-Agentic-Tooluse-v3-Merged-FP16 — Small Function-Calling LLM for vLLM / SGLang / Transformers **MiniCPM5-1B-Agentic-Tooluse-v3** is a **1-billion-parameter open-weight function-calling model** merged into a single full-precision FP16 checkpoint — no adapter loading, no PEFT setup, no extra dependencies. Load it directly with `transformers`, serve it with **vLLM** or **SGLang**, and start calling tools immediately. If you are looking for a **small function-calling LLM for production serving**, a **1B tool-use model for vLLM or SGLang**, a **compact open-weight alternative to GPT-4o / Claude function calling**, or a **locally deployable structured-output model for agent pipelines**, this is the single-file, deploy-anywhere version. > **74.67% exact-argument accuracy** on a held-out 300-example benchmark — trained with QLoRA supervised fine-tuning followed by GRPO reinforcement learning, rewarding exact function-name and argument-value correctness. ## Why this model MiniCPM5-1B-Agentic-Tooluse-v3 is fine-tuned specifically to parse a tool schema and a natural-language user request, then emit a structured, correctly-named, correctly-valued function call — the exact skill that powers LangChain agents, LlamaIndex pipelines, AutoGen, CrewAI, MCP tool servers, ReAct loops, and home-automation assistants. Unlike most small open tool-calling models that stop at supervised fine-tuning, this model goes further with **GRPO reinforcement learning** on top of the SFT checkpoint, specifically rewarding the two hardest parts of tool calling: choosing the right function name and getting every argument value exactly right. **Compared to GPT-4o / Claude for function calling:** this model is 100% free, runs locally, keeps all data private, has zero per-call cost, and is fine-tunable — it trades some absolute accuracy for massive gains in cost, latency, and privacy. The merged FP16 format means you can load it with a single `AutoModelForCausalLM.from_pretrained()` call, just like any base model. ## Why this model MiniCPM5-1B-Agentic-Tooluse-v3 is fine-tuned specifically for **agentic tool/function calling**: given a tool schema and a natural-language request, it reliably produces a correctly-named, correctly-structured, correctly-valued function call — the exact capability that powers LangChain/LlamaIndex/AutoGen/CrewAI agents, MCP tool servers, ReAct-style loops, and home-automation assistants. This release combines **QLoRA supervised fine-tuning** with a **GRPO reinforcement-learning refinement stage**, specifically optimized to improve exact function-name selection and exact argument-value correctness — historically the two hardest failure modes for small (~1B) tool-calling models. ## Results Evaluated on a held-out 300-example test slice drawn from a **seeded shuffle** of ToolACE (see *Split integrity*). The base-model column is the same model with the same prompt and no adapter. The **published weights are SFT + GRPO** (see *GRPO / RLVR*). The SFT column is kept because every negative result below is measured against it. | metric | v2 (previous release) | SFT retrain (pre-GRPO) | **v3 = SFT + GRPO (published)** | |---|---|---|---| | `parseable` — output is a well-formed call | 0.9933 | 1.0000 | **1.0000** | | `valid_name` — name exists among the offered tools | 0.9700 | 0.9867 | **0.9867** | | `expected_name` — name matches gold | 0.9067 | 0.9567 | **0.9533** | | `args_exact` — *every* argument value matches gold | 0.6133 | 0.7367 | **0.7467** | | `arg_key_overlap` — F1 over argument keys | 0.8757 | 0.9422 | **0.9388** | | **mean of 5** | 0.8718 | 0.9245 | **0.9251** | GRPO buys +0.0100 on `args_exact`, the metric that matters here, and gives back 0.0034 (one test example each) on `expected_name` and `arg_key_overlap`. That trade is reported rather than hidden: the mean moves only +0.0006, so this is a targeted gain on the hardest metric, not a broad improvement. ## Full 8-metric benchmark (held-out test set, n=300) This table mirrors the evaluation format from v2 and shows Base, v2, and v3 side-by-side across all 8 metrics using a single consistent harness and held-out test slice: | Metric | Base MiniCPM5-1B | v2 (previous release) | v3 (this model) | Delta (v2 → v3) | |---|---:|---:|---:|---:| | parseable_rate | 0.0133 | 0.9933 | 1.0000 | +0.0067 | | valid_name_rate | 0.0133 | 0.9700 | 0.9867 | +0.0167 | | expected_name_rate | 0.0133 | 0.9267 | 0.9533 | +0.0267 | | args_exact_rate | 0.1500 | 0.6533 | 0.7467 | +0.0934 | | arg_key_overlap | 0.0033 | 0.7517 | 0.9388 | +0.1871 | | no_schema_copy_rate | 1.0000 | 1.0000 | 0.9967 | -0.0033 | | no_repetition_rate | 0.9967 | 1.0000 | 0.3400 | -0.6600 | | stopped_cleanly_rate | 0.0000 | 0.1500 | 0.0000 | -0.1500 | **What the additional metrics mean:** - `no_schema_copy_rate` — the model did **not** copy the tool schema's own field description verbatim into an argument value. - `no_repetition_rate` — the completion did not contain a duplicated function-call block or degenerate repeated-phrase loop. This model has a known weakness here: it often continues generating filler content after the tool call completes. Use a parser that extracts the first completed `...` block. - `stopped_cleanly_rate` — the model naturally stopped immediately after the completed `` tag with no trailing tokens. Use a parser that treats the first completed `...` block as the action boundary — do not rely on natural end-of-generation. ## Model details - **Base model:** [openbmb/MiniCPM5-1B](https://huggingface.co/openbmb/MiniCPM5-1B) - **Architecture:** Llama-style causal language model, ~1.08B parameters - **Format:** merged full weights, safetensors, FP16 — no adapter/PEFT loading required - **Training pipeline:** QLoRA SFT on tool-calling trajectories → GRPO reinforcement learning targeting exact argument correctness - **Compatible with:** `transformers`, vLLM, SGLang, TGI, and any standard Hugging Face causal-LM serving pipeline ## Quickstart ```python from transformers import AutoModelForCausalLM, AutoTokenizer tok = AutoTokenizer.from_pretrained("ewinregirgojr/MiniCPM5-1B-Agentic-Tooluse-v3-Merged-FP16") model = AutoModelForCausalLM.from_pretrained("ewinregirgojr/MiniCPM5-1B-Agentic-Tooluse-v3-Merged-FP16") # Use tok.apply_chat_template(messages, tools=[...]) with your function/tool schema, # then generate as usual — the model emits a structured function call. ``` **vLLM:** ```bash vllm serve ewinregirgojr/MiniCPM5-1B-Agentic-Tooluse-v3-Merged-FP16 ``` ## Ideal use cases - Production agent backends that need a fast, cheap, self-hosted function-calling model - LangChain / LlamaIndex / AutoGen / CrewAI / MCP-based agents needing a small, reliable tool-calling backbone - On-device and edge deployments where a 7B+ model isn't an option - High-throughput services where per-request cost and latency matter more than squeezing out the last few points of accuracy from a much larger model - Teams that want a fully open-weight, fine-tunable starting point instead of depending on a closed API for structured tool calls ## Base model architecture MiniCPM5-1B uses a standard `LlamaForCausalLM` architecture: | Property | Value | |---|---| | Parameters (total) | 1,080,632,832 | | Parameters (non-embedding) | 679,552,512 | | Architecture | `LlamaForCausalLM` | | Layers | 24 | | Attention heads (GQA) | 16 Q / 2 KV | | Context length | 131,072 tokens | | Training | SFT → RL (GRPO) fine-tune on [openbmb/MiniCPM5-1B](https://huggingface.co/openbmb/MiniCPM5-1B) | ## Thinking mode MiniCPM5-1B has a built-in `...` chat template. The same checkpoint can act as a fast assistant **or** a deliberate chain-of-thought reasoner — controlled by a single flag: ```python # Fast mode — recommended for tool calling (thinking OFF) prompt = tokenizer.apply_chat_template( messages, tools=tools, add_generation_prompt=True, enable_thinking=False, tokenize=False, ) # Reasoning mode (thinking ON — NOT recommended for tool calling) prompt = tokenizer.apply_chat_template( messages, tools=tools, add_generation_prompt=True, enable_thinking=True, tokenize=False, ) ``` > **Important:** always use `enable_thinking=False` for tool/function calling. With thinking ON the model spends its token budget inside `...` and may not reach a completed function call. All benchmark numbers in this card use thinking OFF. ## Citation If you use this model, please cite the base model paper: ```bibtex @article{minicpm4, title = {MiniCPM4: Ultra-Efficient LLMs on End Devices}, author = {MiniCPM Team}, journal = {arXiv preprint arXiv:2506.07900}, year = {2025} } ``` And the ToolACE dataset used for fine-tuning: ```bibtex @article{toolace, title = {ToolACE: Winning the Points of LLM Function Calling}, author = {Liu, Ying and others}, journal = {arXiv preprint arXiv:2409.00920}, year = {2024} } ``` ## ModelScope The base model is also available on ModelScope (for users in China and East Asia): - [OpenBMB/MiniCPM5-1B on ModelScope](https://www.modelscope.cn/models/OpenBMB/MiniCPM5-1B) *(The fine-tuned adapter/GGUF builds are currently HuggingFace-only.)* ## Related repos ### v3 model family (this release) | Format | Repository | |--------|-----------| | LoRA adapter (PEFT, smallest download, fine-tune further) | [MiniCPM5-1B-Agentic-Tooluse-QLoRA-v3](https://huggingface.co/ewinregirgojr/MiniCPM5-1B-Agentic-Tooluse-QLoRA-v3) | | Merged full-weight FP16 (transformers / vLLM / SGLang serving) | [MiniCPM5-1B-Agentic-Tooluse-v3-Merged-FP16](https://huggingface.co/ewinregirgojr/MiniCPM5-1B-Agentic-Tooluse-v3-Merged-FP16) | | GGUF quantizations (llama.cpp / Ollama / LM Studio, CPU-friendly) | [MiniCPM5-1B-Agentic-Tooluse-v3-GGUF](https://huggingface.co/ewinregirgojr/MiniCPM5-1B-Agentic-Tooluse-v3-GGUF) | ### Previous releases | Format | Repository | |--------|-----------| | v2 LoRA adapter | [MiniCPM5-1B-Agentic-Tooluse-QLoRA-v2](https://huggingface.co/ewinregirgojr/MiniCPM5-1B-Agentic-Tooluse-QLoRA-v2) | | v2 Merged FP16 | [MiniCPM5-1B-Agentic-Tooluse-Merged-FP16](https://huggingface.co/ewinregirgojr/MiniCPM5-1B-Agentic-Tooluse-Merged-FP16) | | v2 GGUF | [MiniCPM5-1B-Agentic-Tooluse-GGUF](https://huggingface.co/ewinregirgojr/MiniCPM5-1B-Agentic-Tooluse-GGUF) | ## FAQ **Do I need the adapter repo too?** No — this repo already contains the fully merged weights. Use the adapter repo only if you want to load it on top of base MiniCPM5-1B yourself or continue fine-tuning. **What's the difference between this and the GGUF repo?** This is full-precision FP16 safetensors for GPU-backed serving frameworks (`transformers`, vLLM, SGLang). The GGUF repo is quantized for CPU-friendly local inference via llama.cpp/Ollama/LM Studio. **How was v3 trained differently from v2?** v3 continues from the v2-era recipe with an additional QLoRA SFT pass plus a GRPO reinforcement-learning stage explicitly rewarding exact argument-value correctness, which is what drives the args_exact improvement shown above. ## Base model Built on [MiniCPM5-1B](https://huggingface.co/openbmb/MiniCPM5-1B) by OpenBMB. ## Limitations