初始化项目,由ModelHub XC社区提供模型

Model: Harish241412/qwen2.5-1.5b-toolcalling-dpo
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-29 19:28:17 +08:00
commit b8840d7f75
8 changed files with 536 additions and 0 deletions

36
.gitattributes vendored Normal file
View File

@@ -0,0 +1,36 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
tokenizer.json filter=lfs diff=lfs merge=lfs -text

336
README.md Normal file
View File

@@ -0,0 +1,336 @@
---
library_name: transformers
base_model: Qwen/Qwen2.5-1.5B-Instruct
pipeline_tag: text-generation
tags:
- qwen
- qwen2.5
- tool-calling
- function-calling
- dpo
- direct-preference-optimization
- preference-tuning
- when2call
- transformers
- pytorch
license: apache-2.0
datasets:
- nvidia/When2Call
---
# Qwen2.5-1.5B Tool Calling DPO
A DPO fine-tuned version of [Qwen/Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct), trained on NVIDIA's [When2Call](https://huggingface.co/datasets/nvidia/When2Call) preference dataset to improve tool-calling decision making.
The model is trained to better distinguish between requests that **require a tool call**, requests that can be answered directly, and requests that cannot be answered using the available tools.
> This model is based on Qwen2.5-1.5B-Instruct, which is licensed under Apache 2.0. The model was fine-tuned using NVIDIA's When2Call dataset, which is licensed under CC BY 4.0. The dataset license and attribution requirements apply to the use of the When2Call dataset.
## Model Details
| Property | Value |
| ---------------- | --------------------------- |
| Base model | Qwen/Qwen2.5-1.5B-Instruct |
| Parameters | 1.5B |
| Fine-tuning | LoRA + DPO |
| Dataset | NVIDIA When2Call |
| Training split | `train_pref` |
| Final checkpoint | Step 1000 |
| Model format | Merged |
| Framework | PyTorch, Transformers, PEFT |
The model was initially fine-tuned using LoRA and Direct Preference Optimization (DPO). The LoRA adapter was subsequently merged into the base model using `merge_and_unload()` to produce a standalone model.
## Dataset
Training was performed using the **`train_pref` split of NVIDIA's When2Call dataset**.
When2Call is designed specifically to evaluate and train LLMs on decisions about **when (and when not) to call tools**. The dataset includes preference pairs consisting of a chosen and rejected response for a given user request and tool specification.
The `train_pref` split contains **9,000 preference-training examples** with:
* Tool specifications
* User messages
* Chosen responses
* Rejected responses
NVIDIA provides both an SFT dataset (`train_sft`) and a preference dataset (`train_pref`); this model uses the **preference dataset for DPO training**.
The When2Call dataset is synthetic and is licensed under **CC BY 4.0**.
## Training Objective
The objective was to improve the model's **tool-use decision boundary**.
The model learns to distinguish between:
1. **Tool Call** — a tool should be invoked to answer the request.
2. **Request for Information** — the request can be handled without invoking a tool.
3. **Cannot Answer** — the available tools cannot answer the request.
The preference-training setup encourages the model to prefer appropriate responses over incorrect tool-calling behavior.
## Evaluation
The fine-tuned model was evaluated against the original Qwen2.5-1.5B-Instruct model on **300 samples**.
The evaluation uses the same 300-sample LLM-as-a-judge subset provided by When2Call. NVIDIA's dataset contains a larger 3,652-sample MCQ test set and a 300-sample LLM-as-a-judge subset.
### Results
| Metric | Qwen2.5-1.5B-Instruct | Qwen2.5-1.5B + DPO |
| ------------------------ | --------------------: | -----------------: |
| Intent Accuracy | 52.7% | **74.0%** |
| Tool Precision | 40.8% | **72.9%** |
| Tool Recall | **93.0%** | 35.0% |
| Tool F1 | **56.7%** | 47.3% |
| Argument F1 | **71.9%** | 67.7% |
| Unsupported Tool Calls ↓ | 45.0% | **4.3%** |
| Missed Tool Calls ↓ | **2.3%** | 21.7% |
| Throughput | 38.3 tok/s | 26.4 tok/s |
### Key Results
The largest improvement was in **unsupported tool calls**:
**45.0% → 4.3%**
This indicates that DPO substantially reduced inappropriate or hallucinated tool invocations.
Intent accuracy also increased:
**52.7% → 74.0%**
and tool precision increased:
**40.8% → 72.9%**
However, this came with a substantial reduction in tool recall:
**93.0% → 35.0%**
and Tool F1:
**56.7% → 47.3%**
Therefore, the main behavioral change is a shift toward a **more conservative, precision-oriented tool-calling policy**.
## Confusion Matrix
### Qwen2.5-1.5B-Instruct
| Ground Truth \ Prediction | Tool | Request | Refusal |
| ------------------------- | ---: | ------: | ------: |
| Tool Call | 93 | 3 | 4 |
| Request For Information | 76 | 23 | 1 |
| Cannot Answer | 59 | 22 | 19 |
### Qwen2.5-1.5B + DPO
| Ground Truth \ Prediction | Tool | Request | Refusal |
| ------------------------- | ---: | ------: | ------: |
| Tool Call | 35 | 65 | 0 |
| Request For Information | 12 | 86 | 2 |
| Cannot Answer | 1 | 91 | 8 |
The confusion matrix shows that DPO significantly reduced the model's tendency to issue tool calls for requests that should not result in a tool invocation.
## Tool Calling Example
### Tool Definition
```python
tools = [{
"name": "get_stock_price",
"description": "Fetch real-time stock price for a given ticker symbol.",
"parameters": {
"type": "object",
"properties": {
"ticker": {
"type": "string",
"description": "The ticker symbol (e.g., AAPL, NVDA)"
}
},
"required": ["ticker"]
}
}]
```
### User
```text
Can you check Nvidia's current stock price?
```
### Model Output
```text
<tool_call>
{"name": "get_stock_price", "arguments": {"ticker": "NVDA"}}
</tool_call>
```
## Inference
```python
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "YOUR_USERNAME/YOUR_MODEL_NAME"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.float16,
device_map="auto"
)
model.eval()
tools = [{
"name": "get_stock_price",
"description": "Fetch real-time stock price for a given ticker symbol.",
"parameters": {
"type": "object",
"properties": {
"ticker": {
"type": "string",
"description": "The ticker symbol (e.g., AAPL, NVDA)"
}
},
"required": ["ticker"]
}
}]
messages = [
{
"role": "user",
"content": "Can you check Nvidia's current stock price?"
}
]
prompt = tokenizer.apply_chat_template(
messages,
tools=tools,
tokenize=False,
add_generation_prompt=True
)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=128,
do_sample=False,
eos_token_id=tokenizer.eos_token_id,
pad_token_id=tokenizer.pad_token_id or tokenizer.eos_token_id
)
response = tokenizer.decode(
outputs[0][inputs.input_ids.shape[1]:],
skip_special_tokens=False
)
print(response)
```
## Output Post-Processing
During evaluation, model outputs were passed through a lightweight post-processing step to normalize tool-call formatting.
The post-processing:
1. Extracts JSON containing `name` and `arguments`.
2. Removes duplicate or nested `<tool_call>` wrappers.
3. Normalizes the output to:
```text
<tool_call>
{"name": "...", "arguments": {...}}
</tool_call>
```
If no tool-call JSON is detected, the output is treated as a normal conversational response.
The post-processing step is used for **format normalization and evaluation** and does not generate a tool call that the model did not produce.
## Limitations
The main limitation is the precision-recall trade-off introduced by DPO.
The model is considerably better at avoiding unsupported tool calls, but it also misses a larger proportion of valid tool-call opportunities.
Therefore, this model should not be interpreted as universally better than the base model for tool calling. Instead, it demonstrates that preference optimization can strongly shift the **tool-use decision policy** of a small instruction-tuned model.
The evaluation also uses a relatively small 300-sample subset, so additional evaluation on larger and more diverse tool-calling benchmarks would be useful.
## Intended Use
This model is intended for research and experimentation involving:
* Tool calling
* Function calling
* Tool-selection policies
* Preference optimization
* DPO
* Agentic LLM systems
* Small language model alignment
It is **not intended to be considered production-ready** without additional task-specific evaluation.
## Future Work
* Recover tool-call recall while maintaining low unsupported-call rates
* Experiment with DPO hyperparameters
* Improve preference-data construction
* Compare DPO against SFT
* Evaluate larger Qwen models
* Evaluate multi-tool selection
* Evaluate multi-step tool calling
* Improve argument-generation accuracy
* Benchmark inference using vLLM
* Evaluate on larger tool-calling benchmarks
## Base Model
This model is based on:
**Qwen/Qwen2.5-1.5B-Instruct**
Please refer to the base model for its original capabilities, license, and usage restrictions.
## Dataset Citation
This work uses NVIDIA's **When2Call** dataset.
> Ross, Hayley, Ameya Sunil Mahabaleshwarka, and Yoshi Suhara. "When2Call: When (not) to Call Tools." NAACL 2025.
```bibtex
@inproceedings{ross-etal-2025-when2call,
title = "{W}hen2{C}all: When (not) to Call Tools",
author = "Ross, Hayley and
Mahabaleshwarkar, Ameya Sunil and
Suhara, Yoshi",
editor = "Chiruzzo, Luis and
Ritter, Alan and
Wang, Lu",
booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
month = apr,
year = "2025",
address = "Albuquerque, New Mexico",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.naacl-long.174/",
doi = "10.18653/v1/2025.naacl-long.174",
pages = "3391--3409",
ISBN = "979-8-89176-189-6",
abstract = "Leveraging external tools is a key feature for modern Language Models (LMs) to expand their capabilities and integrate them into existing systems. However, existing benchmarks primarily focus on the accuracy of tool calling{---}whether the correct tool is called with the correct parameters{---}and less on evaluating when LMs should (not) call tools. We develop a new benchmark, When2Call, which evaluates tool-calling decision-making: when to generate a tool call, when to ask follow-up questions and when to admit the question can{'}t be answered with the tools provided. We find that state-of-the-art tool-calling LMs show significant room for improvement on When2Call, indicating the importance of this benchmark. We also develop a training set for When2Call and leverage the multiple-choice nature of the benchmark to develop a preference optimization training regime, which shows considerably more improvement than traditional fine-tuning. We release the benchmark and training data as well as evaluation scripts."
}
```
## Acknowledgements
Thanks to the NVIDIA When2Call authors and the Qwen team for releasing the dataset and base model used in this experiment.

54
chat_template.jinja Normal file
View File

@@ -0,0 +1,54 @@
{%- if tools %}
{{- '<|im_start|>system\n' }}
{%- if messages[0]['role'] == 'system' %}
{{- messages[0]['content'] }}
{%- else %}
{{- 'You are Qwen, created by Alibaba Cloud. You are a helpful assistant.' }}
{%- endif %}
{{- "\n\n# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
{%- for tool in tools %}
{{- "\n" }}
{{- tool | tojson }}
{%- endfor %}
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
{%- else %}
{%- if messages[0]['role'] == 'system' %}
{{- '<|im_start|>system\n' + messages[0]['content'] + '<|im_end|>\n' }}
{%- else %}
{{- '<|im_start|>system\nYou are Qwen, created by Alibaba Cloud. You are a helpful assistant.<|im_end|>\n' }}
{%- endif %}
{%- endif %}
{%- for message in messages %}
{%- if (message.role == "user") or (message.role == "system" and not loop.first) or (message.role == "assistant" and not message.tool_calls) %}
{{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
{%- elif message.role == "assistant" %}
{{- '<|im_start|>' + message.role }}
{%- if message.content %}
{{- '\n' + message.content }}
{%- endif %}
{%- for tool_call in message.tool_calls %}
{%- if tool_call.function is defined %}
{%- set tool_call = tool_call.function %}
{%- endif %}
{{- '\n<tool_call>\n{"name": "' }}
{{- tool_call.name }}
{{- '", "arguments": ' }}
{{- tool_call.arguments | tojson }}
{{- '}\n</tool_call>' }}
{%- endfor %}
{{- '<|im_end|>\n' }}
{%- elif message.role == "tool" %}
{%- if (loop.index0 == 0) or (messages[loop.index0 - 1].role != "tool") %}
{{- '<|im_start|>user' }}
{%- endif %}
{{- '\n<tool_response>\n' }}
{{- message.content }}
{{- '\n</tool_response>' }}
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
{{- '<|im_end|>\n' }}
{%- endif %}
{%- endif %}
{%- endfor %}
{%- if add_generation_prompt %}
{{- '<|im_start|>assistant\n' }}
{%- endif %}

61
config.json Normal file
View File

@@ -0,0 +1,61 @@
{
"architectures": [
"Qwen2ForCausalLM"
],
"attention_dropout": 0.0,
"bos_token_id": 151643,
"dtype": "float16",
"eos_token_id": 151645,
"hidden_act": "silu",
"hidden_size": 1536,
"initializer_range": 0.02,
"intermediate_size": 8960,
"layer_types": [
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention"
],
"max_position_embeddings": 32768,
"max_window_layers": 21,
"model_type": "qwen2",
"num_attention_heads": 12,
"num_hidden_layers": 28,
"num_key_value_heads": 2,
"pad_token_id": null,
"rms_norm_eps": 1e-06,
"rope_parameters": {
"rope_theta": 1000000.0,
"rope_type": "default"
},
"sliding_window": null,
"tie_word_embeddings": true,
"transformers_version": "5.0.0",
"use_cache": true,
"use_sliding_window": false,
"vocab_size": 151936
}

14
generation_config.json Normal file
View File

@@ -0,0 +1,14 @@
{
"bos_token_id": 151643,
"do_sample": true,
"eos_token_id": [
151645,
151643
],
"pad_token_id": 151643,
"repetition_penalty": 1.1,
"temperature": 0.7,
"top_k": 20,
"top_p": 0.8,
"transformers_version": "5.0.0"
}

3
model.safetensors Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:247e0b10761b18bbac6b60bbd80f2f05ac302bf19c8dab8c1e559d7004b3a814
size 3087466808

3
tokenizer.json Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3fd169731d2cbde95e10bf356d66d5997fd885dd8dbb6fb4684da3f23b2585d8
size 11421892

29
tokenizer_config.json Normal file
View File

@@ -0,0 +1,29 @@
{
"add_prefix_space": false,
"backend": "tokenizers",
"bos_token": null,
"clean_up_tokenization_spaces": false,
"eos_token": "<|im_end|>",
"errors": "replace",
"extra_special_tokens": [
"<|im_start|>",
"<|im_end|>",
"<|object_ref_start|>",
"<|object_ref_end|>",
"<|box_start|>",
"<|box_end|>",
"<|quad_start|>",
"<|quad_end|>",
"<|vision_start|>",
"<|vision_end|>",
"<|vision_pad|>",
"<|image_pad|>",
"<|video_pad|>"
],
"is_local": true,
"model_max_length": 131072,
"pad_token": "<|endoftext|>",
"split_special_tokens": false,
"tokenizer_class": "Qwen2Tokenizer",
"unk_token": null
}