101 lines
4.7 KiB
Markdown
101 lines
4.7 KiB
Markdown
|
|
---
|
|||
|
|
base_model: meta-llama/Llama-3.2-1B-Instruct
|
|||
|
|
library_name: transformers
|
|||
|
|
tags:
|
|||
|
|
- function-calling
|
|||
|
|
- tool-use
|
|||
|
|
- llama
|
|||
|
|
- lora
|
|||
|
|
- bfcl
|
|||
|
|
license: llama3.2
|
|||
|
|
language:
|
|||
|
|
- en
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# Llama-3.2-1B-Instruct — Function Calling (Refusal-SFT, "Path B")
|
|||
|
|
|
|||
|
|
**Code & writeup:** https://github.com/keitake123/llama-function-calling-study
|
|||
|
|
|
|||
|
|
The [v1 SFT model](https://huggingface.co/Keitsuna123/llama-3.2-1b-fc-sft-full) further fine-tuned on a mix of call-required examples and **targeted refusal examples**, to recover irrelevance detection. **This is the effective fix** in a four-model study — and it succeeded where GRPO did not.
|
|||
|
|
|
|||
|
|
## Results (BFCL v4)
|
|||
|
|
|
|||
|
|
| Category | Base | v1 SFT | GRPO | This model (Path B) |
|
|||
|
|
|---|---|---|---|---|
|
|||
|
|
| irrelevance | 35.8 | 5.8 | 5.4 | **21.7** |
|
|||
|
|
| live_irrelevance | 67.3 | 16.9 | 17.2 | **36.8** |
|
|||
|
|
| simple_python | 75.0 | 77.5 | 77.2 | 78.2 |
|
|||
|
|
| multiple | 50.5 | 74.0 | 74.0 | 72.0 |
|
|||
|
|
| live_simple | 31.8 | 57.0 | 56.2 | 54.3 |
|
|||
|
|
| live_multiple | 7.3 | 38.8 | 39.4 | 35.7 |
|
|||
|
|
| live_relevance | 43.8 | 93.8 | 93.8 | 81.2 |
|
|||
|
|
|
|||
|
|
**Finding:** Supervised refusal examples recovered irrelevance ~4x (5.8→21.7) and more than doubled live_irrelevance (16.9→36.8), at a modest, quantified cost to call-required categories (multiple −2, live_simple −2.7, live_multiple −3, live_relevance −12.5). This is a clean precision/recall tradeoff, and a much larger recovery than GRPO achieved (which showed no gain).
|
|||
|
|
|
|||
|
|
**Takeaway:** for recovering an out-of-distribution behavior in a small model, targeted supervised demonstration outperforms RL reward shaping — because RL requires the target behavior to appear in sampled rollouts, which a strong SFT prior prevents.
|
|||
|
|
|
|||
|
|
## Training details
|
|||
|
|
|
|||
|
|
- **Base:** v1 SFT model · **Method:** LoRA SFT
|
|||
|
|
- **Data:** 1600 examples, 50/50 mix of xLAM call-required + synthetic refusal examples (query + unrelated tools → conversational refusal)
|
|||
|
|
- **Config:** 3 epochs, effective batch 16, lr 1.2e-4 cosine, bf16 · **final eval loss:** 0.17
|
|||
|
|
|
|||
|
|
## Usage
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
from transformers import AutoModelForCausalLM, AutoTokenizer
|
|||
|
|
import torch
|
|||
|
|
|
|||
|
|
tokenizer = AutoTokenizer.from_pretrained("Keitsuna123/llama-3.2-1b-fc-pathb")
|
|||
|
|
model = AutoModelForCausalLM.from_pretrained(
|
|||
|
|
"Keitsuna123/llama-3.2-1b-fc-pathb", torch_dtype=torch.bfloat16, device_map="auto"
|
|||
|
|
)
|
|||
|
|
|
|||
|
|
# Call-required: produces a function call
|
|||
|
|
messages = [{"role": "user", "content": "What's the weather in Tokyo?"}]
|
|||
|
|
# Irrelevant: produces a conversational refusal instead of a hallucinated call
|
|||
|
|
# messages = [{"role": "user", "content": "How do I make bread fluffier?"}]
|
|||
|
|
|
|||
|
|
tools = [{"type": "function", "function": {
|
|||
|
|
"name": "get_weather", "description": "Get the weather for a location",
|
|||
|
|
"parameters": {"type": "object", "properties": {"location": {"type": "string"}}, "required": ["location"]}
|
|||
|
|
}}]
|
|||
|
|
|
|||
|
|
text = tokenizer.apply_chat_template(messages, tools=tools, tokenize=False, add_generation_prompt=True)
|
|||
|
|
inputs = tokenizer(text, return_tensors="pt").to(model.device)
|
|||
|
|
out = model.generate(**inputs, max_new_tokens=128, do_sample=False)
|
|||
|
|
print(tokenizer.decode(out[0][inputs['input_ids'].shape[1]:], skip_special_tokens=True))
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
## Limitations
|
|||
|
|
|
|||
|
|
- Residual irrelevance failures concentrate in **short, factual-seeming queries** ("what's 2+2", "capital of France") where a loosely-related tool exists — the hardest irrelevance sub-case.
|
|||
|
|
- Parallel calls remain near-zero (single-call training only).
|
|||
|
|
|
|||
|
|
## Part of a series
|
|||
|
|
|
|||
|
|
| Model | Description |
|
|||
|
|
|---|---|
|
|||
|
|
| [fc-sft-full](https://huggingface.co/Keitsuna123/llama-3.2-1b-fc-sft-full) | v1 SFT on xLAM |
|
|||
|
|
| [fc-sft-v2-merged](https://huggingface.co/Keitsuna123/llama-3.2-1b-fc-sft-v2-merged) | SFT on xLAM + distilabel (data-scaling ablation) |
|
|||
|
|
| [fc-grpo](https://huggingface.co/Keitsuna123/llama-3.2-1b-fc-grpo) | GRPO for irrelevance recovery (negative result) |
|
|||
|
|
| [fc-pathb](https://huggingface.co/Keitsuna123/llama-3.2-1b-fc-pathb) | Supervised refusal training (this model) |
|
|||
|
|
|
|||
|
|
## References & Acknowledgements
|
|||
|
|
|
|||
|
|
- **Base model:** Llama 3.2 (Meta AI) — [meta-llama/Llama-3.2-1B-Instruct](https://huggingface.co/meta-llama/Llama-3.2-1B-Instruct)
|
|||
|
|
- **Training data:** Salesforce xLAM — [xlam-function-calling-60k](https://huggingface.co/datasets/Salesforce/xlam-function-calling-60k); refusal targets generated for this study
|
|||
|
|
- **Benchmark:** Berkeley Function Calling Leaderboard (BFCL) — [Gorilla project](https://github.com/ShishirPatil/gorilla)
|
|||
|
|
- **Frameworks:** HuggingFace TRL, PEFT, Transformers
|
|||
|
|
|
|||
|
|
## Citation
|
|||
|
|
|
|||
|
|
```bibtex
|
|||
|
|
@misc{taketsuna2026_fc_smallmodel,
|
|||
|
|
title = {Small-Model Function Calling: Comparing SFT, Data Scaling, GRPO, and Supervised Refusal Training},
|
|||
|
|
author = {Taketsuna, Keiichi},
|
|||
|
|
year = {2026},
|
|||
|
|
howpublished = {\url{https://huggingface.co/Keitsuna123}},
|
|||
|
|
note = {Llama-3.2-1B function-calling post-training study on BFCL v4}
|
|||
|
|
}
|
|||
|
|
```
|