97 lines
4.2 KiB
Markdown
97 lines
4.2 KiB
Markdown
---
|
||
base_model: meta-llama/Llama-3.2-1B-Instruct
|
||
library_name: transformers
|
||
tags:
|
||
- function-calling
|
||
- tool-use
|
||
- llama
|
||
- lora
|
||
- bfcl
|
||
license: llama3.2
|
||
language:
|
||
- en
|
||
---
|
||
|
||
# Llama-3.2-1B-Instruct — Function Calling (SFT on xLAM + distilabel)
|
||
|
||
**Code & writeup:** https://github.com/keitake123/llama-function-calling-study
|
||
|
||
Supervised fine-tune of `meta-llama/Llama-3.2-1B-Instruct` on a larger merged dataset (xLAM + Argilla distilabel). This is the **data-scaling ablation** in a four-model function-calling study — testing whether adding ~19k additional synthetic examples to xLAM improves results.
|
||
|
||
## Results (BFCL v4)
|
||
|
||
| Category | Base | v1 (xLAM only) | This model (v2, +distilabel) |
|
||
|---|---|---|---|
|
||
| simple_python | 75.0 | 77.5 | 79.5 |
|
||
| multiple | 50.5 | 74.0 | 77.5 |
|
||
| live_simple | 31.8 | 57.0 | 58.5 |
|
||
| live_multiple | 7.3 | 38.8 | 43.5 |
|
||
| live_relevance | 43.8 | 93.8 | 87.5 |
|
||
| live_irrelevance | 67.3 | 16.9 | 22.2 |
|
||
| irrelevance | 35.8 | 5.8 | 6.2 |
|
||
| parallel | 44.0 | 1.0 | 0.0 |
|
||
| parallel_multiple | 15.0 | 2.0 | 0.0 |
|
||
|
||
**Finding:** Adding more (mixed-quality) data gave marginal call-required gains (+1.5 to +4.7 pts) but **did not fix the irrelevance regression** (still ~6%). Volume is not the fix for irrelevance — see the [Path B model](https://huggingface.co/Keitsuna123/llama-3.2-1b-fc-pathb) for the effective approach. Note: ~12% of the distilabel additions had malformed JSON and were filtered out during preprocessing.
|
||
|
||
## Training details
|
||
|
||
- **Base:** meta-llama/Llama-3.2-1B-Instruct
|
||
- **Data:** xLAM + argilla/apigen (distilabel), filtered to 45k single-call examples
|
||
- **Method:** LoRA (r=16, α=32), 1 epoch, effective batch 16, lr 2e-4 cosine, bf16
|
||
- **Hardware:** 1× A100, ~98 min · **final eval loss:** 0.16
|
||
|
||
## Usage
|
||
|
||
```python
|
||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||
import torch
|
||
|
||
tokenizer = AutoTokenizer.from_pretrained("Keitsuna123/llama-3.2-1b-fc-sft-v2-merged")
|
||
model = AutoModelForCausalLM.from_pretrained(
|
||
"Keitsuna123/llama-3.2-1b-fc-sft-v2-merged", torch_dtype=torch.bfloat16, device_map="auto"
|
||
)
|
||
|
||
messages = [{"role": "user", "content": "What's the weather in Tokyo?"}]
|
||
tools = [{"type": "function", "function": {
|
||
"name": "get_weather", "description": "Get the weather for a location",
|
||
"parameters": {"type": "object", "properties": {"location": {"type": "string"}}, "required": ["location"]}
|
||
}}]
|
||
|
||
text = tokenizer.apply_chat_template(messages, tools=tools, tokenize=False, add_generation_prompt=True)
|
||
inputs = tokenizer(text, return_tensors="pt").to(model.device)
|
||
out = model.generate(**inputs, max_new_tokens=128, do_sample=False)
|
||
print(tokenizer.decode(out[0][inputs['input_ids'].shape[1]:], skip_special_tokens=True))
|
||
```
|
||
|
||
## Limitations
|
||
|
||
- Irrelevance detection still regressed vs. base. Parallel calls near-zero (single-call training only).
|
||
|
||
## Part of a series
|
||
|
||
| Model | Description |
|
||
|---|---|
|
||
| [fc-sft-full](https://huggingface.co/Keitsuna123/llama-3.2-1b-fc-sft-full) | v1 SFT on xLAM |
|
||
| [fc-sft-v2-merged](https://huggingface.co/Keitsuna123/llama-3.2-1b-fc-sft-v2-merged) | SFT on xLAM + distilabel (this model) |
|
||
| [fc-grpo](https://huggingface.co/Keitsuna123/llama-3.2-1b-fc-grpo) | GRPO for irrelevance recovery (negative result) |
|
||
| [fc-pathb](https://huggingface.co/Keitsuna123/llama-3.2-1b-fc-pathb) | Supervised refusal training (effective irrelevance fix) |
|
||
|
||
## References & Acknowledgements
|
||
|
||
- **Base model:** Llama 3.2 (Meta AI) — [meta-llama/Llama-3.2-1B-Instruct](https://huggingface.co/meta-llama/Llama-3.2-1B-Instruct)
|
||
- **Training data:** Salesforce xLAM — [xlam-function-calling-60k](https://huggingface.co/datasets/Salesforce/xlam-function-calling-60k); Argilla APIGen — [argilla/apigen-function-calling](https://huggingface.co/datasets/argilla/apigen-function-calling)
|
||
- **Benchmark:** Berkeley Function Calling Leaderboard (BFCL) — [Gorilla project](https://github.com/ShishirPatil/gorilla)
|
||
- **Frameworks:** HuggingFace TRL, PEFT, Transformers
|
||
|
||
## Citation
|
||
|
||
```bibtex
|
||
@misc{taketsuna2026_fc_smallmodel,
|
||
title = {Small-Model Function Calling: Comparing SFT, Data Scaling, GRPO, and Supervised Refusal Training},
|
||
author = {Taketsuna, Keiichi},
|
||
year = {2026},
|
||
howpublished = {\url{https://huggingface.co/Keitsuna123}},
|
||
note = {Llama-3.2-1B function-calling post-training study on BFCL v4}
|
||
}
|
||
``` |