base_model, library_name, tags, license, language
base_model library_name tags license language
meta-llama/Llama-3.2-1B-Instruct transformers
function-calling
tool-use
llama
lora
bfcl
llama3.2
en

Llama-3.2-1B-Instruct — Function Calling (SFT on xLAM + distilabel)

Code & writeup: https://github.com/keitake123/llama-function-calling-study

Supervised fine-tune of meta-llama/Llama-3.2-1B-Instruct on a larger merged dataset (xLAM + Argilla distilabel). This is the data-scaling ablation in a four-model function-calling study — testing whether adding ~19k additional synthetic examples to xLAM improves results.

Results (BFCL v4)

Category Base v1 (xLAM only) This model (v2, +distilabel)
simple_python 75.0 77.5 79.5
multiple 50.5 74.0 77.5
live_simple 31.8 57.0 58.5
live_multiple 7.3 38.8 43.5
live_relevance 43.8 93.8 87.5
live_irrelevance 67.3 16.9 22.2
irrelevance 35.8 5.8 6.2
parallel 44.0 1.0 0.0
parallel_multiple 15.0 2.0 0.0

Finding: Adding more (mixed-quality) data gave marginal call-required gains (+1.5 to +4.7 pts) but did not fix the irrelevance regression (still ~6%). Volume is not the fix for irrelevance — see the Path B model for the effective approach. Note: ~12% of the distilabel additions had malformed JSON and were filtered out during preprocessing.

Training details

  • Base: meta-llama/Llama-3.2-1B-Instruct
  • Data: xLAM + argilla/apigen (distilabel), filtered to 45k single-call examples
  • Method: LoRA (r=16, α=32), 1 epoch, effective batch 16, lr 2e-4 cosine, bf16
  • Hardware: 1× A100, ~98 min · final eval loss: 0.16

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

tokenizer = AutoTokenizer.from_pretrained("Keitsuna123/llama-3.2-1b-fc-sft-v2-merged")
model = AutoModelForCausalLM.from_pretrained(
    "Keitsuna123/llama-3.2-1b-fc-sft-v2-merged", torch_dtype=torch.bfloat16, device_map="auto"
)

messages = [{"role": "user", "content": "What's the weather in Tokyo?"}]
tools = [{"type": "function", "function": {
    "name": "get_weather", "description": "Get the weather for a location",
    "parameters": {"type": "object", "properties": {"location": {"type": "string"}}, "required": ["location"]}
}}]

text = tokenizer.apply_chat_template(messages, tools=tools, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=128, do_sample=False)
print(tokenizer.decode(out[0][inputs['input_ids'].shape[1]:], skip_special_tokens=True))

Limitations

  • Irrelevance detection still regressed vs. base. Parallel calls near-zero (single-call training only).

Part of a series

Model Description
fc-sft-full v1 SFT on xLAM
fc-sft-v2-merged SFT on xLAM + distilabel (this model)
fc-grpo GRPO for irrelevance recovery (negative result)
fc-pathb Supervised refusal training (effective irrelevance fix)

References & Acknowledgements

Citation

@misc{taketsuna2026_fc_smallmodel,
  title  = {Small-Model Function Calling: Comparing SFT, Data Scaling, GRPO, and Supervised Refusal Training},
  author = {Taketsuna, Keiichi},
  year   = {2026},
  howpublished = {\url{https://huggingface.co/Keitsuna123}},
  note   = {Llama-3.2-1B function-calling post-training study on BFCL v4}
}
Description
Model synced from source: Keitsuna123/llama-3.2-1b-fc-sft-v2-merged
Readme 16 MiB
Languages
Jinja 100%