Supervised fine-tune of meta-llama/Llama-3.2-1B-Instruct on a larger merged dataset (xLAM + Argilla distilabel). This is the data-scaling ablation in a four-model function-calling study — testing whether adding ~19k additional synthetic examples to xLAM improves results.
Results (BFCL v4)
Category
Base
v1 (xLAM only)
This model (v2, +distilabel)
simple_python
75.0
77.5
79.5
multiple
50.5
74.0
77.5
live_simple
31.8
57.0
58.5
live_multiple
7.3
38.8
43.5
live_relevance
43.8
93.8
87.5
live_irrelevance
67.3
16.9
22.2
irrelevance
35.8
5.8
6.2
parallel
44.0
1.0
0.0
parallel_multiple
15.0
2.0
0.0
Finding: Adding more (mixed-quality) data gave marginal call-required gains (+1.5 to +4.7 pts) but did not fix the irrelevance regression (still ~6%). Volume is not the fix for irrelevance — see the Path B model for the effective approach. Note: ~12% of the distilabel additions had malformed JSON and were filtered out during preprocessing.
Training details
Base: meta-llama/Llama-3.2-1B-Instruct
Data: xLAM + argilla/apigen (distilabel), filtered to 45k single-call examples
Hardware: 1× A100, ~98 min · final eval loss: 0.16
Usage
fromtransformersimportAutoModelForCausalLM,AutoTokenizerimporttorchtokenizer=AutoTokenizer.from_pretrained("Keitsuna123/llama-3.2-1b-fc-sft-v2-merged")model=AutoModelForCausalLM.from_pretrained("Keitsuna123/llama-3.2-1b-fc-sft-v2-merged",torch_dtype=torch.bfloat16,device_map="auto")messages=[{"role":"user","content":"What's the weather in Tokyo?"}]tools=[{"type":"function","function":{"name":"get_weather","description":"Get the weather for a location","parameters":{"type":"object","properties":{"location":{"type":"string"}},"required":["location"]}}}]text=tokenizer.apply_chat_template(messages,tools=tools,tokenize=False,add_generation_prompt=True)inputs=tokenizer(text,return_tensors="pt").to(model.device)out=model.generate(**inputs,max_new_tokens=128,do_sample=False)print(tokenizer.decode(out[0][inputs['input_ids'].shape[1]:],skip_special_tokens=True))
Limitations
Irrelevance detection still regressed vs. base. Parallel calls near-zero (single-call training only).
Benchmark: Berkeley Function Calling Leaderboard (BFCL) — Gorilla project
Frameworks: HuggingFace TRL, PEFT, Transformers
Citation
@misc{taketsuna2026_fc_smallmodel,title={Small-Model Function Calling: Comparing SFT, Data Scaling, GRPO, and Supervised Refusal Training},author={Taketsuna, Keiichi},year={2026},howpublished={\url{https://huggingface.co/Keitsuna123}},note={Llama-3.2-1B function-calling post-training study on BFCL v4}}