The v1 SFT model further fine-tuned on a mix of call-required examples and targeted refusal examples, to recover irrelevance detection. This is the effective fix in a four-model study — and it succeeded where GRPO did not.
Results (BFCL v4)
Category
Base
v1 SFT
GRPO
This model (Path B)
irrelevance
35.8
5.8
5.4
21.7
live_irrelevance
67.3
16.9
17.2
36.8
simple_python
75.0
77.5
77.2
78.2
multiple
50.5
74.0
74.0
72.0
live_simple
31.8
57.0
56.2
54.3
live_multiple
7.3
38.8
39.4
35.7
live_relevance
43.8
93.8
93.8
81.2
Finding: Supervised refusal examples recovered irrelevance ~4x (5.8→21.7) and more than doubled live_irrelevance (16.9→36.8), at a modest, quantified cost to call-required categories (multiple −2, live_simple −2.7, live_multiple −3, live_relevance −12.5). This is a clean precision/recall tradeoff, and a much larger recovery than GRPO achieved (which showed no gain).
Takeaway: for recovering an out-of-distribution behavior in a small model, targeted supervised demonstration outperforms RL reward shaping — because RL requires the target behavior to appear in sampled rollouts, which a strong SFT prior prevents.
Config: 3 epochs, effective batch 16, lr 1.2e-4 cosine, bf16 · final eval loss: 0.17
Usage
fromtransformersimportAutoModelForCausalLM,AutoTokenizerimporttorchtokenizer=AutoTokenizer.from_pretrained("Keitsuna123/llama-3.2-1b-fc-pathb")model=AutoModelForCausalLM.from_pretrained("Keitsuna123/llama-3.2-1b-fc-pathb",torch_dtype=torch.bfloat16,device_map="auto")# Call-required: produces a function callmessages=[{"role":"user","content":"What's the weather in Tokyo?"}]# Irrelevant: produces a conversational refusal instead of a hallucinated call# messages = [{"role": "user", "content": "How do I make bread fluffier?"}]tools=[{"type":"function","function":{"name":"get_weather","description":"Get the weather for a location","parameters":{"type":"object","properties":{"location":{"type":"string"}},"required":["location"]}}}]text=tokenizer.apply_chat_template(messages,tools=tools,tokenize=False,add_generation_prompt=True)inputs=tokenizer(text,return_tensors="pt").to(model.device)out=model.generate(**inputs,max_new_tokens=128,do_sample=False)print(tokenizer.decode(out[0][inputs['input_ids'].shape[1]:],skip_special_tokens=True))
Limitations
Residual irrelevance failures concentrate in short, factual-seeming queries ("what's 2+2", "capital of France") where a loosely-related tool exists — the hardest irrelevance sub-case.
Parallel calls remain near-zero (single-call training only).
Training data: Salesforce xLAM — xlam-function-calling-60k; refusal targets generated for this study
Benchmark: Berkeley Function Calling Leaderboard (BFCL) — Gorilla project
Frameworks: HuggingFace TRL, PEFT, Transformers
Citation
@misc{taketsuna2026_fc_smallmodel,title={Small-Model Function Calling: Comparing SFT, Data Scaling, GRPO, and Supervised Refusal Training},author={Taketsuna, Keiichi},year={2026},howpublished={\url{https://huggingface.co/Keitsuna123}},note={Llama-3.2-1B function-calling post-training study on BFCL v4}}