Files
llama-3.2-1b-fc-pathb/README.md
ModelHub XC 439bbe730c 初始化项目,由ModelHub XC社区提供模型
Model: Keitsuna123/llama-3.2-1b-fc-pathb
Source: Original Platform
2026-09-08 19:24:16 +08:00

4.7 KiB
Raw Blame History

base_model, library_name, tags, license, language
base_model library_name tags license language
meta-llama/Llama-3.2-1B-Instruct transformers
function-calling
tool-use
llama
lora
bfcl
llama3.2
en

Llama-3.2-1B-Instruct — Function Calling (Refusal-SFT, "Path B")

Code & writeup: https://github.com/keitake123/llama-function-calling-study

The v1 SFT model further fine-tuned on a mix of call-required examples and targeted refusal examples, to recover irrelevance detection. This is the effective fix in a four-model study — and it succeeded where GRPO did not.

Results (BFCL v4)

Category Base v1 SFT GRPO This model (Path B)
irrelevance 35.8 5.8 5.4 21.7
live_irrelevance 67.3 16.9 17.2 36.8
simple_python 75.0 77.5 77.2 78.2
multiple 50.5 74.0 74.0 72.0
live_simple 31.8 57.0 56.2 54.3
live_multiple 7.3 38.8 39.4 35.7
live_relevance 43.8 93.8 93.8 81.2

Finding: Supervised refusal examples recovered irrelevance ~4x (5.8→21.7) and more than doubled live_irrelevance (16.9→36.8), at a modest, quantified cost to call-required categories (multiple −2, live_simple −2.7, live_multiple −3, live_relevance −12.5). This is a clean precision/recall tradeoff, and a much larger recovery than GRPO achieved (which showed no gain).

Takeaway: for recovering an out-of-distribution behavior in a small model, targeted supervised demonstration outperforms RL reward shaping — because RL requires the target behavior to appear in sampled rollouts, which a strong SFT prior prevents.

Training details

  • Base: v1 SFT model · Method: LoRA SFT
  • Data: 1600 examples, 50/50 mix of xLAM call-required + synthetic refusal examples (query + unrelated tools → conversational refusal)
  • Config: 3 epochs, effective batch 16, lr 1.2e-4 cosine, bf16 · final eval loss: 0.17

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

tokenizer = AutoTokenizer.from_pretrained("Keitsuna123/llama-3.2-1b-fc-pathb")
model = AutoModelForCausalLM.from_pretrained(
    "Keitsuna123/llama-3.2-1b-fc-pathb", torch_dtype=torch.bfloat16, device_map="auto"
)

# Call-required: produces a function call
messages = [{"role": "user", "content": "What's the weather in Tokyo?"}]
# Irrelevant: produces a conversational refusal instead of a hallucinated call
# messages = [{"role": "user", "content": "How do I make bread fluffier?"}]

tools = [{"type": "function", "function": {
    "name": "get_weather", "description": "Get the weather for a location",
    "parameters": {"type": "object", "properties": {"location": {"type": "string"}}, "required": ["location"]}
}}]

text = tokenizer.apply_chat_template(messages, tools=tools, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=128, do_sample=False)
print(tokenizer.decode(out[0][inputs['input_ids'].shape[1]:], skip_special_tokens=True))

Limitations

  • Residual irrelevance failures concentrate in short, factual-seeming queries ("what's 2+2", "capital of France") where a loosely-related tool exists — the hardest irrelevance sub-case.
  • Parallel calls remain near-zero (single-call training only).

Part of a series

Model Description
fc-sft-full v1 SFT on xLAM
fc-sft-v2-merged SFT on xLAM + distilabel (data-scaling ablation)
fc-grpo GRPO for irrelevance recovery (negative result)
fc-pathb Supervised refusal training (this model)

References & Acknowledgements

Citation

@misc{taketsuna2026_fc_smallmodel,
  title  = {Small-Model Function Calling: Comparing SFT, Data Scaling, GRPO, and Supervised Refusal Training},
  author = {Taketsuna, Keiichi},
  year   = {2026},
  howpublished = {\url{https://huggingface.co/Keitsuna123}},
  note   = {Llama-3.2-1B function-calling post-training study on BFCL v4}
}