57 lines
2.2 KiB
Markdown
57 lines
2.2 KiB
Markdown
|
|
---
|
|||
|
|
license: apache-2.0
|
|||
|
|
base_model: Qwen/Qwen2.5-Coder-1.5B-Instruct
|
|||
|
|
tags:
|
|||
|
|
- code
|
|||
|
|
- function-calling
|
|||
|
|
- tool-use
|
|||
|
|
- small-language-model
|
|||
|
|
- small-code
|
|||
|
|
datasets:
|
|||
|
|
- NousResearch/hermes-function-calling-v1
|
|||
|
|
language:
|
|||
|
|
- en
|
|||
|
|
pipeline_tag: text-generation
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# small-code-coder-1.5b-tools
|
|||
|
|
|
|||
|
|
A LoRA fine-tune of **Qwen2.5-Coder-1.5B-Instruct** that teaches the model to emit
|
|||
|
|
**native `<tool_call>` function calls**, so a ≤2B *coder* model can drive an agentic
|
|||
|
|
coding loop.
|
|||
|
|
|
|||
|
|
Built for [**smolcode**](https://gitea.poyner.ai/sean/smolcode) — an SLM-optimized
|
|||
|
|
agentic coding assistant — for the Hugging Face **Build Small** hackathon.
|
|||
|
|
|
|||
|
|
## Why
|
|||
|
|
Out of the box, small Qwen-Coder models describe tool calls as plain-text JSON
|
|||
|
|
instead of emitting the native `<tool_call>` format that runtimes (Ollama,
|
|||
|
|
llama.cpp) parse — which breaks agentic tool-use loops. This fine-tune closes
|
|||
|
|
that gap on a tiny (≤2B, Tiny-Titan-class) model.
|
|||
|
|
|
|||
|
|
## Training
|
|||
|
|
- **Base:** Qwen/Qwen2.5-Coder-1.5B-Instruct
|
|||
|
|
- **Method:** bf16 LoRA (r=16, α=32) on attention + MLP projections, **assistant-only
|
|||
|
|
loss** (loss on tool calls + final answers only).
|
|||
|
|
- **Data:** NousResearch/hermes-function-calling-v1 (breadth) + synthetic smolcode
|
|||
|
|
tool-use trajectories (sharpness on the actual 5 tools), all rendered through the
|
|||
|
|
*same* `apply_chat_template(tools=...)` used at inference — so the training target
|
|||
|
|
is byte-identical to the served prompt.
|
|||
|
|
- **Schedule:** 3 epochs, full 2048 sequence length.
|
|||
|
|
- **Hardware:** trained on Modal (x86/CUDA); served on NVIDIA DGX Spark (GB10).
|
|||
|
|
|
|||
|
|
## Use
|
|||
|
|
Standard Qwen2.5 chat template with `tools=`. The model responds with
|
|||
|
|
`<tool_call>{"name": ..., "arguments": ...}</tool_call>` when a tool is warranted.
|
|||
|
|
|
|||
|
|
## Status — v2
|
|||
|
|
v2 fixes the v1 train/inference template mismatch (v1 hit 0.92 teacher-forced token
|
|||
|
|
accuracy but decoded degenerately because it was trained on a hand-rendered Hermes
|
|||
|
|
ChatML format, not Qwen's `apply_chat_template` output). v2 trains and serves through
|
|||
|
|
one shared template and is gated on a *free-generation* tool-call parse-rate eval
|
|||
|
|
(≥90% on held-out smolcode prompts) before release — see `eval_toolcall.py` in the
|
|||
|
|
smolcode repo.
|
|||
|
|
|
|||
|
|
## License
|
|||
|
|
Apache-2.0 (inherits from the base model).
|