182 lines
7.7 KiB
Markdown
182 lines
7.7 KiB
Markdown
|
|
---
|
|||
|
|
license: apache-2.0
|
|||
|
|
language:
|
|||
|
|
- ar
|
|||
|
|
base_model:
|
|||
|
|
- oddadmix/50M-2048-Emhotob
|
|||
|
|
library_name: transformers
|
|||
|
|
pipeline_tag: text-generation
|
|||
|
|
tags:
|
|||
|
|
- arabic
|
|||
|
|
- tool-calling
|
|||
|
|
- function-calling
|
|||
|
|
- grpo
|
|||
|
|
- rlhf
|
|||
|
|
- distillation
|
|||
|
|
- small-language-model
|
|||
|
|
- slm
|
|||
|
|
- tiny
|
|||
|
|
- llama
|
|||
|
|
- from-scratch
|
|||
|
|
- proof-of-concept
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# Emhotob-50M-GRPO — Arabic Tool-Calling (RL-tuned)
|
|||
|
|
|
|||
|
|
A **~51.8M-parameter** Arabic **tool-calling / function-calling** model, tuned with
|
|||
|
|
**GRPO** (Group Relative Policy Optimization) reinforcement learning on top of a
|
|||
|
|
distilled SFT checkpoint. It is a proof-of-concept for **agentic Arabic behavior at
|
|||
|
|
tiny scale** — deciding *when to call a tool* vs. *when to decline*, entirely from a
|
|||
|
|
model small enough to run on a CPU.
|
|||
|
|
|
|||
|
|
This checkpoint corresponds to the **GRPO-v12** run in the training repo: it is tuned
|
|||
|
|
for **robust abstention** — it declines out-of-scope requests reliably instead of
|
|||
|
|
forcing a tool call.
|
|||
|
|
|
|||
|
|
- **Base model:** [`oddadmix/50M-2048-Emhotob`](https://huggingface.co/oddadmix/50M-2048-Emhotob)
|
|||
|
|
— Llama-architecture, pre-trained **from scratch on ~20B Arabic tokens**, 2048 ctx.
|
|||
|
|
- **Architecture & recipe:** derived from [`SupraLabs/Supra-50M-Base`](https://huggingface.co/SupraLabs/Supra-50M-Base);
|
|||
|
|
a **POC for training Arabic tiny models from scratch**.
|
|||
|
|
- **Format:** ChatML + Hermes-style `<tool_call>{...}</tool_call>`.
|
|||
|
|
|
|||
|
|
> **الملخص بالعربية:** نموذج عربي صغير (~51.8 مليون معامل) لاستدعاء الأدوات (Tool/Function
|
|||
|
|
> Calling)، مبني على معمارية Llama ومدرَّب من الصفر، ثم ضُبط باستخدام التعلّم المعزّز
|
|||
|
|
> **GRPO**. هذه النسخة مُحسَّنة للامتناع عن استدعاء أداة عند عدم توفّر أداة مناسبة، بدلًا من
|
|||
|
|
> اختلاق استدعاء خاطئ. نموذج تجريبي (POC) لسلوك عربي «وكيلي» على نطاق صغير جدًا.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Model details
|
|||
|
|
|
|||
|
|
| | |
|
|||
|
|
|---|---|
|
|||
|
|
| **Parameters** | ~51.8M |
|
|||
|
|
| **Architecture** | Llama (`LlamaForCausalLM`) |
|
|||
|
|
| **Hidden size** | 512 · **Layers** 12 · **Heads** 8 (GQA, 4 KV) · **head_dim** 64 |
|
|||
|
|
| **Vocab size** | 32002 (32000 base + 2 ChatML control tokens) |
|
|||
|
|
| **Context length** | 2048 |
|
|||
|
|
| **Chat format** | ChatML (`<|im_start|>` / `<|im_end|>`) + Hermes `<tool_call>` |
|
|||
|
|
| **Precision** | bfloat16 |
|
|||
|
|
| **License** | Apache-2.0 (matches the base architecture) |
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## How it was trained
|
|||
|
|
|
|||
|
|
The model is the end of a **base → SFT → distillation → RL** pipeline, all in pure
|
|||
|
|
🤗 Transformers (**no TRL** — the GRPO loop is from scratch):
|
|||
|
|
|
|||
|
|
1. **Base:** [`oddadmix/50M-2048-Emhotob`](https://huggingface.co/oddadmix/50M-2048-Emhotob),
|
|||
|
|
pre-trained from scratch on ~20B Arabic tokens (2048 ctx), architecture and training
|
|||
|
|
scripts derived from [`SupraLabs/Supra-50M-Base`](https://huggingface.co/SupraLabs/Supra-50M-Base).
|
|||
|
|
2. **SFT:** supervised fine-tuning on Arabic tool-calling data (ChatML + Hermes
|
|||
|
|
`<tool_call>` format).
|
|||
|
|
3. **Sequence-level distillation (Kim & Rush, 2016):** a 1.2B tool-specialized teacher
|
|||
|
|
([`LiquidAI/LFM2-1.2B-Tool`](https://huggingface.co/LiquidAI/LFM2-1.2B-Tool))
|
|||
|
|
generated targets — including **fluent Arabic abstentions** and hard-negative
|
|||
|
|
(no-matching-tool) examples — that the student learned to imitate. This is the
|
|||
|
|
variant that reliably **declines out-of-scope requests**.
|
|||
|
|
4. **GRPO (Shao et al., 2024 — DeepSeekMath):** on-policy RL with a **verifiable reward**
|
|||
|
|
(the tool-eval scorer itself — no reward model, no LLM judge). For each prompt the
|
|||
|
|
policy samples a group of completions; a group-relative advantage plus a KL leash to
|
|||
|
|
the reference model shapes the call/abstain decision. This checkpoint uses a
|
|||
|
|
**balanced 1:1 call/abstain** batch with a gentle learning rate, which **holds the
|
|||
|
|
high-abstention behavior** while keeping JSON/tool-selection intact.
|
|||
|
|
|
|||
|
|
**What the experiment found (the POC result):** at 50M parameters there is a genuine
|
|||
|
|
**precision ↔ recall frontier** on the call decision — you can maximize valid calls *or*
|
|||
|
|
maximize correct refusals, but not both. GRPO with a perfect verifiable reward moves
|
|||
|
|
*along* this frontier rather than beyond it; this checkpoint deliberately sits at the
|
|||
|
|
**robust-abstention** corner. (A separately-tested 2.5× larger student tied these
|
|||
|
|
metrics, corroborating that the ceiling is capacity, not the training signal.)
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Evaluation
|
|||
|
|
|
|||
|
|
Numbers are reported honestly for this exact checkpoint against the sibling SFT/RL
|
|||
|
|
checkpoints (see the training repo for the full trial log).
|
|||
|
|
|
|||
|
|
**Dialect eval (18 items, Egyptian-dialect prompts):**
|
|||
|
|
|
|||
|
|
| metric | this model | note |
|
|||
|
|
|---|---|---|
|
|||
|
|
| strict success | 12/18 | headline |
|
|||
|
|
| decision (call/answer correct) | 16/18 | |
|
|||
|
|
| **abstention** | **4/4** ✅ | robustly declines out-of-scope requests |
|
|||
|
|
|
|||
|
|
**MSA eval v2 (126 items, higher-resolution):** strict 11/68, decision 49/68,
|
|||
|
|
**abstain 27/46** — the high-abstention end of the decision↔abstain frontier.
|
|||
|
|
Tool-selection (~33%) and argument-exactness (~18%) are the capacity-bound bottlenecks
|
|||
|
|
shared across all checkpoints at this scale.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Usage
|
|||
|
|
|
|||
|
|
The tokenizer uses the `TokenizersBackend` class, which requires **`transformers>=5.12`**.
|
|||
|
|
Build the ChatML prompt manually and pass the tool schema in the system message:
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
import json, torch
|
|||
|
|
from transformers import AutoModelForCausalLM, AutoTokenizer
|
|||
|
|
|
|||
|
|
model_id = "oddadmix/Emhotob-50M-GPRO-Arabic-Final"
|
|||
|
|
tok = AutoTokenizer.from_pretrained(model_id)
|
|||
|
|
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.bfloat16)
|
|||
|
|
|
|||
|
|
tools = [{
|
|||
|
|
"name": "get_weather",
|
|||
|
|
"description": "Get the current weather for a city",
|
|||
|
|
"parameters": {"type": "object",
|
|||
|
|
"properties": {"location": {"type": "string"}},
|
|||
|
|
"required": ["location"]},
|
|||
|
|
}]
|
|||
|
|
|
|||
|
|
system = (
|
|||
|
|
"أنت مساعد يستطيع استدعاء الأدوات. الأدوات المتاحة:\n"
|
|||
|
|
+ json.dumps(tools, ensure_ascii=False)
|
|||
|
|
+ "\nإذا لم تكن هناك أداة مناسبة، اعتذر ولا تختلق استدعاءً."
|
|||
|
|
)
|
|||
|
|
user = "ما حالة الطقس في القاهرة؟"
|
|||
|
|
|
|||
|
|
prompt = (
|
|||
|
|
f"<|im_start|>system\n{system}<|im_end|>\n"
|
|||
|
|
f"<|im_start|>user\n{user}<|im_end|>\n"
|
|||
|
|
f"<|im_start|>assistant\n"
|
|||
|
|
)
|
|||
|
|
ids = tok(prompt, return_tensors="pt").to(model.device)
|
|||
|
|
out = model.generate(
|
|||
|
|
**ids, max_new_tokens=256, do_sample=False,
|
|||
|
|
repetition_penalty=1.2,
|
|||
|
|
eos_token_id=tok.convert_tokens_to_ids("<|im_end|>"),
|
|||
|
|
)
|
|||
|
|
print(tok.decode(out[0][ids.input_ids.shape[1]:], skip_special_tokens=True))
|
|||
|
|
# Expected: <tool_call>{"name": "get_weather", "arguments": {"location": "القاهرة"}}</tool_call>
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
For an out-of-scope request with no matching tool, the model is tuned to **refuse in
|
|||
|
|
Arabic** rather than hallucinate a call.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Intended use & limitations
|
|||
|
|
|
|||
|
|
**Intended use.** Research and demonstration of Arabic tool-calling / agentic behavior at
|
|||
|
|
tiny scale; a CPU-friendly baseline for from-scratch Arabic SLM experiments.
|
|||
|
|
|
|||
|
|
**Limitations.** At ~50M parameters this is a **proof of concept**. Tool-selection with
|
|||
|
|
several confusable tools (~33% correct) and exact argument extraction (~18%) are the
|
|||
|
|
capacity-bound weak spots — validate any tool call before executing it. Training was
|
|||
|
|
predominantly Modern Standard Arabic; dialect prompts are harder. Not for production
|
|||
|
|
agents without a validation/guardrail layer.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Citation & credits
|
|||
|
|
|
|||
|
|
- **Method:** GRPO — Shao et al., 2024 (*DeepSeekMath*); sequence-level KD — Kim & Rush, 2016.
|
|||
|
|
- **Teacher:** [`LiquidAI/LFM2-1.2B-Tool`](https://huggingface.co/LiquidAI/LFM2-1.2B-Tool).
|
|||
|
|
- **Base architecture & training scripts:** [`SupraLabs/Supra-50M-Base`](https://huggingface.co/SupraLabs/Supra-50M-Base) (Apache-2.0).
|
|||
|
|
- **Base Arabic model:** [`oddadmix/50M-2048-Emhotob`](https://huggingface.co/oddadmix/50M-2048-Emhotob) — from-scratch, ~20B Arabic tokens, 2048 ctx.
|