--- license: apache-2.0 language: - ar base_model: - oddadmix/50M-2048-Emhotob library_name: transformers pipeline_tag: text-generation tags: - arabic - tool-calling - function-calling - grpo - rlhf - distillation - small-language-model - slm - tiny - llama - from-scratch - proof-of-concept --- # Emhotob-50M-GRPO — Arabic Tool-Calling (RL-tuned) A **~51.8M-parameter** Arabic **tool-calling / function-calling** model, tuned with **GRPO** (Group Relative Policy Optimization) reinforcement learning on top of a distilled SFT checkpoint. It is a proof-of-concept for **agentic Arabic behavior at tiny scale** — deciding *when to call a tool* vs. *when to decline*, entirely from a model small enough to run on a CPU. This checkpoint corresponds to the **GRPO-v12** run in the training repo: it is tuned for **robust abstention** — it declines out-of-scope requests reliably instead of forcing a tool call. - **Base model:** [`oddadmix/50M-2048-Emhotob`](https://huggingface.co/oddadmix/50M-2048-Emhotob) — Llama-architecture, pre-trained **from scratch on ~20B Arabic tokens**, 2048 ctx. - **Architecture & recipe:** derived from [`SupraLabs/Supra-50M-Base`](https://huggingface.co/SupraLabs/Supra-50M-Base); a **POC for training Arabic tiny models from scratch**. - **Format:** ChatML + Hermes-style `{...}`. > **الملخص بالعربية:** نموذج عربي صغير (~51.8 مليون معامل) لاستدعاء الأدوات (Tool/Function > Calling)، مبني على معمارية Llama ومدرَّب من الصفر، ثم ضُبط باستخدام التعلّم المعزّز > **GRPO**. هذه النسخة مُحسَّنة للامتناع عن استدعاء أداة عند عدم توفّر أداة مناسبة، بدلًا من > اختلاق استدعاء خاطئ. نموذج تجريبي (POC) لسلوك عربي «وكيلي» على نطاق صغير جدًا. --- ## Model details | | | |---|---| | **Parameters** | ~51.8M | | **Architecture** | Llama (`LlamaForCausalLM`) | | **Hidden size** | 512 · **Layers** 12 · **Heads** 8 (GQA, 4 KV) · **head_dim** 64 | | **Vocab size** | 32002 (32000 base + 2 ChatML control tokens) | | **Context length** | 2048 | | **Chat format** | ChatML (`<|im_start|>` / `<|im_end|>`) + Hermes `` | | **Precision** | bfloat16 | | **License** | Apache-2.0 (matches the base architecture) | --- ## How it was trained The model is the end of a **base → SFT → distillation → RL** pipeline, all in pure 🤗 Transformers (**no TRL** — the GRPO loop is from scratch): 1. **Base:** [`oddadmix/50M-2048-Emhotob`](https://huggingface.co/oddadmix/50M-2048-Emhotob), pre-trained from scratch on ~20B Arabic tokens (2048 ctx), architecture and training scripts derived from [`SupraLabs/Supra-50M-Base`](https://huggingface.co/SupraLabs/Supra-50M-Base). 2. **SFT:** supervised fine-tuning on Arabic tool-calling data (ChatML + Hermes `` format). 3. **Sequence-level distillation (Kim & Rush, 2016):** a 1.2B tool-specialized teacher ([`LiquidAI/LFM2-1.2B-Tool`](https://huggingface.co/LiquidAI/LFM2-1.2B-Tool)) generated targets — including **fluent Arabic abstentions** and hard-negative (no-matching-tool) examples — that the student learned to imitate. This is the variant that reliably **declines out-of-scope requests**. 4. **GRPO (Shao et al., 2024 — DeepSeekMath):** on-policy RL with a **verifiable reward** (the tool-eval scorer itself — no reward model, no LLM judge). For each prompt the policy samples a group of completions; a group-relative advantage plus a KL leash to the reference model shapes the call/abstain decision. This checkpoint uses a **balanced 1:1 call/abstain** batch with a gentle learning rate, which **holds the high-abstention behavior** while keeping JSON/tool-selection intact. **What the experiment found (the POC result):** at 50M parameters there is a genuine **precision ↔ recall frontier** on the call decision — you can maximize valid calls *or* maximize correct refusals, but not both. GRPO with a perfect verifiable reward moves *along* this frontier rather than beyond it; this checkpoint deliberately sits at the **robust-abstention** corner. (A separately-tested 2.5× larger student tied these metrics, corroborating that the ceiling is capacity, not the training signal.) --- ## Evaluation Numbers are reported honestly for this exact checkpoint against the sibling SFT/RL checkpoints (see the training repo for the full trial log). **Dialect eval (18 items, Egyptian-dialect prompts):** | metric | this model | note | |---|---|---| | strict success | 12/18 | headline | | decision (call/answer correct) | 16/18 | | | **abstention** | **4/4** ✅ | robustly declines out-of-scope requests | **MSA eval v2 (126 items, higher-resolution):** strict 11/68, decision 49/68, **abstain 27/46** — the high-abstention end of the decision↔abstain frontier. Tool-selection (~33%) and argument-exactness (~18%) are the capacity-bound bottlenecks shared across all checkpoints at this scale. --- ## Usage The tokenizer uses the `TokenizersBackend` class, which requires **`transformers>=5.12`**. Build the ChatML prompt manually and pass the tool schema in the system message: ```python import json, torch from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "oddadmix/Emhotob-50M-GPRO-Arabic-Final" tok = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.bfloat16) tools = [{ "name": "get_weather", "description": "Get the current weather for a city", "parameters": {"type": "object", "properties": {"location": {"type": "string"}}, "required": ["location"]}, }] system = ( "أنت مساعد يستطيع استدعاء الأدوات. الأدوات المتاحة:\n" + json.dumps(tools, ensure_ascii=False) + "\nإذا لم تكن هناك أداة مناسبة، اعتذر ولا تختلق استدعاءً." ) user = "ما حالة الطقس في القاهرة؟" prompt = ( f"<|im_start|>system\n{system}<|im_end|>\n" f"<|im_start|>user\n{user}<|im_end|>\n" f"<|im_start|>assistant\n" ) ids = tok(prompt, return_tensors="pt").to(model.device) out = model.generate( **ids, max_new_tokens=256, do_sample=False, repetition_penalty=1.2, eos_token_id=tok.convert_tokens_to_ids("<|im_end|>"), ) print(tok.decode(out[0][ids.input_ids.shape[1]:], skip_special_tokens=True)) # Expected: {"name": "get_weather", "arguments": {"location": "القاهرة"}} ``` For an out-of-scope request with no matching tool, the model is tuned to **refuse in Arabic** rather than hallucinate a call. --- ## Intended use & limitations **Intended use.** Research and demonstration of Arabic tool-calling / agentic behavior at tiny scale; a CPU-friendly baseline for from-scratch Arabic SLM experiments. **Limitations.** At ~50M parameters this is a **proof of concept**. Tool-selection with several confusable tools (~33% correct) and exact argument extraction (~18%) are the capacity-bound weak spots — validate any tool call before executing it. Training was predominantly Modern Standard Arabic; dialect prompts are harder. Not for production agents without a validation/guardrail layer. --- ## Citation & credits - **Method:** GRPO — Shao et al., 2024 (*DeepSeekMath*); sequence-level KD — Kim & Rush, 2016. - **Teacher:** [`LiquidAI/LFM2-1.2B-Tool`](https://huggingface.co/LiquidAI/LFM2-1.2B-Tool). - **Base architecture & training scripts:** [`SupraLabs/Supra-50M-Base`](https://huggingface.co/SupraLabs/Supra-50M-Base) (Apache-2.0). - **Base Arabic model:** [`oddadmix/50M-2048-Emhotob`](https://huggingface.co/oddadmix/50M-2048-Emhotob) — from-scratch, ~20B Arabic tokens, 2048 ctx.