初始化项目,由ModelHub XC社区提供模型
Model: oddadmix/Emhotob-50M-GRPO-Arabic-Final Source: Original Platform
This commit is contained in:
35
.gitattributes
vendored
Normal file
35
.gitattributes
vendored
Normal file
@@ -0,0 +1,35 @@
|
||||
*.7z filter=lfs diff=lfs merge=lfs -text
|
||||
*.arrow filter=lfs diff=lfs merge=lfs -text
|
||||
*.bin filter=lfs diff=lfs merge=lfs -text
|
||||
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
||||
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
||||
*.ftz filter=lfs diff=lfs merge=lfs -text
|
||||
*.gz filter=lfs diff=lfs merge=lfs -text
|
||||
*.h5 filter=lfs diff=lfs merge=lfs -text
|
||||
*.joblib filter=lfs diff=lfs merge=lfs -text
|
||||
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
||||
*.model filter=lfs diff=lfs merge=lfs -text
|
||||
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
||||
*.npy filter=lfs diff=lfs merge=lfs -text
|
||||
*.npz filter=lfs diff=lfs merge=lfs -text
|
||||
*.onnx filter=lfs diff=lfs merge=lfs -text
|
||||
*.ot filter=lfs diff=lfs merge=lfs -text
|
||||
*.parquet filter=lfs diff=lfs merge=lfs -text
|
||||
*.pb filter=lfs diff=lfs merge=lfs -text
|
||||
*.pickle filter=lfs diff=lfs merge=lfs -text
|
||||
*.pkl filter=lfs diff=lfs merge=lfs -text
|
||||
*.pt filter=lfs diff=lfs merge=lfs -text
|
||||
*.pth filter=lfs diff=lfs merge=lfs -text
|
||||
*.rar filter=lfs diff=lfs merge=lfs -text
|
||||
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
||||
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar filter=lfs diff=lfs merge=lfs -text
|
||||
*.tflite filter=lfs diff=lfs merge=lfs -text
|
||||
*.tgz filter=lfs diff=lfs merge=lfs -text
|
||||
*.wasm filter=lfs diff=lfs merge=lfs -text
|
||||
*.xz filter=lfs diff=lfs merge=lfs -text
|
||||
*.zip filter=lfs diff=lfs merge=lfs -text
|
||||
*.zst filter=lfs diff=lfs merge=lfs -text
|
||||
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
||||
181
README.md
Normal file
181
README.md
Normal file
@@ -0,0 +1,181 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
language:
|
||||
- ar
|
||||
base_model:
|
||||
- oddadmix/50M-2048-Emhotob
|
||||
library_name: transformers
|
||||
pipeline_tag: text-generation
|
||||
tags:
|
||||
- arabic
|
||||
- tool-calling
|
||||
- function-calling
|
||||
- grpo
|
||||
- rlhf
|
||||
- distillation
|
||||
- small-language-model
|
||||
- slm
|
||||
- tiny
|
||||
- llama
|
||||
- from-scratch
|
||||
- proof-of-concept
|
||||
---
|
||||
|
||||
# Emhotob-50M-GRPO — Arabic Tool-Calling (RL-tuned)
|
||||
|
||||
A **~51.8M-parameter** Arabic **tool-calling / function-calling** model, tuned with
|
||||
**GRPO** (Group Relative Policy Optimization) reinforcement learning on top of a
|
||||
distilled SFT checkpoint. It is a proof-of-concept for **agentic Arabic behavior at
|
||||
tiny scale** — deciding *when to call a tool* vs. *when to decline*, entirely from a
|
||||
model small enough to run on a CPU.
|
||||
|
||||
This checkpoint corresponds to the **GRPO-v12** run in the training repo: it is tuned
|
||||
for **robust abstention** — it declines out-of-scope requests reliably instead of
|
||||
forcing a tool call.
|
||||
|
||||
- **Base model:** [`oddadmix/50M-2048-Emhotob`](https://huggingface.co/oddadmix/50M-2048-Emhotob)
|
||||
— Llama-architecture, pre-trained **from scratch on ~20B Arabic tokens**, 2048 ctx.
|
||||
- **Architecture & recipe:** derived from [`SupraLabs/Supra-50M-Base`](https://huggingface.co/SupraLabs/Supra-50M-Base);
|
||||
a **POC for training Arabic tiny models from scratch**.
|
||||
- **Format:** ChatML + Hermes-style `<tool_call>{...}</tool_call>`.
|
||||
|
||||
> **الملخص بالعربية:** نموذج عربي صغير (~51.8 مليون معامل) لاستدعاء الأدوات (Tool/Function
|
||||
> Calling)، مبني على معمارية Llama ومدرَّب من الصفر، ثم ضُبط باستخدام التعلّم المعزّز
|
||||
> **GRPO**. هذه النسخة مُحسَّنة للامتناع عن استدعاء أداة عند عدم توفّر أداة مناسبة، بدلًا من
|
||||
> اختلاق استدعاء خاطئ. نموذج تجريبي (POC) لسلوك عربي «وكيلي» على نطاق صغير جدًا.
|
||||
|
||||
---
|
||||
|
||||
## Model details
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| **Parameters** | ~51.8M |
|
||||
| **Architecture** | Llama (`LlamaForCausalLM`) |
|
||||
| **Hidden size** | 512 · **Layers** 12 · **Heads** 8 (GQA, 4 KV) · **head_dim** 64 |
|
||||
| **Vocab size** | 32002 (32000 base + 2 ChatML control tokens) |
|
||||
| **Context length** | 2048 |
|
||||
| **Chat format** | ChatML (`<|im_start|>` / `<|im_end|>`) + Hermes `<tool_call>` |
|
||||
| **Precision** | bfloat16 |
|
||||
| **License** | Apache-2.0 (matches the base architecture) |
|
||||
|
||||
---
|
||||
|
||||
## How it was trained
|
||||
|
||||
The model is the end of a **base → SFT → distillation → RL** pipeline, all in pure
|
||||
🤗 Transformers (**no TRL** — the GRPO loop is from scratch):
|
||||
|
||||
1. **Base:** [`oddadmix/50M-2048-Emhotob`](https://huggingface.co/oddadmix/50M-2048-Emhotob),
|
||||
pre-trained from scratch on ~20B Arabic tokens (2048 ctx), architecture and training
|
||||
scripts derived from [`SupraLabs/Supra-50M-Base`](https://huggingface.co/SupraLabs/Supra-50M-Base).
|
||||
2. **SFT:** supervised fine-tuning on Arabic tool-calling data (ChatML + Hermes
|
||||
`<tool_call>` format).
|
||||
3. **Sequence-level distillation (Kim & Rush, 2016):** a 1.2B tool-specialized teacher
|
||||
([`LiquidAI/LFM2-1.2B-Tool`](https://huggingface.co/LiquidAI/LFM2-1.2B-Tool))
|
||||
generated targets — including **fluent Arabic abstentions** and hard-negative
|
||||
(no-matching-tool) examples — that the student learned to imitate. This is the
|
||||
variant that reliably **declines out-of-scope requests**.
|
||||
4. **GRPO (Shao et al., 2024 — DeepSeekMath):** on-policy RL with a **verifiable reward**
|
||||
(the tool-eval scorer itself — no reward model, no LLM judge). For each prompt the
|
||||
policy samples a group of completions; a group-relative advantage plus a KL leash to
|
||||
the reference model shapes the call/abstain decision. This checkpoint uses a
|
||||
**balanced 1:1 call/abstain** batch with a gentle learning rate, which **holds the
|
||||
high-abstention behavior** while keeping JSON/tool-selection intact.
|
||||
|
||||
**What the experiment found (the POC result):** at 50M parameters there is a genuine
|
||||
**precision ↔ recall frontier** on the call decision — you can maximize valid calls *or*
|
||||
maximize correct refusals, but not both. GRPO with a perfect verifiable reward moves
|
||||
*along* this frontier rather than beyond it; this checkpoint deliberately sits at the
|
||||
**robust-abstention** corner. (A separately-tested 2.5× larger student tied these
|
||||
metrics, corroborating that the ceiling is capacity, not the training signal.)
|
||||
|
||||
---
|
||||
|
||||
## Evaluation
|
||||
|
||||
Numbers are reported honestly for this exact checkpoint against the sibling SFT/RL
|
||||
checkpoints (see the training repo for the full trial log).
|
||||
|
||||
**Dialect eval (18 items, Egyptian-dialect prompts):**
|
||||
|
||||
| metric | this model | note |
|
||||
|---|---|---|
|
||||
| strict success | 12/18 | headline |
|
||||
| decision (call/answer correct) | 16/18 | |
|
||||
| **abstention** | **4/4** ✅ | robustly declines out-of-scope requests |
|
||||
|
||||
**MSA eval v2 (126 items, higher-resolution):** strict 11/68, decision 49/68,
|
||||
**abstain 27/46** — the high-abstention end of the decision↔abstain frontier.
|
||||
Tool-selection (~33%) and argument-exactness (~18%) are the capacity-bound bottlenecks
|
||||
shared across all checkpoints at this scale.
|
||||
|
||||
---
|
||||
|
||||
## Usage
|
||||
|
||||
The tokenizer uses the `TokenizersBackend` class, which requires **`transformers>=5.12`**.
|
||||
Build the ChatML prompt manually and pass the tool schema in the system message:
|
||||
|
||||
```python
|
||||
import json, torch
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
|
||||
model_id = "oddadmix/Emhotob-50M-GPRO-Arabic-Final"
|
||||
tok = AutoTokenizer.from_pretrained(model_id)
|
||||
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.bfloat16)
|
||||
|
||||
tools = [{
|
||||
"name": "get_weather",
|
||||
"description": "Get the current weather for a city",
|
||||
"parameters": {"type": "object",
|
||||
"properties": {"location": {"type": "string"}},
|
||||
"required": ["location"]},
|
||||
}]
|
||||
|
||||
system = (
|
||||
"أنت مساعد يستطيع استدعاء الأدوات. الأدوات المتاحة:\n"
|
||||
+ json.dumps(tools, ensure_ascii=False)
|
||||
+ "\nإذا لم تكن هناك أداة مناسبة، اعتذر ولا تختلق استدعاءً."
|
||||
)
|
||||
user = "ما حالة الطقس في القاهرة؟"
|
||||
|
||||
prompt = (
|
||||
f"<|im_start|>system\n{system}<|im_end|>\n"
|
||||
f"<|im_start|>user\n{user}<|im_end|>\n"
|
||||
f"<|im_start|>assistant\n"
|
||||
)
|
||||
ids = tok(prompt, return_tensors="pt").to(model.device)
|
||||
out = model.generate(
|
||||
**ids, max_new_tokens=256, do_sample=False,
|
||||
repetition_penalty=1.2,
|
||||
eos_token_id=tok.convert_tokens_to_ids("<|im_end|>"),
|
||||
)
|
||||
print(tok.decode(out[0][ids.input_ids.shape[1]:], skip_special_tokens=True))
|
||||
# Expected: <tool_call>{"name": "get_weather", "arguments": {"location": "القاهرة"}}</tool_call>
|
||||
```
|
||||
|
||||
For an out-of-scope request with no matching tool, the model is tuned to **refuse in
|
||||
Arabic** rather than hallucinate a call.
|
||||
|
||||
---
|
||||
|
||||
## Intended use & limitations
|
||||
|
||||
**Intended use.** Research and demonstration of Arabic tool-calling / agentic behavior at
|
||||
tiny scale; a CPU-friendly baseline for from-scratch Arabic SLM experiments.
|
||||
|
||||
**Limitations.** At ~50M parameters this is a **proof of concept**. Tool-selection with
|
||||
several confusable tools (~33% correct) and exact argument extraction (~18%) are the
|
||||
capacity-bound weak spots — validate any tool call before executing it. Training was
|
||||
predominantly Modern Standard Arabic; dialect prompts are harder. Not for production
|
||||
agents without a validation/guardrail layer.
|
||||
|
||||
---
|
||||
|
||||
## Citation & credits
|
||||
|
||||
- **Method:** GRPO — Shao et al., 2024 (*DeepSeekMath*); sequence-level KD — Kim & Rush, 2016.
|
||||
- **Teacher:** [`LiquidAI/LFM2-1.2B-Tool`](https://huggingface.co/LiquidAI/LFM2-1.2B-Tool).
|
||||
- **Base architecture & training scripts:** [`SupraLabs/Supra-50M-Base`](https://huggingface.co/SupraLabs/Supra-50M-Base) (Apache-2.0).
|
||||
- **Base Arabic model:** [`oddadmix/50M-2048-Emhotob`](https://huggingface.co/oddadmix/50M-2048-Emhotob) — from-scratch, ~20B Arabic tokens, 2048 ctx.
|
||||
32
config.json
Normal file
32
config.json
Normal file
@@ -0,0 +1,32 @@
|
||||
{
|
||||
"architectures": [
|
||||
"LlamaForCausalLM"
|
||||
],
|
||||
"attention_bias": false,
|
||||
"attention_dropout": 0.0,
|
||||
"bos_token_id": 0,
|
||||
"dtype": "bfloat16",
|
||||
"eos_token_id": 32001,
|
||||
"head_dim": 64,
|
||||
"hidden_act": "silu",
|
||||
"hidden_size": 512,
|
||||
"initializer_range": 0.02,
|
||||
"intermediate_size": 1408,
|
||||
"max_position_embeddings": 2048,
|
||||
"mlp_bias": false,
|
||||
"model_type": "llama",
|
||||
"num_attention_heads": 8,
|
||||
"num_hidden_layers": 12,
|
||||
"num_key_value_heads": 4,
|
||||
"pad_token_id": 1,
|
||||
"pretraining_tp": 1,
|
||||
"rms_norm_eps": 1e-06,
|
||||
"rope_parameters": {
|
||||
"rope_theta": 10000,
|
||||
"rope_type": "default"
|
||||
},
|
||||
"tie_word_embeddings": true,
|
||||
"transformers_version": "5.12.1",
|
||||
"use_cache": true,
|
||||
"vocab_size": 32002
|
||||
}
|
||||
10
generation_config.json
Normal file
10
generation_config.json
Normal file
@@ -0,0 +1,10 @@
|
||||
{
|
||||
"_from_model_config": true,
|
||||
"bos_token_id": 0,
|
||||
"eos_token_id": 32001,
|
||||
"output_attentions": false,
|
||||
"output_hidden_states": false,
|
||||
"pad_token_id": 1,
|
||||
"transformers_version": "5.12.1",
|
||||
"use_cache": true
|
||||
}
|
||||
3
model.safetensors
Normal file
3
model.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:0efe905f770bc7e5ad6c637ed70f1e21692195a90eaff1d45fe83f0113de6aa4
|
||||
size 103586664
|
||||
159050
tokenizer.json
Normal file
159050
tokenizer.json
Normal file
File diff suppressed because it is too large
Load Diff
15
tokenizer_config.json
Normal file
15
tokenizer_config.json
Normal file
@@ -0,0 +1,15 @@
|
||||
{
|
||||
"backend": "tokenizers",
|
||||
"bos_token": "<s>",
|
||||
"eos_token": "<|im_end|>",
|
||||
"extra_special_tokens": [
|
||||
"<|im_start|>",
|
||||
"<|im_end|>"
|
||||
],
|
||||
"is_local": true,
|
||||
"local_files_only": false,
|
||||
"model_max_length": 1000000000000000019884624838656,
|
||||
"pad_token": "<pad>",
|
||||
"tokenizer_class": "TokenizersBackend",
|
||||
"unk_token": "<unk>"
|
||||
}
|
||||
Reference in New Issue
Block a user