101 lines
4.5 KiB
Markdown
101 lines
4.5 KiB
Markdown
---
|
|
license: apache-2.0
|
|
base_model: HuggingFaceTB/SmolLM3-3B-Base
|
|
base_model_relation: finetune
|
|
tags:
|
|
- smollm3
|
|
- fine-tuned
|
|
- lora
|
|
- math
|
|
- conversational
|
|
- text-generation-inference
|
|
language:
|
|
- en
|
|
datasets:
|
|
- openai/gsm8k
|
|
---
|
|
|
|
# Al-Khwarizmi-3B
|
|
|
|
*An AI math tutor named after Muhammad al-Khwarizmi, the 9th-century mathematician whose name is the origin of the word "algorithm".*
|
|
|
|
Fine-tune of **HuggingFaceTB/SmolLM3-3B-Base**, trained in two stages — full fine-tuning followed by LoRA — to solve grade-school math word problems with clear, step-by-step reasoning.
|
|
|
|
## Highlights
|
|
|
|
- **87.21% mean token accuracy** on held-out validation data — up from 84.75% after the initial full fine-tune, and 82.4% at the very first checkpoint
|
|
- **Validation loss reduced by ~19%** across the full pipeline (0.678 → 0.462), with training and validation loss tracking closely throughout every stage — no overfitting observed
|
|
- Trained on the complete GSM8K dataset (both `main` and `socratic` reasoning styles) across two LoRA passes, on top of an initial full fine-tune
|
|
- Available in this repo as safetensors. **GGUF version with quantization** (BF16 and Q8_0) for efficient usage on CPU and a smaller size [**available here**](https://huggingface.co/mzoelfakar/Al-Khwarizmi-3B-GGUF).
|
|
|
|

|
|

|
|
|
|
## Training Details
|
|
|
|
**Stage 1 — Full fine-tuning**
|
|
|
|
| | |
|
|
|---|---|
|
|
| Dataset | GSM8K (`main`), 1,000 random samples, 90/10 train/val split |
|
|
| Steps | 450 (1 epoch) |
|
|
| Learning rate | 5e-5, cosine schedule |
|
|
| Final validation loss / accuracy | 0.569 / 84.75% |
|
|
|
|
**Stage 2 — LoRA fine-tuning**
|
|
|
|
| | |
|
|
|---|---|
|
|
| Method | LoRA, r=16, all-linear target modules |
|
|
| Dataset | Full GSM8K — both `main` and `socratic` reasoning styles |
|
|
| Steps | 3,550 (combined across two passes) |
|
|
| Learning rate | 5e-5, cosine schedule |
|
|
| Final validation loss / accuracy | 0.462 / 87.21% |
|
|
|
|
*Run as two consecutive passes over the dataset, with the `main`/`socratic` split swapped between them so every problem was seen in both reasoning styles.*
|
|
|
|
## Limitations
|
|
|
|
Fine-tuned primarily on GSM8K-style problems (single correct numeric answer, grade-school arithmetic/word problems) — performance on more complex, multi-part, or differently-structured math problems is untested. Occasional arithmetic slips on multi-step problems can still occur, consistent with known limitations of models at this scale.
|
|
|
|
## Usage
|
|
|
|
```python
|
|
from transformers import AutoModelForCausalLM, AutoTokenizer
|
|
import torch
|
|
|
|
model_name = "mzoelfakar/Al-Khwarizmi-3B"
|
|
tokenizer = AutoTokenizer.from_pretrained(model_name)
|
|
model = AutoModelForCausalLM.from_pretrained(model_name, dtype=torch.bfloat16, device_map="auto")
|
|
|
|
messages = [
|
|
{"role": "system", "content": "You are a math tutor. Solve problems step by step."},
|
|
{"role": "user", "content": "If a train travels 120 miles in 2 hours, what is its average speed?"}
|
|
]
|
|
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
|
|
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
|
|
outputs = model.generate(**inputs, max_new_tokens=300, temperature=0.7, do_sample=True)
|
|
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
|
|
```
|
|
|
|
### Note on raw output formatting
|
|
|
|
Because this model was fine-tuned on GSM8K (including the `socratic` reasoning style), raw generations may contain training artifacts not meant for direct display:
|
|
|
|
- `<<...>>` — calculator-style intermediate annotations
|
|
- `**` — separator between a sub-question and its calculation in Socratic-style reasoning (not Markdown bold)
|
|
- `#### <answer>` — marker preceding the final numeric answer
|
|
- `*` — used as a multiplication sign (e.g. `8*9`); if two or more appear in the same
|
|
response, Markdown may pair them as emphasis delimiters, causing text between them
|
|
to render in italic with the asterisks hidden
|
|
|
|
If you're piping output through a Markdown renderer or displaying it in a UI, you'll likely want to strip or reformat these first, since `**` in particular can be misread as Markdown bold syntax if left unescaped.
|
|
|
|
## Try it online
|
|
|
|
A live chat demo is available via Colab: [Al-Khwarizmi-3B.ipynb](https://colab.research.google.com/github/mzoelfakar/Al-Khwarizmi-3B/blob/main/Al-Khwarizmi-3B.ipynb)
|
|
|
|
## Credits
|
|
|
|
Fine-tuned by [Mohamed Zoelfakar](https://www.linkedin.com/in/mzoelfakar/), as part of Hugging Face's [smol-course](https://huggingface.co/learn/smol-course/).
|