Files
Al-Khwarizmi-3B/README.md

101 lines
4.5 KiB
Markdown
Raw Permalink Normal View History

---
license: apache-2.0
base_model: HuggingFaceTB/SmolLM3-3B-Base
base_model_relation: finetune
tags:
- smollm3
- fine-tuned
- lora
- math
- conversational
- text-generation-inference
language:
- en
datasets:
- openai/gsm8k
---
# Al-Khwarizmi-3B
*An AI math tutor named after Muhammad al-Khwarizmi, the 9th-century mathematician whose name is the origin of the word "algorithm".*
Fine-tune of **HuggingFaceTB/SmolLM3-3B-Base**, trained in two stages — full fine-tuning followed by LoRA — to solve grade-school math word problems with clear, step-by-step reasoning.
## Highlights
- **87.21% mean token accuracy** on held-out validation data — up from 84.75% after the initial full fine-tune, and 82.4% at the very first checkpoint
- **Validation loss reduced by ~19%** across the full pipeline (0.678 → 0.462), with training and validation loss tracking closely throughout every stage — no overfitting observed
- Trained on the complete GSM8K dataset (both `main` and `socratic` reasoning styles) across two LoRA passes, on top of an initial full fine-tune
- Available in this repo as safetensors. **GGUF version with quantization** (BF16 and Q8_0) for efficient usage on CPU and a smaller size [**available here**](https://huggingface.co/mzoelfakar/Al-Khwarizmi-3B-GGUF).
![Full fine-tune: Training vs Validation Loss](loss_chart.png)
![LoRA fine-tune: Training vs Validation Loss](lora_loss_chart.png)
## Training Details
**Stage 1 — Full fine-tuning**
| | |
|---|---|
| Dataset | GSM8K (`main`), 1,000 random samples, 90/10 train/val split |
| Steps | 450 (1 epoch) |
| Learning rate | 5e-5, cosine schedule |
| Final validation loss / accuracy | 0.569 / 84.75% |
**Stage 2 — LoRA fine-tuning**
| | |
|---|---|
| Method | LoRA, r=16, all-linear target modules |
| Dataset | Full GSM8K — both `main` and `socratic` reasoning styles |
| Steps | 3,550 (combined across two passes) |
| Learning rate | 5e-5, cosine schedule |
| Final validation loss / accuracy | 0.462 / 87.21% |
*Run as two consecutive passes over the dataset, with the `main`/`socratic` split swapped between them so every problem was seen in both reasoning styles.*
## Limitations
Fine-tuned primarily on GSM8K-style problems (single correct numeric answer, grade-school arithmetic/word problems) — performance on more complex, multi-part, or differently-structured math problems is untested. Occasional arithmetic slips on multi-step problems can still occur, consistent with known limitations of models at this scale.
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_name = "mzoelfakar/Al-Khwarizmi-3B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, dtype=torch.bfloat16, device_map="auto")
messages = [
{"role": "system", "content": "You are a math tutor. Solve problems step by step."},
{"role": "user", "content": "If a train travels 120 miles in 2 hours, what is its average speed?"}
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=300, temperature=0.7, do_sample=True)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```
### Note on raw output formatting
Because this model was fine-tuned on GSM8K (including the `socratic` reasoning style), raw generations may contain training artifacts not meant for direct display:
- `<<...>>` — calculator-style intermediate annotations
- `**` — separator between a sub-question and its calculation in Socratic-style reasoning (not Markdown bold)
- `#### <answer>` — marker preceding the final numeric answer
- `*` — used as a multiplication sign (e.g. `8*9`); if two or more appear in the same
response, Markdown may pair them as emphasis delimiters, causing text between them
to render in italic with the asterisks hidden
If you're piping output through a Markdown renderer or displaying it in a UI, you'll likely want to strip or reformat these first, since `**` in particular can be misread as Markdown bold syntax if left unescaped.
## Try it online
A live chat demo is available via Colab: [Al-Khwarizmi-3B.ipynb](https://colab.research.google.com/github/mzoelfakar/Al-Khwarizmi-3B/blob/main/Al-Khwarizmi-3B.ipynb)
## Credits
Fine-tuned by [Mohamed Zoelfakar](https://www.linkedin.com/in/mzoelfakar/), as part of Hugging Face's [smol-course](https://huggingface.co/learn/smol-course/).