Files
qwen2.5-1.5b-exam-tutor/README.md
ModelHub XC 7420692989 初始化项目,由ModelHub XC社区提供模型
Model: ptvnck/qwen2.5-1.5b-exam-tutor
Source: Original Platform
2026-08-23 07:47:16 +08:00

178 lines
7.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
license: apache-2.0
base_model: unsloth/Qwen2.5-1.5B-Instruct
language:
- en
- ru
library_name: transformers
pipeline_tag: text-generation
tags:
- qwen2
- unsloth
- trl
- sft
- lora
- education
- tutoring
- conversational
datasets:
- ptvnck/TutoringDialogs
- eth-nlped/mathdial
model-index:
- name: qwen2.5-1.5b-exam-tutor
results: []
---
<div align="center">
<span style="font-size:44px; font-weight:bold">Qwen2.5-1.5B Exam Tutor</span>
**Tutoring assistant for preparing to exams, fine-tuned to help you *think*, not just get answers.**
[![Base Model](https://img.shields.io/badge/base-Qwen2.5--1.5B--Instruct-blue)](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
[![License](https://img.shields.io/badge/license-Apache%202.0-green)](https://www.apache.org/licenses/LICENSE-2.0)
[![Made with Unsloth](https://img.shields.io/badge/made%20with-Unsloth%20%2B%20TRL-orange)](https://github.com/unslothai/unsloth)
[![Method](https://img.shields.io/badge/method-LoRA%20SFT-purple)]()
</div>
---
## Overview
`qwen2.5-1.5b-exam-tutor` is a LoRA fine-tune of [`Qwen2.5-1.5B-Instruct`](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct), trained to behave like a **patient human tutor** rather than an answer-dispensing machine. Instead of solving a problem outright, the model is trained to ask guiding questions, probe for misconceptions, and walk the student toward the solution themselves — the same pattern a good teacher uses during exam prep.
This model is the **first stage** of a larger personal-assistant project for exam preparation, which also includes a RAG pipeline over practice problems and a FastAPI serving layer accelerated with vLLM/Ollama.
> This is an educational / portfolio project, not a production system. Scope, dataset size, and evaluation depth are intentionally sized for a learning exercise — see [Limitations](#-limitations--scope) below.
## Model Details
| | |
|---|---|
| **Base model** | [`unsloth/Qwen2.5-1.5B-Instruct`](https://huggingface.co/unsloth/Qwen2.5-1.5B-Instruct) |
| **Fine-tuning method** | LoRA (rank 16), full precision (no 4-bit quantization) |
| **Frameworks** | [Unsloth](https://github.com/unslothai/unsloth) + [TRL](https://github.com/huggingface/trl) `SFTTrainer` |
| **Weights format** | Merged 16-bit safetensors (adapter merged into base) |
| **Language** | English / Russian |
| **License** | Apache 2.0 (inherited from base model) |
## Intended Use
- Conversational tutoring for exam preparation: math word problems, conceptual explanations, step-by-step reasoning practice.
- Designed to be embedded as the generation backend of a larger RAG + FastAPI tutoring assistant (see [Roadmap](#-roadmap)).
- **Not intended** as a general-purpose assistant, factual knowledge base, or replacement for a real teacher — the model's job is to *guide*, its factual accuracy on niche topics is not separately verified.
## Training Data
650 student ↔ tutor dialogues, combined from two sources:
| Source | Dialogues used | Notes |
|---|---|---|
| [`ptvnck/TutoringDialogs`](https://huggingface.co/datasets/ptvnck/TutoringDialogs) | 500 | Synthetically generated, manually curated tutoring dialogues across mixed subjects |
| [`eth-nlped/mathdial`](https://huggingface.co/datasets/eth-nlped/mathdial) | 150 | Filtered (dialogues with >11 turns) and reformatted subset, added specifically to cover math word-problem tutoring, which was underrepresented in the primary dataset |
Data was split 85/15 into train/validation (≈552 / 98 examples), formatted with the tokenizer's ChatML template, and capped at 2500 tokens (covering the 99th percentile of dialogue length with no truncation).
## Training Procedure
<details>
<summary><b>LoRA configuration</b></summary>
| Parameter | Value |
|---|---|
| Rank (`r`) | 16 |
| Alpha (`lora_alpha`) | 32 |
| Target modules | `q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj` |
| Dropout | 0.1 |
| Bias | none |
| Gradient checkpointing | Unsloth-optimized |
</details>
<details>
<summary><b>Optimization hyperparameters</b></summary>
| Parameter | Value |
|---|---|
| Effective batch size | 12 (4 × grad. accumulation 3) |
| Epochs | 4 (best checkpoint auto-selected) |
| Learning rate | 2e-4, cosine schedule |
| Warmup | 10% of total steps |
| Optimizer | AdamW (torch) |
| Precision | fp16 |
| Loss masking | Response-only (`train_on_responses_only`) — loss computed exclusively on tutor turns |
| Hardware | 1× NVIDIA T4 (Google Colab) |
</details>
**Response-only loss masking.** Only the tutor's turns contribute to the training loss; the student's turns are masked out. This keeps the adapter's limited capacity focused entirely on learning *how to tutor*, rather than also learning to imitate the student side of the conversation.
## Results
| Epoch | Training Loss | Validation Loss |
|---|---|---|
| 1 | 1.655 | 1.724 |
| 2 | 1.389 | 1.516 |
| **3** | **1.050** | **1.506** ← best |
| 4 | 0.703 | 1.580 |
- **Best validation loss:** 1.506 (epoch 3) → **perplexity ≈ 4.51**
- Training loss keeps decreasing through epoch 4, while validation loss starts rising after epoch 3 — a clear sign of overfitting setting in on the final epoch, expected given the modest dataset size (~550 training examples).
- `load_best_model_at_end=True` automatically restored the epoch-3 checkpoint as the final model, so the released weights are **not** the last-epoch weights, but the best-validation checkpoint.
Full training curves (loss, LR schedule) were tracked with Weights & Biases.
## How to Use
**With 🤗 Transformers:**
```python
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("ptvnck/qwen2.5-1.5b-exam-tutor")
model = AutoModelForCausalLM.from_pretrained("ptvnck/qwen2.5-1.5b-exam-tutor")
messages = [
{"role": "user", "content": "I need to solve 2x + 5 = 15 but I don't know where to start."}
]
inputs = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, tokenize=True,
return_dict=True, return_tensors="pt"
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=200, temperature=0.7)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))
```
**With Unsloth (2x faster inference):**
```python
from unsloth import FastLanguageModel
model, tokenizer = FastLanguageModel.from_pretrained(
model_name="ptvnck/qwen2.5-1.5b-exam-tutor",
max_seq_length=2048,
)
FastLanguageModel.for_inference(model)
```
**With vLLM:**
```bash
pip install vllm
vllm serve "ptvnck/qwen2.5-1.5b-exam-tutor"
```
## Limitations & Scope
- Trained on 650 dialogues — sufficient to learn a tutoring *pattern*, but not a broad knowledge base. Expect a strong grasp of *conversational tutoring style*, and a shallower grasp of niche subject-matter facts.
- No dedicated generation-quality evaluation (human eval / LLM-as-judge) was run as part of this stage — this is deferred to the RAG + FastAPI integration stage of the project, where end-to-end assistant responses will be evaluated in context rather than in isolation.
- Not safety-tuned beyond what the base `Qwen2.5-1.5B-Instruct` already provides.
## Acknowledgements
- Base model: [Qwen2.5-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct) by the Qwen team
- Training accelerated with [Unsloth](https://github.com/unslothai/unsloth)
- Trained using Hugging Face [TRL](https://github.com/huggingface/trl)
- `mathdial` subset: [eth-nlped/mathdial](https://huggingface.co/datasets/eth-nlped/mathdial)