173 lines
4.1 KiB
Markdown
173 lines
4.1 KiB
Markdown
|
|
---
|
|||
|
|
license: apache-2.0
|
|||
|
|
language:
|
|||
|
|
- en
|
|||
|
|
tags:
|
|||
|
|
- reasoning
|
|||
|
|
- chain-of-thought
|
|||
|
|
- qwen
|
|||
|
|
- tiny
|
|||
|
|
- whirlwindai
|
|||
|
|
pipeline_tag: text-generation
|
|||
|
|
datasets:
|
|||
|
|
- WhirlwindAI/Soft-CoT-1K
|
|||
|
|
library_name: transformers
|
|||
|
|
base_model:
|
|||
|
|
- Qwen/Qwen2.5-0.5B-Instruct
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
<div align="center">
|
|||
|
|
|
|||
|
|
<img src="https://capsule-render.vercel.app/api?type=waving&height=220&color=gradient&customColorList=12,19,24,30&text=Qwen-R1-0.5B&fontSize=48&fontColor=ffffff&animation=twinkling"/>
|
|||
|
|
|
|||
|
|
<img src="https://readme-typing-svg.demolab.com?font=Space+Grotesk&weight=700&size=27&duration=2300&pause=1200&color=A855F7¢er=true&vCenter=true&width=850&lines=Qwen-R1-0.5B;Reason+First.+Answer+Second.;Chain-of-Thought+on+a+Tiny+Model." />
|
|||
|
|
|
|||
|
|
<img src="https://img.shields.io/badge/Parameters-0.5B-A855F7?style=for-the-badge">
|
|||
|
|
<img src="https://img.shields.io/badge/Base-Qwen2.5--0.5B-7C3AED?style=for-the-badge">
|
|||
|
|
<img src="https://img.shields.io/badge/Trained%20On-Soft--CoT--1K-06B6D4?style=for-the-badge">
|
|||
|
|
<img src="https://img.shields.io/badge/License-Apache--2.0-22C55E?style=for-the-badge">
|
|||
|
|
|
|||
|
|
</div>
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# 💡 The Idea
|
|||
|
|
|
|||
|
|
<div align="center">
|
|||
|
|
|
|||
|
|
> **Good answers come from good thinking.**
|
|||
|
|
|
|||
|
|
Qwen-R1-0.5B is a fine-tuned version of Qwen2.5-0.5B-Instruct trained to **reason before it answers** using explicit `<thinking>` tags.
|
|||
|
|
|
|||
|
|
</div>
|
|||
|
|
|
|||
|
|
Instead of jumping straight to the answer, this model generates its reasoning first — making it more transparent, more reliable, and easier to debug.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# 🧠 How It Works
|
|||
|
|
|
|||
|
|
Every response is structured as:
|
|||
|
|
|
|||
|
|
```
|
|||
|
|
User: {question}
|
|||
|
|
Assistant: <thinking>{reasoning}</thinking>
|
|||
|
|
{answer}
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
The model learns to:
|
|||
|
|
1. **Think** – generate step-by-step reasoning
|
|||
|
|
2. **Answer** – provide the final response
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# 📊 Training Details
|
|||
|
|
|
|||
|
|
| Property | Value |
|
|||
|
|
|----------|-------|
|
|||
|
|
| Base Model | Qwen2.5-0.5B-Instruct |
|
|||
|
|
| Dataset | WhirlwindAI/Soft-CoT-1K |
|
|||
|
|
| Examples | 1,355 |
|
|||
|
|
| Method | QLoRA (4-bit) |
|
|||
|
|
| Epochs | 3 |
|
|||
|
|
| Learning Rate | 2e-4 |
|
|||
|
|
| LoRA Rank | 16 |
|
|||
|
|
| LoRA Alpha | 32 |
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# 🚀 Quick Start
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
from transformers import AutoTokenizer, AutoModelForCausalLM
|
|||
|
|
|
|||
|
|
model_name = "WhirlwindAI/Qwen-R1-0.5B"
|
|||
|
|
|
|||
|
|
tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
|
|||
|
|
model = AutoModelForCausalLM.from_pretrained(model_name, trust_remote_code=True)
|
|||
|
|
|
|||
|
|
prompt = "User: What is 2+2?\nAssistant: <thinking>"
|
|||
|
|
inputs = tokenizer(prompt, return_tensors="pt")
|
|||
|
|
outputs = model.generate(**inputs, max_new_tokens=50, do_sample=True, temperature=0.7)
|
|||
|
|
|
|||
|
|
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# 📋 Sample Output
|
|||
|
|
|
|||
|
|
```
|
|||
|
|
User: What is the capital of France?
|
|||
|
|
Assistant: <thinking>Paris is the capital of France.</thinking>
|
|||
|
|
Paris
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# 📈 Performance
|
|||
|
|
|
|||
|
|
The model was evaluated on 10 out-of-distribution questions:
|
|||
|
|
|
|||
|
|
| Category | Performance |
|
|||
|
|
|----------|-------------|
|
|||
|
|
| Format (thinking tags) | ✅ Excellent |
|
|||
|
|
| General Knowledge | ✅ Good |
|
|||
|
|
| Creative Reasoning | ✅ Good |
|
|||
|
|
| Math/Logic | ⚠️ Needs improvement |
|
|||
|
|
| Physics/Science | ⚠️ Needs improvement |
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# 🔬 What It Learned
|
|||
|
|
|
|||
|
|
| Strength | Weakness |
|
|||
|
|
|----------|----------|
|
|||
|
|
| ✅ Consistent `<thinking>` format | ❌ Sometimes hallucinates facts |
|
|||
|
|
| ✅ Generates reasoning before answering | ❌ Struggles with multi-step math |
|
|||
|
|
| ✅ Retains general knowledge | ❌ Physics reasoning needs more data |
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# 🧪 Test It Yourself
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
questions = [
|
|||
|
|
"What is the capital of France?",
|
|||
|
|
"Explain entropy like I'm 5.",
|
|||
|
|
"Write a short poem about a robot.",
|
|||
|
|
]
|
|||
|
|
|
|||
|
|
for q in questions:
|
|||
|
|
prompt = f"User: {q}\nAssistant: <thinking>"
|
|||
|
|
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
|
|||
|
|
outputs = model.generate(**inputs, max_new_tokens=80)
|
|||
|
|
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# 📜 Citation
|
|||
|
|
|
|||
|
|
```bibtex
|
|||
|
|
@model{qwenr1_2026,
|
|||
|
|
title={Qwen-R1-0.5B},
|
|||
|
|
author={WhirlwindAI},
|
|||
|
|
year={2026},
|
|||
|
|
publisher={Hugging Face}
|
|||
|
|
}
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
<div align="center">
|
|||
|
|
|
|||
|
|
### 🌪️ WhirlwindAI
|
|||
|
|
|
|||
|
|
**Efficient Models • Practical Research • Open AI**
|
|||
|
|
|
|||
|
|
<br>
|
|||
|
|
|
|||
|
|
<img src="https://capsule-render.vercel.app/api?type=waving&height=140§ion=footer&color=0:A855F7,100:06B6D4"/>
|
|||
|
|
|
|||
|
|
</div>
|