Files
Qwen_2.5_3B_GRPO_Reasoning_…/README.md

141 lines
4.4 KiB
Markdown
Raw Normal View History

---
license: cc-by-nc-2.0
base_model:
- Qwen/Qwen2.5-3B-Instruct
tags:
- Reasoning
- GRPO
- DeepSeek
- CoT
- finetune
- toggleable-reasoning
language:
- en
pipeline_tag: text-generation
---
# Qwen2.5-3B Toggleable-Reasoning (GRPO)
A fine-tune of **Qwen2.5-3B-Instruct** with a **runtime-toggleable reasoning mode**: with a short system prompt the model
thinks inside `<reasoning>` tags before answering inside `<answer>` tags; without it, the model answers directly like a normal
instruct model. One checkpoint, two behaviours, switched at inference time.
Trained and released by [Aaryan Kapoor](https://huggingface.co/AaryanK) as an independent research project.
**Released February 16, 2025.**
DOI: [10.57967/hf/5366](https://doi.org/10.57967/hf/5366)
---
## Toggleable reasoning
| System prompt | Behaviour |
|---|---|
| Included (below) | Explicit `<reasoning>` … `</reasoning>` block, then `<answer>` … `</answer>` |
| Omitted | Standard conversational response, no visible reasoning |
System prompt that enables reasoning:
```
Respond in the following format:
<reasoning>
...
</reasoning>
<answer>
...
</answer>
```
The switch is purely prompt-driven — no extra tokens, adapters, or generation flags are required. The model was trained so that
the reasoning format is triggered by this instruction and remains dormant otherwise, which keeps default responses short and
lets applications opt into chain-of-thought only where it pays off.
---
## Training
- **Base model:** `Qwen/Qwen2.5-3B-Instruct`
- **Method:** Group Relative Policy Optimization (GRPO) — the reinforcement-learning objective introduced in
[DeepSeekMath](https://arxiv.org/abs/2402.03300) and used to train DeepSeek-R1. Group-relative advantages remove the need for
a separate value model, which is what makes RL on reasoning practical at small scale.
- **Rewards:** rule-based rewards for answer correctness plus format rewards for well-formed `<reasoning>` / `<answer>` structure.
- **Efficiency:** parameter-efficient fine-tuning (QLoRA / PEFT) so the full run fits on a single consumer GPU.
- **Goal:** reproduce the "aha-moment" behaviour of R1-style training — the model learning to reflect on and revise its own
intermediate steps — in a 3B model, while keeping the ability to answer plainly when reasoning is not requested.
---
## Files
- Safetensors / PyTorch weights (transformers-compatible)
- GGUF quantizations for [llama.cpp](https://github.com/ggml-org/llama.cpp): **F16**, **Q8_0**, **Q5_K_M**
---
## Usage
### transformers
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "AaryanK/Qwen_2.5_3B_GRPO_Reasoning_XIOSERV"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype="auto", device_map="auto")
REASONING_PROMPT = (
"Respond in the following format:\n<reasoning>\n...\n</reasoning>\n<answer>\n...\n</answer>"
)
messages = [
{"role": "system", "content": REASONING_PROMPT}, # drop this line for direct answers
{"role": "user", "content": "A bat and a ball cost $1.10 in total. The bat costs $1.00 more than the ball. How much does the ball cost?"},
]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=512)
print(tok.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True))
```
### llama.cpp
```bash
llama-cli -m <model>_Q5_K_M.gguf \
--system-prompt "Respond in the following format:
<reasoning>
...
</reasoning>
<answer>
...
</answer>" \
-p "A bat and a ball cost \$1.10 in total. The bat costs \$1.00 more than the ball. How much does the ball cost?"
```
Omit `--system-prompt` for standard instruct-style answers.
---
## Limitations
- 3B parameters: reasoning helps most on short math, logic, and structured tasks; it does not turn the model into a frontier reasoner.
- Reasoning traces are not guaranteed to be faithful to the final answer.
- Inherits the base model's knowledge cutoff, biases, and safety behaviour.
- Released under **CC BY-NC 2.0** (non-commercial).
---
## Citation
```bibtex
@misc{kapoor2025toggleable,
author = {Kapoor, Aaryan},
title = {Qwen2.5-3B Toggleable-Reasoning (GRPO)},
year = {2025},
month = feb,
doi = {10.57967/hf/5366},
url = {https://huggingface.co/AaryanK/Qwen_2.5_3B_GRPO_Reasoning_XIOSERV},
publisher = {Hugging Face}
}
```
Questions and feedback: open a discussion in the Community tab.