初始化项目,由ModelHub XC社区提供模型
Model: AaryanK/Qwen_2.5_3B_GRPO_Reasoning_XIOSERV Source: Original Platform
This commit is contained in:
140
README.md
Normal file
140
README.md
Normal file
@@ -0,0 +1,140 @@
|
||||
---
|
||||
license: cc-by-nc-2.0
|
||||
base_model:
|
||||
- Qwen/Qwen2.5-3B-Instruct
|
||||
tags:
|
||||
- Reasoning
|
||||
- GRPO
|
||||
- DeepSeek
|
||||
- CoT
|
||||
- finetune
|
||||
- toggleable-reasoning
|
||||
language:
|
||||
- en
|
||||
pipeline_tag: text-generation
|
||||
---
|
||||
|
||||
# Qwen2.5-3B Toggleable-Reasoning (GRPO)
|
||||
|
||||
A fine-tune of **Qwen2.5-3B-Instruct** with a **runtime-toggleable reasoning mode**: with a short system prompt the model
|
||||
thinks inside `<reasoning>` tags before answering inside `<answer>` tags; without it, the model answers directly like a normal
|
||||
instruct model. One checkpoint, two behaviours, switched at inference time.
|
||||
|
||||
Trained and released by [Aaryan Kapoor](https://huggingface.co/AaryanK) as an independent research project.
|
||||
**Released February 16, 2025.**
|
||||
|
||||
DOI: [10.57967/hf/5366](https://doi.org/10.57967/hf/5366)
|
||||
|
||||
---
|
||||
|
||||
## Toggleable reasoning
|
||||
|
||||
| System prompt | Behaviour |
|
||||
|---|---|
|
||||
| Included (below) | Explicit `<reasoning>` … `</reasoning>` block, then `<answer>` … `</answer>` |
|
||||
| Omitted | Standard conversational response, no visible reasoning |
|
||||
|
||||
System prompt that enables reasoning:
|
||||
|
||||
```
|
||||
Respond in the following format:
|
||||
<reasoning>
|
||||
...
|
||||
</reasoning>
|
||||
<answer>
|
||||
...
|
||||
</answer>
|
||||
```
|
||||
|
||||
The switch is purely prompt-driven — no extra tokens, adapters, or generation flags are required. The model was trained so that
|
||||
the reasoning format is triggered by this instruction and remains dormant otherwise, which keeps default responses short and
|
||||
lets applications opt into chain-of-thought only where it pays off.
|
||||
|
||||
---
|
||||
|
||||
## Training
|
||||
|
||||
- **Base model:** `Qwen/Qwen2.5-3B-Instruct`
|
||||
- **Method:** Group Relative Policy Optimization (GRPO) — the reinforcement-learning objective introduced in
|
||||
[DeepSeekMath](https://arxiv.org/abs/2402.03300) and used to train DeepSeek-R1. Group-relative advantages remove the need for
|
||||
a separate value model, which is what makes RL on reasoning practical at small scale.
|
||||
- **Rewards:** rule-based rewards for answer correctness plus format rewards for well-formed `<reasoning>` / `<answer>` structure.
|
||||
- **Efficiency:** parameter-efficient fine-tuning (QLoRA / PEFT) so the full run fits on a single consumer GPU.
|
||||
- **Goal:** reproduce the "aha-moment" behaviour of R1-style training — the model learning to reflect on and revise its own
|
||||
intermediate steps — in a 3B model, while keeping the ability to answer plainly when reasoning is not requested.
|
||||
|
||||
---
|
||||
|
||||
## Files
|
||||
|
||||
- Safetensors / PyTorch weights (transformers-compatible)
|
||||
- GGUF quantizations for [llama.cpp](https://github.com/ggml-org/llama.cpp): **F16**, **Q8_0**, **Q5_K_M**
|
||||
|
||||
---
|
||||
|
||||
## Usage
|
||||
|
||||
### transformers
|
||||
|
||||
```python
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
|
||||
repo = "AaryanK/Qwen_2.5_3B_GRPO_Reasoning_XIOSERV"
|
||||
tok = AutoTokenizer.from_pretrained(repo)
|
||||
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype="auto", device_map="auto")
|
||||
|
||||
REASONING_PROMPT = (
|
||||
"Respond in the following format:\n<reasoning>\n...\n</reasoning>\n<answer>\n...\n</answer>"
|
||||
)
|
||||
|
||||
messages = [
|
||||
{"role": "system", "content": REASONING_PROMPT}, # drop this line for direct answers
|
||||
{"role": "user", "content": "A bat and a ball cost $1.10 in total. The bat costs $1.00 more than the ball. How much does the ball cost?"},
|
||||
]
|
||||
inputs = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
|
||||
out = model.generate(inputs, max_new_tokens=512)
|
||||
print(tok.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True))
|
||||
```
|
||||
|
||||
### llama.cpp
|
||||
|
||||
```bash
|
||||
llama-cli -m <model>_Q5_K_M.gguf \
|
||||
--system-prompt "Respond in the following format:
|
||||
<reasoning>
|
||||
...
|
||||
</reasoning>
|
||||
<answer>
|
||||
...
|
||||
</answer>" \
|
||||
-p "A bat and a ball cost \$1.10 in total. The bat costs \$1.00 more than the ball. How much does the ball cost?"
|
||||
```
|
||||
|
||||
Omit `--system-prompt` for standard instruct-style answers.
|
||||
|
||||
---
|
||||
|
||||
## Limitations
|
||||
|
||||
- 3B parameters: reasoning helps most on short math, logic, and structured tasks; it does not turn the model into a frontier reasoner.
|
||||
- Reasoning traces are not guaranteed to be faithful to the final answer.
|
||||
- Inherits the base model's knowledge cutoff, biases, and safety behaviour.
|
||||
- Released under **CC BY-NC 2.0** (non-commercial).
|
||||
|
||||
---
|
||||
|
||||
## Citation
|
||||
|
||||
```bibtex
|
||||
@misc{kapoor2025toggleable,
|
||||
author = {Kapoor, Aaryan},
|
||||
title = {Qwen2.5-3B Toggleable-Reasoning (GRPO)},
|
||||
year = {2025},
|
||||
month = feb,
|
||||
doi = {10.57967/hf/5366},
|
||||
url = {https://huggingface.co/AaryanK/Qwen_2.5_3B_GRPO_Reasoning_XIOSERV},
|
||||
publisher = {Hugging Face}
|
||||
}
|
||||
```
|
||||
|
||||
Questions and feedback: open a discussion in the Community tab.
|
||||
Reference in New Issue
Block a user