36 lines
1.0 KiB
Markdown
36 lines
1.0 KiB
Markdown
|
|
---
|
||
|
|
base_model: Qwen/Qwen3-0.6B
|
||
|
|
language:
|
||
|
|
- en
|
||
|
|
library_name: transformers
|
||
|
|
license: apache-2.0
|
||
|
|
pipeline_tag: text-generation
|
||
|
|
tags:
|
||
|
|
- metacognitive-behavioral-tuning
|
||
|
|
- multi-hop-qa
|
||
|
|
- reasoning
|
||
|
|
- sft
|
||
|
|
- grpo
|
||
|
|
- MBT-R
|
||
|
|
---
|
||
|
|
|
||
|
|
# Qwen3-0.6B · MBT-R
|
||
|
|
|
||
|
|
MBT-R (main table) — **final checkpoint (SFT → GRPO)**. Base: [`Qwen/Qwen3-0.6B`](https://huggingface.co/Qwen/Qwen3-0.6B).
|
||
|
|
Paper: *Metacognitive Behavioral Tuning of Large Language Models for Multi-Hop Question Answering*.
|
||
|
|
|
||
|
|
- **Method**: MBT-R (Refinement): the student's own reasoning traces are rewritten into the 5-phase structure for SFT, then GRPO.
|
||
|
|
- **Base model**: `Qwen/Qwen3-0.6B`
|
||
|
|
- **Training**: SFT (LR 1e-4, BS 128, HotpotQA) → GRPO
|
||
|
|
- **Benchmarks**: HotpotQA (ID), MuSiQue / 2WikiMultiHopQA (OOD)
|
||
|
|
|
||
|
|
## Usage
|
||
|
|
|
||
|
|
```python
|
||
|
|
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||
|
|
|
||
|
|
repo = "metacognitive-behavioral-tuning/Qwen3-0.6B-MBT-R"
|
||
|
|
tok = AutoTokenizer.from_pretrained(repo)
|
||
|
|
model = AutoModelForCausalLM.from_pretrained(repo, dtype="bfloat16", device_map="auto")
|
||
|
|
```
|