Files
Qwen3-0.6B-MBT-R/README.md
ModelHub XC 47c380e84a 初始化项目,由ModelHub XC社区提供模型
Model: metacognitive-behavioral-tuning/Qwen3-0.6B-MBT-R
Source: Original Platform
2026-08-05 20:05:22 +08:00

36 lines
1.0 KiB
Markdown

---
base_model: Qwen/Qwen3-0.6B
language:
- en
library_name: transformers
license: apache-2.0
pipeline_tag: text-generation
tags:
- metacognitive-behavioral-tuning
- multi-hop-qa
- reasoning
- sft
- grpo
- MBT-R
---
# Qwen3-0.6B · MBT-R
MBT-R (main table) — **final checkpoint (SFT → GRPO)**. Base: [`Qwen/Qwen3-0.6B`](https://huggingface.co/Qwen/Qwen3-0.6B).
Paper: *Metacognitive Behavioral Tuning of Large Language Models for Multi-Hop Question Answering*.
- **Method**: MBT-R (Refinement): the student's own reasoning traces are rewritten into the 5-phase structure for SFT, then GRPO.
- **Base model**: `Qwen/Qwen3-0.6B`
- **Training**: SFT (LR 1e-4, BS 128, HotpotQA) → GRPO
- **Benchmarks**: HotpotQA (ID), MuSiQue / 2WikiMultiHopQA (OOD)
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "metacognitive-behavioral-tuning/Qwen3-0.6B-MBT-R"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, dtype="bfloat16", device_map="auto")
```