初始化项目,由ModelHub XC社区提供模型

Model: metacognitive-behavioral-tuning/Qwen3-0.6B-MBT-R
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-05 20:05:22 +08:00
commit 47c380e84a
12 changed files with 151919 additions and 0 deletions

35
README.md Normal file
View File

@@ -0,0 +1,35 @@
---
base_model: Qwen/Qwen3-0.6B
language:
- en
library_name: transformers
license: apache-2.0
pipeline_tag: text-generation
tags:
- metacognitive-behavioral-tuning
- multi-hop-qa
- reasoning
- sft
- grpo
- MBT-R
---
# Qwen3-0.6B · MBT-R
MBT-R (main table) — **final checkpoint (SFT → GRPO)**. Base: [`Qwen/Qwen3-0.6B`](https://huggingface.co/Qwen/Qwen3-0.6B).
Paper: *Metacognitive Behavioral Tuning of Large Language Models for Multi-Hop Question Answering*.
- **Method**: MBT-R (Refinement): the student's own reasoning traces are rewritten into the 5-phase structure for SFT, then GRPO.
- **Base model**: `Qwen/Qwen3-0.6B`
- **Training**: SFT (LR 1e-4, BS 128, HotpotQA) → GRPO
- **Benchmarks**: HotpotQA (ID), MuSiQue / 2WikiMultiHopQA (OOD)
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "metacognitive-behavioral-tuning/Qwen3-0.6B-MBT-R"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, dtype="bfloat16", device_map="auto")
```