初始化项目,由ModelHub XC社区提供模型
Model: metacognitive-behavioral-tuning/Qwen3-0.6B-MBT-R Source: Original Platform
This commit is contained in:
35
README.md
Normal file
35
README.md
Normal file
@@ -0,0 +1,35 @@
|
||||
---
|
||||
base_model: Qwen/Qwen3-0.6B
|
||||
language:
|
||||
- en
|
||||
library_name: transformers
|
||||
license: apache-2.0
|
||||
pipeline_tag: text-generation
|
||||
tags:
|
||||
- metacognitive-behavioral-tuning
|
||||
- multi-hop-qa
|
||||
- reasoning
|
||||
- sft
|
||||
- grpo
|
||||
- MBT-R
|
||||
---
|
||||
|
||||
# Qwen3-0.6B · MBT-R
|
||||
|
||||
MBT-R (main table) — **final checkpoint (SFT → GRPO)**. Base: [`Qwen/Qwen3-0.6B`](https://huggingface.co/Qwen/Qwen3-0.6B).
|
||||
Paper: *Metacognitive Behavioral Tuning of Large Language Models for Multi-Hop Question Answering*.
|
||||
|
||||
- **Method**: MBT-R (Refinement): the student's own reasoning traces are rewritten into the 5-phase structure for SFT, then GRPO.
|
||||
- **Base model**: `Qwen/Qwen3-0.6B`
|
||||
- **Training**: SFT (LR 1e-4, BS 128, HotpotQA) → GRPO
|
||||
- **Benchmarks**: HotpotQA (ID), MuSiQue / 2WikiMultiHopQA (OOD)
|
||||
|
||||
## Usage
|
||||
|
||||
```python
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
|
||||
repo = "metacognitive-behavioral-tuning/Qwen3-0.6B-MBT-R"
|
||||
tok = AutoTokenizer.from_pretrained(repo)
|
||||
model = AutoModelForCausalLM.from_pretrained(repo, dtype="bfloat16", device_map="auto")
|
||||
```
|
||||
Reference in New Issue
Block a user