--- base_model: Qwen/Qwen3-0.6B language: - en library_name: transformers license: apache-2.0 pipeline_tag: text-generation tags: - metacognitive-behavioral-tuning - multi-hop-qa - reasoning - sft - grpo - gpt-oss-distill --- # Qwen3-0.6B ยท gpt-oss-distill gpt-oss-distill (appendix) โ€” **final checkpoint (SFT โ†’ GRPO)**. Base: [`Qwen/Qwen3-0.6B`](https://huggingface.co/Qwen/Qwen3-0.6B). Paper: *Metacognitive Behavioral Tuning of Large Language Models for Multi-Hop Question Answering*. - **Method**: gpt-oss-distill: naive distillation of teacher (gpt-oss-120b) raw traces via SFT, then GRPO. - **Base model**: `Qwen/Qwen3-0.6B` - **Training**: SFT (LR 1e-4, BS 128, HotpotQA) โ†’ GRPO - **Benchmarks**: HotpotQA (ID), MuSiQue / 2WikiMultiHopQA (OOD) ## Usage ```python from transformers import AutoModelForCausalLM, AutoTokenizer repo = "metacognitive-behavioral-tuning/Qwen3-0.6B-gpt-oss-distill" tok = AutoTokenizer.from_pretrained(repo) model = AutoModelForCausalLM.from_pretrained(repo, dtype="bfloat16", device_map="auto") ```