MBT-S (main table) — final checkpoint (SFT → GRPO). Base: Qwen/Qwen3-4B.
Paper: Metacognitive Behavioral Tuning of Large Language Models for Multi-Hop Question Answering.
Method: MBT-S (Synthesis): base is SFT'd on gpt-oss-120b-synthesized 5-phase metacognitive traces, then GRPO.
Base model: Qwen/Qwen3-4B
Training: SFT (LR 1e-4, BS 128, HotpotQA) → GRPO
Benchmarks: HotpotQA (ID), MuSiQue / 2WikiMultiHopQA (OOD)