GRPO (main table, base+GRPO) — final checkpoint (SFT → GRPO). Base: Qwen/Qwen3-1.7B.
Paper: Metacognitive Behavioral Tuning of Large Language Models for Multi-Hop Question Answering.
Method: GRPO: base Qwen3 trained directly with GRPO (no SFT stage).
Base model: Qwen/Qwen3-1.7B
Training: SFT (LR 1e-4, BS 128, HotpotQA) → GRPO
Benchmarks: HotpotQA (ID), MuSiQue / 2WikiMultiHopQA (OOD)