gpt-oss-distill (appendix) — final checkpoint (SFT → GRPO). Base: Qwen/Qwen3-1.7B.
Paper: Metacognitive Behavioral Tuning of Large Language Models for Multi-Hop Question Answering.
Method: gpt-oss-distill: naive distillation of teacher (gpt-oss-120b) raw traces via SFT, then GRPO.
Base model: Qwen/Qwen3-1.7B
Training: SFT (LR 1e-4, BS 128, HotpotQA) → GRPO
Benchmarks: HotpotQA (ID), MuSiQue / 2WikiMultiHopQA (OOD)