Model: metacognitive-behavioral-tuning/Qwen3-4B-gpt-oss-distill Source: Original Platform
36 lines
1.0 KiB
Markdown
36 lines
1.0 KiB
Markdown
---
|
|
base_model: Qwen/Qwen3-4B
|
|
language:
|
|
- en
|
|
library_name: transformers
|
|
license: apache-2.0
|
|
pipeline_tag: text-generation
|
|
tags:
|
|
- metacognitive-behavioral-tuning
|
|
- multi-hop-qa
|
|
- reasoning
|
|
- sft
|
|
- grpo
|
|
- gpt-oss-distill
|
|
---
|
|
|
|
# Qwen3-4B · gpt-oss-distill
|
|
|
|
gpt-oss-distill (appendix) — **final checkpoint (SFT → GRPO)**. Base: [`Qwen/Qwen3-4B`](https://huggingface.co/Qwen/Qwen3-4B).
|
|
Paper: *Metacognitive Behavioral Tuning of Large Language Models for Multi-Hop Question Answering*.
|
|
|
|
- **Method**: gpt-oss-distill: naive distillation of teacher (gpt-oss-120b) raw traces via SFT, then GRPO.
|
|
- **Base model**: `Qwen/Qwen3-4B`
|
|
- **Training**: SFT (LR 1e-4, BS 128, HotpotQA) → GRPO
|
|
- **Benchmarks**: HotpotQA (ID), MuSiQue / 2WikiMultiHopQA (OOD)
|
|
|
|
## Usage
|
|
|
|
```python
|
|
from transformers import AutoModelForCausalLM, AutoTokenizer
|
|
|
|
repo = "metacognitive-behavioral-tuning/Qwen3-4B-gpt-oss-distill"
|
|
tok = AutoTokenizer.from_pretrained(repo)
|
|
model = AutoModelForCausalLM.from_pretrained(repo, dtype="bfloat16", device_map="auto")
|
|
```
|