This repository provides a merged 16-bit model produced by fine-tuning
Qwen/Qwen3-4B-Instruct-2507 with QLoRA (4-bit, Unsloth) and then merging the LoRA adapter
into the base model weights.
This model is fully self-contained — no adapter loading required.
Intended as the base model for subsequent DPO training.