license, language, library_name, pipeline_tag, tags, base_model, datasets
license language library_name pipeline_tag tags base_model datasets
apache-2.0
en
gguf text-generation
gguf
llama.cpp
quantized
qwen3
reasoning
unsloth
text-generation
ermiaazarkhalili/Qwen3-4B-SFT-Claude-Opus-Reasoning-Unsloth
ermiaazarkhalili/claude-reasoning-distillation

Qwen3-4B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF

GGUF quantizations of Qwen3-4B-SFT-Claude-Opus-Reasoning-Unsloth for CPU and edge inference with llama.cpp, Ollama, LM Studio, and other GGUF runtimes.

This model is a fine-tune of Qwen3-4B (Unsloth 4-bit) for reasoning distillation (chain-of-thought), trained with Unsloth on claude-reasoning-distillation.

Available Quantizations

Quant Size Recommended Use File
Q2_K 1.67 GB Smallest, lowest quality — quick tests / very constrained devices qwen3-4b-sft-claude-opus-reasoning-unsloth.q2_k.gguf
Q3_K_M 2.08 GB Small, acceptable quality qwen3-4b-sft-claude-opus-reasoning-unsloth.q3_k_m.gguf
Q4_K_M 2.50 GB Recommended — best size/quality balance qwen3-4b-sft-claude-opus-reasoning-unsloth.q4_k_m.gguf
Q5_K_M 2.89 GB High quality, larger qwen3-4b-sft-claude-opus-reasoning-unsloth.q5_k_m.gguf
Q6_K 3.31 GB Very high quality, near-fp16 qwen3-4b-sft-claude-opus-reasoning-unsloth.q6_k.gguf
Q8_0 4.28 GB Near-lossless, largest qwen3-4b-sft-claude-opus-reasoning-unsloth.q8_0.gguf

Q4_K_M is the recommended default for most users.

Usage

Download a single quant

pip install -U "huggingface_hub[cli]"
hf download ermiaazarkhalili/Qwen3-4B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF \
  --include "*q4_k_m*.gguf" --local-dir ./Qwen3-4B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF

llama.cpp

# build: https://github.com/ggerganov/llama.cpp
./llama-cli -m ./Qwen3-4B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF/qwen3-4b-sft-claude-opus-reasoning-unsloth.q2_k.gguf \
  -p "Solve step by step: What is the sum of the first 10 prime numbers?" -n 512

Ollama

ollama run hf.co/ermiaazarkhalili/Qwen3-4B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF:Q4_K_M "Solve step by step: What is the sum of the first 10 prime numbers?"

Training Outcome

Metric Value
SLURM Job ID 36885896
Runtime 21m 03s
Final Training Loss 0.998
Peak VRAM 14.01 GB
GPU H100 80GB HBM3 (MIG 3g.40gb)

License

APACHE-2.0 — see the base model for full terms.

Acknowledgments

Description
Model synced from source: ermiaazarkhalili/Qwen3-4B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF
Readme 26 KiB