license, language, library_name, pipeline_tag, tags, base_model, datasets
license language library_name pipeline_tag tags base_model datasets
apache-2.0
en
gguf text-generation
gguf
llama.cpp
quantized
granite
reasoning
unsloth
text-generation
ermiaazarkhalili/Granite-4.1-3B-SFT-Claude-Opus-Reasoning-Unsloth
ermiaazarkhalili/claude-reasoning-distillation

Granite-4.1-3B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF

GGUF quantizations of Granite-4.1-3B-SFT-Claude-Opus-Reasoning-Unsloth for CPU and edge inference with llama.cpp, Ollama, LM Studio, and other GGUF runtimes.

This model is a fine-tune of Granite 4.1-3B for reasoning distillation (chain-of-thought), trained with Unsloth on claude-reasoning-distillation.

Available Quantizations

Quant Size Recommended Use File
Q2_K 1.37 GB Smallest, lowest quality — quick tests / very constrained devices granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q2_k.gguf
Q3_K_M 1.73 GB Small, acceptable quality granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q3_k_m.gguf
Q4_K_M 2.10 GB Recommended — best size/quality balance granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q4_k_m.gguf
Q5_K_M 2.44 GB High quality, larger granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q5_k_m.gguf
Q6_K 2.80 GB Very high quality, near-fp16 granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q6_k.gguf
Q8_0 3.62 GB Near-lossless, largest granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q8_0.gguf

Q4_K_M is the recommended default for most users.

Usage

Download a single quant

pip install -U "huggingface_hub[cli]"
hf download ermiaazarkhalili/Granite-4.1-3B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF \
  --include "*q4_k_m*.gguf" --local-dir ./Granite-4.1-3B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF

llama.cpp

# build: https://github.com/ggerganov/llama.cpp
./llama-cli -m ./Granite-4.1-3B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF/granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q2_k.gguf \
  -p "Solve step by step: What is the sum of the first 10 prime numbers?" -n 512

Ollama

ollama run hf.co/ermiaazarkhalili/Granite-4.1-3B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF:Q4_K_M "Solve step by step: What is the sum of the first 10 prime numbers?"

Training Outcome

Metric Value
SLURM Job ID 38330896
Runtime 29m 27s
Final Training Loss 0.8932
Peak VRAM 10.18 GB
GPU H100 80GB HBM3 (MIG 3g.40gb)

License

APACHE-2.0 — see the base model for full terms.

Acknowledgments

Description
Model synced from source: ermiaazarkhalili/Granite-4.1-3B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF
Readme 26 KiB