commit a08851bdd122d5593e537693ef3a928c626652a3 Author: ModelHub XC Date: Sat Jul 25 08:04:11 2026 +0800 初始化项目,由ModelHub XC社区提供模型 Model: ermiaazarkhalili/Granite-4.1-3B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF Source: Original Platform diff --git a/.gitattributes b/.gitattributes new file mode 100644 index 0000000..2942885 --- /dev/null +++ b/.gitattributes @@ -0,0 +1,41 @@ +*.7z filter=lfs diff=lfs merge=lfs -text +*.arrow filter=lfs diff=lfs merge=lfs -text +*.bin filter=lfs diff=lfs merge=lfs -text +*.bz2 filter=lfs diff=lfs merge=lfs -text +*.ckpt filter=lfs diff=lfs merge=lfs -text +*.ftz filter=lfs diff=lfs merge=lfs -text +*.gz filter=lfs diff=lfs merge=lfs -text +*.h5 filter=lfs diff=lfs merge=lfs -text +*.joblib filter=lfs diff=lfs merge=lfs -text +*.lfs.* filter=lfs diff=lfs merge=lfs -text +*.mlmodel filter=lfs diff=lfs merge=lfs -text +*.model filter=lfs diff=lfs merge=lfs -text +*.msgpack filter=lfs diff=lfs merge=lfs -text +*.npy filter=lfs diff=lfs merge=lfs -text +*.npz filter=lfs diff=lfs merge=lfs -text +*.onnx filter=lfs diff=lfs merge=lfs -text +*.ot filter=lfs diff=lfs merge=lfs -text +*.parquet filter=lfs diff=lfs merge=lfs -text +*.pb filter=lfs diff=lfs merge=lfs -text +*.pickle filter=lfs diff=lfs merge=lfs -text +*.pkl filter=lfs diff=lfs merge=lfs -text +*.pt filter=lfs diff=lfs merge=lfs -text +*.pth filter=lfs diff=lfs merge=lfs -text +*.rar filter=lfs diff=lfs merge=lfs -text +*.safetensors filter=lfs diff=lfs merge=lfs -text +saved_model/**/* filter=lfs diff=lfs merge=lfs -text +*.tar.* filter=lfs diff=lfs merge=lfs -text +*.tar filter=lfs diff=lfs merge=lfs -text +*.tflite filter=lfs diff=lfs merge=lfs -text +*.tgz filter=lfs diff=lfs merge=lfs -text +*.wasm filter=lfs diff=lfs merge=lfs -text +*.xz filter=lfs diff=lfs merge=lfs -text +*.zip filter=lfs diff=lfs merge=lfs -text +*.zst filter=lfs diff=lfs merge=lfs -text +*tfevents* filter=lfs diff=lfs merge=lfs -text +granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q2_k.gguf filter=lfs diff=lfs merge=lfs -text +granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q3_k_m.gguf filter=lfs diff=lfs merge=lfs -text +granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q4_k_m.gguf filter=lfs diff=lfs merge=lfs -text +granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q5_k_m.gguf filter=lfs diff=lfs merge=lfs -text +granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q6_k.gguf filter=lfs diff=lfs merge=lfs -text +granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q8_0.gguf filter=lfs diff=lfs merge=lfs -text diff --git a/README.md b/README.md new file mode 100644 index 0000000..fe1dc83 --- /dev/null +++ b/README.md @@ -0,0 +1,86 @@ +--- +license: apache-2.0 +language: + - en +library_name: gguf +pipeline_tag: text-generation +tags: + - gguf + - llama.cpp + - quantized + - granite + - reasoning + - unsloth + - text-generation +base_model: ermiaazarkhalili/Granite-4.1-3B-SFT-Claude-Opus-Reasoning-Unsloth +datasets: + - ermiaazarkhalili/claude-reasoning-distillation +--- + +# Granite-4.1-3B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF + +GGUF quantizations of [Granite-4.1-3B-SFT-Claude-Opus-Reasoning-Unsloth](https://huggingface.co/ermiaazarkhalili/Granite-4.1-3B-SFT-Claude-Opus-Reasoning-Unsloth) for **CPU and edge inference** with [llama.cpp](https://github.com/ggerganov/llama.cpp), [Ollama](https://ollama.com), LM Studio, and other GGUF runtimes. + +This model is a fine-tune of [Granite 4.1-3B](https://huggingface.co/ibm-granite/granite-4.1-3b) for **reasoning distillation (chain-of-thought)**, trained with [Unsloth](https://github.com/unslothai/unsloth) on [claude-reasoning-distillation](https://huggingface.co/datasets/ermiaazarkhalili/claude-reasoning-distillation). + +- **Full-precision model:** [Granite-4.1-3B-SFT-Claude-Opus-Reasoning-Unsloth](https://huggingface.co/ermiaazarkhalili/Granite-4.1-3B-SFT-Claude-Opus-Reasoning-Unsloth) +- **Base model:** [Granite 4.1-3B](https://huggingface.co/ibm-granite/granite-4.1-3b) +- **Parameters:** 3B + +## Available Quantizations + +| Quant | Size | Recommended Use | File | +|-------|------|-----------------|------| +| `Q2_K` | 1.37 GB | Smallest, lowest quality — quick tests / very constrained devices | [`granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q2_k.gguf`](https://huggingface.co/ermiaazarkhalili/Granite-4.1-3B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF/blob/main/granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q2_k.gguf) | +| `Q3_K_M` | 1.73 GB | Small, acceptable quality | [`granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q3_k_m.gguf`](https://huggingface.co/ermiaazarkhalili/Granite-4.1-3B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF/blob/main/granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q3_k_m.gguf) | +| `Q4_K_M` | 2.10 GB | Recommended — best size/quality balance | [`granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q4_k_m.gguf`](https://huggingface.co/ermiaazarkhalili/Granite-4.1-3B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF/blob/main/granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q4_k_m.gguf) | +| `Q5_K_M` | 2.44 GB | High quality, larger | [`granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q5_k_m.gguf`](https://huggingface.co/ermiaazarkhalili/Granite-4.1-3B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF/blob/main/granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q5_k_m.gguf) | +| `Q6_K` | 2.80 GB | Very high quality, near-fp16 | [`granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q6_k.gguf`](https://huggingface.co/ermiaazarkhalili/Granite-4.1-3B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF/blob/main/granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q6_k.gguf) | +| `Q8_0` | 3.62 GB | Near-lossless, largest | [`granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q8_0.gguf`](https://huggingface.co/ermiaazarkhalili/Granite-4.1-3B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF/blob/main/granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q8_0.gguf) | + +`Q4_K_M` is the recommended default for most users. + +## Usage + +### Download a single quant + +```bash +pip install -U "huggingface_hub[cli]" +hf download ermiaazarkhalili/Granite-4.1-3B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF \ + --include "*q4_k_m*.gguf" --local-dir ./Granite-4.1-3B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF +``` + +### llama.cpp + +```bash +# build: https://github.com/ggerganov/llama.cpp +./llama-cli -m ./Granite-4.1-3B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF/granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q2_k.gguf \ + -p "Solve step by step: What is the sum of the first 10 prime numbers?" -n 512 +``` + +### Ollama + +```bash +ollama run hf.co/ermiaazarkhalili/Granite-4.1-3B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF:Q4_K_M "Solve step by step: What is the sum of the first 10 prime numbers?" +``` + +## Training Outcome + +| Metric | Value | +|--------|-------| +| SLURM Job ID | `38330896` | +| Runtime | 29m 27s | +| Final Training Loss | 0.8932 | +| Peak VRAM | 10.18 GB | +| GPU | H100 80GB HBM3 (MIG 3g.40gb) | + +## License + +APACHE-2.0 — see the [base model](https://huggingface.co/ibm-granite/granite-4.1-3b) for full terms. + +## Acknowledgments + +- [Unsloth](https://github.com/unslothai/unsloth) for 2x faster fine-tuning +- [llama.cpp](https://github.com/ggerganov/llama.cpp) for GGUF quantization +- Base model developers (ibm-granite) +- [Compute Canada / DRAC](https://alliancecan.ca/) for HPC resources diff --git a/granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q2_k.gguf b/granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q2_k.gguf new file mode 100644 index 0000000..bc84673 --- /dev/null +++ b/granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q2_k.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:189f6691a788162742feb3dece125d86026879ef148a52bc87523450028f5fc1 +size 1371437568 diff --git a/granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q3_k_m.gguf b/granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q3_k_m.gguf new file mode 100644 index 0000000..6ce600c --- /dev/null +++ b/granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q3_k_m.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:49aeb056d259a8866022a5fe2c25f03da4499afd2ad0cfd50034adde946cc6a4 +size 1725577728 diff --git a/granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q4_k_m.gguf b/granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q4_k_m.gguf new file mode 100644 index 0000000..1cd8e1d --- /dev/null +++ b/granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q4_k_m.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:356ca23a98d16b4967031934ffdbba6eb63df38e282c586a84d2a02a67791e9d +size 2099501568 diff --git a/granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q5_k_m.gguf b/granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q5_k_m.gguf new file mode 100644 index 0000000..55b92b5 --- /dev/null +++ b/granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q5_k_m.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:71e5d9e92673039a7b755825afbd919a558d63e2d1ee32dd2bc86ec246094630 +size 2437011968 diff --git a/granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q6_k.gguf b/granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q6_k.gguf new file mode 100644 index 0000000..e860904 --- /dev/null +++ b/granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q6_k.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:73ef07ca2ecbbf94dbf24f70e94a75c72a13b62fe79cbedd57bbe071c5e8b66e +size 2795616768 diff --git a/granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q8_0.gguf b/granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q8_0.gguf new file mode 100644 index 0000000..c8cc51e --- /dev/null +++ b/granite-4.1-3b-sft-claude-opus-reasoning-unsloth.q8_0.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:058f5089e19866fc29deda8aff96fb5ced6767a7c5891fad3e554187ec880f53 +size 3619691008