commit a0b0c6e4c96bceade260021ce7602ee0d18a85af Author: ModelHub XC Date: Sat Jul 25 08:27:09 2026 +0800 初始化项目,由ModelHub XC社区提供模型 Model: ermiaazarkhalili/Qwen3-8B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF Source: Original Platform diff --git a/.gitattributes b/.gitattributes new file mode 100644 index 0000000..9b80461 --- /dev/null +++ b/.gitattributes @@ -0,0 +1,41 @@ +*.7z filter=lfs diff=lfs merge=lfs -text +*.arrow filter=lfs diff=lfs merge=lfs -text +*.bin filter=lfs diff=lfs merge=lfs -text +*.bz2 filter=lfs diff=lfs merge=lfs -text +*.ckpt filter=lfs diff=lfs merge=lfs -text +*.ftz filter=lfs diff=lfs merge=lfs -text +*.gz filter=lfs diff=lfs merge=lfs -text +*.h5 filter=lfs diff=lfs merge=lfs -text +*.joblib filter=lfs diff=lfs merge=lfs -text +*.lfs.* filter=lfs diff=lfs merge=lfs -text +*.mlmodel filter=lfs diff=lfs merge=lfs -text +*.model filter=lfs diff=lfs merge=lfs -text +*.msgpack filter=lfs diff=lfs merge=lfs -text +*.npy filter=lfs diff=lfs merge=lfs -text +*.npz filter=lfs diff=lfs merge=lfs -text +*.onnx filter=lfs diff=lfs merge=lfs -text +*.ot filter=lfs diff=lfs merge=lfs -text +*.parquet filter=lfs diff=lfs merge=lfs -text +*.pb filter=lfs diff=lfs merge=lfs -text +*.pickle filter=lfs diff=lfs merge=lfs -text +*.pkl filter=lfs diff=lfs merge=lfs -text +*.pt filter=lfs diff=lfs merge=lfs -text +*.pth filter=lfs diff=lfs merge=lfs -text +*.rar filter=lfs diff=lfs merge=lfs -text +*.safetensors filter=lfs diff=lfs merge=lfs -text +saved_model/**/* filter=lfs diff=lfs merge=lfs -text +*.tar.* filter=lfs diff=lfs merge=lfs -text +*.tar filter=lfs diff=lfs merge=lfs -text +*.tflite filter=lfs diff=lfs merge=lfs -text +*.tgz filter=lfs diff=lfs merge=lfs -text +*.wasm filter=lfs diff=lfs merge=lfs -text +*.xz filter=lfs diff=lfs merge=lfs -text +*.zip filter=lfs diff=lfs merge=lfs -text +*.zst filter=lfs diff=lfs merge=lfs -text +*tfevents* filter=lfs diff=lfs merge=lfs -text +qwen3-8b-sft-claude-opus-reasoning-unsloth.q2_k.gguf filter=lfs diff=lfs merge=lfs -text +qwen3-8b-sft-claude-opus-reasoning-unsloth.q3_k_m.gguf filter=lfs diff=lfs merge=lfs -text +qwen3-8b-sft-claude-opus-reasoning-unsloth.q4_k_m.gguf filter=lfs diff=lfs merge=lfs -text +qwen3-8b-sft-claude-opus-reasoning-unsloth.q5_k_m.gguf filter=lfs diff=lfs merge=lfs -text +qwen3-8b-sft-claude-opus-reasoning-unsloth.q6_k.gguf filter=lfs diff=lfs merge=lfs -text +qwen3-8b-sft-claude-opus-reasoning-unsloth.q8_0.gguf filter=lfs diff=lfs merge=lfs -text diff --git a/README.md b/README.md new file mode 100644 index 0000000..49a93db --- /dev/null +++ b/README.md @@ -0,0 +1,86 @@ +--- +license: apache-2.0 +language: + - en +library_name: gguf +pipeline_tag: text-generation +tags: + - gguf + - llama.cpp + - quantized + - qwen3 + - reasoning + - unsloth + - text-generation +base_model: ermiaazarkhalili/Qwen3-8B-SFT-Claude-Opus-Reasoning-Unsloth +datasets: + - ermiaazarkhalili/claude-reasoning-distillation +--- + +# Qwen3-8B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF + +GGUF quantizations of [Qwen3-8B-SFT-Claude-Opus-Reasoning-Unsloth](https://huggingface.co/ermiaazarkhalili/Qwen3-8B-SFT-Claude-Opus-Reasoning-Unsloth) for **CPU and edge inference** with [llama.cpp](https://github.com/ggerganov/llama.cpp), [Ollama](https://ollama.com), LM Studio, and other GGUF runtimes. + +This model is a fine-tune of [Qwen3-8B (Unsloth 4-bit)](https://huggingface.co/unsloth/qwen3-8b-unsloth-bnb-4bit) for **reasoning distillation (chain-of-thought)**, trained with [Unsloth](https://github.com/unslothai/unsloth) on [claude-reasoning-distillation](https://huggingface.co/datasets/ermiaazarkhalili/claude-reasoning-distillation). + +- **Full-precision model:** [Qwen3-8B-SFT-Claude-Opus-Reasoning-Unsloth](https://huggingface.co/ermiaazarkhalili/Qwen3-8B-SFT-Claude-Opus-Reasoning-Unsloth) +- **Base model:** [Qwen3-8B (Unsloth 4-bit)](https://huggingface.co/unsloth/qwen3-8b-unsloth-bnb-4bit) +- **Parameters:** 8B + +## Available Quantizations + +| Quant | Size | Recommended Use | File | +|-------|------|-----------------|------| +| `Q2_K` | 3.28 GB | Smallest, lowest quality — quick tests / very constrained devices | [`qwen3-8b-sft-claude-opus-reasoning-unsloth.q2_k.gguf`](https://huggingface.co/ermiaazarkhalili/Qwen3-8B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF/blob/main/qwen3-8b-sft-claude-opus-reasoning-unsloth.q2_k.gguf) | +| `Q3_K_M` | 4.12 GB | Small, acceptable quality | [`qwen3-8b-sft-claude-opus-reasoning-unsloth.q3_k_m.gguf`](https://huggingface.co/ermiaazarkhalili/Qwen3-8B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF/blob/main/qwen3-8b-sft-claude-opus-reasoning-unsloth.q3_k_m.gguf) | +| `Q4_K_M` | 5.03 GB | Recommended — best size/quality balance | [`qwen3-8b-sft-claude-opus-reasoning-unsloth.q4_k_m.gguf`](https://huggingface.co/ermiaazarkhalili/Qwen3-8B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF/blob/main/qwen3-8b-sft-claude-opus-reasoning-unsloth.q4_k_m.gguf) | +| `Q5_K_M` | 5.85 GB | High quality, larger | [`qwen3-8b-sft-claude-opus-reasoning-unsloth.q5_k_m.gguf`](https://huggingface.co/ermiaazarkhalili/Qwen3-8B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF/blob/main/qwen3-8b-sft-claude-opus-reasoning-unsloth.q5_k_m.gguf) | +| `Q6_K` | 6.73 GB | Very high quality, near-fp16 | [`qwen3-8b-sft-claude-opus-reasoning-unsloth.q6_k.gguf`](https://huggingface.co/ermiaazarkhalili/Qwen3-8B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF/blob/main/qwen3-8b-sft-claude-opus-reasoning-unsloth.q6_k.gguf) | +| `Q8_0` | 8.71 GB | Near-lossless, largest | [`qwen3-8b-sft-claude-opus-reasoning-unsloth.q8_0.gguf`](https://huggingface.co/ermiaazarkhalili/Qwen3-8B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF/blob/main/qwen3-8b-sft-claude-opus-reasoning-unsloth.q8_0.gguf) | + +`Q4_K_M` is the recommended default for most users. + +## Usage + +### Download a single quant + +```bash +pip install -U "huggingface_hub[cli]" +hf download ermiaazarkhalili/Qwen3-8B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF \ + --include "*q4_k_m*.gguf" --local-dir ./Qwen3-8B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF +``` + +### llama.cpp + +```bash +# build: https://github.com/ggerganov/llama.cpp +./llama-cli -m ./Qwen3-8B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF/qwen3-8b-sft-claude-opus-reasoning-unsloth.q2_k.gguf \ + -p "Solve step by step: What is the sum of the first 10 prime numbers?" -n 512 +``` + +### Ollama + +```bash +ollama run hf.co/ermiaazarkhalili/Qwen3-8B-SFT-Claude-Opus-Reasoning-Unsloth-GGUF:Q4_K_M "Solve step by step: What is the sum of the first 10 prime numbers?" +``` + +## Training Outcome + +| Metric | Value | +|--------|-------| +| SLURM Job ID | `36885901` | +| Runtime | 40m 30s | +| Final Training Loss | 0.8753 | +| Peak VRAM | 14.23 GB | +| GPU | H100 80GB HBM3 (MIG 3g.40gb) | + +## License + +APACHE-2.0 — see the [base model](https://huggingface.co/unsloth/qwen3-8b-unsloth-bnb-4bit) for full terms. + +## Acknowledgments + +- [Unsloth](https://github.com/unslothai/unsloth) for 2x faster fine-tuning +- [llama.cpp](https://github.com/ggerganov/llama.cpp) for GGUF quantization +- Base model developers (unsloth) +- [Compute Canada / DRAC](https://alliancecan.ca/) for HPC resources diff --git a/qwen3-8b-sft-claude-opus-reasoning-unsloth.q2_k.gguf b/qwen3-8b-sft-claude-opus-reasoning-unsloth.q2_k.gguf new file mode 100644 index 0000000..2a274a1 --- /dev/null +++ b/qwen3-8b-sft-claude-opus-reasoning-unsloth.q2_k.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:657d46831171f130f2f66eb028b6a1f7855415819357f1da64929655c0fdaa9d +size 3281732800 diff --git a/qwen3-8b-sft-claude-opus-reasoning-unsloth.q3_k_m.gguf b/qwen3-8b-sft-claude-opus-reasoning-unsloth.q3_k_m.gguf new file mode 100644 index 0000000..c1ec6b7 --- /dev/null +++ b/qwen3-8b-sft-claude-opus-reasoning-unsloth.q3_k_m.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:9c76eaaf16b31be4daac93ff5c2d9b5f1ef011ad6c6d232fdddbb730c432192c +size 4124161216 diff --git a/qwen3-8b-sft-claude-opus-reasoning-unsloth.q4_k_m.gguf b/qwen3-8b-sft-claude-opus-reasoning-unsloth.q4_k_m.gguf new file mode 100644 index 0000000..7746f33 --- /dev/null +++ b/qwen3-8b-sft-claude-opus-reasoning-unsloth.q4_k_m.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:4dcbdfd372562533dc9b0615fc4f7290a0c95ccbe7385a9d160fe05615141f4e +size 5027783872 diff --git a/qwen3-8b-sft-claude-opus-reasoning-unsloth.q5_k_m.gguf b/qwen3-8b-sft-claude-opus-reasoning-unsloth.q5_k_m.gguf new file mode 100644 index 0000000..0d43836 --- /dev/null +++ b/qwen3-8b-sft-claude-opus-reasoning-unsloth.q5_k_m.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:08475cc348df2e196fbac076a9d5f24bbdef24ef76a3de30c13001df956653d9 +size 5851112640 diff --git a/qwen3-8b-sft-claude-opus-reasoning-unsloth.q6_k.gguf b/qwen3-8b-sft-claude-opus-reasoning-unsloth.q6_k.gguf new file mode 100644 index 0000000..d8e7936 --- /dev/null +++ b/qwen3-8b-sft-claude-opus-reasoning-unsloth.q6_k.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:18ab76828da81c808840c7f505ea9e47046a8cc4b0ef30d23a887eab7dbf44f9 +size 6725899456 diff --git a/qwen3-8b-sft-claude-opus-reasoning-unsloth.q8_0.gguf b/qwen3-8b-sft-claude-opus-reasoning-unsloth.q8_0.gguf new file mode 100644 index 0000000..5d9f756 --- /dev/null +++ b/qwen3-8b-sft-claude-opus-reasoning-unsloth.q8_0.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:8e21da36f0a58102a4083e25bf00677224a51df4c942b1244a62ec8c00b4b838 +size 8709518528