commit 51c0997282f2a4a12efef1455bcbef57dcebb545 Author: ModelHub XC Date: Sat Aug 29 12:48:17 2026 +0800 初始化项目,由ModelHub XC社区提供模型 Model: HAD653/qwen3-1.7b-magistral-math-gguf Source: Original Platform diff --git a/.gitattributes b/.gitattributes new file mode 100644 index 0000000..ea698ed --- /dev/null +++ b/.gitattributes @@ -0,0 +1,57 @@ +*.7z filter=lfs diff=lfs merge=lfs -text +*.arrow filter=lfs diff=lfs merge=lfs -text +*.bin filter=lfs diff=lfs merge=lfs -text +*.bz2 filter=lfs diff=lfs merge=lfs -text +*.ckpt filter=lfs diff=lfs merge=lfs -text +*.ftz filter=lfs diff=lfs merge=lfs -text +*.gz filter=lfs diff=lfs merge=lfs -text +*.h5 filter=lfs diff=lfs merge=lfs -text +*.joblib filter=lfs diff=lfs merge=lfs -text +*.lfs.* filter=lfs diff=lfs merge=lfs -text +*.mlmodel filter=lfs diff=lfs merge=lfs -text +*.model filter=lfs diff=lfs merge=lfs -text +*.msgpack filter=lfs diff=lfs merge=lfs -text +*.npy filter=lfs diff=lfs merge=lfs -text +*.npz filter=lfs diff=lfs merge=lfs -text +*.onnx filter=lfs diff=lfs merge=lfs -text +*.ot filter=lfs diff=lfs merge=lfs -text +*.parquet filter=lfs diff=lfs merge=lfs -text +*.pb filter=lfs diff=lfs merge=lfs -text +*.pickle filter=lfs diff=lfs merge=lfs -text +*.pkl filter=lfs diff=lfs merge=lfs -text +*.pt filter=lfs diff=lfs merge=lfs -text +*.pth filter=lfs diff=lfs merge=lfs -text +*.rar filter=lfs diff=lfs merge=lfs -text +*.safetensors filter=lfs diff=lfs merge=lfs -text +saved_model/**/* filter=lfs diff=lfs merge=lfs -text +*.tar.* filter=lfs diff=lfs merge=lfs -text +*.tar filter=lfs diff=lfs merge=lfs -text +*.tflite filter=lfs diff=lfs merge=lfs -text +*.tgz filter=lfs diff=lfs merge=lfs -text +*.wasm filter=lfs diff=lfs merge=lfs -text +*.xz filter=lfs diff=lfs merge=lfs -text +*.zip filter=lfs diff=lfs merge=lfs -text +*.zst filter=lfs diff=lfs merge=lfs -text +*tfevents* filter=lfs diff=lfs merge=lfs -text +Qwen3-1.7B-Base.F16.gguf filter=lfs diff=lfs merge=lfs -text +Qwen3-1.7B-Base.Q8_0.gguf filter=lfs diff=lfs merge=lfs -text +Qwen3-1.7B-Base.Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text +ggml-vocab-aquila.gguf filter=lfs diff=lfs merge=lfs -text +ggml-vocab-baichuan.gguf filter=lfs diff=lfs merge=lfs -text +ggml-vocab-bert-bge.gguf filter=lfs diff=lfs merge=lfs -text +ggml-vocab-command-r.gguf filter=lfs diff=lfs merge=lfs -text +ggml-vocab-deepseek-coder.gguf filter=lfs diff=lfs merge=lfs -text +ggml-vocab-deepseek-llm.gguf filter=lfs diff=lfs merge=lfs -text +ggml-vocab-falcon.gguf filter=lfs diff=lfs merge=lfs -text +ggml-vocab-gpt-2.gguf filter=lfs diff=lfs merge=lfs -text +ggml-vocab-gpt-neox.gguf filter=lfs diff=lfs merge=lfs -text +ggml-vocab-llama-bpe.gguf filter=lfs diff=lfs merge=lfs -text +ggml-vocab-llama-spm.gguf filter=lfs diff=lfs merge=lfs -text +ggml-vocab-mpt.gguf filter=lfs diff=lfs merge=lfs -text +ggml-vocab-nomic-bert-moe.gguf filter=lfs diff=lfs merge=lfs -text +ggml-vocab-phi-3.gguf filter=lfs diff=lfs merge=lfs -text +ggml-vocab-qwen2.gguf filter=lfs diff=lfs merge=lfs -text +ggml-vocab-refact.gguf filter=lfs diff=lfs merge=lfs -text +ggml-vocab-starcoder.gguf filter=lfs diff=lfs merge=lfs -text +Qwen3-1.7B-Magistral-Math-F16.gguf filter=lfs diff=lfs merge=lfs -text +Qwen3-1.7B-Magistral-Math-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text diff --git a/Qwen3-1.7B-Base.Q4_K_M.gguf b/Qwen3-1.7B-Base.Q4_K_M.gguf new file mode 100644 index 0000000..904bb2a --- /dev/null +++ b/Qwen3-1.7B-Base.Q4_K_M.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:3b54656dd5d3f10c8fd4d7457f3ef02227eebcf86528dd022f1d718b54f014c0 +size 1107404896 diff --git a/Qwen3-1.7B-Magistral-Math-F16.gguf b/Qwen3-1.7B-Magistral-Math-F16.gguf new file mode 100644 index 0000000..61a03b5 --- /dev/null +++ b/Qwen3-1.7B-Magistral-Math-F16.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:995984e86cbf479c7fbdb0a574f64abd08c27f3c3beec6991a1b2b3c7075ed19 +size 3447345248 diff --git a/Qwen3-1.7B-Magistral-Math-Q8_0.gguf b/Qwen3-1.7B-Magistral-Math-Q8_0.gguf new file mode 100644 index 0000000..1c43c70 --- /dev/null +++ b/Qwen3-1.7B-Magistral-Math-Q8_0.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:c5009347f0dd04635b83cedbf6488ccda306c1e8c5b542ab2c303f23d48aed7d +size 1834422368 diff --git a/README.md b/README.md new file mode 100644 index 0000000..001d97b --- /dev/null +++ b/README.md @@ -0,0 +1,306 @@ +--- +language: +- en +license: apache-2.0 +tags: +- gguf +- qwen3 +- unsloth +- math +- mathematical-reasoning +- chain-of-thought +- cot +- gsm8k +- openmath +- instruction-tuning +- quantization +base_model: HAD653/qwen3-1.7b-magistral-math +datasets: +- HAD653/GSM8K-OpenMath-MathReason-13k +pipeline_tag: text-generation + +model-index: +- name: Qwen3-1.7B-Magistral-Math-GGUF + results: [] + quantized_from: + - HAD653/qwen3-1.7b-magistral-math +--- + +# Qwen3-1.7B Magistral Math (GGUF) + +[![License: Apache-2.0](https://img.shields.io/badge/License-Apache--2.0-brightgreen.svg)](https://www.apache.org/licenses/LICENSE-2.0) +![Model: Qwen3-1.7B](https://img.shields.io/badge/Model-Qwen3--1.7B-blue.svg) +![Format: GGUF](https://img.shields.io/badge/Format-GGUF-orange.svg) +![Domain: Math Reasoning](https://img.shields.io/badge/Domain-Math%20Reasoning-purple.svg) +![Quantizations: F16, Q8_0, Q4_K_M](https://img.shields.io/badge/Quants-F16%20%7C%20Q8_0%20%7C%20Q4_K_M-lightgrey.svg) + +--- + +## TL;DR + +This is a **math-focused fine-tune** of [`unsloth/Qwen3-1.7B-Base`](https://huggingface.co/unsloth/Qwen3-1.7B-Base), +exported to **GGUF** (F16 / Q8_0 / Q4_K_M) with **Unsloth**. + +- **Goal:** small 1.7B model specialized for **grade-school & early high-school math reasoning**. +- **Data:** [`HAD653/GSM8K-OpenMath-MathReason-13k`](https://huggingface.co/datasets/HAD653/GSM8K-OpenMath-Magistral-13k) – 13.9k math word problems with structured chain-of-thought. +- **Format:** answers always follow the same pattern: + + ```text + Problem: + ... + + Reasoning: + ... + + Answer: + + ```` + +* **Best use:** GSM8K-style problems, OpenMath-style word problems, step-by-step reasoning with a **single numeric final answer**. + +--- + +## Model Description + +* **Base model:** [`unsloth/Qwen3-1.7B-Base`](https://huggingface.co/unsloth/Qwen3-1.7B-Base) (Apache-2.0) +* **Architecture:** Qwen3 dense causal LM, ~1.7B params, 28 layers, GQA attention, **32k context**. +* **Type:** decoder-only LLM, text generation. +* **This repo:** inference-only **GGUF weights** for llama.cpp / LM Studio / Ollama / text-generation-webui. + +### Available files + +From the **Files** tab: + +* `Qwen3-1.7B-Magistral-Math-F16.gguf` – highest quality, requires the most VRAM. +* `Qwen3-1.7B-Magistral-Math-Q8_0.gguf` – 8-bit quantization. +* `Qwen3-1.7B-Magistral-Math-Q4_K_M.gguf` – 4-bit K-quant, best for smaller GPUs. + +> These files contain **fine-tuned math weights**, exported via `model.save_pretrained_gguf` after full BF16 training. + +--- + +## Training Data + +This model is fine-tuned on: + +* **Dataset:** [`HAD653/GSM8K-OpenMath-MathReason-13k`](https://huggingface.co/datasets/HAD653/GSM8K-OpenMath-MathReason-13k) +* **Size:** 13,857 examples. +* **Fields:** + + * `question`: natural language math word problem. + * `cot`: structured solution with three blocks: + + * `Problem:` + * `Reasoning:` + * `Answer:` + * `final_answer`: canonical numeric answer (string). + +The dataset focuses on **easy–medium difficulty**: +basic arithmetic, fractions, percentages, rate problems, simple algebra, and simple combinatorics – the kind of tasks a **1–3B model can genuinely master**. + +--- + +## Training Setup (Summary) + +Fine-tuning was done with **Unsloth + TRL** on a single **RTX 4090**, using **full BF16 fine-tuning** (no LoRA). + +Main hyperparameters: + +* **Base:** `unsloth/Qwen3-1.7B-Base` +* **Sequence length:** 2048 +* **Batching:** `per_device_train_batch_size = 2`, `gradient_accumulation_steps = 8` +* **Effective batch size:** ≈ 16 sequences +* **Epochs:** 2 +* **Optimizer / schedule:** + + * `learning_rate = 7e-5` + * linear scheduler, `warmup_ratio = 0.05` + * `weight_decay = 0.01` +* **Precision & memory:** + + * `dtype = bfloat16` + * `gradient_checkpointing = True` + +### Supervision format + +The training text for each sample is: + +```text +### Instruction: +{question} + +### Response: +{cot} +``` + +where `` is the tokenizer EOS token. +Adding `eos_token` at the end of each sample teaches the model **when to stop**, which greatly reduces “Answer: 36 / Answer: 36 / …” loops during inference. + +--- + +## Prompting & Templates + +### Recommended system prompt (optional but useful) + +```text +You are a math reasoning assistant. + +For every question, answer in exactly this format: + +Problem: + + +Reasoning: + + +Answer: + + +Do not add any extra commentary before or after the answer. +Do not repeat the answer multiple times. +Stop after writing the final answer. +``` + +### Inference template (matches training) + +Single-turn format: + +```text +### Instruction: +{question} + +### Response: +``` + +The model will then generate: + +```text +Problem: +... + +Reasoning: +... + +Answer: + +``` + +### Stop strings + +On top of the EOS token, you can add **stop strings** in your UI: + +* `### Instruction:` +* `### Response:` + +Many frontends (LM Studio, text-generation-webui, KoboldCpp, etc.) let you configure these so the model stops cleanly when it tries to start the next turn. + +--- + +## Quantization & Hardware Tips + +The three variants in this repo roughly behave as follows (ballpark): + +* **`Q4_K_M` (~1.1 GB)** – best for: + + * 4–6 GB GPUs or pure CPU inference. + * Fast experimentation / local tools / “math assistant on a laptop”. +* **`Q8_0` (~1.8 GB)** – good compromise: + + * 8–12 GB GPUs. + * Often slightly more stable than Q4 on harder problems. +* **`F16` (~3.5 GB)** – highest fidelity: + + * 12+ GB GPUs (4090, 4080, 4070 12GB, A4000 etc.). + * Recommended if VRAM allows and you care about maximum accuracy. + +As a rule of thumb, choose a file that is **1–2 GB smaller than your available VRAM**. + +--- + +## Usage Examples + +### llama.cpp + +Once you have built `llama.cpp`, you can run the model like this (replace with your path): + +```bash +./llama-cli \ + -m Qwen3-1.7B-Magistral-Math-Q4_K_M.gguf \ + -p "### Instruction: +Albert buys 2 large pizzas and 2 small pizzas. A large pizza has 16 slices and a small pizza has 8 slices. If he eats it all, how many pieces does he eat that day? + +### Response: +" \ + -n 256 \ + --temp 0.1 \ + --top-p 0.9 \ + --repeat-penalty 1.05 +``` + +Suggested decoding for math: + +* `temperature`: 0.0–0.2 +* `top_p`: 0.9 +* `repeat_penalty`: 1.05–1.1 +* `top_k`: 20–40 (optional tweak) + +### LM Studio / other UIs + +Set the **prompt template** to: + +```text +### Instruction: +{{prompt}} + +### Response: +``` + +Add stop strings: + +* `### Instruction:` +* `### Response:` + +and keep temperature low for math benchmarks. + +--- + +## Intended Uses & Limitations + +### Intended uses + +* Solving **GSM8K-style** and **OpenMath-style** word problems. +* Training / evaluating **small-scale math reasoning pipelines**. +* Serving as a **local math tutor** for grade-school / early high-school algebra & arithmetic. + +### Limitations + +* Not a general chat/instruction model; it is **biased toward math**. +* CoT is learned from **synthetic teacher traces**, not human-written solutions. +* Not suitable for **high-stakes educational or decision-making** without human oversight. +* Performance on very hard competition math (Olympiad-level, deep proofs) will be limited – the training data explicitly focuses on **easy–medium difficulty**. + +Users are responsible for ensuring there is no data leakage if they evaluate on GSM8K/OpenMath-derived benchmarks. + +--- + +## Acknowledgements + +* Base model: [Qwen/Qwen3-1.7B-Base](https://huggingface.co/unsloth/Qwen3-1.7B-Base) and the Qwen / Unsloth teams. +* Unsloth for fast fine-tuning and GGUF export. +* Training data: [`HAD653/GSM8K-OpenMath-MathReason-13k`](https://huggingface.co/datasets/HAD653/GSM8K-OpenMath-MathReason-13k). + +--- + +## Citation + +If you use this model in your work, please cite: + +```bibtex +@misc{had653_qwen3_magistral_math_gguf_2025, + author = {HAD653}, + title = {Qwen3-1.7B Magistral Math (GGUF): A 1.7B Math Reasoning Model with Magistral Chain-of-Thought}, + year = {2025}, + howpublished = {\url{https://huggingface.co/HAD653/qwen3-1.7b-magistral-math-gguf}}, + note = {Fine-tuned on GSM8K + OpenMath MathReason 13k, exported to GGUF (F16 / Q8\_0 / Q4\_K\_M).} +} +``` \ No newline at end of file