From ecaabdcad3443a11497dd75457b7cd342337a3a5 Mon Sep 17 00:00:00 2001 From: ModelHub XC Date: Tue, 1 Sep 2026 12:58:16 +0800 Subject: [PATCH] =?UTF-8?q?=E5=88=9D=E5=A7=8B=E5=8C=96=E9=A1=B9=E7=9B=AE?= =?UTF-8?q?=EF=BC=8C=E7=94=B1ModelHub=20XC=E7=A4=BE=E5=8C=BA=E6=8F=90?= =?UTF-8?q?=E4=BE=9B=E6=A8=A1=E5=9E=8B?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Model: build-small-hackathon/mind-of-tashi-micro-grpo-gguf Source: Original Platform --- .gitattributes | 37 ++++++++++++++ README.md | 75 ++++++++++++++++++++++++++++ mind-of-tashi-micro-grpo-Q4_K_M.gguf | 3 ++ mind-of-tashi-micro-grpo-f16.gguf | 3 ++ 4 files changed, 118 insertions(+) create mode 100644 .gitattributes create mode 100644 README.md create mode 100644 mind-of-tashi-micro-grpo-Q4_K_M.gguf create mode 100644 mind-of-tashi-micro-grpo-f16.gguf diff --git a/.gitattributes b/.gitattributes new file mode 100644 index 0000000..b856114 --- /dev/null +++ b/.gitattributes @@ -0,0 +1,37 @@ +*.7z filter=lfs diff=lfs merge=lfs -text +*.arrow filter=lfs diff=lfs merge=lfs -text +*.bin filter=lfs diff=lfs merge=lfs -text +*.bz2 filter=lfs diff=lfs merge=lfs -text +*.ckpt filter=lfs diff=lfs merge=lfs -text +*.ftz filter=lfs diff=lfs merge=lfs -text +*.gz filter=lfs diff=lfs merge=lfs -text +*.h5 filter=lfs diff=lfs merge=lfs -text +*.joblib filter=lfs diff=lfs merge=lfs -text +*.lfs.* filter=lfs diff=lfs merge=lfs -text +*.mlmodel filter=lfs diff=lfs merge=lfs -text +*.model filter=lfs diff=lfs merge=lfs -text +*.msgpack filter=lfs diff=lfs merge=lfs -text +*.npy filter=lfs diff=lfs merge=lfs -text +*.npz filter=lfs diff=lfs merge=lfs -text +*.onnx filter=lfs diff=lfs merge=lfs -text +*.ot filter=lfs diff=lfs merge=lfs -text +*.parquet filter=lfs diff=lfs merge=lfs -text +*.pb filter=lfs diff=lfs merge=lfs -text +*.pickle filter=lfs diff=lfs merge=lfs -text +*.pkl filter=lfs diff=lfs merge=lfs -text +*.pt filter=lfs diff=lfs merge=lfs -text +*.pth filter=lfs diff=lfs merge=lfs -text +*.rar filter=lfs diff=lfs merge=lfs -text +*.safetensors filter=lfs diff=lfs merge=lfs -text +saved_model/**/* filter=lfs diff=lfs merge=lfs -text +*.tar.* filter=lfs diff=lfs merge=lfs -text +*.tar filter=lfs diff=lfs merge=lfs -text +*.tflite filter=lfs diff=lfs merge=lfs -text +*.tgz filter=lfs diff=lfs merge=lfs -text +*.wasm filter=lfs diff=lfs merge=lfs -text +*.xz filter=lfs diff=lfs merge=lfs -text +*.zip filter=lfs diff=lfs merge=lfs -text +*.zst filter=lfs diff=lfs merge=lfs -text +*tfevents* filter=lfs diff=lfs merge=lfs -text +mind-of-tashi-micro-grpo-f16.gguf filter=lfs diff=lfs merge=lfs -text +mind-of-tashi-micro-grpo-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text diff --git a/README.md b/README.md new file mode 100644 index 0000000..bc054d3 --- /dev/null +++ b/README.md @@ -0,0 +1,75 @@ +--- +license: apache-2.0 +base_model: build-small-hackathon/mind-of-tashi-micro-grpo +tags: + - gguf + - llama-cpp + - qwen3moe + - reasoning + - game + - bilingual + - grpo +language: + - en + - hi + - sa +pipeline_tag: text-generation +--- + +# The Mind of Tashi — micro student (GRPO, GGUF) + +The GRPO-trained student exported to **GGUF** for llama.cpp. Drop-in +replacement for the SFT GGUF in the playable +[Space](https://huggingface.co/spaces/build-small-hackathon/mind-of-tashi) +after an A/B (winning the game is not enough — the mind-scroll prose must +hold up). Transformers source: +[`…/mind-of-tashi-micro-grpo`](https://huggingface.co/build-small-hackathon/mind-of-tashi-micro-grpo). + +> **Build status:** this GGUF is **built at push time** from the GRPO +> checkpoint — it does not exist as a by-product of training. Use the exact +> same recipe as the SFT GGUF. + +## Files (after build) + +| File | Approx size | Use | +|---|---|---| +| `mind-of-tashi-micro-grpo-Q4_K_M.gguf` | ~256 MB | deployed candidate | +| `mind-of-tashi-micro-grpo-f16.gguf` | ~786 MB | zero-loss reference | + +## Build recipe (no compiled binary needed) + +1. Download the GRPO transformers checkpoint **with `chat_template.jinja`** + (a missing template silently yields a garbage GGUF). +2. `python convert_hf_to_gguf.py --outtype f16` → f16 GGUF. +3. Quantise via the `llama-cpp-python` C binding: + ```python + import ctypes, llama_cpp + p = llama_cpp.llama_model_quantize_default_params() + p.ftype = 15 # LLAMA_FTYPE_MOSTLY_Q4_K_M + llama_cpp.llama_model_quantize(b"in-f16.gguf", b"out-Q4_K_M.gguf", ctypes.byref(p)) + ``` +4. Grade via the format gate through `llama-cpp-python` (the real deploy path); + ship Q4 if it clears ≥15/20 and stays within ~5 ladder points of f16. + +### ⚠️ `norm_topk_prob` — required for llama.cpp + +Inherited `norm_topk_prob=true` from SFT; llama.cpp's `qwen3moe` graph +hardcodes `norm_w=true` and a mismatched checkpoint produces garbage on every +llama.cpp runtime. (See the SFT GGUF card.) + +## Usage + +```python +from llama_cpp import Llama + +llm = Llama.from_pretrained( + repo_id="build-small-hackathon/mind-of-tashi-micro-grpo-gguf", + filename="mind-of-tashi-micro-grpo-Q4_K_M.gguf", + n_ctx=4096, n_gpu_layers=0, logits_all=True, +) +``` + +## Part of the bundle + +Game Space · self-play dataset · SFT model + GGUF · OpenEnv gym · +GRPO model + **GGUF (this)** — all under `build-small-hackathon/mind-of-tashi-*`. diff --git a/mind-of-tashi-micro-grpo-Q4_K_M.gguf b/mind-of-tashi-micro-grpo-Q4_K_M.gguf new file mode 100644 index 0000000..b565971 --- /dev/null +++ b/mind-of-tashi-micro-grpo-Q4_K_M.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:26f6b1b6c48aecc0218e903680b334eb4b86a447f72330551d67c067c9a45a57 +size 255625344 diff --git a/mind-of-tashi-micro-grpo-f16.gguf b/mind-of-tashi-micro-grpo-f16.gguf new file mode 100644 index 0000000..29c9d57 --- /dev/null +++ b/mind-of-tashi-micro-grpo-f16.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:d19d66671e4e01c5adf9ccb5ca9b55822aeeff5eb8dba27507c493fe555992fa +size 786273920