commit f485cd2be3f6e704e67784ef3429a2e742008e8e Author: ModelHub XC Date: Sat Jul 18 05:48:09 2026 +0800 初始化项目,由ModelHub XC社区提供模型 Model: Mungert/gemma-3-1b-it-gguf Source: Original Platform diff --git a/.gitattributes b/.gitattributes new file mode 100644 index 0000000..cdd2797 --- /dev/null +++ b/.gitattributes @@ -0,0 +1,48 @@ +*.7z filter=lfs diff=lfs merge=lfs -text +*.arrow filter=lfs diff=lfs merge=lfs -text +*.bin filter=lfs diff=lfs merge=lfs -text +*.bz2 filter=lfs diff=lfs merge=lfs -text +*.ckpt filter=lfs diff=lfs merge=lfs -text +*.ftz filter=lfs diff=lfs merge=lfs -text +*.gz filter=lfs diff=lfs merge=lfs -text +*.h5 filter=lfs diff=lfs merge=lfs -text +*.joblib filter=lfs diff=lfs merge=lfs -text +*.lfs.* filter=lfs diff=lfs merge=lfs -text +*.mlmodel filter=lfs diff=lfs merge=lfs -text +*.model filter=lfs diff=lfs merge=lfs -text +*.msgpack filter=lfs diff=lfs merge=lfs -text +*.npy filter=lfs diff=lfs merge=lfs -text +*.npz filter=lfs diff=lfs merge=lfs -text +*.onnx filter=lfs diff=lfs merge=lfs -text +*.ot filter=lfs diff=lfs merge=lfs -text +*.parquet filter=lfs diff=lfs merge=lfs -text +*.pb filter=lfs diff=lfs merge=lfs -text +*.pickle filter=lfs diff=lfs merge=lfs -text +*.pkl filter=lfs diff=lfs merge=lfs -text +*.pt filter=lfs diff=lfs merge=lfs -text +*.pth filter=lfs diff=lfs merge=lfs -text +*.rar filter=lfs diff=lfs merge=lfs -text +*.safetensors filter=lfs diff=lfs merge=lfs -text +saved_model/**/* filter=lfs diff=lfs merge=lfs -text +*.tar.* filter=lfs diff=lfs merge=lfs -text +*.tar filter=lfs diff=lfs merge=lfs -text +*.tflite filter=lfs diff=lfs merge=lfs -text +*.tgz filter=lfs diff=lfs merge=lfs -text +*.wasm filter=lfs diff=lfs merge=lfs -text +*.xz filter=lfs diff=lfs merge=lfs -text +*.zip filter=lfs diff=lfs merge=lfs -text +*.zst filter=lfs diff=lfs merge=lfs -text +*tfevents* filter=lfs diff=lfs merge=lfs -text +google_gemma-3-1b-it-q4_1.gguf filter=lfs diff=lfs merge=lfs -text +google_gemma-3-1b-it-bf16-q8.gguf filter=lfs diff=lfs merge=lfs -text +google_gemma-3-1b-it-q6_k.gguf filter=lfs diff=lfs merge=lfs -text +google_gemma-3-1b-it-q4_k.gguf filter=lfs diff=lfs merge=lfs -text +google_gemma-3-1b-it-q4_0.gguf filter=lfs diff=lfs merge=lfs -text +google_gemma-3-1b-it-q5_k_s.gguf filter=lfs diff=lfs merge=lfs -text +google_gemma-3-1b-it-bf16.gguf filter=lfs diff=lfs merge=lfs -text +google_gemma-3-1b-it-iq4_nl.gguf filter=lfs diff=lfs merge=lfs -text +google_gemma-3-1b-it-q5_k.gguf filter=lfs diff=lfs merge=lfs -text +google_gemma-3-1b-it-q8.gguf filter=lfs diff=lfs merge=lfs -text +google_gemma-3-1b-it-q3_k_s.gguf filter=lfs diff=lfs merge=lfs -text +google_gemma-3-1b-it-q4_k_s.gguf filter=lfs diff=lfs merge=lfs -text +google_gemma-3-1b-it-f16-q8.gguf filter=lfs diff=lfs merge=lfs -text diff --git a/README.md b/README.md new file mode 100644 index 0000000..5b5502f --- /dev/null +++ b/README.md @@ -0,0 +1,164 @@ +--- +license: gemma +pipeline_tag: text-generation +tags: +- gemma +--- + +# Gemma-3 1B Instruct GGUF Models + + +**Note llama-quantize was not able to fully quantize the ggufs for k quants as the tensor dimensions of some weights where not divisible by 256. fallback quants where used.** + +## **Choosing the Right Model Format** + +Selecting the correct model format depends on your **hardware capabilities** and **memory constraints**. + +### **BF16 (Brain Float 16) – Use if BF16 acceleration is available** +- A 16-bit floating-point format designed for **faster computation** while retaining good precision. +- Provides **similar dynamic range** as FP32 but with **lower memory usage**. +- Recommended if your hardware supports **BF16 acceleration** (check your device’s specs). +- Ideal for **high-performance inference** with **reduced memory footprint** compared to FP32. + +📌 **Use BF16 if:** +✔ Your hardware has native **BF16 support** (e.g., newer GPUs, TPUs). +✔ You want **higher precision** while saving memory. +✔ You plan to **requantize** the model into another format. + +📌 **Avoid BF16 if:** +❌ Your hardware does **not** support BF16 (it may fall back to FP32 and run slower). +❌ You need compatibility with older devices that lack BF16 optimization. + +--- + +### **F16 (Float 16) – More widely supported than BF16** +- A 16-bit floating-point **high precision** but with less of range of values than BF16. +- Works on most devices with **FP16 acceleration support** (including many GPUs and some CPUs). +- Slightly lower numerical precision than BF16 but generally sufficient for inference. + +📌 **Use F16 if:** +✔ Your hardware supports **FP16** but **not BF16**. +✔ You need a **balance between speed, memory usage, and accuracy**. +✔ You are running on a **GPU** or another device optimized for FP16 computations. + +📌 **Avoid F16 if:** +❌ Your device lacks **native FP16 support** (it may run slower than expected). +❌ You have memory limtations. + +--- + +### **Quantized Models (Q4_K, Q6_K, Q8, etc.) – For CPU & Low-VRAM Inference** +Quantization reduces model size and memory usage while maintaining as much accuracy as possible. +- **Lower-bit models (Q4_K)** → **Best for minimal memory usage**, may have lower precision. +- **Higher-bit models (Q6_K, Q8_0)** → **Better accuracy**, requires more memory. + +📌 **Use Quantized Models if:** +✔ You are running inference on a **CPU** and need an optimized model. +✔ Your device has **low VRAM** and cannot load full-precision models. +✔ You want to reduce **memory footprint** while keeping reasonable accuracy. + +📌 **Avoid Quantized Models if:** +❌ You need **maximum accuracy** (full-precision models are better for this). +❌ Your hardware has enough VRAM for higher-precision formats (BF16/F16). + +--- + +### **Summary Table: Model Format Selection** + +| Model Format | Precision | Memory Usage | Device Requirements | Best Use Case | +|--------------|------------|---------------|----------------------|---------------| +| **BF16** | Highest | High | BF16-supported GPU/CPUs | High-speed inference with reduced memory | +| **F16** | High | High | FP16-supported devices | GPU inference when BF16 isn’t available | +| **Q4_K** | Low | Very Low | CPU or Low-VRAM devices | Best for memory-constrained environments | +| **Q6_K** | Medium Low | Low | CPU with more memory | Better accuracy while still being quantized | +| **Q8** | Medium | Moderate | CPU or GPU with enough VRAM | Best accuracy among quantized models | + + +## **Included Files & Details** + +### `google_gemma-3-1b-it-bf16.gguf` +- Model weights preserved in **BF16**. +- Use this if you want to **requantize** the model into a different format. +- Best if your device supports **BF16 acceleration**. + +### `google_gemma-3-1b-it-f16.gguf` +- Model weights stored in **F16**. +- Use if your device supports **FP16**, especially if BF16 is not available. + +### `google_gemma-3-1b-it-bf16-q8.gguf` +- **Output & embeddings** remain in **BF16**. +- All other layers quantized to **Q8_0**. +- Use if your device supports **BF16** and you want a quantized version. + +### `google_gemma-3-1b-it-f16-q8.gguf` +- **Output & embeddings** remain in **F16**. +- All other layers quantized to **Q8_0**. + +### `google_gemma-3-1b-it-q4_k.gguf` +- **Output & embeddings** quantized to **Q8_0**. +- All other layers quantized to **Q4_K**. +- Good for **CPU inference** with limited memory. + +### `google_gemma-3-1b-it-q4_k_s.gguf` +- Smallest **Q4_K** variant, using less memory at the cost of accuracy. +- Best for **very low-memory setups**. + +### `google_gemma-3-1b-it-q6_k.gguf` +- **Output & embeddings** quantized to **Q8_0**. +- All other layers quantized to **Q6_K** . + + +### `google_gemma-3-1b-it-q8.gguf` +- Fully **Q8** quantized model for better accuracy. +- Requires **more memory** but offers higher precision. + + +# Gemma 3 model card + +**Model Page**: [Gemma](https://ai.google.dev/gemma/docs/core) + +**Resources and Technical Documentation**: + +* [Gemma 3 Technical Report][g3-tech-report] +* [Responsible Generative AI Toolkit][rai-toolkit] +* [Gemma on Kaggle][kaggle-gemma] +* [Gemma on Vertex Model Garden][vertex-mg-gemma3] + +**Terms of Use**: [Terms][terms] + +**Authors**: Google DeepMind + +## Model Information + +Summary description and brief definition of inputs and outputs. + +### Description + +Gemma is a family of lightweight, state-of-the-art open models from Google, +built from the same research and technology used to create the Gemini models. +Gemma 3 1B model handles text only. + +### Inputs and outputs + +- **Input:** + - Text string, such as a question, a prompt, or a document to be summarized + - Total input context 32K tokens for the 1B size + +- **Output:** + - Generated text in response to the input, such as an answer to a + question, analysis of image content, or a summary of a document + - Total output context of 8192 tokens + + +# 🚀 If you find these models useful + +Please give like a click ❤️ . Also I’d really appreciate it if you could test my Network Monitor Assistant at 👉 [Network Monitor Assitant](https://readyforquantum.com). +💬 Click the **chat icon** (bottom right of the main and dashboard pages) . Choose a LLM; toggle between the LLM Types TurboLLM -> FreeLLM -> TestLLM. + +### What I'm Testing +I'm experimenting with **function calling** against my network monitoring service. Using small open source models. I am into the question "How small can it go and still function". +🟡 **TestLLM** – Runs **Phi-4-mini-instruct** using phi-4-mini-q4_0.gguf , llama.cpp on 6 threads of a Cpu VM (Should take about 15s to load. Inference speed is quite slow and it only processes one user prompt at a time—still working on scaling!). If you're curious, I'd be happy to share how it works! . + +### The other Available AI Assistants +🟢 **TurboLLM** – Uses **gpt-4o-mini** Fast! . Note: tokens are limited since OpenAI models are pricey, but you can [Login](https://readyforquantum.com) or [Download](https://readyforquantum.com/download/?utm_source=huggingface&utm_medium=referral&utm_campaign=huggingface_repo_readme) the Quantum Network Monitor agent to get more tokens, Alternatively use the TestLLM . +🔵 **HugLLM** – Runs **open-source Hugging Face models** Fast, Runs small models (≈8B) hence lower quality, Get 2x more tokens (subject to Hugging Face API availability) \ No newline at end of file diff --git a/google_gemma-3-1b-it-bf16-q8.gguf b/google_gemma-3-1b-it-bf16-q8.gguf new file mode 100644 index 0000000..e1792a9 --- /dev/null +++ b/google_gemma-3-1b-it-bf16-q8.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:375bceb9776d8ee1ab7ccac5177b2c8cfc2e1770a8573a7e9fa53c25aa32f2b9 +size 1352422112 diff --git a/google_gemma-3-1b-it-bf16.gguf b/google_gemma-3-1b-it-bf16.gguf new file mode 100644 index 0000000..821ea2b --- /dev/null +++ b/google_gemma-3-1b-it-bf16.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:7c5ed9b8b4c1a68e3e4677b56e615df7eab3d9757874d7c96b210d53266de527 +size 2006573568 diff --git a/google_gemma-3-1b-it-f16-q8.gguf b/google_gemma-3-1b-it-f16-q8.gguf new file mode 100644 index 0000000..cbcafdf --- /dev/null +++ b/google_gemma-3-1b-it-f16-q8.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:7ae47ddde581cab0b9ce60615124598ce117e6fc902c36a926b7ce1be3956925 +size 1352422112 diff --git a/google_gemma-3-1b-it-iq4_nl.gguf b/google_gemma-3-1b-it-iq4_nl.gguf new file mode 100644 index 0000000..d287604 --- /dev/null +++ b/google_gemma-3-1b-it-iq4_nl.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:17897a9cd660436eab0450476977f4c338a5712c9a66f9ef45ee0001aceacb54 +size 721863392 diff --git a/google_gemma-3-1b-it-q3_k_s.gguf b/google_gemma-3-1b-it-q3_k_s.gguf new file mode 100644 index 0000000..7343dbc --- /dev/null +++ b/google_gemma-3-1b-it-q3_k_s.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:0f4ed79efe7598ad2849a28989877cf39f1bde93117f5a57c46921b85bd0c2ff +size 688856288 diff --git a/google_gemma-3-1b-it-q4_0.gguf b/google_gemma-3-1b-it-q4_0.gguf new file mode 100644 index 0000000..5a24e69 --- /dev/null +++ b/google_gemma-3-1b-it-q4_0.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:f15940d368523acb14ede9a4c21b6cf75f510b788186f19f2c2e6828c382dcea +size 720425696 diff --git a/google_gemma-3-1b-it-q4_1.gguf b/google_gemma-3-1b-it-q4_1.gguf new file mode 100644 index 0000000..94785fa --- /dev/null +++ b/google_gemma-3-1b-it-q4_1.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:e251c16da6e376a8acc29812447a0fc8e075096989815e4aeea6d39cfe3422da +size 764035808 diff --git a/google_gemma-3-1b-it-q4_k.gguf b/google_gemma-3-1b-it-q4_k.gguf new file mode 100644 index 0000000..c7be868 --- /dev/null +++ b/google_gemma-3-1b-it-q4_k.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:4c8ef53ada1f6647cd27ab8d4ddb052690d4e3a94cf0fea934bd25312bac8aed +size 806058464 diff --git a/google_gemma-3-1b-it-q4_k_s.gguf b/google_gemma-3-1b-it-q4_k_s.gguf new file mode 100644 index 0000000..a27b8be --- /dev/null +++ b/google_gemma-3-1b-it-q4_k_s.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:879405c451d6937c9f9ec02f7ba1248919f528c570d69bcf662515a317c86f03 +size 780993248 diff --git a/google_gemma-3-1b-it-q5_k.gguf b/google_gemma-3-1b-it-q5_k.gguf new file mode 100644 index 0000000..5802ba2 --- /dev/null +++ b/google_gemma-3-1b-it-q5_k.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:8d1ad26f175f3dec47c88cf03b49d779d5846ab11b1a00c55ae2c1d6a7c2a9fe +size 851345888 diff --git a/google_gemma-3-1b-it-q5_k_s.gguf b/google_gemma-3-1b-it-q5_k_s.gguf new file mode 100644 index 0000000..2c429c3 --- /dev/null +++ b/google_gemma-3-1b-it-q5_k_s.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:99eee8e4fd1ee0d7fb517dadc900af7bf49b3dfb95118f1b1cdf2bc9d103c998 +size 836399840 diff --git a/google_gemma-3-1b-it-q6_k.gguf b/google_gemma-3-1b-it-q6_k.gguf new file mode 100644 index 0000000..43c36b3 --- /dev/null +++ b/google_gemma-3-1b-it-q6_k.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:14feb67da83beaa00558b8aec00fb3f290d5b807ce94ab774c1dcd61effbbf68 +size 1011738848 diff --git a/google_gemma-3-1b-it-q8.gguf b/google_gemma-3-1b-it-q8.gguf new file mode 100644 index 0000000..60b3019 --- /dev/null +++ b/google_gemma-3-1b-it-q8.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:75c8fe93256fdeb5f56d86ea3eb7e69ff436f650ce470335e4101d70b3a58a9a +size 1069306368