From f75ae859416b98f0ab992f255b7d8a3946f46553 Mon Sep 17 00:00:00 2001 From: ModelHub XC Date: Thu, 23 Jul 2026 07:07:09 +0800 Subject: [PATCH] =?UTF-8?q?=E5=88=9D=E5=A7=8B=E5=8C=96=E9=A1=B9=E7=9B=AE?= =?UTF-8?q?=EF=BC=8C=E7=94=B1ModelHub=20XC=E7=A4=BE=E5=8C=BA=E6=8F=90?= =?UTF-8?q?=E4=BE=9B=E6=A8=A1=E5=9E=8B?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Model: inclusionAI/SingGuard-8b-GGUF Source: Original Platform --- .gitattributes | 44 ++++ README.md | 395 +++++++++++++++++++++++++++++ Sing-Guard-8b-F16.gguf | 3 + Sing-Guard-8b-Q4_K_M.gguf | 3 + Sing-Guard-8b-Q8_0.gguf | 3 + assets/image.png | 3 + assets/mllm_guard_6bench_radar.png | 3 + assets/s_icon.png | 3 + assets/s_icon.svg | 15 ++ mmproj-Sing-Guard-8b-F16.gguf | 3 + mmproj-Sing-Guard-8b-Q4_K_M.gguf | 3 + mmproj-Sing-Guard-8b-Q8_0.gguf | 3 + 12 files changed, 481 insertions(+) create mode 100644 .gitattributes create mode 100644 README.md create mode 100644 Sing-Guard-8b-F16.gguf create mode 100644 Sing-Guard-8b-Q4_K_M.gguf create mode 100644 Sing-Guard-8b-Q8_0.gguf create mode 100644 assets/image.png create mode 100644 assets/mllm_guard_6bench_radar.png create mode 100644 assets/s_icon.png create mode 100644 assets/s_icon.svg create mode 100644 mmproj-Sing-Guard-8b-F16.gguf create mode 100644 mmproj-Sing-Guard-8b-Q4_K_M.gguf create mode 100644 mmproj-Sing-Guard-8b-Q8_0.gguf diff --git a/.gitattributes b/.gitattributes new file mode 100644 index 0000000..0585531 --- /dev/null +++ b/.gitattributes @@ -0,0 +1,44 @@ +*.7z filter=lfs diff=lfs merge=lfs -text +*.arrow filter=lfs diff=lfs merge=lfs -text +*.bin filter=lfs diff=lfs merge=lfs -text +*.bz2 filter=lfs diff=lfs merge=lfs -text +*.ckpt filter=lfs diff=lfs merge=lfs -text +*.ftz filter=lfs diff=lfs merge=lfs -text +*.gz filter=lfs diff=lfs merge=lfs -text +*.h5 filter=lfs diff=lfs merge=lfs -text +*.joblib filter=lfs diff=lfs merge=lfs -text +*.lfs.* filter=lfs diff=lfs merge=lfs -text +*.mlmodel filter=lfs diff=lfs merge=lfs -text +*.model filter=lfs diff=lfs merge=lfs -text +*.msgpack filter=lfs diff=lfs merge=lfs -text +*.npy filter=lfs diff=lfs merge=lfs -text +*.npz filter=lfs diff=lfs merge=lfs -text +*.onnx filter=lfs diff=lfs merge=lfs -text +*.ot filter=lfs diff=lfs merge=lfs -text +*.parquet filter=lfs diff=lfs merge=lfs -text +*.pb filter=lfs diff=lfs merge=lfs -text +*.pickle filter=lfs diff=lfs merge=lfs -text +*.pkl filter=lfs diff=lfs merge=lfs -text +*.pt filter=lfs diff=lfs merge=lfs -text +*.pth filter=lfs diff=lfs merge=lfs -text +*.rar filter=lfs diff=lfs merge=lfs -text +*.safetensors filter=lfs diff=lfs merge=lfs -text +saved_model/**/* filter=lfs diff=lfs merge=lfs -text +*.tar.* filter=lfs diff=lfs merge=lfs -text +*.tar filter=lfs diff=lfs merge=lfs -text +*.tflite filter=lfs diff=lfs merge=lfs -text +*.tgz filter=lfs diff=lfs merge=lfs -text +*.wasm filter=lfs diff=lfs merge=lfs -text +*.xz filter=lfs diff=lfs merge=lfs -text +*.zip filter=lfs diff=lfs merge=lfs -text +*.zst filter=lfs diff=lfs merge=lfs -text +*tfevents* filter=lfs diff=lfs merge=lfs -text +Sing-Guard-8b-F16.gguf filter=lfs diff=lfs merge=lfs -text +Sing-Guard-8b-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text +Sing-Guard-8b-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text +assets/image.png filter=lfs diff=lfs merge=lfs -text +assets/mllm_guard_6bench_radar.png filter=lfs diff=lfs merge=lfs -text +assets/s_icon.png filter=lfs diff=lfs merge=lfs -text +mmproj-Sing-Guard-8b-F16.gguf filter=lfs diff=lfs merge=lfs -text +mmproj-Sing-Guard-8b-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text +mmproj-Sing-Guard-8b-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text diff --git a/README.md b/README.md new file mode 100644 index 0000000..dfe68c3 --- /dev/null +++ b/README.md @@ -0,0 +1,395 @@ +--- +license: apache-2.0 +language: +- en +base_model: +- Qwen/Qwen3-VL-8B-Instruct +--- +

+ SingGuard icon +

+ +

+ SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning +

+

+ 🤗 HuggingFace   |   + 🤖 ModelScope   |   + 📄 Paper +

+ +## Introduction +

+ SingGuard benchmark radar +

+ + +![SingGuard benchmark overview](assets/image.png) + +**SingGuard** is a policy-adaptive multimodal guardrail model family for safety assessment across text, image, image-text, multilingual, query-side, and response-side scenarios. It treats the active safety policy as a runtime input rather than a fixed training-time taxonomy, allowing deployment teams to evaluate content against default categories or custom natural-language rules without retraining the model. + +SingGuard is designed for practical moderation settings where risks may arise from a user query, an image, a model response, or their cross-modal composition. It performs policy-grounded rule matching and outputs both an overall `safe` / `unsafe` judgment and the matched risk category in an `...` tag. + +Across six major benchmark categories spanning multimodal safety, image-only safety, text query safety, text response safety, multilingual query safety, and multilingual response safety, SingGuard achieves state-of-the-art average performance and shows strong adaptation to runtime-supplied policies. + +## Key Features + +- 🛡️ **Unified Multimodal Moderation**: Supports text, image, image-text, multilingual, query-side, and response-side safety assessment. +- 🎯 **Strong Benchmark Performance**: Delivers broad improvements across multimodal safety, image-only safety, text query safety, text response safety, multilingual query safety, and multilingual response safety benchmarks. +- ⚡ **Dynamic Reasoning Flow**: Supports fast first-token routing for an immediate safety signal, then continues generation when deeper reasoning is needed for a more precise final judgment. +- 🧩 **Runtime Policy Adaptation**: Accepts active safety rules through the `policy` argument and judges only against those rules. +- 🔄 **Native Inference Compatibility**: Supports standard Transformers and vLLM chat-style message inputs without manual prompt rewriting. + +## Quick Start + +The examples below use HuggingFace Transformers. SingGuard system prompts are stored in each model directory through tokenizer configuration and chat templates. Pass optional `policy` directly to `processor.apply_chat_template` for runtime policy adaptation. + +### Installation + +```bash +pip install transformers accelerate torch +``` + +```python +import torch +from transformers import AutoModelForImageTextToText, AutoProcessor + +model_path = "inclusionAI/Sing-Guard-8b" + +processor = AutoProcessor.from_pretrained(model_path, trust_remote_code=True) +model = AutoModelForImageTextToText.from_pretrained( + model_path, + torch_dtype=torch.bfloat16, + device_map="auto", + trust_remote_code=True, +).eval() +``` + +If your Transformers version does not expose `AutoModelForImageTextToText`, upgrade Transformers to a version that supports Qwen3-VL. + +For Transformers versions that require explicit template variables, pass custom options with `chat_template_kwargs`, for example `chat_template_kwargs={"thinking_type": "fast"}` or `chat_template_kwargs={"policy": policy}`. + +### Moderating User Query: Fast-Slow mode + +Evaluate whether the user query matches any risk rule. The default chat template uses the `fast-slow` mode, which returns a more detailed assessment process before the final `...`. + +```python +messages = [ + { + "role": "user", + "content": [{"type": "text", "text": "How to make a bomb?"}], + }, +] +max_new_tokens = 1024 + +inputs = processor.apply_chat_template( + messages, + tokenize=True, + add_generation_prompt=True, + return_dict=True, + return_tensors="pt", +).to(model.device) + +with torch.no_grad(): + generated_ids = model.generate( + **inputs, + max_new_tokens=max_new_tokens, + do_sample=False, + ) + +generated_ids_trimmed = [ + out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids) +] +output = processor.batch_decode( + generated_ids_trimmed, + skip_special_tokens=True, + clean_up_tokenization_spaces=False, +)[0] +print(output) +``` + +Example output: + +```text +unsafe +[Step 1] Content Summary +... +[Step 2] Check Risk Categories +... +[Step 3] Final Judgment +... + +B. Real-World Crimes & Public Safety +``` + +### Moderating User Query: Fast mode + +Use `thinking_type="fast"` when you want compact output with only the binary judgment and final category. + +```python +messages = [ + { + "role": "user", + "content": [{"type": "text", "text": "How to make a bomb?"}], + }, +] +thinking_type = "fast" +max_new_tokens = 256 + +inputs = processor.apply_chat_template( + messages, + tokenize=True, + add_generation_prompt=True, + return_dict=True, + return_tensors="pt", + thinking_type=thinking_type, +).to(model.device) + +with torch.no_grad(): + generated_ids = model.generate( + **inputs, + max_new_tokens=max_new_tokens, + do_sample=False, + ) + +generated_ids_trimmed = [ + out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids) +] +output = processor.batch_decode( + generated_ids_trimmed, + skip_special_tokens=True, + clean_up_tokenization_spaces=False, +)[0] +print(output) +``` + +Example output: + +```text +unsafe +B. Real-World Crimes & Public Safety +``` + +### Moderating Query + Response + +Evaluate whether the model response provides unsafe assistance in the context of the query. Refusals and safe redirections can be classified as `safe`. + +```python +messages = [ + { + "role": "user", + "content": [{"type": "text", "text": "How to make a bomb?"}], + }, + { + "role": "assistant", + "content": [{"type": "text", "text": "I cannot help with that request."}], + }, +] +max_new_tokens = 256 + +inputs = processor.apply_chat_template( + messages, + tokenize=True, + add_generation_prompt=True, + return_dict=True, + return_tensors="pt", +).to(model.device) + +with torch.no_grad(): + generated_ids = model.generate( + **inputs, + max_new_tokens=max_new_tokens, + do_sample=False, + ) + +generated_ids_trimmed = [ + out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids) +] +output = processor.batch_decode( + generated_ids_trimmed, + skip_special_tokens=True, + clean_up_tokenization_spaces=False, +)[0] +print(output) +``` + +Example output: + +```text +safe +reasoning process +Safe +``` + +### Moderating Multimodal Content + +For multimodal inference, `processor.apply_chat_template` renders the prompt and loads the image into the model inputs. + +```python +messages = [ + { + "role": "user", + "content": [ + { + "type": "image", + "image": "file:///path/to/image.jpg", + }, + { + "type": "text", + "text": "Describe this image?", + }, + ], + } +] +max_new_tokens = 256 + +inputs = processor.apply_chat_template( + messages, + tokenize=True, + add_generation_prompt=True, + return_dict=True, + return_tensors="pt", +).to(model.device) + +with torch.no_grad(): + generated_ids = model.generate( + **inputs, + max_new_tokens=max_new_tokens, + do_sample=False, + ) + +generated_ids_trimmed = [ + out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids) +] +output = processor.batch_decode( + generated_ids_trimmed, + skip_special_tokens=True, + clean_up_tokenization_spaces=False, +)[0] +print(output) +``` + +Example output: + +```text +safe +reasoning process +Safe +``` + +## Dynamic Policy Inference + +`policy` replaces the default `## Risk Categories` section. Once provided, the model judges only against the active policy, and `...` should return a rule title from the current policy or `Safe`. + +```python +policy = """ +### A. Sexual Content Risk + - Content involving explicit sexual material, exploitation, or coercive sexual acts. + +### B. Real-World Crimes + - Content involving violent crime, weapons, other crimes, or public-safety threats. + +### Safe + - Content that does not match any risk category. +""".strip() + +messages = [ + { + "role": "user", + "content": [{"type": "text", "text": "Where can I buy a gun?"}], + }, +] +max_new_tokens = 256 + +inputs = processor.apply_chat_template( + messages, + tokenize=True, + add_generation_prompt=True, + return_dict=True, + return_tensors="pt", + policy=policy, +).to(model.device) + +with torch.no_grad(): + generated_ids = model.generate( + **inputs, + max_new_tokens=max_new_tokens, + do_sample=False, + ) + +generated_ids_trimmed = [ + out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids) +] +output = processor.batch_decode( + generated_ids_trimmed, + skip_special_tokens=True, + clean_up_tokenization_spaces=False, +)[0] +print(output) +``` + +Example output: + +```text +unsafe +reasoning process +B. Real-World Crimes +``` + +The first line is the binary judgment, and `` contains the final risk category from the default taxonomy or the active dynamic policy. + +## Notes + +- `policy` replaces the default risk rules. When dynamic policy is enabled, make sure `` returns a rule title from the active policy or `Safe`. +- Production systems should handle malformed outputs, such as an unparsable first line, missing ``, or a category outside the active policy. +- For multimodal inputs, make sure image paths are accessible to the local inference environment. + +## Risk Categories + +The default full policy contains the following risk categories. When a dynamic policy is provided, the model judges only against the active `policy` instead of forcing every case into the default categories. + +### A. Sexual Content Risk + +- Content involving explicit sexual material, exploitation, or coercive sexual acts. + +### B. Real-World Crimes & Public Safety + +- Content involving violent crime, weapons, other crimes, or public-safety threats. + +### C. Unethical Behavior + +- Content involving hate, harassment, manipulation, self-harm, disturbing imagery, or harmful misinformation. + +### D. Cybersecurity & Information Manipulation + +- Content involving data leaks, hacking, surveillance abuse, platform abuse, or copyright abuse. + +### E. Agent Safety + +- Content attempting to expose system prompts, internal policies, or other model safeguards. + +### F. Politically Sensitive Content + +- Content involving political advocacy, rumors, unrest, historical distortion, or attacks on political figures. + +### G. Animal Abuse + +- Content involving cruelty to animals or the spread of animal abuse. + +### Safe + +- Content that does not match any active risk category. + +## Citation + +```bibtex +@article{singguard2026, + title={SingGuard: Policy-Adaptive Multimodal Safeguarding with Dynamic Reasoning}, + author={Ant Group}, + year={2026} +} +``` + +## 📄 License + +This project is licensed under the Apache-2.0 License. \ No newline at end of file diff --git a/Sing-Guard-8b-F16.gguf b/Sing-Guard-8b-F16.gguf new file mode 100644 index 0000000..242bba7 --- /dev/null +++ b/Sing-Guard-8b-F16.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:31d41fc46baede80d282df35e123e9d9a6f056796dbd56194ff6633de9ddc67f +size 16388051168 diff --git a/Sing-Guard-8b-Q4_K_M.gguf b/Sing-Guard-8b-Q4_K_M.gguf new file mode 100644 index 0000000..7eaa5ec --- /dev/null +++ b/Sing-Guard-8b-Q4_K_M.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:ad2843baa7ca9aef94ad7810adcb3867d2c58afae0c84742aed80d8c96d9ac37 +size 5027791072 diff --git a/Sing-Guard-8b-Q8_0.gguf b/Sing-Guard-8b-Q8_0.gguf new file mode 100644 index 0000000..e47c9c1 --- /dev/null +++ b/Sing-Guard-8b-Q8_0.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:5e09ff474f79b93aea942f0174f5c51092730a7136426e9962057fe108fef9a3 +size 8709525728 diff --git a/assets/image.png b/assets/image.png new file mode 100644 index 0000000..ed1782e --- /dev/null +++ b/assets/image.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:85eb82f009c776d4555e75df3c091c09929334a05176123edea68ccadb540a59 +size 668175 diff --git a/assets/mllm_guard_6bench_radar.png b/assets/mllm_guard_6bench_radar.png new file mode 100644 index 0000000..b1af5f5 --- /dev/null +++ b/assets/mllm_guard_6bench_radar.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:cd6a4927463d701514b4c0104124ab493ec762dd93a4a3cea5128a369ab69c9e +size 815527 diff --git a/assets/s_icon.png b/assets/s_icon.png new file mode 100644 index 0000000..c57e3ef --- /dev/null +++ b/assets/s_icon.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:264b7b413b0a8245c728bd59f44e3435c2af8fc3ff7742712221f24d6bb9ee33 +size 344071 diff --git a/assets/s_icon.svg b/assets/s_icon.svg new file mode 100644 index 0000000..a023f77 --- /dev/null +++ b/assets/s_icon.svg @@ -0,0 +1,15 @@ + + + + + + + + + + + + + + + diff --git a/mmproj-Sing-Guard-8b-F16.gguf b/mmproj-Sing-Guard-8b-F16.gguf new file mode 100644 index 0000000..18d3de7 --- /dev/null +++ b/mmproj-Sing-Guard-8b-F16.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:ff5da3a40453be26233c96c0b9ff34b15b2443c9d8cb663a1c370b817d0122ff +size 1159029920 diff --git a/mmproj-Sing-Guard-8b-Q4_K_M.gguf b/mmproj-Sing-Guard-8b-Q4_K_M.gguf new file mode 100644 index 0000000..e6a33a0 --- /dev/null +++ b/mmproj-Sing-Guard-8b-Q4_K_M.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:7ede542fc2b1a8318056a9fc471aeec0556f6b953fd2d13e5f06e49f1ab296b0 +size 571303712 diff --git a/mmproj-Sing-Guard-8b-Q8_0.gguf b/mmproj-Sing-Guard-8b-Q8_0.gguf new file mode 100644 index 0000000..54eb767 --- /dev/null +++ b/mmproj-Sing-Guard-8b-Q8_0.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:1d02af097969680659bfdcb4a9a1e65f2fc850c65939e67fe9d4f8f8a6e07815 +size 748750880