Files
PolicyShiftGuard-3B/README.md
ModelHub XC 5818578519 初始化项目,由ModelHub XC社区提供模型
Model: PolicyShiftGuard/PolicyShiftGuard-3B
Source: Original Platform
2026-09-01 21:44:36 +08:00

1.6 KiB

base_model, datasets, license, pipeline_tag, library_name, tags
base_model datasets license pipeline_tag library_name tags
Qwen/Qwen2.5-VL-3B-Instruct
PolicyShiftBench/PolicyShiftBench
apache-2.0 image-text-to-text transformers
vision-language
image-safety
guardrails
policy-conditioned
qwen2.5-vl

PolicyShiftGuard-3B

📚 Paper | 💻 GitHub | 🏠 Project Page

PolicyShiftGuard-3B is a policy-conditioned image guardrail model based on Qwen2.5-VL-3B. It is trained to decide whether an image violates a supplied policy bundle and to return a structured safe/unsafe decision with the violated risk category when applicable.

Expected Output Format

true | <two-digit risk category id> | <short reason>
false | <short reason>

Training Data

This checkpoint is trained with the PolicyShiftBench public data release:

  • Dataset: PolicyShiftBench/PolicyShiftBench
  • Main evaluation splits: ID/adaptive branch and OOD/shift branch
  • Training stages: randomized policy SFT followed by boundary-pair policy adaptation

Intended Use

Use this model for research on policy-conditioned multimodal safety, adaptive image moderation, and robustness under policy shifts. The model should be evaluated with explicit policy bundles rather than as a fixed universal safety classifier.

Limitations

This is a research checkpoint. It may fail under policies, languages, visual domains, or deployment settings not represented in the benchmark. Outputs should not be treated as legal or compliance advice.