ModelHub XC 5818578519 初始化项目,由ModelHub XC社区提供模型
Model: PolicyShiftGuard/PolicyShiftGuard-3B
Source: Original Platform
2026-09-01 21:44:36 +08:00

base_model, datasets, license, pipeline_tag, library_name, tags
base_model datasets license pipeline_tag library_name tags
Qwen/Qwen2.5-VL-3B-Instruct
PolicyShiftBench/PolicyShiftBench
apache-2.0 image-text-to-text transformers
vision-language
image-safety
guardrails
policy-conditioned
qwen2.5-vl

PolicyShiftGuard-3B

📚 Paper | 💻 GitHub | 🏠 Project Page

PolicyShiftGuard-3B is a policy-conditioned image guardrail model based on Qwen2.5-VL-3B. It is trained to decide whether an image violates a supplied policy bundle and to return a structured safe/unsafe decision with the violated risk category when applicable.

Expected Output Format

true | <two-digit risk category id> | <short reason>
false | <short reason>

Training Data

This checkpoint is trained with the PolicyShiftBench public data release:

  • Dataset: PolicyShiftBench/PolicyShiftBench
  • Main evaluation splits: ID/adaptive branch and OOD/shift branch
  • Training stages: randomized policy SFT followed by boundary-pair policy adaptation

Intended Use

Use this model for research on policy-conditioned multimodal safety, adaptive image moderation, and robustness under policy shifts. The model should be evaluated with explicit policy bundles rather than as a fixed universal safety classifier.

Limitations

This is a research checkpoint. It may fail under policies, languages, visual domains, or deployment settings not represented in the benchmark. Outputs should not be treated as legal or compliance advice.

Description
Model synced from source: PolicyShiftGuard/PolicyShiftGuard-3B
Readme 2 MiB
Languages
Jinja 100%