43 lines
1.6 KiB
Markdown
43 lines
1.6 KiB
Markdown
|
|
---
|
||
|
|
base_model: Qwen/Qwen2.5-VL-3B-Instruct
|
||
|
|
datasets:
|
||
|
|
- PolicyShiftBench/PolicyShiftBench
|
||
|
|
license: apache-2.0
|
||
|
|
pipeline_tag: image-text-to-text
|
||
|
|
library_name: transformers
|
||
|
|
tags:
|
||
|
|
- vision-language
|
||
|
|
- image-safety
|
||
|
|
- guardrails
|
||
|
|
- policy-conditioned
|
||
|
|
- qwen2.5-vl
|
||
|
|
---
|
||
|
|
|
||
|
|
# PolicyShiftGuard-3B
|
||
|
|
|
||
|
|
[📚 Paper](https://huggingface.co/papers/2607.05910) | [💻 GitHub](https://github.com/ssmisya/PolicyShiftGuard) | [🏠 Project Page](https://policyshiftguard.github.io/)
|
||
|
|
|
||
|
|
PolicyShiftGuard-3B is a policy-conditioned image guardrail model based on Qwen2.5-VL-3B. It is trained to decide whether an image violates a supplied policy bundle and to return a structured safe/unsafe decision with the violated risk category when applicable.
|
||
|
|
|
||
|
|
## Expected Output Format
|
||
|
|
|
||
|
|
```text
|
||
|
|
true | <two-digit risk category id> | <short reason>
|
||
|
|
false | <short reason>
|
||
|
|
```
|
||
|
|
|
||
|
|
## Training Data
|
||
|
|
|
||
|
|
This checkpoint is trained with the PolicyShiftBench public data release:
|
||
|
|
|
||
|
|
- Dataset: `PolicyShiftBench/PolicyShiftBench`
|
||
|
|
- Main evaluation splits: ID/adaptive branch and OOD/shift branch
|
||
|
|
- Training stages: randomized policy SFT followed by boundary-pair policy adaptation
|
||
|
|
|
||
|
|
## Intended Use
|
||
|
|
|
||
|
|
Use this model for research on policy-conditioned multimodal safety, adaptive image moderation, and robustness under policy shifts. The model should be evaluated with explicit policy bundles rather than as a fixed universal safety classifier.
|
||
|
|
|
||
|
|
## Limitations
|
||
|
|
|
||
|
|
This is a research checkpoint. It may fail under policies, languages, visual domains, or deployment settings not represented in the benchmark. Outputs should not be treated as legal or compliance advice.
|