ModelHub XC b64825235a 初始化项目,由ModelHub XC社区提供模型
Model: CodeGoat24/UnifiedReward-2.0-qwen3vl-8b
Source: Original Platform
2026-09-04 06:05:17 +08:00

license, base_model
license base_model
mit
Qwen/Qwen3-VL-8B-Instruct

Model Summary

UnifiedReward-2.0-qwen3vl-8b is the first unified reward model based on Qwen/Qwen3-VL-8B-Instruct for multimodal understanding and generation assessment, enabling both pairwise ranking and pointwise scoring, which can be employed for vision model preference alignment.

For further details, please refer to the following resources:

🏁 Compared with Current Reward Models

Reward Model Method Image Generation Image Understanding Video Generation Video Understanding
PickScore Point √
HPS Point √
ImageReward Point √
LLaVA-Critic Pair/Point √
IXC-2.5-Reward Pair/Point √ √
VideoScore Point √
LiFT Point √
VisionReward Point √ √
VideoReward Point √
UnifiedReward (Ours) Pair/Point √ √ √ √

Citation

@article{unifiedreward,
  title={Unified reward model for multimodal understanding and generation},
  author={Wang, Yibin and Zang, Yuhang and Li, Hao and Jin, Cheng and Wang, Jiaqi},
  journal={arXiv preprint arXiv:2503.05236},
  year={2025}
}
Description
Model synced from source: CodeGoat24/UnifiedReward-2.0-qwen3vl-8b
Readme 13 MiB
Languages
Jinja 100%