Files
ViGaL-7B/README.md
ModelHub XC 7d0832a17f 初始化项目,由ModelHub XC社区提供模型
Model: yunfeixie/ViGaL-7B
Source: Original Platform
2026-09-03 16:36:18 +08:00

1.7 KiB

base_model, language, license, pipeline_tag, tags, library_name
base_model language license pipeline_tag tags library_name
Qwen/Qwen2.5-VL-7B-Instruct
en
apache-2.0 image-text-to-text
transformers
multimodal
transformers

Model Overview

We present Visual Game Learning (ViGaL), a novel post-training paradigm where multimodal large language models (MLLMs) develop out-of-domain generalization of multimodal reasoning through playing arcade-like games.

ViGaL-7B demonstrates that training a 7B-parameter MLLM via reinforcement learning on simple arcade-like games like Snake significantly enhances its downstream performance on multimodal math benchmarks like MathVista, and on multi-discipline questions like MMMU, without seeing any worked solutions, equations, or diagrams during RL, suggesting the capture of transferable reasoning skills.

Resources

For details of our approach and performance comparison, please see our paper.

For details of training and evaluation, please see our code repo.

| 🚀 Project Page | 📖 Paper | 🔗 GitHub | 🤗 Training Data | 🤗 Model |

Citation

If you feel this model useful, please give us a free cite:

@article{xie2025play,
  title     = {Play to Generalize: Learning to Reason Through Game Play},
  author    = {Xie, Yunfei and Ma, Yinsong and Lan, Shiyi and Yuille, Alan and Xiao, Junfei and Wei, Chen},
  journal   = {arXiv preprint arXiv:2506.08011},
  year      = {2025},
}