license, library_name, pipeline_tag, tags
license library_name pipeline_tag tags
fair-noncommercial-research-license transformers image-text-to-text
qwen3-vl
video-language-model
region-understanding
motion-captioning

MotionAtlas-4B

This repository contains the MotionAtlas-4B model for MotionAtlas: Detailed Region Captioning for Motion-Centric Videos.

TL; DR: MotionAtlas shifts motion captioning from global video descriptions to region-aware motion captions, enabling precise evaluation with MotionAtlas-Bench and scalable training with MotionAtlas-Data. The model is designed for detailed motion-centric video understanding over referred regions.

Usage

For detailed usage of this model, please refer to our GitHub repo and project page.

Description
Model synced from source: maxLWSv2/MotionAtlas-4B
Readme 13 MiB
Languages
Jinja 100%