21 lines
462 B
Markdown
21 lines
462 B
Markdown
---
|
|
license: apache-2.0
|
|
language:
|
|
- en
|
|
base_model:
|
|
- Qwen/Qwen2.5-VL-7B-Instruct
|
|
pipeline_tag: visual-question-answering
|
|
---
|
|
|
|
# Video-MTR
|
|
|
|
Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding.
|
|
|
|
This checkpoint extends Qwen2.5-VL-7B-Instruct with a multi-turn frame-retrieval
|
|
policy trained via PPO, with an 80-frame video input budget.
|
|
|
|
## References
|
|
|
|
* [Paper](https://arxiv.org/abs/2508.20478)
|
|
* [Code](https://github.com/Xyuan13/Video-MTR)
|