ReVA dataset advances remote sensing video question answering worldwide
ReVA: A Scene-Centric Dataset Beyond Repetition for Remote Sensing Video Question Answering
Computer Vision and Pattern Recognition
Summary
Understanding videos taken from drones over cities is hard because questions often repeat and images don’t show change over time. The authors created ReVA, a large collection of drone videos with lots of unique questions and answers about what’s happening in the scenes. They tested many video AI models on ReVA and found that current technology still struggles with understanding these complex videos. ReVA helps improve AI’s ability to reason about changes and scenes in remote sensing videos.
What this means in practice
- •For geospatial analysts: Use the ReVA dataset to train AI systems that better understand changes and events in drone videos over cities.
- •For surveillance technology developers: Develop video AI tools that improve scene and temporal reasoning in aerial footage using ReVA’s diverse questions and answers.
Authors
Zhen Yao, Likai Wang, Yuming Yang, Zhihao Zheng, Bo Lang, Qiuyu Tang, Jialu Sheng, Jingqi Xu, Yuehai Yang, Jumal Barker, Xiaowen Ying, Mooi Choo Chuah
Abstract
Multimodal Large Language Models (MLLMs) have demonstrated remarkable advances in remote sensing. However, existing remote sensing multimodal reasoning benchmarks exhibit two critical limitations: they rely on (i) template-driven questions, which causes repetitive questions; and (ii) static images that fail to capture the inherent temporal nature of drone/UAV videos. This leaves systematic evaluation of remote sensing video reasoning largely unexplored. To address this gap, we introduce ReVA, a new dataset for remote sensing video question answering, designed to assess spatiotemporal, scene-centric, and reasoning-oriented capabilities of MLLMs. ReVA comprises 2,438 drone videos spanning 18 cities worldwide (580K frames) and 22K high-quality question-answer pairs across 11 challenging QA tasks. We develop a semi-automatic annotation pipeline that leverages Text LLMs and MLLMs for question-answer generation with human verification. We evaluate 23 proprietary and open-source Video LLMs on ReVA, exposing fundamental limitations of current models. These findings position ReVA as a critical benchmark toward better remote sensing video understanding and temporal reasoning capabilities for real-world deployments. Our code and dataset are available at: https://github.com/zyaocoder/ReVA