ReVA dataset advances remote sensing video question answering worldwide

ReVA: A Scene-Centric Dataset Beyond Repetition for Remote Sensing Video Question Answering

Computer Vision and Pattern Recognition

Summary

Understanding videos taken from drones over cities is hard because questions often repeat and images don’t show change over time. The authors created ReVA, a large collection of drone videos with lots of unique questions and answers about what’s happening in the scenes. They tested many video AI models on ReVA and found that current technology still struggles with understanding these complex videos. ReVA helps improve AI’s ability to reason about changes and scenes in remote sensing videos.

What this means in practice

Authors

Zhen Yao, Likai Wang, Yuming Yang, Zhihao Zheng, Bo Lang, Qiuyu Tang, Jialu Sheng, Jingqi Xu, Yuehai Yang, Jumal Barker, Xiaowen Ying, Mooi Choo Chuah

Abstract

Multimodal Large Language Models (MLLMs) have demonstrated remarkable advances in remote sensing. However, existing remote sensing multimodal reasoning benchmarks exhibit two critical limitations: they rely on (i) template-driven questions, which causes repetitive questions; and (ii) static images that fail to capture the inherent temporal nature of drone/UAV videos. This leaves systematic evaluation of remote sensing video reasoning largely unexplored. To address this gap, we introduce ReVA, a new dataset for remote sensing video question answering, designed to assess spatiotemporal, scene-centric, and reasoning-oriented capabilities of MLLMs. ReVA comprises 2,438 drone videos spanning 18 cities worldwide (580K frames) and 22K high-quality question-answer pairs across 11 challenging QA tasks. We develop a semi-automatic annotation pipeline that leverages Text LLMs and MLLMs for question-answer generation with human verification. We evaluate 23 proprietary and open-source Video LLMs on ReVA, exposing fundamental limitations of current models. These findings position ReVA as a critical benchmark toward better remote sensing video understanding and temporal reasoning capabilities for real-world deployments. Our code and dataset are available at: https://github.com/zyaocoder/ReVA