Knowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMs

2026-07-17Computer Vision and Pattern Recognition

Computer Vision and Pattern Recognition
AI summary

The authors created a new benchmark called UAV-DualCog to test how well multimodal large language models (MLLMs) can think about both a drone's own position and its surroundings using multiple views and time-based data. Unlike previous tests, their benchmark includes images and videos and requires detailed spatial and temporal reasoning, not just simple answers. They built a large and varied dataset using 3D scene information and found that current models struggle with tasks like understanding viewpoints, precise location, and timing. Their benchmark is easy for humans but hard for existing AI, and it can also be used to help train better drone-aware models.

Multimodal Large Language ModelsUAV (Unmanned Aerial Vehicle)Dual-CognitionSpatio-Temporal ReasoningViewpoint TransformationSpatial GroundingTemporal Interval LocalizationSemantic Point CloudsBenchmark Dataset
Authors
Like Liu, Zhengzheng Xu, Haitao He, Hongzhe Li, Shuchang Zhang, Dian Shao
Abstract
Multimodal large language models have achieved strong performance across diverse vision-language tasks, yet their capabilities in UAV scenarios remain insufficiently explored. Recent UAV-oriented benchmarks have begun to evaluate MLLMs in aerial scenarios, but they typically focus on scene understanding, event recognition, or navigation completion, rather than jointly assessing the dual-cognition capability required for UAV agents: reasoning about both the UAV's own state and the external environment in multiview spatio-temporal contexts. To address this gap, we present UAV-DualCog, a benchmark for aerial multiview spatio-temporal reasoning built on this dual-cognition perspective. UAV-DualCog includes both image and video tasks to jointly evaluate self-state and environment-state reasoning, while requiring spatial or temporal grounding beyond discrete answer prediction. We also develop an automated pipeline that constructs data from scene-level semantic point clouds, yielding a scalable benchmark with diverse scenes, hundreds of landmarks, and thousands of QA samples. Extensive evaluations show that current MLLMs remain far from reliable in UAV dual cognition. Self-state reasoning, viewpoint transformation, precise spatial grounding, and temporal interval localization are persistent bottlenecks, and additional validation with thinking/frontier models and a human baseline confirms that the benchmark is understandable to humans but challenging for existing models. We further construct UAV-DualCog-Train from disjoint scenes and show through a lightweight optimization probe that it provides useful structured supervision, suggesting its value not only as an evaluation benchmark but also as a data resource for advancing MLLM-based UAV agents. Project website and supplementary materials: https://uav-dualcog.lozumi.com