UAV memories combined on ground to improve low-altitude answers

Memory in the Sky: Low-Altitude Question Answering with Multi-Agent Memory Aggregation

RoboticsInformation Theory

Summary

Answering questions about what drones see over time is hard because their memories are spread out and communication can be limited. The authors propose a way to score and pick the most useful drone memories to send to a ground server, even without looking inside complex AI systems. They also develop methods to pick which drones to use and how much power to send data using new formulas. Their approach works well in both simulated and real-world tests, helping robots understand and navigate environments better.

What this means in practice

  • For drone fleet operators: Select and power UAVs to maximize useful memory data sent to ground stations for long-term environmental monitoring.
  • For mobile robot developers: Improve robot navigation and environment understanding using aggregated memories from multiple UAVs.

Authors

Chengyang Li, Yujie Wan, Shuai Wang, Kejiang Ye, Weijie Yuan, Boyu Zhou, Yik-Chung Wu, Chengzhong Xu, Huseyin Arslan

Abstract

This paper studies low-altitude question answering (LAQA), in which distributed unmanned aerial vehicle (UAV) memories are aggregated at a ground server to answer questions about observations over a long horizon. Unlike conventional resource allocation based on sensing, communication, control, or computation metrics, LAQA requires an explicit measure of memory value. We propose a generative adversarial exam (GAE) that uses forward simulation to evaluate memory retrieval and exam scores to quantify memory quality. This enables the downstream QA value of candidate memories to be measured and optimized without accessing the internal mechanisms of the black-box captioning, retrieval, and reasoning pipeline. Building on this metric, we develop a memory-centric (MemCen) framework that jointly selects UAVs and allocates transmit power to maximize memory quality under communication constraints. In the noise-limited regime, we derive a QoM-aware capped water-filling law that explicitly connects task utility with physical-layer power allocation. We further develop penalty successive optimization (PSO) and learning to memorize (L2M) solvers. MemCen achieves QA accuracies of 92.4% and 84.0% in CARLA Town04 and Town05 under static and dynamic communication conditions, respectively. In real-world experiments, MemCen achieves 88.5% QA accuracy on the panoramic multi-agent system (PMAS) benchmark. Finally, UAV-to-robot-dog demonstrations further validate the practical utility of the acquired memories for environmental understanding and navigation.