Vision language action model improves robot manipulation of moving objects

D$^2$-VLA: Dual-Memory Dual-Frequency Vision-Language-Action Model For Long Dynamic Manipulation

Computer Vision and Pattern Recognition

Summary

Robots that handle tasks over a long time need to remember things they saw earlier, even when those things are no longer visible. The authors present a new robot control system called D²-VLA that keeps two kinds of memory and updates these differently to better handle moving objects and remember past information. Their method helps robots do complex tasks more successfully than previous approaches, even when objects move and key visual clues are no longer in view. They also test their approach on multiple robot benchmarks and real robots with improved results.

What this means in practice

  • For robotics engineers: Build robots that better remember past visual cues to handle long and dynamic manipulation tasks involving moving objects.
  • For industrial automation teams: Deploy improved robot control policies that increase success rates for complex assembly or sorting tasks requiring memory of previous observations.

Authors

Zijian Ye, Chengqi Wei, Wei Huang, Anlin Zheng, Chunyu Zou, Liangyu Wu, Zikang Zhao, Zhenjie Peng, Yushuo Yang, Shuman Zhao, Zhongrui Wang, Xiaojuan Qi

Abstract

Long-horizon manipulation requires robots to remember cues that are no longer in view while responding to moving objects. Yet vision-language-action (VLA) policies often rely on the latest observation, and refreshing their visual context typically requires another costly vision-language model (VLM) pass. We present D$^2$-VLA, which combines dual memory and dual-frequency control at the KV-cache interface of a pretrained VLA. D$^2$-VLA uses block-wise causal KV caching to encode observations incrementally and, guided by distinct temporal attention patterns, constructs separate historical KV read views for the VLM and action expert. Between periodic VLM updates, a gated adapter incorporates fresh visual features into the latest history-conditioned KV block, while a short fast-memory queue supports action replanning. We introduce DOMINO-Long, a ten-task benchmark requiring robots to use earlier visual cues when manipulating moving objects. D$^2$-VLA achieves complete-task success rates of 29.3\% on DOMINO, compared with 9.6\% for $π_{0.5}$ and 17.2\% for PUMA, and 60.0\% on DOMINO-Long, compared with 35.4\% and 20.6\%, respectively. It improves success rates on eight real-robot tasks and reaches 97.5\% on LIBERO-Long and 74.3\% on RoboTwin 2.0.