Robot memory improves manipulation by remembering count and time

ReCAT: Remember, Count, and Time: Structured Recurrent Memory for Robot Manipulation

RoboticsArtificial IntelligenceComputer Vision and Pattern RecognitionMachine Learning

Summary

Robots often need to remember things from earlier to do tasks well, like spotting something they saw before, keeping track of repeated steps, or knowing how much time has passed. The authors created ReCAT, a system that helps robots use past information effectively by combining different memory and attention methods. This system helps robots perform better on tasks requiring memory, such as recalling places, counting events, and estimating time passed. Tests show ReCAT outperforms other approaches on many robot manipulation challenges, both in simulation and on real robots.

What this means in practice

  • For robotics software engineers: Improve robot manipulation skills by integrating structured recurrent memory to recall past observations, count repeated actions, and estimate elapsed time in complex tasks.
  • For automation system integrators: Build industrial robots that handle multi-step tasks requiring memory of previous steps, enabling more accurate and autonomous operation.

Authors

Pankhuri Vanjani, Mostafa Hatab, Can Mizrakli, Vaisakh Shaj, Zhuoyue Li, Moritz Reuss, Rudolf Lioutikov

Abstract

Memory-dependent manipulation requires robots to make decisions using information that is no longer available to their current sensors, such as recalling an earlier visual cue, tracking task progress, counting repeated events, or estimating elapsed time. We present ReCAT, a language-conditioned policy with structured recurrent memory. An instruction-conditioned encoder forms features from the current observation. A recurrent memory integrates the observation stream through Mamba-2 layers and one causal attention layer. A flow-matching Transformer decoder reads the current and the historical representation through separate cross-attention in every block. ReCAT reaches 95.3\% average success on LIBERO and 62.4\% on RMBench, with the best or tied-best result on six of nine tasks. On three real-robot tasks probing spatial recall, event counting, and interval timing, the best ReCAT variant reaches 66.7\% average success, against 8.3\% for the strongest short-history baseline. Controlled comparisons within ReCAT show that the observation encoder and every-block memory conditioning are needed for this performance. They also show that update rules developed for efficient sequence modeling behave differently as robot memory: additive updates have the highest observed success on counting and timing, and delta-rule updates on spatial recall. Project website is at https://intuitive-robots.github.io/ReCAT