Object memory system improves robot manipulation over long tasks

Where Memory Belongs: Ledger, an Object Ledger for Memory-Augmented VLAs

Robotics

Summary

Robots need memory to handle tasks over long periods, like remembering where they put things or how many steps they took. The paper shows that different kinds of memory should live in different places: short-term memory stays inside the robot's main brain, while long-term memory lives outside in a special record. The authors created Ledger, a system that keeps track of objects and events separately from the robot's main policy. This approach helped robots perform better on tests of remembering objects and actions.

What this means in practice

  • For robotics engineers: Improve robot manipulation by integrating distinct short-term and long-term memory systems for better object tracking and action planning over extended tasks.
  • For industrial automation teams: Develop robots that maintain accurate object state and history outside core control software, enabling more reliable manipulation in complex environments.

Authors

Tanguy Dieudonné, Jack B. Jedlicki, Heng Yang

Abstract

Memory is essential for long-horizon, partially observed robotic manipulation: a robot must remember which object was placed in a drawer, whose cup it moved, or how many action cycles have elapsed. Recent vision-language-action (VLA) models embed memory directly inside the policy, but benchmarks show no single in-policy mechanism covers all spatio-temporal dimensions, trailing oracle methods by a wide margin. We argue that memory type dictates where memory should reside: short-term perceptual memory (repetition, timing, retracing) belongs inside the policy, while long-term object memory (persistent spatial state, containment, event history) belongs outside as an explicit, readable record. We present Ledger, a harness that realizes this split over a single fine-tuned $π_{0.5}$ policy by pairing an in-policy frame-sampling memory with an external spatio-temporal object memory, the ledger, built from a SAM3 tracker and a VLM captioner of the demonstration and read by an LLM planner that decides at step boundaries. On RoboMME, Ledger reaches the highest four-suite average among the evaluated methods, 64.3% (vs. 45.9% for the strongest prior method under identical evaluation), leading object reference (60.7% vs. 40.3%) and object permanence (86.7% vs. 56.2%) using a single set of weights. Choosing the memory source at runtime, from the instruction and the record, removes the need for a task-level router.