Memory systems affect robot navigation success with changing conditions
MemTransfer: Benchmarking Memory Beyond Matched Experience in Embodied Decision-Making
Robotics
Summary
Robots use memory to remember past experiences, but having memory doesn’t always mean they can use it well when things change. The authors tested six different ways robots can remember and use past experience to navigate in a simulated warehouse. They found some memory types work well when conditions stay the same but struggle when starting points or routes change. More examples help some memory methods perform better in new situations, but not all benefit equally. This study shows it’s important to test both how much information robots store and how they actually use it during decisions.
What this means in practice
- •For autonomous robot developers: Choose and evaluate memory systems for robots that navigate changing real-world environments to improve robustness to new starting points and route changes.
- •For warehouse automation teams: Design navigation controls that effectively incorporate past experience when conditions like starting locations or available routes vary in a warehouse.
Authors
Haiming Tang, Xianjie Dai, Gujie Shao, Zuyi Guo, Jingguang Li, Kailang Ma, Yihong Tang, Heye Huang
Abstract
Memory lets an embodied agent reuse past experience, yet retaining useful information does not ensure that the agent can apply it when conditions change. We present MemTransfer, a benchmark comparing six memory representations, a working-memory baseline and five representations of past experience, under a shared frozen vision-language-model policy. It comprises 100 navigation cases across ten task types in a simulated warehouse, with expert demonstrations supplying the history. Three comparisons vary the starting pose, route availability, and amount and task relevance of history. With one demonstration per task, Full-context and Episodic memory reach 95.3% and 100.0% success at the original demonstration start, but lose 48-49 percentage points at a new test start. Summary changes little between these two test starts, yet with four demonstrations per task it retains a smaller fraction of its unchanged-route success after blocking (39.3%) than Working memory (44.8%) or the two trajectory memories (56-58%). At the new test start, increasing from one to four relevant demonstrations raises Episodic success by 14.3 percentage points, while the other evaluated representations gain no more than 1.3 percentage points. Replacing half of the relevant histories with other-task experience lowers success for both trajectory memories. These results show that robustness to one kind of mismatch does not imply robustness to another, motivating evaluation of both stored information and its use at decision time.