World models learn to recall helpful memories from multiple cues

Learning What to Recall: Adaptive Multi-Cue Episodic Memory for World Models

Machine LearningArtificial IntelligenceComputer Vision and Pattern Recognition

Summary

Predicting what will happen next can require remembering things from far in the past. But it’s tricky to know which past memories are useful and which ways of finding them work best. The authors created a method called Future-Aware Recall (FAR) that learns which memories to bring back by checking how useful they are for predicting the future. FAR also figures out which types of clues, like time or visual similarity, are trustworthy when recalling memories. This approach works better than fixed methods and adapts as situations change.

What this means in practice

  • For robotics developers: Improve robots' ability to predict future events by selectively recalling past observations using multiple cues.
  • For game ai programmers: Create non-player characters that better anticipate player actions by adaptively recalling useful past experiences.

Authors

Beomsu Kim, Chieh-Hsin Lai, Bac Nguyen, Amir Bar, Jong Chul Ye, Yuki Mitsufuji

Abstract

World models predict future observations from current experience and actions, yet prediction can depend on observations seen far in the past. Episodic memory preserves past observations for later recall; however, as memory accumulates, it raises a fundamental question: which memories are useful for the current prediction, and which available retrieval cues should be trusted to find them? This is challenging because fixed criteria based on recency, pose overlap, or visual similarity can be unreliable across environments and queries. We propose Future-Aware Recall (FAR), a framework that learns episodic recall from future-aware predictive supervision and adaptive multi-cue scoring. During training, FAR measures predictive utility by the conditional log-likelihood of the realized future given recalled context, approximated by negative diffusion prediction loss, and uses it to train a retriever that remains future-blind at inference. The retriever learns cue-specific relevance and automatically determines which available retrieval cues, such as time, pose, vision, and audio, to trust for each query when selecting memories. Across three complementary settings, FAR outperforms hand-designed recall even with the same retrieval cues, automatically adapts which available cues to trust, and recalls the right history as the world changes. Together, these results establish FAR as a flexible, principled approach to episodic memory access in world models.