Memory system for AI agents improves itself before errors happen
Remember Before You're Asked: MemDream for Self-Probing Memory Evolution
Artificial Intelligence
Summary
Language models that act as agents need memory to keep track of information over long tasks, but usually they only fix memory mistakes after they cause problems. The authors propose MemDream, a method where the system pretends to test its own memory when it's not actively working, finding and fixing issues early. This approach uses three roles working together to analyze and improve memory proactively, resulting in better answers and performance on tests compared to older methods. MemDream helps AI agents remember better before real mistakes occur.
What this means in practice
- •For virtual assistant developers: Improve long-term personalized memory in conversational AI by proactively detecting and fixing memory errors before user queries expose them.
- •For customer support automation teams: Enhance chatbot memory reliability by periodically testing and refining stored information to reduce retrieval mistakes during customer interactions.
Authors
Mingfei Lu, Mengjia Wu, Runsong Jia, Zhe Luo, Yi Zhang
Abstract
Memory is essential for enabling LLM-based agents to maintain coherent, personalized behavior over long-horizon interactions. However, existing memory systems share a fundamental limitation: they never proactively test their own memory, repairing it only after real queries expose weaknesses. This reactive paradigm means every retrieval failure corresponds to a real interaction in which the cost has already been paid. We propose MemDream, a framework that enables self-probing memory evolution for LLM agents. Our framework periodically enters offline dream cycles where three specialized agents (Dreamer, Analyst, Consolidator) collaboratively probe, diagnose, and repair the memory graph before failures occur. A policy trained via Group Relative Policy Optimization learns which repair operations produce durable retrieval improvements, while a soft decay mechanism provides reversible forgetting driven by the same anticipatory signal. Experiments on LoCoMo and MemoryAgentBench demonstrate that MemDream improves answer F1 by 4.5 points on LoCoMo and achieves a 9.1-point higher overall score on MAB over the strongest reactive-evolution baselines.