Memory system improves long term chat with less token use

Remember by Asking: Retrieval-Induced Memory Evolution for LLM Agents

Computation and LanguageArtificial Intelligence

Summary

Keeping track of conversations over many sessions is hard for language AI because important details can get lost when everything is compressed at once. The authors present RIME, which builds memory step-by-step by asking itself helpful questions to find key parts of past talks. This method keeps memory updated over time and uses the stored facts first when answering, but can also retrieve more context if needed. Their tests show RIME works better and uses fewer resources than other memory methods.

What this means in practice

  • For chatbot developers: Enhance chatbot memory across multiple user sessions by selectively integrating key conversation details for improved response accuracy with less computational cost.
  • For customer support teams: Provide conversational AI tools that maintain relevant previous interaction details more effectively, helping agents reference past cases without processing full histories.
  • For personal assistant app builders: Create AI assistants that remember and recall user preferences over long periods with efficient memory use, improving personalized interactions without expensive processing.$Commercial implications: Enables marketable personal assistant apps that offer smarter long-term memory capabilities, differentiating products through sustained, context-aware user engagement.

Authors

Wanqi Zhou, Jiawei Lu, Yang Wang, Zhaolong Xing, Zhen Chen, Ai Han, Haoyue Shi

Abstract

Long-term memory is essential for language agents to maintain coherent and effective behavior over extended, multi-session interactions. Existing memory systems mainly use retrieval at read time, while write-time memory formation still relies on direct extraction or compression. However, when future information needs are unknown, compressing an entire interaction in one pass can overlook locally important details that may matter later. To this end, we introduce RIME, a retrieval-induced memory framework that shifts memory construction from monolithic compression toward evidence-centered integration. RIME uses generic self-questions to retrieve focused dialogue evidence and grounds memory formation in both the retrieved evidence and relevant historical memories, which are jointly reconciled into an evolving memory bank with temporal and provenance information. At inference time, compressed memory serves as the primary rather than the sole source of evidence: when it cannot support an answer, RIME retrieves relevant source dialogue together with its local context to recover information omitted during memory formation, without resorting to full-history processing. Extensive experiments on LoCoMo with Qwen3-235B-A22B and GPT-5.6 Sol show that RIME consistently achieves the best performance across all three quality metrics among the compared methods, while requiring substantially fewer query-time LLM tokens.