Agent memory works as well selecting raw turns as extracting facts

When Does Selection Replace Extraction? A Pre-Registered Test of Agent Memory with a Typed Decision Model

Artificial IntelligenceComputation and LanguageInformation Retrieval

Summary

This paper looks at how computers remember information in conversations. It asks whether it’s better to pick important parts of the chat directly or to first turn those parts into facts using fancy language models. The authors found that just choosing the right parts of the conversation works just as well as making new fact summaries, but costs far less time and computing. They also discovered that rearranging which parts to keep helps a little more when you have only a few pieces to work with. This helps explain why previous studies disagreed on the best method.

What this means in practice

  • For chatbot developers: Use typed decision models to select raw conversation turns efficiently for memory, reducing costs without hurting chatbot answer quality.
  • For customer support teams: Improve memory retrieval in support chat systems by selecting key conversation parts rather than generating extra fact summaries, speeding up processing.

Authors

Rishabh Sharma, Rishika Lall

Abstract

Does conversational memory need LLM-extracted facts, or is selecting the right raw turns enough? Published results disagree. Extraction-based systems report gains from distilled facts. Recent studies find raw history with good ranking does as well, but disagree about whether ranking matters. We ran a pre-registered study on held-out LoCoMo conversations and LongMemEval. At a tight budget on LoCoMo, raw turns selected by a single call to Jev, a typed decision model, are non-inferior to an LLM-extraction memory (one-sided 95% bound -3.0 points against a -5-point margin). Blind human grading narrows the margin but does not change the result. Raw turns cost 3,061 times less to write, and the result holds with a second answer model. Within this study, reranking's gain shrinks as the budget grows. It adds 17.4 points on LoCoMo and 9.1 on LongMemEval when three of 30 candidates are kept. At generous budgets it adds 1.5 and 1.1, and extraction systems are more accurate. This suggests why published results disagree. At matched context, Jev selects as accurately as an LLM reranker (non-inferiority bound -2.0) at a third of the latency, and more accurately than a multi-call graph traversal. Reranking lowers correct abstention. Plans, code and graded answers are released.