RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States

2026-08-03Machine Learning

Machine LearningComputation and Language
AI summary

The authors address problems in learning-based memory systems for language model agents, where feedback gets spread thin over growing interaction histories and irrelevant memories can get wrongly rewarded. They propose a method called Reduced-Order Memory Reinforcement Learning (RoMeRL) that compresses this growing utility space into a fixed set of memory states organized by outcome type and memory changes. This approach helps concentrate feedback on meaningful memories and prevents reward errors from persisting. Tests show RoMeRL improves task success, makes feedback more efficient, reduces memory use, and lowers the need for calls to large language models.

Language Model (LLM)Memory Reinforcement LearningTrajectory-indexed UtilitiesReward ContaminationReduced-Order ParameterizationSemantic CoordinatesFeedback DensitySelf-evolving AgentsTask PerformanceCold-Q Ratio
Authors
Yi Yang, Zhennan Chen, Yihong Zhuang, Tiehan Fan, Yinan Chen, Jian Li, Jian Yang, Ying Tai
Abstract
Learning-based memory systems for self-evolving LLM agents face two tightly coupled challenges. First, trajectory-indexed utilities grow with the interaction history, thereby dispersing limited feedback over an ever-expanding state space. Second, because trajectory-level rewards are jointly assigned to co-retrieved memories, irrelevant experiences may receive misleading utility updates and consequently enter the memory-reward trap. To address these challenges, we introduce Reduced-Order Memory Reinforcement Learning (RoMeRL), which represents the growing trajectory-indexed utility space using a fixed-dimensional per-task memory state factorized by outcome polarity and memory dynamics. RoMeRL incorporates new experiences through a fixed set of semantic coordinates whose contents are updated or replaced over time, thereby concentrating feedback over a bounded utility support. Theoretically, we show that this reduced-order parameterization increases the average feedback received by each utility coordinate and characterize the steady-state occupancy of erroneous coordinates under a generic coordinate-transition model. Empirically, across ALFWorld and LifelongAgentBench, RoMeRL improves task performance, reduces the Cold-Q ratio by 80.0%, increases feedback density by approximately 6.0 times, reduces the maintained memory size by 84.4%, and cuts LLM calls by 21.1%. These results show that reduced-order utility states support efficient self-evolving agent memory while limiting persistent reward contamination. Code is available at: https://github.com/YOUNG-fnxm/RoMeRL