Counterfactual memory improves language agent task performance consistently
COUNTERMEM: World-Model Verified Counter-Factual Memory for Language Agents
Artificial Intelligence
Summary
Memory systems in language agents usually learn from what actually happened, but don’t imagine what could have been. The authors introduce COUNTERMEM, a framework that lets an agent consider alternative actions after a failure by simulating and verifying outcomes. This approach helps the agent remember better and choose smarter next steps without changing its core language model. Tested across various tasks and language models, COUNTERMEM shows clear improvements in success rates and efficiency.
What this means in practice
- •For conversational ai developers: Improve dialogue agents by storing and using verified alternative action outcomes to reduce errors and enhance response accuracy.
- •For automated reasoning engineers: Use verified counterfactual feedback to optimize agent decisions in environments requiring proof checking or solver integration.
Authors
Hongji Pu, Ruixiang Tang, Yongfeng Zhang
Abstract
Existing agent memory frameworks mainly create memory through an agent's interaction with the factual world, e.g., remembering feedback from actions taken to improve performance on future tasks. However, these frameworks seldom ask the "what if" question during memory construction: what if a different action had been taken, would the feedback have changed, and how could this feedback become useful memory? Obtaining such feedback directly in an active environment can be expensive and can alter the state needed for comparison. In this work, we introduce COUNTERMEM, a reinforcement-learning framework for constructing and using verified counterfactual memory across tasks. After a failed action, COUNTERMEM evaluates local alternatives from a copy or reset of the original state using executable world models, such as tests, proof checkers, and solvers. It stores improvements with the original and corrected actions, checked outcomes, and conditions for reuse. A learned memory-use policy selects a retrieved record or skips memory to balance task success and interaction cost, while the base LLM remains fixed. Both memory and policy are frozen during held-out evaluation. We evaluate COUNTERMEM on 12 benchmark settings across six domains. With gpt-oss-120b, COUNTERMEM improves both ReAct and Reflexion on all 12 benchmarks across six domains, averaging a gain of 12.6 percentage points over their unaugmented versions. In the four-domain comparison across two backbones, task-run tokens decrease by 7.7-42.0%, excluding offline selector-training costs. Further analyses show that removing verification or persistent storage weakens the gains, while applying verified corrections to unsuitable decisions can reverse them. Code will be released upon acceptance.