Agent memory often breaks when AI models get upgraded
Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability
Artificial IntelligenceComputation and LanguageInformation Retrieval
Summary
When AI agents get new versions of their brains, their memory doesn’t always carry over well. The researchers found that if the memory is stored in a strict, organized format, it transfers smoothly. But when memory is compressed into notes or stored in chunks with embeddings (special number forms), accuracy drops a lot after the upgrade. They also show that keeping the original memory helps fix problems, while just trying to repair the compressed memory usually fails. This means AI systems need careful testing and handling when their models change.
memory portabilitymodel upgradeembeddingretrieval-augmented generationknowledge graphmemory compressionlong context readingmemory repair
Authors
Ankit Goyal, Jaideep Ray
Abstract
Model upgrades are routine; memory migrations are not. An agent can keep the same memory store and still forget: a new model may interpret old notes differently, mixed embedding versions may break retrieval, and repair may fail without the original evidence. We compare memory as the same history is preserved verbatim for long-context reading (LC-RAW), divided into chunks for retrieval-augmented generation (RAG), compressed by a model into natural-language notes (NOTES), or normalized into a fixed-schema knowledge graph (KG-fixed). The study uses 48 synthetic histories with randomized answer codes, exact scoring, and two open-weight models with sub 10 billion parameters. Our measurements show that fixed-schema structures transfer reliably, with KG-fixed accuracy changing by only $+0.0004 \pm 0.0020$ following a writer swap. Conversely, compressed NOTES exhibit high model coupling, with accuracy shifting asymmetrically by $+9.91$ or $-13.28$ percentage points depending on the specific migration direction. In RAG systems, partial embedding migrations using a 50/50 mixed index capture only a 4.96-point accuracy improvement, forfeiting the majority of the 11.90-point gain achieved through full re-embedding. Diagnostic decomposition attributes 80% ($0.467 \pm 0.014$) of the NOTES accuracy deficit to information lost during initial construction, whereas retrieval failures drive 81% ($0.364 \pm 0.012$) of the RAG deficit. Finally, store-only repair of NOTES fails to reach a 90% performance recovery target in all 48 test cases, whereas retaining the raw source history enables successful recovery in 34 of 48 cases for one tested direction. These findings highlight the necessity of direction-specific migration testing, strict embedding space isolation, and the retention of source histories for memory repair.