Muscle Memory for Agents: Compile not Merely Retrieve

2026-08-10Multiagent Systems

Multiagent Systems
AI summary

The authors point out that current methods for memory in language model (LLM) assistants usually involve storing and retrieving information, but this may not be the best way to personalize user interactions. They propose a new approach called Muscle Memory, which creates specialized agents tailored to common user intents by compiling patterns from past conversations. Their system goes through steps to gather, analyze, enhance, and test these specialized agents, leading to better personalization with minimal loss in accuracy. The authors also explain why this compilation method works better for some tasks and discuss what this means for future memory designs.

LLM agentsmemory architecturespersonalizationspecialist agentscompilationretrievalmulti-turn conversationuser intenttrigger matchingevaluation metrics
Authors
Pouya Ghiasnezhad Omran, Soujanya Lanka, Qin Zhang, Tanya Dixit
Abstract
Memory for LLM agents has converged on a single architectural pattern: store experience as text, embeddings, reflections, or rules; retrieve at inference time; let a general-purpose orchestrator interpret what to do. This paper argues that the pattern is the wrong default for personalization. We position Muscle Memory - the practice of compiling recurring user intent into purpose-built specialist agents - as a distinct memory paradigm from retrieval, and we argue that compilation is a better fit for the workloads where current assistants impose a multi-turn tax on their users: making them repeatedly correct format, depth, and scope to obtain a domain-appropriate answer. We support the position with a reference implementation and empirical evidence. The implementation is a four-phase pipeline (Harvest $\rightarrow$ Analyze $\rightarrow$ Augment $\rightarrow$ Evaluate) that mines conversational history, separates behavioral from task patterns, and emits quality-gated executable compiled specialists with two-stage trigger matching. On 90 held-out scenarios across five user personas, the augmented assistant wins 32 of 36 cases where a specialist fires, an 88.9% win rate, with a +2.05 personalization gain and only a $-0.28$ accuracy cost on a 1-4 scale. We discuss why compilation is better suited than retrieval in this regime, what the result implies for the broader memory design space, and what open problems remain.