Stashbird cuts memory tokens for ai agents with indexed updates
Stashbird: Efficient Speaker-Indexed Memory for Conversational Agents
Artificial IntelligenceMachine Learning
Summary
AI assistants need to remember past conversations accurately, including who said what and when information changes. The authors created Stashbird, a system that organizes memory with clear links to original conversations and allows easy updates or deletions. This system uses much fewer resources to process and recall information than some existing methods, while maintaining similar or better accuracy. Stashbird is tested across multiple benchmarks involving long-term conversational memory.
What this means in practice
- •For ai assistant developers: Build conversational AI that remembers and updates past talks efficiently with minimal prompt tokens and accurate recall.
- •For customer support platform engineers: Integrate better memory handling in chatbots to manage multi-party conversations and update customer info as it changes.
Authors
Chidera Biringa, Lucas Yannul, Xiaowen Wang, Marco Ayala, Nicholas Yi, Alex Moyse, Nishant Manchanda, Vivek Gupta
Abstract
AI agents require memory that preserves information across user-agent exchanges, user-to-user conversations, and group conversations with or without agent participation, while supporting updates as evidence changes or is removed. We present Stashbird, an agent memory system that links source episodes to derived memory state through explicit provenance. Stashbird organizes memory into episodic records, semantic relations, community summaries, and persisted graph state, with lifecycle operations for incremental updates and episode-level deletion. We evaluate question-answering accuracy and model-facing workload across four long-term memory benchmarks. On LoCoMo, Stashbird uses 76.4x fewer ingestion prompt tokens than Graphiti. Compared with reproduced Hindsight on the same benchmark, it uses 8.1x fewer retrieval prompt tokens, with accuracy 1.6 percentage points lower. It achieves higher accuracy than Hindsight on LongMemEval-S and GroupMemBench and comparable accuracy on EverMemBench.