Tessera cuts delays in memory-heavy AI model requests by up to 3.6 times
Tessera: Demand-Driven KV Cache Management for Retrieval-Augmented LLM Serving
Distributed, Parallel, and Cluster Computing
Summary
Large language models (LLMs) often use extra information from documents or memory to answer questions better. However, reusing parts of previous computations to save time is tricky because this useful content can appear in different places or orders, making simple caching ineffective. The authors designed Tessera, a system that smartly manages caching by knowing what pieces of information are really needed upfront and sharing that knowledge across servers to prepare answers faster. Tessera speeds up response times significantly while keeping answer quality high, especially when many users ask questions at once.
What this means in practice
- •For cloud ai service providers: Serve LLM-based applications with faster response times by managing memory caching more efficiently under heavy query loads.$Commercial implications: Enables cloud providers to build faster AI text generation APIs that handle complex memory usage and large-scale concurrency, creating competitive hosting services.
- •For software engineers building ai assistants: Improve speed and scalability of retrieval-augmented LLM-powered assistants by coordinating cache and request routing across multiple servers.
Authors
Fei Fang, Chung-Hsiang Lo, Yi Liu, Yifan Hua, Chen Qian
Abstract
RAG and retrieval-based agent memory both inject retrieved content into LLM prompts, as document chunks and recalled memory records, respectively. The same content can recur across requests at different prompt positions or after different preceding contexts, preventing reuse through conventional prefix caching. Our characterization finds that records recurring outside the matching prefix account for over 70% of injected memory tokens in agent-memory workloads. Composable KV-reuse methods enable reuse in such cases, but online serving introduces a management problem: a recurring unit's KV states may not yet exist, may have been evicted, or may reside on another node. We present Tessera, a disaggregated serving system that makes retrieval the control plane for KV reuse. By exposing the context units needed before model execution, retrieval allows Tessera to combine current demand with retrieval history, KV residency, and generation load to coordinate cache management and request routing. Generation nodes concurrently prepare locally cached, remotely cached, and missing states, while retaining newly computed states off the request's critical path. Across RAG and agent-memory workloads, Tessera lowers mean TTFT by up to 3.6x over SGLang and LMCache with EPIC at matched request rates, and sustains low TTFT at rates where the baselines saturate, while matching the answer quality of the underlying composition policy.