CacheRepair speeds up language model retrieval with smarter cache fusion

CacheRepair: Learning to Repair Cross-Chunk Context in RAG for KV Cache Fusion

Machine Learning

Summary

When a language model answers questions using multiple documents, it processes pieces of text called chunks. Usually, it stores information from each chunk separately, which makes combining them less accurate and slower. The authors created CacheRepair, a small network that fixes missing links between chunks by learning how caches differ when chunks are processed together. This method makes answering faster without losing quality and works well across many models and datasets.

What this means in practice

  • For machine learning engineers: Accelerate multi-document question answering systems by efficiently repairing key-value caches during chunk fusion to reduce latency without degrading answer quality.
  • For search infrastructure teams: Improve response times in retrieval-augmented language model services by integrating CacheRepair to maintain cross-chunk context in cached representations.
  • For automated customer support developers: Speed up retrieval-based chatbots by applying CacheRepair to enhance multi-document context handling for faster and more accurate replies.$Commercial implications: Enables more responsive and higher-quality AI chatbots that can be sold to enterprises requiring efficient automated customer interactions.

Authors

Genglin Wang, Wangsong Yin, Yeerzhati Abudunuer, Haoxuan Xu, Guoliang Xing, Zhenyu Yan

Abstract

Multi-document retrieval-augmented generation (RAG) requires a language model to process multiple retrieved text chunks before answering a question. Precomputing each chunk's KV cache independently and concatenating the caches when the chunks are retrieved can accelerate this step. However, the assembled cache lacks cross-chunk attention information, reducing answer quality. Selective recomputation methods recover the missing cross-chunk context by rerunning the target LLM on selected tokens, incurring substantial online computation. We introduce CacheRepair, a lightweight network that learns the difference between independently computed KV caches and those produced by processing the chunks together. The network combines compressed KV features with token embeddings and uses attention that is bidirectional within each chunk and flows from earlier to later chunks. Each repair block receives the compressed cache features, and the predicted residual is added to every document token's cache. Each repair network is trained for a specific frozen target LLM on a generic retrieval corpus and reused across downstream datasets. Our analysis shows that repair reduces KV errors both near chunk boundaries and throughout chunk interiors. Evaluation across three target LLMs and four downstream datasets places CacheRepair on the measured answer-quality-latency Pareto frontier in eleven of twelve model-dataset combinations. Reported time to first token (TTFT) includes online cache transfer and repair. Across all twelve combinations, the largest repairers achieve 1.69-4.61$\times$ speedups in median TTFT over full prefill and improve mean F1 by 2.1-26.1 percentage points over direct cache reuse.