Papers for

automated customer support developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

CacheRepair speeds up language model retrieval with smarter cache fusion

CacheRepair: Learning to Repair Cross-Chunk Context in RAG for KV Cache Fusion

Abstract: Multi-document retrieval-augmented generation (RAG) requires a language model to process multiple retrieved text chunks before answering a question. Precomputing each chunk's KV cache independently and concatenating the caches when the chunks are retrieved can accelerate this step. However, the assembled cache lacks cross-chunk attention information, reducing answer quality. Selective recomputation methods recover the missing cross-chunk context by rerunning the target LLM on selected tokens, incurring substantial online computation. We introduce CacheRepair, a lightweight network that learns the difference between independently computed KV caches and those produced by processing the chunks together. The network combines compressed KV features with token embeddings and uses attention that is bidirectional within each chunk and flows from earlier to later chunks. Each repair block receives the compressed cache features, and the predicted residual is added to every document token's cache. Each repair network is trained for a specific frozen target LLM on a generic retrieval corpus and reused across downstream datasets. Our analysis shows that repair reduces KV errors both near chunk boundaries and throughout chunk interiors. Evaluation across three target LLMs and four downstream datasets places CacheRepair on the measured answer-quality-latency Pareto frontier in eleven of twelve model-dataset combinations. Reported time to first token (TTFT) includes online cache transfer and repair. Across all twelve combinations, the largest repairers achieve 1.69-4.61$\times$ speedups in median TTFT over full prefill and improve mean F1 by 2.1-26.1 percentage points over direct cache reuse.

Mon 28 SeptMachine Learning
The gist
When a language model answers questions using multiple documents, it processes pieces of text called chunks. Usually, it stores information from each chunk separately, which makes combining them less accurate and slower. The authors created CacheRepair, a small network that fixes missing links between chunks by learning how caches differ when chunks are processed together. This method makes answering faster without losing quality and works well across many models and datasets.
Open → 2609.35139v1

Agents reuse historical action credit to reduce tool interactions

When Does Action Credit Need Updating?

Abstract: Tool-using agents are continually updated with new interaction data. After each policy update, however, previously estimated action credits may become stale. Recomputing them from scratch can require many additional tool calls and environment interactions, making repeated updates increasingly expensive. We ask a simple question: when does historical action credit actually need to be updated? Our key observation is that a change in action value does not necessarily imply a change in the decision. Historical credit can still be useful as long as policy-induced drift is too small to overturn the existing action ranking. Building on this idea, we introduce pairwise branch sensitivity to capture how strongly a policy update affects the downstream regions that distinguish two candidate actions. We then derive a first-order anchored credit-transport estimator that updates historical credit using old interventional trajectories, and propose a Decision-Sufficient Credit Gate (DSC-Gate) that chooses whether to reuse, transport, or resample credit. Experiments show that branch sensitivity explains credit drift substantially better than global policy distance. With sufficient historical data, credit transport reduces estimation error, while its benefit to decision making is concentrated on updates that affect action-distinguishing branches. On a fully independent test set, DSC-Gate changes mean regret by only +0.00004 relative to a gap-based gate while reducing mean new tool steps from 472 to 286, a 39.4% reduction. We observe the same pattern after a real tool-agent parameter update. Overall, our results show that agents do not need to recompute action credit after every policy update: much of the historical evidence can be reused or cheaply corrected, reducing the additional interaction required to keep action decisions up to date.

Thu 24 SeptArtificial Intelligence
The gist
For agents that use tools, updating decisions after each change can be slow and costly if they recalculate old information every time. The authors found that agents often do not need to redo all their previous assessments because small updates usually do not change which option is best. They created a way to measure when old information stays useful and a method to update it efficiently. Their system cuts down the number of times agents have to redo tool actions by nearly 40% while keeping decision quality nearly the same.
Open → 2609.29007v1