Papers for

search infrastructure teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

CacheRepair speeds up language model retrieval with smarter cache fusion

CacheRepair: Learning to Repair Cross-Chunk Context in RAG for KV Cache Fusion

Abstract: Multi-document retrieval-augmented generation (RAG) requires a language model to process multiple retrieved text chunks before answering a question. Precomputing each chunk's KV cache independently and concatenating the caches when the chunks are retrieved can accelerate this step. However, the assembled cache lacks cross-chunk attention information, reducing answer quality. Selective recomputation methods recover the missing cross-chunk context by rerunning the target LLM on selected tokens, incurring substantial online computation. We introduce CacheRepair, a lightweight network that learns the difference between independently computed KV caches and those produced by processing the chunks together. The network combines compressed KV features with token embeddings and uses attention that is bidirectional within each chunk and flows from earlier to later chunks. Each repair block receives the compressed cache features, and the predicted residual is added to every document token's cache. Each repair network is trained for a specific frozen target LLM on a generic retrieval corpus and reused across downstream datasets. Our analysis shows that repair reduces KV errors both near chunk boundaries and throughout chunk interiors. Evaluation across three target LLMs and four downstream datasets places CacheRepair on the measured answer-quality-latency Pareto frontier in eleven of twelve model-dataset combinations. Reported time to first token (TTFT) includes online cache transfer and repair. Across all twelve combinations, the largest repairers achieve 1.69-4.61$\times$ speedups in median TTFT over full prefill and improve mean F1 by 2.1-26.1 percentage points over direct cache reuse.

Mon 28 SeptMachine Learning
The gist
When a language model answers questions using multiple documents, it processes pieces of text called chunks. Usually, it stores information from each chunk separately, which makes combining them less accurate and slower. The authors created CacheRepair, a small network that fixes missing links between chunks by learning how caches differ when chunks are processed together. This method makes answering faster without losing quality and works well across many models and datasets.
Open → 2609.35139v1

Embedding subspaces enable flexible multi-goal recommendations at scale

Embedding Subspace Partitioning for Dynamic Multi-Objective Retrieval

Abstract: Modern industrial recommender systems must optimize across competing objectives, balancing semantic relevance with business metrics such as engagement and revenue. While bi-encoders dominate large-scale retrieval due to their efficiency, they collapse these heterogeneous signals into a single static embedding space. This design creates a fundamental limitation: once trained, the retriever cannot adapt to shifting objective priorities at serving time without retraining. Moreover, joint optimization with multi-objective losses often induces interference between objectives, leading to suboptimal trade-offs. We propose Embedding Subspace Partitioning (ESP), a retrieval framework that decomposes the embedding into task-aware subspaces and replaces the single dot product with a weighted sum of per-subspace similarities, whose weights are tunable at serving time. For Transformer bi-encoders, ESP uses the model's native end-of-sequence token as a segment delimiter, with segment-aware attention masking and position encoding resets to guarantee subspace isolation in a single forward pass. Serving is performed via GPU-accelerated exhaustive kNN over one concatenated index, eliminating the need for per-objective Approximate Nearest Neighbor (ANN) infrastructure required by multi-head approaches. We evaluate ESP on an open-source benchmark built from MS MARCO. A single ESP model traces a broad Pareto frontier, consistently outperforming strong multi-task baselines across diverse operating points. In LinkedIn's job matching platform (70M+ weekly users), ESP enabled dynamic retrieval reconfiguration and delivered significant key business metric lifts.

Thu 24 SeptInformation Retrieval
The gist
Recommender systems often need to balance different goals like showing relevant items and increasing business value, but traditional methods mix all goals into one fixed understanding. The authors propose splitting the understanding into separate parts, each focused on a specific goal, so the system can adjust priorities on the fly without retraining. Their method works efficiently on large-scale systems by doing one search over combined data rather than multiple searches. Tests show this approach better balances goals and improved user engagement in a large job matching platform.
Open → 2609.30601v1

Hybrid gpu and cpu system improves personalized search at trillion scale

Hybrid GPU-CPU Retrieval for Personalized Search at Ultra-Large Scale

Abstract: Embedding-based retrieval on user-generated content at the trillion-document scale exposes a sharp conflict between two production demands: deep, expressive personalization for queries with rich user intent, and broad coverage of a massive inventory under fixed latency and resource budgets. We characterize this as the personalization-scale paradox: hosting the full serving inventory in GPU memory is too resource intensive, while CPU compute cannot execute the same interaction-heavy model on the latency-critical path. We present a hybrid GPU-CPU co-serving system that resolves the paradox through orchestration rather than a new model class. A high-depth GPU pathway fuses retrieval and interaction pre-ranking over a curated online pool on the order of a billion documents, while a high-breadth CPU pathway searches an independently selected online inventory roughly twenty times larger with lightweight personalized scoring. Either or both pathways can run per request; candidates are deduplicated before shared downstream ranking. The system is deployed in production. A full-system A/B test against the legacy CPU-only configuration improves model-scored relevance and substantive engagement, while separate pathway experiments show positive value at their own deployment scopes. Retrieval logs show that the pathways contribute structurally distinct candidates, production serving measurements characterize their latency, and a matched capacity plan quantifies the economic rationale for assigning modeling depth to GPUs and inventory breadth to CPUs. Together, these results validate a practical, independently evolvable depth-breadth architecture for ultra-large-scale personalized search.

Fri 18 SeptInformation RetrievalDistributed, Parallel, and Cluster ComputingMachine Learning
The gist
Searching billions to trillions of documents to find what each user wants is very hard because powerful methods require a lot of memory but must respond quickly. The authors describe a system that uses both GPUs and CPUs together: GPUs do deep and detailed searching on a smaller set, while CPUs quickly scan a much larger set with simpler methods. This mix helps balance speed, resource use, and personalization, and it runs in production with better results than previous CPU-only approaches.
Open → 2609.21281v1

AI privacy improves by auditing vector retrieval process safety

Beyond Private Training: The New Landscape of AI Privacy

Abstract: Retrieval-augmented systems increasingly rely on vector indexes that may retain deleted items in their search graph. Existing deletion interfaces can prevent deleted identifiers from appearing in returned results while still computing distances to their embeddings during graph traversal. We formalize this distinction as output safety versus traversal safety, and introduce TSD-AUDIT, a framework for auditing and enforcing traversal-safe deletion in graph-based approximate nearest-neighbor retrieval. On Faiss IndexHNSWFlat, native filtering leaves the number of distance computations unchanged relative to unfiltered search; at a 70% deletion rate, trace-faithful replay detects deleted-vector scoring in all 100 audited queries. Code inspection of hnswlib's mark_deleted path reveals the same scoring-before-liveness pattern. TSD-AUDIT enforces an alive-before-scoring invariant, repairs connectivity using only live candidates, and emits per-query scored-trace certificates that an independent verifier can check against the deletion snapshot. Under region-targeted deletion, TSD-AUDIT improves Recall@10 over native filtering by 4.3--42.2 percentage points across deletion fractions from 0.5 to 0.9, while remaining comparable under random deletion. These results show that output-only deletion audits can miss process-level exposure: auditing deletion in vector retrieval requires accounting for the vectors scored during search, not only the identifiers returned.

Wed 16 SeptCryptography and SecurityInformation Retrieval
The gist
When AI systems look up information stored as vectors, deleted items may still affect how results are found even if they don’t appear in outputs. The authors studied this hidden exposure risk and made a tool called TSD-AUDIT to check and fix it. Their tool ensures deleted vectors aren’t even considered during search, not just hidden afterward, which keeps deleted data truly private. This improves search accuracy after deleting parts of the data, showing that just filtering results is not enough to protect privacy.
Open → 2609.19456v1