Papers for

search infrastructure engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Gating skips reranker calls for faster multi-hop retrieval with small accuracy loss

Per-Query Gating of LLM Rerankers for Multi-Hop Retrieval

Abstract: LLM rerankers add of the order of \$0.2-0.3 per 1,000 queries and about a second of tail latency on top of a graph-augmented dense pipeline such as HippoRAG2, and on three multi-hop benchmarks they improve final-hop top-K coverage on seven of nine (dataset, K) cells, by up to +34.8 pp. We ask whether a learned per-query gate can skip the reranker where it will not help, using only features available before the LLM call (27 score and lexical statistics of the two retrieval lists plus a PCA of a small query embedding) with an executable fallback. Every choice, including the fallback and the threshold, is made inside the training fold and applied once to held-out queries, and harmful skips (the rerank would have found the target, the fallback did not) are reported next to the aggregate coverage. Across nine cells on 2WikiMultiHopQA, MuSiQue and HotpotQA the gate skips 51% of calls at an average held-out LastHop@K cost of 1.2 pp; four cells meet a pre-registered 1 pp rule, harmful skips occur in eight (190 harmful against 136 beneficial), and a random gate at the same skip rate loses 2 to 11 pp on the high-lift cells. A second rule sets each cell's threshold from a pre-specified budget on the expected harmful-skip rate over Platt-calibrated harm probabilities (ECE 0.025 after calibration, 0.094 before): at a 1 pp budget the gate skips 42% at -0.8 pp with 66 harmful skips and six cells within 1 pp, but realised harm exceeds the promise in six cells (mean 1.45 vs 0.83 pp), a selection optimism we quantify; a 0.5 pp budget realises about 1 pp. The harm probabilities are calibrated but barely discriminative (AUC 0.16 to 0.70). An earlier version reported 73% "lossless" savings; that figure rested on an oracle fallback and a wrong MuSiQue target, and we document both.

Sat 19 SeptInformation RetrievalArtificial IntelligenceComputation and Language
The gist
Large language model (LLM) rerankers improve search accuracy but slow down the system and add cost. The authors develop a gate that decides for each query whether the reranker is needed, based on quick-to-get features before running the costly reranker. This gate skips about half of reranker calls, saving time and money with only a small drop in accuracy. The authors carefully measure harm caused when the gate skips useful reranking and adjust thresholds to control this risk. The gate’s predictive power is limited, so some harmful skips remain despite calibration.
Open 2609.22880v1

Mixture-of-experts language models improve search speed and accuracy

Mixture-of-Experts Language Models Can Be Strong and Efficient Retrievers

Abstract: Recent work has shown that fine-tuning decoder-only large language models (LLMs) for retrieval yields strong first-stage retrievers, with effectiveness improving as backbones grow in size. However, every query and document must pass through the full model, so encoding cost increases with model size. Mixture-of-Experts (MoE) LLMs activate only a subset of parameters per token and are widely used to scale generative models, yet remain underexplored as retrievers. We systematically study MoE backbones for retrieval by training MoE and dense LLMs from several families using the same procedure, evaluating them across diverse datasets, and measuring query encoding time under the same serving configuration. We show that MoE retrievers outperform dense retrievers with comparable active parameter counts by up to 3.0 nDCG@10 points on BEIR. One of our strongest MoE retrievers matches an 8B dense retriever with 59% fewer active parameters and 18% lower query encoding time. We further show that the number of experts used for query encoding can be reduced without retraining or re-indexing, retaining more than 99% of retrieval effectiveness while reducing query encoding time by up to 26%. Recent rerankers provide only modest additional gains over strong MoE first stages, which often match or exceed the reranked configurations we evaluate. Together, these results show that MoE LLMs can be strong and efficient first-stage retrievers.

Fri 11 SeptInformation RetrievalComputation and Language
The gist
Searching through large amounts of text quickly and accurately is hard because big language models take a lot of computing power. The authors studied a special kind of model called mixture-of-experts (MoE) that only uses part of its brain when answering, making it faster. They found these MoE models can find relevant information better and more efficiently than regular models of similar size. Also, you can reduce how much of the model is active during search without losing much accuracy. This means MoE models can make searching smarter and quicker at the same time.
Open 2609.13486v1

Compact binary codes improve document search efficiency and accuracy

Matryoshka Hash Representations for Model-Aware Compact Semantic Retrieval

Abstract: Retrieval-augmented generation (RAG) depends on dense retrieval: each document is stored as a learned vector, and a query is answered by finding its nearest neighbors in that vector space. Keeping one full-precision vector per document is the dominant index cost at corpus scale, so retrieval systems replace each vector with a short code of a few bytes---a step called quantization. Standard quantizers such as product quantization (PQ) pick the code that reconstructs the original vector most closely. A single code is even more useful if it serves several byte budgets at once: when its short prefixes are each directly searchable, a deployment can set its efficiency--quality operating point without re-encoding the corpus. But training all prefixes under one objective makes the early bits a compromise across budgets---short codes improve while the full-width code degrades. Quantization to low-bit representation, such as binary codes, further sharpens the conflict. We introduce Matryoshka Hash Representations (MHR), a two-stage procedure that separates full-width training from prefix organization. MHR first learns a longer binary code, then freezes the model and trains additional zero-initialized residual code adaptors for directly searchable prefixes. Documents are stored at one bit per coordinate, while queries keep continuous logits like PQ to attain sufficient expressivity. We implement the search process with FAISS FastScan. Trained on MS MARCO and zero-shot transferred to seven BEIR datasets, MHR reaches .5561 NDCG@10 and .6535 Recall@100 at 32 bytes, surpassing the best baseline of the same budget. The advantage is more pronounced in lower budgets. The same code also strengthens two common pipelines: shortlisting candidates for full-precision reranking, and pruning a low-storage graph index such as LEANN.

Mon 7 SeptInformation RetrievalArtificial IntelligenceMachine Learning
The gist
Searching large collections of documents can be slow and costly when storing detailed information about each document. The authors introduce a new method called Matryoshka Hash Representations that stores document data in compact binary codes which can be searched quickly without losing much accuracy. This method first learns a longer binary code, then creates shorter subsections of that code that can be searched efficiently. Their approach works better than current techniques, especially when very little storage space is available.
Open 2609.07276v1