Papers for

search system developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Rubric calibration improves accuracy of large language model document rankings

Rubric-Calibrated Preferences: Cross-Query Calibration of LLM Judgments via Item Response Theory

Abstract: Rerankers decide which documents users and LLMs see, yet their standard metric, nDCG, relies on human relevance labels that are costly, sparse, noisy, and discretely graded. As rerankers approach each other in quality, nDCG on these labels therefore increasingly fails to separate them. LLM judges could supply dense labels. Relative judgments within one query tell even close candidates apart, yet their scores share no scale across queries. Absolute grades share one scale but are too coarse to distinguish documents of similar relevance. We propose Rubric-Calibrated Preferences (RCP), which combine both kinds of judgment. A listwise Bradley-Terry tournament orders each query's documents, and a rubric of yes/no criteria of increasing stringency provides an absolute standard. Item Response Theory (IRT), which scores test-takers based on their answers to common questions, then uses the shared criteria to put all queries' tournament scores on one scale. RCP's retrieval metric, RCP-nDCG, replaces nDCG's discrete labels with the resulting calibrated relevance probabilities. Against blind grades from 46 external annotators, calibration raises the correlation between a query's mean score and its mean human grade from 0.538 to 0.795. The probabilities rank a useful document above a non-useful one with probability 0.910 (AUC, chance 0.5), versus 0.651 for the benchmark labels. When the annotators' grades prefer one of two rerankers and exactly one metric agrees, that metric is RCP-nDCG in 72.4% of 185 comparisons (chance about 53%). On TREC-DL, RCP-nDCG sides with NIST assessors' grades on every reranker pair that these grades separate significantly. RCP-nDCG also resolves many of nDCG's ties and separates 1.9 times as many reranker pairs on NanoBEIR. Rubric calibration thus turns relative LLM judgments into dense relevance labels that are comparable across queries and agree with human judgment.

Mon 28 SeptInformation Retrieval
The gist
Deciding which documents show up first in search results is tricky because human labels that judge relevance are limited and inconsistent. The authors created a method called Rubric-Calibrated Preferences that combines detailed comparisons within searches and clear yes/no standards across searches to better rate document relevance. This method uses a testing theory to put all these ratings on the same scale, improving agreement with human judgments and distinguishing closely ranked results more effectively. Their new metric, RCP-nDCG, better separates good documents from less useful ones compared to traditional methods.
Open → 2609.35739v1

Agentic model development improves video discovery retrieval at scale

Verify, Don't Trust: Agentic Model Development for Video Discovery Retrieval at Scale

Abstract: Large language model (LLM) agents can propose, implement, and evaluate model changes. Autoresearch loops demonstrate this capability through minutes-scale iterations on a self-contained program. Online autoresearch instead spans asynchronous systems, hours-long variants, and weeks-long campaigns that can influence a product. A completed run can still support an invalid conclusion when a code change is a no-op, data windows leak, evaluator semantics drift, or the two arms traverse different serving funnels. We present EvoPilot, a human-gated method for long-horizon online autoresearch. Role-specific agents execute each round through a versioned domain skill and typed adapter. Durable records preserve experiments and failures; deterministic checks enforce recorded lessons. We study a 37-day campaign for the retrieval system that powers Video Deep Dive (VDD), an online experience for discovering follow-on videos after a user opens a seed video. The campaign covered seven directions and used an hourly refreshed index of hundreds of millions of videos. Earlier manual experiments had not established a benefit from an interaction head. A primitive autoresearch attempt revisited the direction but incorrectly attributed an offline hit-rate decline of 22 percentage points to the head. We then introduced EvoPilot. Its human-gated verification traced the drop to a pre-existing evaluation defect that produced output depths of 3,000 and 600. After repair, a matched comparison measured an offline improvement of 3.20 percentage points. Post-study replay and mutation tests rejected invalid comparisons while admitting valid counterparts. Durable state recovered an interrupted round, and artifact reuse avoided approximately five GPU-hours. Separately, a seven-day randomized online evaluation estimated a 0.66% relative increase in the VDD slice of Good Search Result Rate for Retention (GSRR).

Fri 18 SeptInformation RetrievalArtificial IntelligenceMachine Learning
The gist
It can be hard to trust changes made to complicated video search systems because mistakes or hidden problems can lead to wrong conclusions about what works. The authors created EvoPilot, a method where specialized automated agents propose and test improvements while humans verify results to avoid errors. This was tested over a month on a large video retrieval system and found real improvements that earlier automated trials missed due to evaluation bugs. Their approach keeps detailed records and reproducible checks, helping systems recover from interruptions and avoid wasteful compute.
Open → 2609.21257v1

Optimal chunk order reduces computation in large language model serving

Prefix Sharing Is a Sorting Problem

Abstract: LLM serving reuses KV cache by exact prefix match, so when a prompt is assembled from a set of reusable pieces -- retrieved passages, tool definitions, few-shot exemplars -- the order chosen for those pieces determines how much computation can be shared. Every deployed system fixes that order by a single global convention. We prove this is optimal only when requests contain at most two pieces, and asymptotically wrong in general. Our main result is a structure theorem: the minimum prefix-trie cost equals min_H sum_x w(x) t_x(H) over binary hierarchies H on the requests, where t_x(H) is the canonical decomposition size of the set of requests needing chunk x. Choosing chunk orders is therefore equivalent to choosing one hierarchy over requests. The identity yields an O(3^m) exact algorithm, identifies the two-chunk case as minimum vertex cover, and shows that on the leave-one-out family the optimum is the minimum external path length of a binary tree -- the merge-sort recursion -- so a global order pays Theta(n^2) against a true cost of Theta(n log n). Agglomerative clustering by common intersection is a tight 1/2-approximation for the achievable saving. On BM25 retrieval traces over three BEIR corpora the resulting layout reduces prefill by 17-36% against production RAG ordering, and the margin widens with retrieval depth as the theory predicts. Serving requests in the hierarchy's DFS order finally lets a cache holding one request's context attain the unbounded-cache optimum exactly, so cache capacity and reorder window act as substitutes.

Sat 12 SeptData Structures and AlgorithmsComputation and Language
The gist
Large language model systems reuse previous computations when a new prompt starts with the same pieces as before. How these pieces are ordered affects how much work can be saved. The authors show that simply using one fixed order isn’t best except in very simple cases, and they map the problem to sorting and tree structures. Their approach can reduce unnecessary computation by reorganizing the pieces intelligently, leading to more efficient responses in practice.
Open → 2609.13692v1