Query adaptive indexing improves retrieval in expert archives

ORDER: Task-Conditioned Routing for Retrieval-Augmented Generation

Artificial Intelligence

Summary

Many systems that find and generate answers from large collections use the same setup for all questions, even though different questions need different ways of organizing and searching the data. The authors propose ORDER, a method that changes how the data is divided and searched depending on the type of question asked. First, it groups similar questions and learns the best way to split and filter the underlying information for each group. Then, when a new question comes in, it routes it to the right way of searching based on its similarity to the groups. This approach helps get better answers from complex, historical archives with diverse information.

What this means in practice

  • For digital archive managers: Adapt indexing and search strategies dynamically based on query type to improve retrieval quality from diverse historical collections.
  • For enterprise search teams: Deploy query-conditioned retrieval pipelines to better handle heterogeneous internal document searches with varied metadata and content structures.

Authors

Aurélien Pellet, Julien Perez, Marie Puren

Abstract

Retrieval-Augmented Generation (RAG) pipelines typically rely on a fixed indexing and retrieval configuration determined at preprocessing time. This one-size-fits-all design is ill-suited to domain-expert settings, where heterogeneous queries require different chunking granularities, metadata constraints, and source-selection strategies. As a result, configurations that are effective for one family of queries often perform poorly for others. In this paper, we introduce ORDER (Optimal Routing for Dynamic Evidence Retrieval), a query-conditioned RAG framework that jointly adapts indexing and retrieval to the incoming query. Our approach first discovers semantic clusters over a given set of questions associated to a corpus and learns, for each cluster, a chunking strategy together with a suited metadata filtering and reranking configuration. At inference time, queries are routed to the appropriate pre-built index through nearest-centroid assignment. To further improve retrieval, we propose a supervised query router (QRe) that predicts which collections are most likely to contain relevant evidence, coupled with a Uniform Multi-source Sampler (UMS) that allocates the retrieval budget evenly across the selected sources. We evaluate our framework on large-scale, heterogeneous historical archives and show that conditioning both indexing and retrieval on the query consistently outperforms both naive baselines and strong state-of-the-art RAG systems in complex expert-domain environments.