Papers for

digital library engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

ConRAG improves retrieval of multi-step entity connections with less computing cost

ConRAG: Lightweight inference of multi-hop relations

Abstract: Understanding how two entities are connected often requires tracing multi-hop relations across documents to identify intermediate entities and supporting evidence that explain a connection. This is a task that appears frequently in scientific research and other knowledge-intensive analyses. We formalise this setting as multi-hop relation inference: given two known endpoint entities, we aim to recover the bridge entities and evidence-grounded reasoning chains that connect them across a document corpus, and to generate an explanation grounded in the retrieved evidence. Existing multi-hop RAG systems typically seek an unknown answer entity rather than explicitly recovering the connection between two known endpoints and graph-based approaches often rely on costly LLM-extracted knowledge graphs that limit scalability to large document collections. We introduce ConRAG, which builds a lightweight entity-document graph from entity co-occurrence and LLM-based entity filtering. Its connective retrieval infers and semantically ranks paths between two endpoints. On MuSiQue and 2WikiMultiHopQA, ConRAG consistently improves bridge entity and reasoning chain recovery over strong RAG baselines, while reducing graph-indexing token cost by up to roughly 1.5 orders of magnitude. Our results show that endpoint-constrained path retrieval provides an effective and index-efficient approach to evidence-grounded relation discovery.

Mon 28 SeptMachine Learning
The gist
Sometimes to understand how two things are related, you need to find other things in between that connect them. The authors call this multi-hop relation inference. They created ConRAG, a method that makes it easier and faster to find these connecting steps and explain the link between two known things using lots of documents. ConRAG uses a simple graph showing where entities appear together and filters them smartly to find the best path. It works better and uses much less computing power than older methods on two test sets.
Open → 2609.35193v1

Visual document retrieval improves by picking tokens after query arrives

Query-Aware Token Budgeting for Efficient Late-Interaction Visual Document Retrieval

Abstract: Late-interaction visual document retrievers preserve fine-grained page evidence by storing many token embeddings per page, but the resulting storage and query-time interaction costs make large-scale deployment expensive. Pooling document tokens before indexing offers a natural remedy, yet static pooling must decide which visual evidence to preserve before the query is known. We study an alternative: a heavily compressed hot-path index generates candidates, after which query-aware token budgeting operates on the original token sets of the shortlisted pages. We formulate this stage-two selection as a budgeted MaxSim coverage problem, show that a clipped version is monotone submodular, and compare coverage-only, cluster-guided, token-wise, and marginal-gain policies. On ten ViDoRe tasks with ColModernVBERT, direct static pooling reduces macro normalized discounted cumulative gain at rank five from 0.6309 without compression to 0.4738 at a thirty-two-fold pool factor. Under the same candidate-generation regime and a pool-factor-eight-equivalent reranking budget, token top-k recovers 93.93 percent of the full-token score, while greedy marginal-gain selection recovers 98.39 percent. Held-out and leave-one-dataset-out evaluations yield positive greedy improvements over token top-k on every dataset. The latency analysis reveals two useful operating points: token top-k for interactive retrieval and the naive greedy implementation as a quality upper envelope. Together, these results show that late-interaction visual retrieval benefits from query-aware allocation rather than query-agnostic pooling alone.

Mon 7 SeptInformation RetrievalArtificial Intelligence
The gist
Finding relevant images of documents quickly is hard because pages store lots of detailed information, which slows things down and uses lots of storage space. The authors study a two-step method: first, they use a small index to pick some promising pages, then they carefully choose which pieces of those pages to look at based on the actual search query. This method chooses tokens (small units of information) smartly, which keeps the search accurate but much faster. Their tests show this adaptive approach works better than just compressing data without considering the query.
Open → 2609.07262v1