ConRAG improves retrieval of multi-step entity connections with less computing cost

ConRAG: Lightweight inference of multi-hop relations

Machine Learning

Summary

Sometimes to understand how two things are related, you need to find other things in between that connect them. The authors call this multi-hop relation inference. They created ConRAG, a method that makes it easier and faster to find these connecting steps and explain the link between two known things using lots of documents. ConRAG uses a simple graph showing where entities appear together and filters them smartly to find the best path. It works better and uses much less computing power than older methods on two test sets.

What this means in practice

  • For knowledge management teams: Quickly discover connections and explanations linking concepts across large document collections for research or analysis.
  • For digital library engineers: Build scalable search tools that retrieve multi-step relation paths between known topics without high indexing costs.

Authors

Kilian Bänziger, Sonia Laguna, Markus Kreft, Robert Jakob, Kevin O'Sullivan, Lasse B. Strand, Julia E. Vogt

Abstract

Understanding how two entities are connected often requires tracing multi-hop relations across documents to identify intermediate entities and supporting evidence that explain a connection. This is a task that appears frequently in scientific research and other knowledge-intensive analyses. We formalise this setting as multi-hop relation inference: given two known endpoint entities, we aim to recover the bridge entities and evidence-grounded reasoning chains that connect them across a document corpus, and to generate an explanation grounded in the retrieved evidence. Existing multi-hop RAG systems typically seek an unknown answer entity rather than explicitly recovering the connection between two known endpoints and graph-based approaches often rely on costly LLM-extracted knowledge graphs that limit scalability to large document collections. We introduce ConRAG, which builds a lightweight entity-document graph from entity co-occurrence and LLM-based entity filtering. Its connective retrieval infers and semantically ranks paths between two endpoints. On MuSiQue and 2WikiMultiHopQA, ConRAG consistently improves bridge entity and reasoning chain recovery over strong RAG baselines, while reducing graph-indexing token cost by up to roughly 1.5 orders of magnitude. Our results show that endpoint-constrained path retrieval provides an effective and index-efficient approach to evidence-grounded relation discovery.