Scientific knowledge graphs improve distributed vector search efficiency
COMPASS: Steering Distributed Vector Search with Scientific Knowledge Graphs
DatabasesDistributed, Parallel, and Cluster Computing
Summary
Searching large collections of data with vector databases can be slow because data is split randomly across many parts, causing queries to check everything. The authors show that using scientific knowledge graphs, which capture relationships between facts, can better group data and guide queries to fewer relevant parts. Their approach, COMPASS, checks less data while still finding the important connections between pieces of information. This leads to much faster searching without losing important results.
What this means in practice
- •For biomedical data engineers: Improve speed and accuracy of searching large biomedical knowledge bases by grouping data with scientific relationships for targeted queries.
- •For distributed database developers: Build vector search systems that use knowledge graph structure to partition and query data efficiently, reducing latency and load.
Authors
Song Young Oh, Amal Gueroudji, Seth Ockerman, Rob Latham, Orcun Yildiz, Ian Foster, Kyle Chard, Robert Ross
Abstract
Vector databases use hashing to partition data across "shards," logical units for distributed execution. This placement, however, destroys semantic locality, forcing each query into scatter-gather limited by the slowest shard. Vector-space clustering can help, but scientific evidence is often connected by factual relations that do not align with embedding distance. We present COMPASS, a framework that uses a knowledge graph (KG) to determine data placement and query-time shard selection. COMPASS detects communities, splits oversized communities, inserts embeddings by subject entity, and routes queries to a small set of shards. Across four biomedical KGs, our method searches only 13-18% of the corpus while preserving broadcast recall and recovering up to 2.6x more multi-hop evidence than an embedding-based baseline. On 15 HPC nodes, COMPASS sustains 7.9x higher throughput with lower tail latency than hash-based broadcast. These results show that KG structure provides a compact complement to embedding geometry for scalable vector search.