Federated knowledge graphs enable multi-hop question answering without sharing raw data
FedV-KGQA in Practice: Design Lessons and an Interactive Prototype
Artificial IntelligenceComputation and LanguageInformation Retrieval
Summary
Question answering systems that use knowledge graphs usually need access to all the information in one place. The authors study what happens when different organizations hold parts of the knowledge graph with unique relationships but share common entity identifiers. They show that combining knowledge locally enriched and embedded in separate silos can nearly match the accuracy of a system that sees everything. They also find that focusing on linking questions to key entities and enriching the data matters more than using complex embedding models. The authors provide lessons from comparing different settings and an interactive prototype that demonstrates the approach.
What this means in practice
- •For enterprise data engineers: Combine information from different business units without moving sensitive relation data to answer complex queries involving multiple departments.
- •For customer support platform developers: Integrate fragmented knowledge bases from partner companies to provide better multi-step answers to customer questions without exposing raw data.
Authors
Md Saikat Islam Khan Bappy, Oshani Seneviratne
Abstract
Knowledge graph question answering usually assumes that one system can reach the whole graph. In practice, facts are often held by organizations that share entity identifiers but own disjoint relation types, so no single party sees a complete reasoning chain. This poster presents the empirical findings of FedV-KGQA on multi-hop question answering over such vertically partitioned graphs. Each silo enriches its local graph and trains a knowledge graph embedding on its own triples. A server then concatenates the silo-specific entity views, anchors the projected question at the topic entity, and ranks candidates by similarity. Raw triples and relation embeddings never leave a silo. Comparing the FedV-KGQA experiments with one another yields three results. First, federated fusion recovers most of the centralized accuracy, while a single silo recovers little. Second, anchoring and enrichment matter more than the choice of embedding model. Third, the cheapest encoder depends on the target accuracy rather than on parameter count. This poster paper contributes that cross-experiment comparison, four design lessons drawn from it, and an interactive prototype that runs real inference and traces the full pipeline, per question, on released checkpoints.