Constrained Entity Selection under Partial Knowledge for LLM-Based Knowledge Graph QA
2026-08-25 • Artificial Intelligence
Artificial Intelligence
AI summaryⓘ
The authors looked at how large language models (LLMs) can answer questions using knowledge graphs but sometimes give wrong answers. Instead of fully translating questions into complex queries, they propose a method where the LLM suggests possible answers and then checks those answers using simple rules derived from the question. Their method accounts for incomplete knowledge graphs by marking constraints as satisfied, violated, or unknown, which helps avoid wrongly discarding good answers. They tested this on a biomedical knowledge graph and found their approach improved accuracy without losing correct answers.
large language modelsknowledge graphquestion answeringsemantic parsingSPARQLsymbolic constraintsopen-world assumptionHetionetconstraint satisfactioncandidate filtering
Authors
Emanuel Kitzelmann
Abstract
Large language models are increasingly used for knowledge graph question answering (KGQA), but can fail to correctly ground answers in the underlying graph. Current approaches to LLM-based KGQA either rely on full semantic parsing into executable queries such as SPARQL, which is brittle in practice due to complex schemas or incompleteness of real-world KGs, or on LLM-reasoning and answer generation over KGs, which can be more robust but lacks formal guarantees. In this work, we study a complementary setting in which \emph{candidate} answers are generated by an LLM-based system and subsequently verified using lightweight symbolic constraints derived from the question. We introduce \emph{Constrained Entity Selection under Partial Knowledge (CES-PK)}, a problem formulation that focuses on eliminating invalid answers and providing symbolic support for valid ones without requiring construction of executable logical forms. To account for incomplete KGs, we employ a three-valued constraint semantics (\emph{satisfied, violated, unknown}) that avoids incorrect rejections under open-world assumptions. To demonstrate the effects of our method, we instantiate this framework over the Hetionet biomedical knowledge graph and evaluate the impact of type, relation, and exclusion constraints. Experiments show that precision improves by filtering invalid candidates, while recall is preserved due to retaining candidates whose constraints are not explicitly violated. Satisfied constraints provide additional positive symbolic evidence to rank remaining candidates.