Knowledge graph method measures large language model context understanding
Do LLMs Understand Context? A Knowledge Graph-Based Evaluation Framework
Artificial IntelligenceMachine Learning
Summary
Large language models can produce impressive answers but it’s unclear if they truly understand the information around them or just guess based on patterns. The authors developed a way to check if these models really get the context by comparing their answers to knowledge graphs, which show facts and their connections. They created a special score to see how well the model's answers match the structure and meaning of these facts. This method also helps spot exactly where models make reasoning mistakes. Their tests on multiple benchmarks showed better results than previous tools.
What this means in practice
- •For qa system developers: Evaluate and improve question answering models by measuring how well they comprehend and reason over contextual knowledge.
- •For natural language processing engineers: Diagnose specific reasoning errors in language models to enhance factual accuracy and context grounding in generated text.
Authors
Subavarshana Arumugam, Mamta Nallaretnam, Kithuni Wickramasinghe, Chamath Gunapala, Pragatheeswaran Vipulanandan, Kamal Premaratne, Uthayasanker Thayasivam
Abstract
While large language models (LLMs) have achieved remarkable linguistic capabilities, a profound question lingers at their core: do these models truly comprehend context or simply excel at pattern matching on an unprecedented scale? Contextual understanding in LLMs refers to the ability to correctly extract relevant information from a given context, integrate it into a coherent internal representation, and reason over it to produce factually consistent and contextually grounded responses. However, traditional methods such as BiLingual Evaluation Understudy (BLEU) and perplexity simply measure surface-level performance. This reveals a critical gap in question answering (QA), where responses must be contextually grounded rather than simply being memorized associations. To fill this void, we propose a novel knowledge graph (KG) based evaluation framework for LLM contextual understanding in QA. Central to this is Semantic Structural Similarity for KGs (S3KG), a hybrid similarity measure combining structural and semantic signals into a single score. In addition, a diagnostic analysis framework is developed to identify and categorize reasoning errors at the triplet level, enabling fine-grained analysis of model failures. Together, across nine benchmarks, S3KG achieves F1 gains of up to $+7.6$ points over the strongest baseline and AUROC up to $0.973$.