Large language models vary in admitting what they do not know
Do Large Language Models Know What They Don't Know II? A Fully Behavioral, Non-Cognitive Measure of Epistemic Honesty
Artificial Intelligence
Summary
Large language models often give confident answers, but it's unclear if they recognize when they don't know something. The authors created a way to measure how honestly these models admit their knowledge limits, using a new score called the Epistemic Honesty Quotient (EHQ). They tested 15 models on 3,000 questions designed to challenge their knowledge boundaries and found big differences in how models expressed uncertainty or avoided making false claims. This shows that looking just at right or wrong answers doesn’t capture the full picture of how models handle their own knowledge gaps.
Large Language ModelsEpistemic HonestyKnowledge BoundariesUncertainty EstimationBehavioral MetricsModel CalibrationBenchmarkingFalse Information Detection
Authors
Ali Şenol, H. Russell Bernard, Huan Liu
Abstract
Large Language Models (LLMs) are frequently confident, eloquent, and well versed. A natural question arises: do they know what they don't know? To answer this question, we borrow the concept of epistemic honesty and develop a novel metric to systematically evaluate whether an LLM appropriately acknowledges the boundaries of its knowledge. In this work, we introduce the Epistemic Honesty Quotient (EHQ), which reports three observable sub-scores across two operational axes (epistemic restraint and substantive-answer calibration), and construct EHQ-3000, a 3,000-question benchmark spanning Fabricated Entity, Post-Cutoff Event, Hyper-Niche True, and Context-Conditioned Questions. From a frozen registry of 21 model API routes, 15 completed the protocol after endpoint and eligibility checks; 14 entered the confirmatory analysis because severe provider-side truncation made one route's score indeterminate. The study reveals substantial variation across models, including a difference that can not be explained by their capability to extract explicitly available information. Composite EHQ ranges from 0.31 to 0.81 across the analysed panel, despite near-ceiling performance on the document-grounded capability probe. The two restraint criteria overlap strongly under the present category composition, whereas substantive-answer calibration varies across models and does not reliably co-vary with restraint; however, the small panel leaves substantial uncertainty. Thus, EHQ reveals behavioral differences that are not visible to conventional correctness-based assessment, while also showing why dataset composition, provider behavior, and confidence elicitation must remain part of the interpretation.