Semantically quantify uncertainty with better factual equivalence measures
Semantic Uncertainty Quantification Needs Factual Equivalence
Computation and LanguageArtificial Intelligence
Summary
Figuring out how much uncertainty there is in answers from AI language models often means checking if multiple answers agree. The authors found that the key problem is how these answers are compared, especially whether they really mean the same fact. They created a simple method that trains an AI to better focus on whether answers truly match in fact, not just in wording. This new method consistently improves uncertainty measurement across many tests and can even help sharpen how confident we feel about parts of a single answer.
What this means in practice
- •For language model developers: Improve uncertainty estimates in language model outputs by integrating a fact-focused comparison operator to better measure answer equivalence.
- •For vision-language model engineers: Enhance uncertainty detection in vision-language models by using a simple contrastive encoder trained to capture factual equivalence in generated answers.
Authors
Joseph Hoche, Quentin Guimard, Gianni Franchi
Abstract
Semantic uncertainty quantification for large language models rests on a common template: sample several answers, measure how much they agree, and treat disagreement as uncertainty. We first formalize this template as two separate roles: an operator that compares two answers, and an aggregator that combines all pairwise comparisons into a scalar. Existing methods differ almost entirely in how they aggregate, while taking the operator off the shelf, typically an NLI model or a generic sentence encoder. We show that this reliance on off-the-shelf operators is the primary bottleneck of semantic UQ: they do not accurately measure factual equivalence of multiple answers to the same question. We resolve this with a deliberately simple recipe: a single encoder trained contrastively to isolate the targeted fact, utilizing synthetic data generated by an LLM and dataset both disjoint from all evaluation settings. Integrating the resulting operator into existing methods improves performance on 120 of 126 evaluation settings (95%) spanning 18 model dataset combinations across language and vision-language models. The best variant reaches 0.76 mean AUROC against 0.68 for the strongest baseline, while replacing the quadratic cross-encoder comparisons of entailment-based operators with one encoder pass per answer. The uniformity of the improvement supports the view that the operator, not the aggregator, is the limiting factor. The same operator also improves single generation token-level estimators: the norm it assigns to each token measures how much that token bears on the answer, and reweighting token log-likelihoods accordingly sharpens the estimate.