Large language models improve confidence by combining hidden signals and verbal reasoning

When Words Fall Short: Iterative Synergy Between Verbalized Reasoning and Hidden Features for LLM Confidence Estimation

Computation and Language

Summary

Knowing how sure a language model is about its answers is important for trust. The authors found that models contain hidden confidence clues that can be better than just asking the model to say how sure it is. They created a new way that mixes these hidden clues with the model’s own explanations, improving how well confidence is estimated. This approach works well on different data and models, helping to better understand when the model’s answers can be trusted.

What this means in practice

  • For ai developers: Enhance AI systems’ reliability by integrating richer internal confidence signals and verbal reasoning to better judge answer trustworthiness.
  • For software quality assurance teams: Improve automated testing of language-based systems by using more accurate confidence estimates to flag uncertain outputs.

Authors

Yekun Xu, Ante Wang, Jingyi Ren, Xuanyi Chen, Weizhi Ma, Yang Liu

Abstract

Confidence estimation is crucial for developing trustworthy large language models (LLMs), with most methods following estimator-based or verbalization-based paradigms. While recent research increasingly focuses on improving verbalized self-reports of confidence, we challenge the prevailing view that this approach surpasses independent confidence estimators. Our empirical study shows that a dedicated confidence estimator can substantially outperform verbalized confidence, indicating that LLMs' internal representations contain richer confidence signals. Building on this finding, we propose Iterative Policy-Estimator Training (IPoET), a framework that synergizes the complementary strengths of verbalized reasoning traces and informative representations. IPoET alternates policy optimization with estimator updating, integrating estimator-derived confidence feedback into policy learning and refreshing the estimator on new policy rollouts. Experiments across diverse datasets and Qwen and Llama backbones demonstrate that, by iteratively exploiting richer hidden features and adapting to the evolving policy distribution, IPoET consistently outperforms both estimator- and verbalization-based baselines in-domain and achieves superior or comparable results across all out-of-domain metrics. For more details, refer to https://github.com/xyk829/ipoet.