Confidence measures of large language models struggle to spot their own errors
On the Limits of Metacognitive Monitoring in LLMs
Artificial Intelligence
Summary
Deciding if an answer is wrong is important for making good choices. The authors checked how well large language models can judge when their own answers might be mistakes. They found that even when these models solve hard problems very accurately, their confidence scores often can't clearly tell a right answer from a wrong one. Trying to review answers or compare with another model helps a little but doesn’t solve the problem for hard questions. So, models are still not very good at knowing when they are wrong.
What this means in practice
- •For ai safety teams: Improve AI risk assessment by understanding limits of confidence-based error detection in language models solving complex problems.
- •For automated tutoring developers: Design better feedback systems by accounting for unreliable confidence signals from language models when validating student answers.
Authors
Dongqi Han, Yifan Yang, Dongsheng Li
Abstract
Reliable decisions depend on recognizing when an answer may be wrong. In biological cognition, metacognitive monitoring can dissociate from task performance, raising the question of how closely solving and judging are linked in language models. Here we study the confidence reports of four frontier models across 15 benchmarks. High task accuracy can coexist with weak error discrimination: a model solves 97% of competition mathematics problems while its answer-time confidence ranks correct answers above errors barely better than chance. Confidence separates correct answers from errors more effectively on questions solved by a separate reference model, while review brings limited improvement on reference-hard questions. Aggregate discrimination also rewards ranking correct answers on easy questions above errors on hard ones, which question-only forecasts already do well. Cross-evaluation helps most where the evaluator answered correctly, and errors shared by the two models usually retain high confidence. Hard questions and shared errors remain difficult targets for prompted self-review and peer oversight, even in models with strong problem-solving performance.