Detecting early false confidence in large language models reasoning

When Confidence Rises Too Early: Detecting Shortcut Reasoning via Premature Answer Commitment

Computation and Language

Summary

Large language models sometimes jump to answers too quickly by using shortcuts rather than careful thinking. This can make their explanations seem plausible but actually be unfaithful. The authors introduce a new way to watch how confident the model is at each step of its thinking. They find that models using shortcuts become overly confident very early. Their method, called Distributional Answer Commitment Score (DACS), helps spot this early overconfidence without needing the right answer in advance. This improves detecting shortcut reasoning in math and code problems and helps preference systems avoid it.

What this means in practice

Authors

Zhaohan Zhang, Junjie Liu, Chengzhengxu Li, Chen Shen, Xiaoming Liu, Chao Shen, Jieping Ye, Ziquan Liu, Ioannis Patras

Abstract

The reasoning trajectory of a Large Language Model (LLM) is often treated as a verbalized description of its internal reasoning. However, such trajectories can be unfaithful: a model may rely on shortcuts to reach an answer and then post-rationalize the decision with a seemingly coherent chain of thought. Detecting this shortcut reasoning is challenging because existing monitors and verifiers mainly inspect textual traces or final outcomes, rather than how the model's belief in its answer develops during generation. We introduce ConfLens, a framework that tracks the evolution of confidence in the final answer throughout reasoning. Across three shortcut reasoning settings, we observe a common pattern of premature confidence, where shortcut samples become highly confident in the final answer at early reasoning stages. Existing confidence estimation methods, however, show limited generalizability, reliability, or efficiency for detecting this behavior. We therefore propose the Distributional Answer Commitment Score (DACS), a distributional confidence estimator that measures the entropy of the model's probability distribution over answer commitment at each reasoning step. DACS captures how concentrated the model's answer belief is without requiring ground-truth answers or task-specific verifiers. We further convert ConfLens detection results into interpretable signals for reward models to reduce their preference for shortcut reasoning. Experiments on mathematical and code reasoning tasks show that ConfLens with DACS improves shortcut reasoning detection by over 4.3% F1 compared with strong baselines and reduces the mismatch between faithfulness and correctness in reward model preferences.