Reasoning models improve answer confidence with smooth calibration
Beyond Verbalized Confidence: Calibrating Reasoners with Differentiable Readouts
Machine LearningArtificial IntelligenceComputation and Language
Summary
Many AI systems that solve problems say how sure they are about their answers, but they often claim to be more confident than they should be. The authors show that current methods which guess confidence as text make the system's certainty jump around and hard to train properly. They created a new method called CREDO that picks confidence values directly from the AI's internal signals in a smooth way, leading to better and more reliable confidence scores. This helps the system know when it might be wrong and when to skip answering, improving overall performance.
What this means in practice
- •For software engineers: Build AI assistants that provide better confidence estimates to decide when to ask for human help or skip uncertain answers.
- •For automated code reviewers: Improve tools that suggest code fixes by weighting suggestions with calibrated confidence for safer code changes.
Authors
Chenxiao Fan, Chongming Gao, Gangyi Zhang, Leyang Shen, Yaxin Gong, Jiamin Wang, Jiakai Wang, Dong Wang, Yang Liu, Fuli Feng, Xiangnan He
Abstract
Reinforcement learning with verifiable rewards (RLVR) trains reasoning models to produce correct answers, but does not ensure that their stated confidence is calibrated. The resulting models are systematically overconfident. Recent methods train calibration inside the RLVR loop by having the model state a numerical confidence alongside its answer, but they all obtain the confidence by sampling it as text. This choice imposes two costs: a sampled confidence introduces variance and in practice collapses to a handful of distinct values, and sampling makes the confidence non-differentiable, forcing the calibration loss through a scalar reward. We propose CREDO (Confidence REaDOut) to replace sampling with a deterministic readout. While RLVR optimizes correctness, CREDO reads the confidence from a dedicated token pair in the model's output distribution and trains it by differentiable regression. CREDO further turns the trained confidence into a signal for accuracy, weighting rollouts by how far confidence and outcome disagree, so that accuracy and calibration improve together. Across mathematical and code reasoning, CREDO attains the best accuracy and calibration, and the gains extend to abstention and selective prediction.