Calibrating reasoning models improves confidence estimates greatly
Calibration, Not Answer Selection: Distilling Internal Confidence in Reasoning Models
Computation and Language
Summary
Models that answer questions often say how confident they are, but this confidence is usually too high and doesn't match the true chance of being right. The authors found a better way to measure confidence by checking the model’s internal reasoning steps, which gives more accurate confidence levels. They also developed a method to teach the model to trust this better confidence measure, improving its predictions without needing complex new training techniques. This helps models not just pick answers, but also know how sure they should be about them.
What this means in practice
- •For ai application developers: Improve the reliability of AI answers by using internal confidence signals to better weigh responses on factual question answering tasks.
- •For automated customer support teams: Increase trust in AI-generated answers by providing well-calibrated confidence scores that reflect true answer correctness.
Authors
Yadong Xi, Rongsheng Zhang, Tangjie Lv, Ziyang Luo, Ruochen Zhao
Abstract
Reinforcement learning with binary correctness rewards trains correctness, not calibrated confidence. The confidence that reasoning models verbalize is systematically overconfident, and the problem is not merely one of scale: verbalized confidence tracks how willing a model is to commit to an answer, not how likely the answer is to be right. Post-hoc rescaling therefore fits one distribution but rarely transfers. We look inside the model instead. On factual question answering, a linear probe on the hidden state between the chain of thought and the answer is substantially better calibrated: its expected calibration error is 5 to 38 times lower than that of the verbalized score across four benchmarks and two model families. However, when used to pick among N sampled answers, that same probe nearly ties majority voting yet falls far short of the oracle. Internal states answer "how certain am I" well and "which answer is right" poorly, so the signal should be reported as a confidence rather than used to select answers. As a result, we introduce probe-guided self-distillation (Probe-SD): score a model's own sampled traces with the probe, overwrite the confidence each trace states, and finetune the base checkpoint of the same family, so nothing but the model itself remains at test time. On Qwen3-14B, Probe-SD cuts ECE from 0.178 to 0.024 in-domain and from 0.542 to 0.113 out-of-domain, where it also beats post-hoc recalibration and self-consistency distillation. The resulting confidence is well-calibrated and useful for weighted voting, behaviors previously attributed to online RL, here obtained with supervised finetuning alone.