Papers for

automated code reviewers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Reasoning models improve answer confidence with smooth calibration

Beyond Verbalized Confidence: Calibrating Reasoners with Differentiable Readouts

Abstract: Reinforcement learning with verifiable rewards (RLVR) trains reasoning models to produce correct answers, but does not ensure that their stated confidence is calibrated. The resulting models are systematically overconfident. Recent methods train calibration inside the RLVR loop by having the model state a numerical confidence alongside its answer, but they all obtain the confidence by sampling it as text. This choice imposes two costs: a sampled confidence introduces variance and in practice collapses to a handful of distinct values, and sampling makes the confidence non-differentiable, forcing the calibration loss through a scalar reward. We propose CREDO (Confidence REaDOut) to replace sampling with a deterministic readout. While RLVR optimizes correctness, CREDO reads the confidence from a dedicated token pair in the model's output distribution and trains it by differentiable regression. CREDO further turns the trained confidence into a signal for accuracy, weighting rollouts by how far confidence and outcome disagree, so that accuracy and calibration improve together. Across mathematical and code reasoning, CREDO attains the best accuracy and calibration, and the gains extend to abstention and selective prediction.

Mon 28 SeptMachine LearningArtificial IntelligenceComputation and Language
The gist
Many AI systems that solve problems say how sure they are about their answers, but they often claim to be more confident than they should be. The authors show that current methods which guess confidence as text make the system's certainty jump around and hard to train properly. They created a new method called CREDO that picks confidence values directly from the AI's internal signals in a smooth way, leading to better and more reliable confidence scores. This helps the system know when it might be wrong and when to skip answering, improving overall performance.
Open → 2609.34857v1