Large reasoning models get better uncertainty estimates without internals
Jailbreaks for Black-Box Uncertainty Quantification in Large Reasoning Models
Artificial IntelligenceComputation and LanguageMachine Learning
Summary
Large reasoning models (LRMs) can be too confident in their answers, which is risky if we can't see their internal scores. The authors find that usual ways to guess how uncertain these models are don't work well because the models give less varied answers after fine-tuning. They introduce a new way to 'jailbreak' the model's answers to get more realistic uncertainty estimates without needing internal access. Their method, called J4U, shows better uncertainty predictions across several datasets and models.
What this means in practice
- •For ai deployment engineers: Improve trustworthiness by enhancing uncertainty estimation for large language models without accessing internal model scores.
- •For automated customer service teams: Better detect when AI-generated answers might be unreliable by using improved uncertainty measures in black-box models.
Authors
Lucas Biechy, Cédric Eichler, Adrien Boiret, Nicolas Anciaux
Abstract
While Large Reasoning Models (LRMs) excel at complex reasoning, alignment through reinforcement learning often induces systemic overconfidence. In production environments, where logits may be unavailable, robust black-box uncertainty quantification (UQ) is essential for trustworthiness and safety. Focusing on question-answering for LRMs, we show that existing black-box methods, such as paraphrase-based self-consistency and confidence verbalization, offer little to no improvement over simple repeated sampling, suggesting that alignment suppresses useful output variability. We introduce prompt-level relaxation operators that broaden the model's effective output distribution by approximating the effect of an optimal policy obtained with a stronger KL-regularization parameter, hence closer to the reference model. Theoretically, we demonstrate that relaxation improves calibration. We propose Jailbreak for Uncertainty (J4U), a jailbreak-derived technique for UQ that empirically reproduces the behavioral signatures predicted by our relaxation theory. Across 3 datasets and 4 LRMs, including a closed-source production model, J4U's improvement over repeated sampling achieves statistical significance in up to 6 times more LRM-dataset-metric settings than the strongest black-box UQ state-of-the-art baseline we evaluate, with average ECE reductions up to 5 times larger. These results provide a practical tool for UQ in black-box LRM deployment.