Jev system reveals hidden uncertainty improves probability estimates
Jev thinks "I don't know'', but doesn't say it: Introducing Sys1Cal-v1 Dataset for Probability Calibration
Artificial IntelligenceLogic in Computer Science
Summary
It can be hard for AI systems to give true probabilities for whether something is true or false. The authors created a special set of questions where they know the exact chance an answer is true. They tested Jev, a system that gives probabilities, and found it acts like there’s a hidden 'I don’t know' option besides true or false. By accounting for this hidden option, Jev’s probability estimates got much more accurate. This shows that some AI models might be quietly unsure but don’t openly say so.
What this means in practice
- •For ai system developers: Improve probability outputs in AI systems by modeling an explicit 'I don’t know' option to enhance calibration and accuracy of predictions.
- •For data validation teams: Use the Sys1Cal-v1 dataset to benchmark and validate the probability calibration of binary decision systems.
Authors
Riccardo Porcedda
Abstract
The appearance of Jev marked the era of System One Models, foundation models that return structured decisions with probability distributions rather than text. Aside from low cost and great speed, Jev's central promise is that these probabilities are calibrated: such claim is not backed by any public test and available external benchmarks evaluate confidence calibration, not whether every returned option probability has the right numerical meaning. To tackle this issue, we introduce Sys1Cal-v1, a dataset of True/False questions about a proposition $A$ for which the exact probability $P(A)$ is known by construction. Each item is queried through the three Jev primitives - Noul, Choice and Score - and evaluated by total variation distance from the ground-truth distribution, which can be used to estimate a soft accuracy of System One Models. We showcase the utility of Sys1Cal-v1 as a benchmark dataset by evaluating Jev and SemIf, an open-source Choice-style baseline. In this work, however, we focus even more deeply on Jev, by studying the calibration of its Score and Choice answers. In particular, we discover a peculiar behaviour that can be explained by assuming that Jev suppresses a third truth value, going beyond True and False. In other words, in \texttt{Choice} answers, $P(A)$ and $P(\neg A)$ are presented as if $P(A)+P(\neg A)=1$, while a term $P(U)\neq0$ is missing in the sum. Recovering $P(U)$ leads to an improvement of median soft accuracy in \texttt{Choice} answers from $0.771$ to $0.978$, suggesting that, even in binary decisions, Jev wants to answer with a third option:``I don't know''.