Papers for

data validation teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Jev system reveals hidden uncertainty improves probability estimates

Jev thinks "I don't know'', but doesn't say it: Introducing Sys1Cal-v1 Dataset for Probability Calibration

Abstract: The appearance of Jev marked the era of System One Models, foundation models that return structured decisions with probability distributions rather than text. Aside from low cost and great speed, Jev's central promise is that these probabilities are calibrated: such claim is not backed by any public test and available external benchmarks evaluate confidence calibration, not whether every returned option probability has the right numerical meaning. To tackle this issue, we introduce Sys1Cal-v1, a dataset of True/False questions about a proposition $A$ for which the exact probability $P(A)$ is known by construction. Each item is queried through the three Jev primitives - Noul, Choice and Score - and evaluated by total variation distance from the ground-truth distribution, which can be used to estimate a soft accuracy of System One Models. We showcase the utility of Sys1Cal-v1 as a benchmark dataset by evaluating Jev and SemIf, an open-source Choice-style baseline. In this work, however, we focus even more deeply on Jev, by studying the calibration of its Score and Choice answers. In particular, we discover a peculiar behaviour that can be explained by assuming that Jev suppresses a third truth value, going beyond True and False. In other words, in \texttt{Choice} answers, $P(A)$ and $P(\neg A)$ are presented as if $P(A)+P(\neg A)=1$, while a term $P(U)\neq0$ is missing in the sum. Recovering $P(U)$ leads to an improvement of median soft accuracy in \texttt{Choice} answers from $0.771$ to $0.978$, suggesting that, even in binary decisions, Jev wants to answer with a third option:``I don't know''.

Mon 28 SeptArtificial IntelligenceLogic in Computer Science
The gist
It can be hard for AI systems to give true probabilities for whether something is true or false. The authors created a special set of questions where they know the exact chance an answer is true. They tested Jev, a system that gives probabilities, and found it acts like there’s a hidden 'I don’t know' option besides true or false. By accounting for this hidden option, Jev’s probability estimates got much more accurate. This shows that some AI models might be quietly unsure but don’t openly say so.
Open → 2609.35342v1

Typed decision models fail to align probabilities with probability rules

Beyond Calibration: Do a Typed-Decision Model's Probabilities Obey the Probability Axioms?

Abstract: Typed-decision models such as TypeSafe's Jev answer a declared yes/no or multiple-choice question about a state with a probability instead of text, and their evaluations report accuracy and calibration. Neither requires that the probabilities a model gives to logically related questions fit together item by item, which is what a system that acts on those probabilities needs. We test this property, coherence, with a battery of logically linked questions that needs no labels. For 160 items from ChaosNLI and PubMedQA, each with three mutually exclusive labels, we ask whether the label is X, whether it is not X, whether it is one of the other two labels, and which label applies. On 480 negation pairs, Jev's probabilities for "the label is X" and "the label is not X" miss summing to one by 0.064 on average (95% CI 0.055 to 0.072). Qwen3.8-27B, run from its official BF16 weights, misses by 0.293 with first-token probabilities and by 0.122 with verbalized probabilities. The gap to the first-token readout persists on pairs where both systems give similar probabilities, without double-negation labels, and after averaging Jev's repeated calls. Jev is not coherent either: its violations are about five times its repeat noise, and it over-endorses statements about single labels, so that its three single-label probabilities sum to 1.14 on average. The two systems also fail differently. Qwen3.8-27B's first-token readout under-endorses the complement of a label whether or not the question contains "not", rejecting both a statement and its negation in 196 of 480 pairs, and it does not become more coherent where it is more confident, whereas Jev's violations concentrate where its answer is uncertain. Because the checks need no labels, they expose biases that appear only when question forms are compared, and inconsistencies within items.

Sun 27 SeptComputation and LanguageArtificial IntelligenceMachine Learning
The gist
Probability-based AI models are tested to see if their answers about yes/no or multiple-choice questions add up logically, as they mathematically should. The authors find that models like TypeSafe's Jev and Qwen3.8-27B give probability answers that often don't fit together correctly, such as the chance something is true plus the chance it is not true not adding up to 100%. These inconsistencies happen even when the models seem confident, pointing to hidden biases and logical gaps in how they handle probabilities. The tests used don't need labeled data, which helps reveal subtle problems in probability reasoning within these models.
Open → 2609.33209v1