Typed decision models fail to align probabilities with probability rules
Beyond Calibration: Do a Typed-Decision Model's Probabilities Obey the Probability Axioms?
Computation and LanguageArtificial IntelligenceMachine Learning
Summary
Probability-based AI models are tested to see if their answers about yes/no or multiple-choice questions add up logically, as they mathematically should. The authors find that models like TypeSafe's Jev and Qwen3.8-27B give probability answers that often don't fit together correctly, such as the chance something is true plus the chance it is not true not adding up to 100%. These inconsistencies happen even when the models seem confident, pointing to hidden biases and logical gaps in how they handle probabilities. The tests used don't need labeled data, which helps reveal subtle problems in probability reasoning within these models.
What this means in practice
- •For ai developers: Improve AI systems by checking their probability outputs for logical consistency to enhance decision-making reliability.
- •For data validation teams: Use label-free coherence tests to detect and diagnose hidden biases and inconsistencies in AI model probabilities.
Authors
Keyi Li, Yihao He, Quanyi Li
Abstract
Typed-decision models such as TypeSafe's Jev answer a declared yes/no or multiple-choice question about a state with a probability instead of text, and their evaluations report accuracy and calibration. Neither requires that the probabilities a model gives to logically related questions fit together item by item, which is what a system that acts on those probabilities needs. We test this property, coherence, with a battery of logically linked questions that needs no labels. For 160 items from ChaosNLI and PubMedQA, each with three mutually exclusive labels, we ask whether the label is X, whether it is not X, whether it is one of the other two labels, and which label applies. On 480 negation pairs, Jev's probabilities for "the label is X" and "the label is not X" miss summing to one by 0.064 on average (95% CI 0.055 to 0.072). Qwen3.8-27B, run from its official BF16 weights, misses by 0.293 with first-token probabilities and by 0.122 with verbalized probabilities. The gap to the first-token readout persists on pairs where both systems give similar probabilities, without double-negation labels, and after averaging Jev's repeated calls. Jev is not coherent either: its violations are about five times its repeat noise, and it over-endorses statements about single labels, so that its three single-label probabilities sum to 1.14 on average. The two systems also fail differently. Qwen3.8-27B's first-token readout under-endorses the complement of a label whether or not the question contains "not", rejecting both a statement and its negation in 196 of 480 pairs, and it does not become more coherent where it is more confident, whereas Jev's violations concentrate where its answer is uncertain. Because the checks need no labels, they expose biases that appear only when question forms are compared, and inconsistencies within items.