Papers for

automated quality inspectors

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Model uncertainty does not match human ambiguity in vision tasks

Does Model Uncertainty Track Human Ambiguity? Evidence from Multi-Annotator Vision Benchmarks

Abstract: Human-model alignment is critical for trustworthy AI-assisted decision-making systems. Yet, most work evaluates model predictions against single ground-truth labels, overlooking that humans themselves often disagree on labels, a signal of genuine ambiguity. We investigate whether models struggle on the same instances that humans find difficult. We measure this on two vision datasets (FER+ and CIFAR-10H) where multiple human annotations per image capture human disagreement patterns. We evaluate eight pretrained models across three architectures (ResNet, EfficientNet, MobileNetV3) in two parts: first, whether model uncertainty (softmax confidence, entropy) correlates with human disagreement, and second, whether predictive multiplicity measures (inter-model disagreement, Jensen-Shannon divergence) do. We find that it does not: alignment is weak in both dimensions. At the discrete label level, 50.4% of CIFAR-10H images and 33.5% of FER+ images receive multiple valid classifications from humans, while the models converge on only one. These instances represent a critical failure case where humans perceive ambiguity and would request expert review, yet models decide confidently. At the continuous score level, single-model uncertainty correlates weakly with human disagreement ($ρ= 0.24--0.55$), and predictive multiplicity provides only modest improvement. Widely-used uncertainty quantification methods do not reliably identify instances humans find ambiguous. Model uncertainty should not be treated as a trustworthy signal by default for decision-making in high-stakes scenarios.

Mon 28 SeptArtificial Intelligence
The gist
Humans often disagree when labeling images because some pictures are truly unclear or ambiguous. This paper shows that popular AI models, even when uncertain, do not flag the same confusing images as humans do. The researchers tested many AI models on datasets with multiple human labels and found that models tend to be confidently wrong on images where humans see ambiguity. Therefore, relying on AI model uncertainty alone might be risky in situations needing human-like judgment.
Open → 2609.34506v1