Learning from uncertain labels improves prediction uncertainty estimates
Epistemic Learning from Imprecise Annotation
Machine Learning
Summary
When labels in data are unclear or imprecise, traditional methods often guess one exact answer, which can hide uncertainty about what the true label might be. The authors propose a new way to teach machines that considers all plausible labels as a set, allowing the model to express uncertainty more honestly. They develop a method called POCC that trains two versions of a classifier to cover the worst- and best-case label possibilities, giving a useful score of uncertainty. This helps the model make more reliable predictions, especially when the training data labels are not fully precise.
What this means in practice
- •For machine learning engineers: Design classifiers that provide uncertainty scores reflecting annotation ambiguity for more robust real-world predictions.
- •For automated decision system developers: Improve selective classification by integrating models that better handle uncertain or noisy training labels.
Authors
Kaizheng Wang, Siu Lun Chau
Abstract
Imprecise annotations may support several plausible labelling distributions, yet learning methods often resolve this ambiguity into a single predictive distribution. This can obscure what the annotation evidence leaves unresolved. We introduce epistemic learning from credal supervision, a framework that uses convex sets of plausible labelling distributions, called credal sets, as supervision and learns sets of predictive distributions. We instantiate the framework with the pessimistic--optimistic credal classifier (POCC), which combines a shared backbone with two classification heads trained to minimise worst-case and best-case losses over the supervision sets. Their outputs define a predictive credal set whose spread provides an uncertainty score. We also show how credal labels can be obtained through a simple relaxation of existing probabilistic labels, reducing commitment to their precise probability assignments. This construction admits closed-form inner optimisation under cross-entropy loss, enabling efficient training. Assuming the supervision sets contain the true conditional label distributions, and other regularity assumptions, we establish a finite-sample generalisation bound for the averaged predictor with an explicit penalty for supervision imprecision. We evaluate POCC using human annotator disagreement and teacher predictions, alongside label smoothing as a controlled proxy for annotation imprecision. Across these settings, POCC achieves a favourable balance of predictive accuracy, calibration, and uncertainty-based selective classification versus competitive baselines.