Learning from uncertain labels improves prediction uncertainty estimates

Epistemic Learning from Imprecise Annotation

Machine Learning

Summary

When labels in data are unclear or imprecise, traditional methods often guess one exact answer, which can hide uncertainty about what the true label might be. The authors propose a new way to teach machines that considers all plausible labels as a set, allowing the model to express uncertainty more honestly. They develop a method called POCC that trains two versions of a classifier to cover the worst- and best-case label possibilities, giving a useful score of uncertainty. This helps the model make more reliable predictions, especially when the training data labels are not fully precise.

What this means in practice

Authors

Kaizheng Wang, Siu Lun Chau

Abstract

Imprecise annotations may support several plausible labelling distributions, yet learning methods often resolve this ambiguity into a single predictive distribution. This can obscure what the annotation evidence leaves unresolved. We introduce epistemic learning from credal supervision, a framework that uses convex sets of plausible labelling distributions, called credal sets, as supervision and learns sets of predictive distributions. We instantiate the framework with the pessimistic--optimistic credal classifier (POCC), which combines a shared backbone with two classification heads trained to minimise worst-case and best-case losses over the supervision sets. Their outputs define a predictive credal set whose spread provides an uncertainty score. We also show how credal labels can be obtained through a simple relaxation of existing probabilistic labels, reducing commitment to their precise probability assignments. This construction admits closed-form inner optimisation under cross-entropy loss, enabling efficient training. Assuming the supervision sets contain the true conditional label distributions, and other regularity assumptions, we establish a finite-sample generalisation bound for the averaged predictor with an explicit penalty for supervision imprecision. We evaluate POCC using human annotator disagreement and teacher predictions, alongside label smoothing as a controlled proxy for annotation imprecision. Across these settings, POCC achieves a favourable balance of predictive accuracy, calibration, and uncertainty-based selective classification versus competitive baselines.