Informative Label Missingness in Multiclass Classification Information Geometry and Excess Risk

2026-08-31Machine Learning

Machine Learning
AI summary

The authors studied how missing labels in data, when the pattern of missingness itself holds information, can affect the performance of classifiers that try to sort things into multiple categories. They created a theory that breaks down how much information is lost due to missing labels and how much is gained from the missing-label pattern. Their work shows that sometimes using partial information about missing labels can actually improve classification accuracy in certain ways. They also used mathematical tools and experiments to explore when these effects happen and how strong they are.

multiclass classificationmissing labelslikelihood theoryBayes boundaryFisher informationquadratic discriminant analysisexcess riskmissing completely at randominformation decompositiongeneralized eigenvalue criterion
Authors
Fariborz Setoudehtazang, Geoffrey J. McLachlan
Abstract
Informative label missingness can change the usual efficiency ordering between completely and partially labelled classifiers because the pattern of missing labels may itself carry information about the classification model. We develop a general likelihood-based theory for this phenomenon in parametric multiclass classification. An efficient-information decomposition separates information lost through unavailable class memberships from information contributed by the missing-label mechanism. We then derive a quadratic expansion of plug-in excess risk over the active pairwise faces of the multiclass Bayes boundary, showing that classification efficiency depends on how information gains and losses align with directions that perturb the decision boundary. This yields a classification-weighted generalized-eigenvalue criterion under which informative partial classification may have smaller asymptotic classification risk without globally dominating complete classification in Fisher information. Near missing completely at random, with the marginal missing-label proportion fixed, redistribution of missing labels changes lost class-label information at first order, whereas efficient information from the missingness pattern appears only at second order. Three-class quadratic discriminant calculations, finite-sample experiments, and a semi-synthetic multiclass application illustrate the resulting regime-dependent behaviour.