Entropy-Shapley method reveals how features drive multivariate uncertainty
A Hierarchy of Entropy-Shapley Games for Multivariate Predictive Uncertainty
Machine Learning
Summary
Predictive models often guess multiple related outcomes at once, and understanding which inputs cause uncertainty in these guesses is tricky because the outputs influence each other. The authors introduce a new way to measure how each input feature affects both the individual uncertainties and the connections between outputs, using a special kind of game based on entropy. Their method can tell if an input mostly changes an output’s own uncertainty or how different outputs relate to each other. This helps improve decisions and model understanding where knowing uncertainty matters.
What this means in practice
- •For risk management teams: Identify specific input factors causing joint uncertainties to improve risk-informed decision-making in probabilistic forecasting models.
- •For machine learning engineers: Diagnose and improve multivariate models by pinpointing features affecting output dependence rather than just marginal uncertainty.
Authors
Niklas Koenen, Claudia Battistin, Jeriek Van den Abeele, Martin Jullum
Abstract
Modern probabilistic machine learning models increasingly produce multivariate outputs with complex dependence structure, from multi-step time-series forecasts to sample path predictions. Understanding which input features drive the predictive uncertainty is important for risk-aware decisions, model diagnostics, and deciding whether the uncertainty should be mitigated or hedged against. This attribution problem requires a choice of how dependencies between output components are treated. Existing approaches reduce the output to a scalar through aggregation or projection before attribution, thereby obscuring whether features affect marginal uncertainty, dependence structure, or both, while component-wise analyses can miss dependence effects entirely. We close this gap by introducing a hierarchy of three entropy-based Shapley games that make this output-side choice explicit for any ordered multivariate outcome, ranging from per-component marginal entropy to fully joint entropy. The hierarchy isolates a cross-component attribution term that captures how each feature shifts the dependence between output components, a quantity invisible to component-wise methods. We establish a chain-rule decomposition of the joint attribution and characterize the cross-component term through conditional total correlation, providing both closed-form and sample-based estimators. Finally, we demonstrate how the framework captures differences in learned joint structure across probabilistic models from distributional regression to a zero-shot time series foundation model.