How Does Bayesian Causal Discovery Fail? Characterising Structural Consequences in Linear Gaussian Networks under Latent Confounding
2026-07-10 • Artificial Intelligence
Artificial Intelligence
AI summaryⓘ
The authors study how Bayesian causal discovery, which tries to find cause-effect relationships while showing uncertainty, behaves when there is hidden confounding (an unseen factor influencing two variables). They focus on simple linear models where exactly two observed variables share a hidden influence. They find there is a critical level of correlation between these variables above which the method wrongly favors adding a direct connection that doesn’t really exist. This threshold gets lower with more data, meaning bigger datasets might make the mistake more likely. The authors also identify two different ways the method can fail based on the local graph structure around the confounded variables.
Bayesian causal discoverylatent confoundingdirected acyclic graphepistemic uncertaintyposterior distributionlinear Gaussian causal modelscorrelation thresholdscore functionidentifiabilitygraph structure
Authors
Debargha Ghosh, Silja Renooij, Anna Kononova
Abstract
Bayesian causal discovery is widely used for its ability to quantify epistemic uncertainty over directed acyclic graphs (DAGs) through posterior inference. However, its behaviour under latent confounding remains poorly understood, as existing work typically notes that confounding breaks identifiability without characterising how the posterior distribution over DAGs responds. In this work, we analyse posterior behaviour under latent confounding in linear Gaussian causal models, focusing on additive latent confounding between exactly two observed variables. We derive a critical correlation threshold above which the score function favours graphs with a spurious edge between the confounded variables, and show that this threshold decreases with sample size -- more data lowers the correlation required for the spurious edge to be favoured. Beyond this threshold, we characterize two distinct posterior failure regimes determined by the local structure around the confounded variables. Our findings are supported by exact posterior computations on multiple graph structures, demonstrating both the predicted failure regimes.