Biomedical data clusters improved by causal and flexible dependence modeling
Copula Adapted Directed Acyclic Graph for Cluster Representation of Biomedical Data
Machine Learning
Summary
Biomedical data can be complicated and often lacks clear labels, making it hard for computers to group similar patients or conditions accurately. The authors developed a new way to look at the data by combining methods that model complex relationships and find cause-effect links between features. This approach helps group data more accurately without needing labels, providing clear visualizations and insights about what features matter. Their method outperformed many other techniques on multiple biomedical datasets.
What this means in practice
- •For hospital data teams: Improve patient or sample grouping without labels by leveraging complex feature dependencies and causal structures in biomedical datasets.
- •For biomedical software developers: Build bioinformatics tools that provide explainable cluster results and causal feature relationships for unlabeled biomedical data.
Authors
Heranga K. Rathnasekara, Norou Diawara, Manar D. Samad
Abstract
Diagnostic errors and mislabeling are common in biomedicine, which compromise the reliability of predictive models and data-driven outcomes. Stratifying unlabeled biomedical data based on complex relationships between features eliminates the need for data labels and overcomes the limitations of supervised learning. Traditional clustering methods assume restrictive data distributions, making them suboptimal for capturing complex dependencies in high-dimensional biomedical data. This paper introduces a novel cluster-friendly data presentation framework that integrates the non-Gaussian and non-linear feature dependence of copula models with an ensemble of causal structure discovery (CSD) methods based on Directed Acyclic Graphs (DAGs). While copulas model flexible multivariate distributions by relaxing assumptions related to multivariate normality, linear dependence, and symmetric relationships, an ensemble of DAG-based CSD methods identifies stable causal relationships between features. When clustered using K-means, the new data representation obtained by the proposed copula-adapted DAG (CopDAG) ranks first among the 12 methods in normalized clustering accuracy and adjusted Rand index across 16 biomedical datasets. Our CopDAG method predicts ground-truth class labels directly from feature relationships without data annotations and supervised learning, while also providing cluster visualizations and explainable causal structures of the biomedical data features.