ToxLens: A Reproducible Graph-Learning Framework for Leakage-Aware, Uncertainty-Calibrated Molecular Toxicity Prediction
2026-08-31 • Machine Learning
Machine Learning
AI summaryⓘ
The authors developed ToxLens, a new computer method that predicts the toxicity of chemicals using graph-learning techniques. They designed it carefully to avoid data leakage, making the test results more reliable than usual. Their method combines different data features and statistical tools to improve prediction accuracy across 11 toxicity tests. They also used explainable AI tools to identify important chemical substructures related to toxicity. Overall, their approach worked better than simpler baseline methods and provided useful insights into toxic chemical features.
Molecular toxicity predictionGraph learningData leakageMatthews correlation coefficientConformal predictionSHAP valuesTox21 assaysAmes mutagenicityhERG inhibitionMonte Carlo dropout
Authors
Magnus H. Strømme, Alex G. C. de Sá, David B. Ascher
Abstract
Molecular toxicity prediction is increasingly used to prioritise compounds before experimental testing, but conventional benchmark performance can overstate practical utility when structurally related molecules occur across training and test folds. We introduce ToxLens, a reproducible multi-task graph-learning framework for 11 toxicity endpoints spanning Ames mutagenicity, acute oral toxicity, hERG inhibition, and Tox21 nuclear-receptor and stress-response assays. The workflow combines conservative chemical curation, sphere-exclusion filtering, a leakage-aware UMAP-HDBSCAN split, parallel graph and global-feature encoders joined by late concatenation, temperature-scaled Monte Carlo dropout with conformal-style prediction sets, applicability-domain analysis, and SHAP-guided toxicophore discovery with occlusion controls. On the leakage-controlled test fold, a five-seed soft-voting ensemble achieved a Matthews correlation coefficient score of 0.44, an area under the receiver operating characteristic curve score of 0.83, and an area under the precision-recall curve score of 0.58. It exceeded four ECFP4-based shallow baselines on all 11 endpoints under the same split and validation-based threshold-selection protocol. Controlled ablations showed that the global pathway was important, whereas late concatenation outperformed the tested gated and feature-wise linear modulation fusion variants. Conformal-style prediction sets revealed substantial endpoint-specific variation in set efficiency, and discrimination and calibration improved with similarity to the training domain. Retraining on fixed published Tox21 Challenge and TDA folds produced competitive, but not uniformly state-of-the-art, performance. SHAP-guided occlusion and consensus subgraph mining yielded model-derived structural hypotheses, 44 of which contained at least one occurrence that passed the predefined counterfactual criteria.