Trustworthy AI in Digital Health: A Comprehensive Review of Robustness and Explainability
2026-08-03 • Artificial Intelligence
Artificial IntelligenceMachine Learning
AI summaryⓘ
The authors review how to make AI systems trustworthy, focusing on healthcare applications. They emphasize two key aspects: robustness (how well AI works in different situations) and explainability (how well humans can understand AI decisions). The paper organizes existing methods and challenges, especially in areas like intensive care and neonatal health. It also covers new techniques to handle limited or changing data and ways to explain AI outputs. Overall, the authors aim to help researchers create safer and clearer AI tools for digital health.
trustworthy AIrobustnessexplainabilitydigital healthdistributional shiftsfeature attributioncounterfactual explanationslarge language modelsevaluation metricsAI lifecycle
Authors
Abdullah Mamun, Shovito Barua Soumma, Hassan Ghasemzadeh
Abstract
Ensuring trust in AI systems is essential for the safe and ethical integration of machine learning systems into high-stakes domains such as digital health. Key dimensions, including robustness, explainability, fairness, accountability, and privacy, need to be addressed throughout the AI lifecycle, from problem formulation and data collection to model deployment and human interaction. While various contributions address different aspects of trustworthy AI, a focused synthesis on robustness and explainability, especially tailored to the healthcare context, remains limited. This review addresses that need by organizing recent advancements into an accessible framework, highlighting both technical and practical considerations. We present a structured overview of methods, challenges, and solutions, aiming to support researchers and practitioners in developing reliable and explainable AI solutions for digital health. This review article is organized into three main parts. First, we introduce the pillars of trustworthy AI and discuss the technical and ethical challenges, particularly in the context of digital health. Second, we explore application-specific trust considerations across domains such as intensive care, neonatal health, and metabolic health, highlighting how robustness and explainability support trust. Lastly, we present recent advancements in techniques aimed at improving robustness under data scarcity and distributional shifts, as well as explainable AI methods ranging from feature attribution to gradient-based interpretations and counterfactual explanations. This paper is further enriched with detailed discussions of the contributions toward robustness and explainability in digital health, the development of trustworthy AI systems in the era of LLMs, and various evaluation metrics for measuring trust and related parameters such as validity, fidelity, and diversity.