Related Experiment Videos
Avoiding common failures in AI for health and medicine
Olawale Salaudeen1, Haoran Zhang1, Eileen Kim2
1Department of Electrical Engineering and Computer Science, Massachussets Institute of Technology, Cambridge, MA, USA.
Abstract:
The widespread adoption of artificial intelligence (AI) in healthcare necessitates reliable AI systems. Reliability refers to clinically acceptable performance across contexts such as patient and clinician populations and over time under real-world deployment conditions. This review synthesizes common reliability failure modes in predictive and generative AI systems, which include erroneous model outputs, clinically unjustified performance differences across patient populations, and performance degradation under changing deployment conditions. We examine approaches to addressing these failures and critically assess the empirical support and practical limits of current solutions, in both predictive and generative AI settings. We argue that current technical solutions are often insufficient to address the challenges that lead to unreliable AI, thus motivating the need for lifecycle-aware evaluation, continuous monitoring, and governance.