Related Experiment Video
Updated: Jan 16, 2026

Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches
Published on: December 15, 2023
Data Leakage in Deep Learning for Alzheimer's Disease Diagnosis: A Scoping Review of Methodological Rigor and
Vanessa M Young1,2, Samantha Gates1, Layla Y Garcia1
1Glenn Biggs Institute for Alzheimer's and Neurodegenerative Diseases, University of Texas Health Science at San Antonio, San Antonio, TX 78229, USA.
Abstract:
Background: Deep-learning models for Alzheimer's disease (AD) diagnosis frequently report revolutionary accuracies exceeding 95% yet consistently fail in clinical translation. This scoping review investigates whether methodological flaws, particularly data leakage, systematically inflates performance metrics, and examines the broader landscape of validation practices that impact clinical readiness. Methods: We conducted a scoping review following PRISMA-ScR guidelines, with protocol pre-registered in the Open Science Framework (OSF osf.io/2s6e9). We searched PubMed, Scopus, and CINAHL databases through May 2025 for studies employing deep learning for AD diagnosis. We developed a novel three-tier risk stratification framework to assess data leakage potential and systematically extracted data on validation practices, interpretability methods, and performance metrics. Results: From 2368 identified records, 44 studies met inclusion criteria, with 90.9% published between 2020-2023. We identified a striking inverse relationship between methodological rigor and reported accuracy. Studies with confirmed subject-wise data splitting reported accuracies of 66-90%, while those with high data leakage risk claimed 95-99% accuracy. Direct comparison within a single study demonstrated a 28-percentage point accuracy drop (from 94% to 66%) when proper validation was implemented. Only 15.9% of studies performed external validation, and 79.5% failed to control for confounders. While interpretability methods like Gradient-weighted Class Activation Mapping (Grad-CAM) were used in 18.2% of studies, clinical validation of these explanations remained largely absent. Encouragingly, high-risk methodologies decreased from 66.7% (2016-2019) to 9.5% (2022-2023). Conclusions: Data leakage and associated methodological flaws create a pervasive illusion of near-perfect performance in AD deep-learning research. True accuracy ranges from 66-90% when properly validated-comparable to existing clinical methods but far from revolutionary. The disconnect between technical implementation of interpretability methods and their clinical validation represents an additional barrier. These findings reveal fundamental challenges that must be addressed through adoption of a "methodological triad": proper data splitting, external validation, and confounder control.
More Related Videos
12:50Lesion Explorer: A Video-guided, Standardized Protocol for Accurate and Reliable MRI-derived Volumetrics in Alzheimer's Disease and Normal Elderly
Published on: April 14, 2014
09:38Generalized Psychophysiological Interaction PPI Analysis of Memory Related Connectivity in Individuals at Genetic Risk for Alzheimer's Disease
Published on: November 14, 2017