From Volumetrics to 3D Tensors: A Multi-Cohort Evaluation of Machine Learning and Deep Learning for Alzheimer's
Firuz Kamalov1, Anthony Ibrahim2
1Department of Electrical Engineering, Canadian University Dubai, Dubai, UAE. firuz@cud.ac.ae.
Abstract:
Automated Alzheimer's disease (AD) classification from structural MRI typically employs either feature-engineered machine learning (ML) or end-to-end 3D deep learning (DL). Addressing a lack of rigorous statistical comparison and multi-dataset validation in current literature, this study evaluates classical ML algorithms utilizing volumetric and voxel-based morphometry (VBM) features against 3D CNNs processing raw MRI tensors. Using the ADNI and OASIS datasets, models underwent internal 4-fold cross-validation and zero-shot cross-cohort testing to assess predictive performance, domain shift resilience, and clinical sensitivity. Welch's t-tests confirmed that 16 of 19 macro-regional brain aspects differed significantly between AD and cognitively normal subjects ( ), led by the medial-temporal lobe. While DL architectures achieved marginal numerical superiority in peak accuracy on ADNI (DenseNet121: accuracy , F1 , ROC-AUC ), they exhibited higher variance and lacked statistical significance compared to optimized ML baselines (XGBoost-VOL: F1 ; for all 3D CNNs). On the highly imbalanced OASIS dataset, VBM-enhanced Logistic Regression matched top DL models in overall F1-score ( ) while delivering statistically superior AD recall ( ; versus baseline), exceeding every 3D CNN. Although both paradigms generalized robustly across cohorts (only a 2-4% zero-shot F1 reduction; ROC-AUC up to 0.9506), the findings highlight a crucial clinical trade-off: 3D CNNs autonomously extract complex spatial features, but simpler, interpretable ML models provide superior inferential stability and diagnostic sensitivity, making them highly viable for real-world deployment.
