Related Experiment Video
Updated: Aug 10, 2026

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Quantifying generalization error in machine learning prediction of cognitive decline
Roya Melanie Hüppi1, Nicolas Langer2, Bruno Hebling Vieira2
1Methods of Plasticity Research, Department of Psychology, University of Zurich, 8050, Zurich, Zurich, Switzerland; Neuroscience Center Zurich (ZNZ), University of Zurich & ETH Zurich, 8057, Zurich, Zurich, Switzerland; Department of Adult Psychiatry and Psychotherapy, Psychiatric University Clinic Zurich and University of Zurich, 8032, Zurich, Zurich, Switzerland.
Background:
Predicting cognitive decline as a continuum, from healthy age-related decline to mild cognitive impairment and dementia, enables more precise individual-level predictions. However, the practical value of such models for early intervention and prevention depends on their ability to generalize to independent cohorts, a property that is often not evaluated.
Objectives:
This study investigated whether adding structural magnetic resonance imaging (MRI) to non-brain data improved machine learning predictions of continuous cognitive decline and analyzed the models' generalizability.
Design:
Multi-target random forest regression models predicted annual decline in the Clinical Dementia Rating Scale Sum of Boxes (CDR-SOB) and Mini-Mental State Examination (MMSE) using non-brain data, structural MRI data, or their combination from the Alzheimer's Disease Neuroimaging Initiative (ADNI; N = 1237) and Open Access Series of Imaging Studies (OASIS-3; N = 662) datasets. Cross-site generalizability was evaluated.
Setting:
Data from ADNI and OASIS-3 were used for this study.
Participants:
A total of 1899 participants who had demographic, clinical, and brain imaging data from a baseline session and clinical data from at least 2 follow-up sessions were included.
Measurements:
Baseline non-brain (demographics, clinical and neuropsychological scores, information on APOE genotype, cognitive diagnosis, health, and number of sessions before baseline) and/or structural MRI data were used to predict the yearly rate of change in CDR-SOB and MMSE scores.
Results:
Including structural MRI data improved prediction of CDR-SOB and MMSE change, reaching respective R2 values of .41 and .33 in ADNI and .42 and .33 in OASIS-3. Model performance for across-dataset predictions was reduced (R2 between .18 and .35), unexplained by distributional shifts of target variables. Models using only top predictive features performed similarly to full models when tested externally (R2 between .18 and .34), suggesting predictor redundancy.
Conclusions:
Incorporating structural MRI data enhances within-dataset prediction of continuous cognitive decline, allowing for more precise individual-level prediction and advancing towards precision medicine. Even though external validation remains limited, quantifying the generalizability gap is a crucial step towards the responsible use of ML models in clinical intervention and prevention.
Related Concept Videos
Regression Toward the Mean
Detection of Gross Error: The Q Test
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
The...
