Related Experiment Video
Updated: Oct 5, 2026

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Explainable machine learning for identifying mild cognitive impairment in older adults with chronic diseases: Model
1School of Martial Arts, Shandong Sport University, Jinan, 250102, Shandong, China.
Objective:
This study aimed to develop and validate an interpretable machine-learning classification model for identifying MCI at the time of assessment among older adults with chronic conditions. We also evaluated model performance and the stability of TreeSHAP-based feature explanations across survey waves.
Methods:
Data were obtained from the CHARLS. A total of 8222 participants from the 2018 wave were stratified by MCI status and randomly divided into a development set (n = 5754) and an internal validation set (n = 2468). Data from the 2020 wave (n = 5558) were used for temporal validation. LASSO regression was performed exclusively in the development set for feature selection, resulting in 27 retained predictors. Using the same feature set, seven classification models were developed: logistic regression, support vector machine, multilayer perceptron, random forest, XGBoost, LightGBM, and k-nearest neighbors. Encoding, standardization, SMOTENC resampling, and hyperparameter tuning were performed within the training folds of 10-fold cross-validation to minimize information leakage, while the original class distributions were retained in both validation sets. Model performance was assessed using the AUROC, AUPRC, classification metrics, Brier score, calibration analysis, and decision curve analysis. TreeSHAP was then applied to the prespecified XGBoost model to characterize feature contributions and assess the stability of global feature importance across survey waves.
Results:
In the 2018 internal validation set, the prespecified XGBoost model achieved an AUROC of 0.692 (95% CI: 0.660-0.724), an AUPRC of 0.282 (95% CI: 0.242-0.328), a sensitivity of 0.511, a specificity of 0.746, and a Brier score of 0.109. In the 2020 temporal validation set, the corresponding values were 0.702 for AUROC, 0.305 for AUPRC, 0.390 for sensitivity, 0.843 for specificity, and 0.137 for the Brier score. These findings indicate that discrimination remained broadly stable across survey waves, although sensitivity and calibration were limited. Decision curve analysis showed a positive net benefit for XGBoost mainly across low-to-moderate threshold probabilities. TreeSHAP analysis identified number of living children, sleep duration, educational attainment, activities of daily living (ADL), and age as the strongest contributors to model predictions. Global feature-importance rankings were highly consistent between the 2018 internal validation set and the 2020 temporal validation set (Spearman ρ = 0.996).
Conclusion:
This study developed an interpretable machine-learning framework for identifying current MCI status among older adults with chronic conditions. The prespecified XGBoost model showed moderate and relatively stable discrimination in both internal and temporal validation, while its TreeSHAP-based feature explanations were highly stable across survey waves. The model may serve as an adjunctive screening and preliminary risk-stratification tool before formal cognitive assessment, but it should not replace clinical diagnosis of MCI or be used to predict incident MCI or future cognitive decline. Further recalibration and external validation in independent populations are required before implementation.