Related Experiment Video
Updated: Aug 14, 2026

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Explainable XGBoost model and nomogram for risk factor identification and risk prediction in cerebral small vessel
Xi Zhu1, Xuhui Liu2, Xujie Wang3
1Department of Neurology, The Fifth Affiliated Hospital of Xinjiang Medical University, Ürümqi, China.
Insights
This study developed an interpretable machine learning model to predict cerebral small vessel disease (CSVD) risk using common clinical data. The model shows high accuracy, aiding early intervention for this common vascular disorder.
Area of Science:
- Neurology
- Cardiovascular Disease
- Artificial Intelligence in Medicine
Background:
- Cerebral small vessel disease (CSVD) is a prevalent vascular disorder linked to cognitive decline and poor prognosis.
- Early identification of high-risk individuals for CSVD is challenging due to complex pathophysiology.
- This study focused on developing a predictive model for CSVD occurrence.
Purpose of the Study:
- To develop and validate an interpretable machine learning (ML) model for predicting cerebral small vessel disease (CSVD).
- To identify key predictors for CSVD risk stratification.
- To create a clinically applicable tool for early CSVD risk assessment.
Main Methods:
- Retrospective analysis of 1,640 adult patients.
- Feature selection using LASSO regression and logistic regression.
- Comparison of six ML algorithms (XGBoost, SVM, etc.), with XGBoost selected as optimal.
- Model interpretation using SHAP and nomogram construction.
Main Results:
- Ten predictors identified: blood glucose, hypertension history, systolic blood pressure, age, triglycerides, stroke history, cystatin C, C-reactive protein, homocysteine, and BMI.
- XGBoost model achieved high performance (AUC 0.968 training, 0.938 validation).
- The model demonstrated clinical utility via calibration plots and DCA, with strong prognostic discrimination.
Conclusions:
- An interpretable XGBoost-based ML model was validated for early CSVD risk stratification.
- The model utilizes routinely collected, low-cost variables, making it suitable for resource-limited settings.
- Future work includes prospective validation and EHR integration for decision support.
Background:
Cerebral small vessel disease (CSVD) is a common, clinically significant vascular disorder that frequently leads to cognitive impairment, dementia, and poor overall prognosis. Owing to its complex hemodynamic characteristics and multifactorial pathophysiology, early identification of individuals at high risk for CSVD remains a clinical challenge. This study aimed to develop and validate an interpretable machine learning (ML) model for predicting the occurrence of CSVD.
Methods:
We retrospectively enrolled 1,640 adult patients treated at the Fifth Affiliated Hospital of Xinjiang Medical University between September 2019 and December 2024. Twenty-three candidate variables (demographics, vitals, biomarkers, comorbidities) were evaluated. Feature selection was performed using least absolute shrinkage and selection operator (LASSO) regression, followed by stepwise backward elimination in multivariable logistic regression. Six supervised ML algorithms (DT, KNN, LR, LightGBM, XGBoost, SVM) were compared. Performance was assessed using ROC curves, calibration plots, and decision curve analysis (DCA). The optimal model was interpreted using SHapley Additive exPlanations (SHAP), and a bedside clinical nomogram was constructed.
Results:
Ten independent predictors were identified: blood glucose, history of hypertension, systolic blood pressure, age, triglycerides, history of stroke, cystatin C, C-reactive protein, homocysteine, and body mass index. Among all models, XGBoost demonstrated the best performance, with an AUC of 0.968 in the training cohort and 0.938 in the validation cohort. Calibration plots and DCA confirmed its clinical utility. The derived nomogram demonstrated strong prognostic discrimination (p < 0.0001). The XGBoost model achieved an accuracy of 88.0%, sensitivity of 80.9%, specificity of 93.8%, and an F1 score of 0.86, corresponding to a 5.4-percentage-point gain in AUC over logistic regression. Ten-fold cross-validation confirmed this ranking, with a mean AUC of 0.934 ± 0.016.
Conclusions:
We validated an interpretable XGBoost-based ML model that facilitates early risk stratification and targeted interventions for CSVD. Because the model relies only on routinely collected, low-cost variables and open-source software, it is readily transferable to resource-limited settings; future work will focus on prospective, multicentre external validation and on embedding the nomogram into electronic-health-record decision support.