Related Experiment Video
Updated: May 12, 2026

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Multidimensional machine learning for early neurological deterioration prediction in acute ischemic stroke
1Shinhan University of Physical Education, Uijeongbu-si, Gyeonggi-do, Republic of Korea.
Objective:
This study aimed to develop and validate a multidimensional clinical feature-based machine learning model for accurately predicting the risk of early neurological deterioration (END) in patients with acute ischemic stroke (AIS).
Methods:
A total of 338 AIS patients were randomly divided into a training set (n = 236) and a validation set (n = 102). Five core predictors were identified from multiple clinical and pathological indicators: admission National Institutes of Health Stroke Scale (NIHSS) score, admission blood glucose, infarct core volume, collateral circulation status, and neutrophil-to-lymphocyte ratio (NLR). In the training set, univariate analysis was first performed to screen prognosis-related factors. After variable compression via least absolute shrinkage and selection operator (LASSO) regression, multivariate logistic regression was employed to determine independent risk factors for poor prognosis. Using Python, three prediction models-Random Forest (RF), Gradient Boosting Machine (GBM), and K-Nearest Neighbors (KNN)-were constructed. Model performance was evaluated by the area under the receiver operating characteristic curve (AUC), and the optimal model was selected.
Results:
No statistically significant differences were observed in baseline characteristics between the training and validation sets (P > 0.05). Multivariate logistic regression revealed that admission NIHSS score, blood glucose, infarct core volume, and NLR were independent risk factors (P < 0.05), while collateral circulation status was an independent protective factor (P < 0.05). The RF model demonstrated superior predictive performance, with AUC values of 0.779 (training set) and 0.775 (validation set), significantly outperforming KNN (0.727, 0.741) and GBM (0.736, 0.665).
Conclusion:
The multidimensional model provides a potential practical tool for early clinical identification of high-risk END patients and timely intervention.