Related Experiment Video
Updated: Jul 31, 2026

Automated Midline Shift and Intracranial Pressure Estimation based on Brain CT Images
Published on: April 13, 2013
Prediction of NIHSS Scores and Acute Ischemic Stroke Severity Using a Cross-attention Vision Transformer Model with
Pahati Tuxunjiang1, Chencui Huang2, Zhen Zhou2
1The First Affiliated Hospital of Xinjiang Medical University, Department of Radiology, Urumqi, Xinjiang Uygur Autonomous Region, PR China (P.T., H.H.K., W.Z., R.X., A.A., Y.S., Y.W.).
Rationale And Objectives:
This study aimed to develop and evaluate models for classifying the severity of neurological impairment in acute ischemic stroke (AIS) patients using multimodal MRI data.
Methods:
A retrospective cohort of 1227 AIS patients was collected and categorized into mild (NIHSS<5) and moderate-to-severe (NIHSS≥5) stroke groups based on NIHSS scores. Eight baseline models were constructed for performance comparison, including a clinical model, radiomics models using DWI or multiple MRI sequences, and deep learning (DL) models with varying fusion strategies (early fusion, later fusion, full cross-fusion, and DWI-centered cross-fusion). All DL models were based on the Vision Transformer (ViT) framework. Model performance was evaluated using metrics such as AUC and ACC, and robustness was assessed through subgroup analyses and visualization using Grad-CAM.
Results:
Among the eight models, the DL model using DWI as the primary sequence with cross-fusion of other MRI sequences (Model 8) achieved the best performance. In the test cohort, Model 8 demonstrated an AUC of 0.914, ACC of 0.830, and high specificity (0.818) and sensitivity (0.853). Subgroup analysis shows that model 8 is robust in most subgroups with no significant prediction difference (p > 0.05), and the AUC value consistently exceeds 0.900. A significant predictive difference was observed in the BMI group (p < 0.001). The results of external validation showed that the AUC values of the model 8 in center 2 and center 3 reached 0.910 and 0.912, respectively. Visualization using Grad-CAM emphasized the infarct core as the most critical region contributing to predictions, with consistent feature attention across DWI, T1WI, T2WI, and FLAIR sequences, further validating the interpretability of the model.
Conclusion:
A ViT-based DL model with cross-modal fusion strategies provides a non-invasive and efficient tool for classifying AIS severity. Its robust performance across subgroups and interpretability make it a promising tool for personalized management and decision-making in clinical practice.

