Related Experiment Video
Updated: Jun 12, 2026

Computerized Adaptive Testing System of Functional Assessment of Stroke
Published on: January 7, 2019
Clinical evaluation of an automated Alberta Stroke Program Early Computed Tomography Score (ASPECTS)-scoring system
Ömer Bagcilar1, Alexander Rau1, Stephan Rau2
1Department of Neuroradiology, Medical Center-University of Freiburg, Faculty of Medicine, University of Freiburg, Freiburg, Germany.
Background:
The Alberta Stroke Program Early Computed Tomography Score (ASPECTS) is widely established to assess early ischemic changes on non-contrast computed tomography (NCCT) and guide treatment decisions in acute stroke. While automated ASPECTS tools are increasingly available, independent validation and the potential influence on human ratings remain important. The aim of this study was to evaluate agreement between an automated ASPECTS scoring system and expert readers and to examine whether software assistance was associated with a systematic shift in human ASPECTS scoring.
Methods:
We implemented an automated ASPECTS scoring system, incorporating image normalization, anatomical registration, and regional intensity analysis through hemispheric comparisons of net water uptake (NWU) values. A total of 224 cases were retrospectively analyzed. We included two expert unassisted readers (U1-2; reference reader group) and four software-assisted readers (A1-4). Agreement was assessed using intraclass correlation coefficient (ICC), mean difference (MD), mean absolute difference (MAD), and the distribution of absolute score differences (Δ=0/1/2/≥3). Unassisted vs assisted scores were compared using a Wilcoxon rank-sum test. Analyses were performed using two bootstrapped stratifications (approximately uniform and clinically representative).
Results:
Inter-rater agreement within the unassisted and assisted reader groups was high (uniform: unassisted ICC 0.96, MAD 0.65; assisted ICC 0.95, MAD 0.66; clinical: unassisted ICC 0.91, MAD 0.83; assisted ICC 0.88, MAD 0.89). Agreement between unassisted and assisted readings was similarly high (uniform: ICC 0.96, MAD 0.60, MD -0.02; clinical: ICC 0.90, MAD 0.80, MD -0.05), with no significant differences between unassisted and assisted scores (Wilcoxon: uniform P=0.82, clinical P=0.61). The automated output showed good agreement with human ratings, though consistently lower than inter-reader agreement (uniform: ICC 0.89, MAD 0.97-1.01, MD -0.20 to -0.22; clinical: ICC 0.76-0.77, MAD 1.18-1.27, MD -0.16 to -0.21). Sensitivity analyses supported an NWU threshold of approximately 7%.
Conclusions:
The automated system demonstrated good-to-moderate agreement with expert ratings and was not associated with a systematic group-level shift in ASPECTS scoring. It may support more standardized ASPECTS evaluation, particularly in settings with limited access to expert readers, while maintaining the autonomy of clinical judgment.
