Related Experiment Video
Updated: Jun 3, 2025

Objective Nociceptive Assessment in Ventilated ICU Patients: A Feasibility Study Using Pupillometry and the Nociceptive Flexion Reflex
Published on: July 4, 2018
External validation of AI-based scoring systems in the ICU: a systematic review and meta-analysis
Patrick Rockenschaub1,2,3, Ela Marie Akay1, Benjamin Gregory Carlisle4
1CLAIM - Charité Lab for AI in Medicine, Charité - Universitätsmedizin Berlin, Berlin, Germany.
Background:
Machine learning (ML) is increasingly used to predict clinical deterioration in intensive care unit (ICU) patients through scoring systems. Although promising, such algorithms often overfit their training cohort and perform worse at new hospitals. Thus, external validation is a critical - but frequently overlooked - step to establish the reliability of predicted risk scores to translate them into clinical practice. We systematically reviewed how regularly external validation of ML-based risk scores is performed and how their performance changed in external data.
Methods:
We searched MEDLINE, Web of Science, and arXiv for studies using ML to predict deterioration of ICU patients from routine data. We included primary research published in English before December 2023. We summarised how many studies were externally validated, assessing differences over time, by outcome, and by data source. For validated studies, we evaluated the change in area under the receiver operating characteristic (AUROC) attributable to external validation using linear mixed-effects models.
Results:
We included 572 studies, of which 84 (14.7%) were externally validated, increasing to 23.9% by 2023. Validated studies made disproportionate use of open-source data, with two well-known US datasets (MIMIC and eICU) accounting for 83.3% of studies. On average, AUROC was reduced by -0.037 (95% CI -0.052 to -0.027) in external data, with more than 0.05 reduction in 49.5% of studies.
Discussion:
External validation, although increasing, remains uncommon. Performance was generally lower in external data, questioning the reliability of some recently proposed ML-based scores. Interpretation of the results was challenged by an overreliance on the same few datasets, implicit differences in case mix, and exclusive use of AUROC.
More Related Videos
04:54Author Spotlight: IntelliSleepScorer — A High-Accuracy, Accessible GUI Software for Automated Sleep Stage Scoring in Mice and its Application in Psychiatric Research
Published on: November 8, 2024
10:38Observational Study Protocol for Repeated Clinical Examination and Critical Care Ultrasonography Within the Simple Intensive Care Studies
Published on: January 16, 2019
Related Concept Videos
Data Validation
Nursing assessment guides are generally based on holistic models rather than medical...
Imaging Studies for Cardiovascular System VI: Calcium -Scoring CT
Pre-Procedural Guidelines for Assessing Blood Pressure