Related Experiment Video
Updated: Jun 3, 2025

Objective Nociceptive Assessment in Ventilated ICU Patients: A Feasibility Study Using Pupillometry and the Nociceptive Flexion Reflex
Published on: July 4, 2018
External validation of AI-based scoring systems in the ICU: a systematic review and meta-analysis
Patrick Rockenschaub1,2,3, Ela Marie Akay1, Benjamin Gregory Carlisle4
1CLAIM - Charité Lab for AI in Medicine, Charité - Universitätsmedizin Berlin, Berlin, Germany.
External validation of machine learning (ML) risk scores for intensive care unit (ICU) patients is uncommon, with performance often decreasing in new datasets. This highlights concerns about the reliability of current ML models in clinical practice.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Clinical Prediction Models
Background:
- Machine learning (ML) models are increasingly used for predicting clinical deterioration in intensive care units (ICUs).
- However, these models often exhibit poor performance in new hospital settings due to overfitting.
- External validation is crucial for assessing the reliability of ML-based risk scores before clinical implementation.
Purpose of the Study:
- To systematically review the frequency of external validation for ML-based risk scores in ICU patient deterioration prediction.
- To evaluate the impact of external validation on the performance of these ML models.
Main Methods:
- A systematic literature search was conducted across MEDLINE, Web of Science, and arXiv for studies predicting ICU patient deterioration using ML.
- Included studies were published in English before December 2023.
- The frequency of external validation was summarized, and changes in performance (Area Under the Receiver Operating Characteristic curve - AUROC) were analyzed using linear mixed-effects models.
Main Results:
- Out of 572 included studies, only 14.7% underwent external validation, rising to 23.9% by 2023.
- External validation was more common in studies using open-source data, particularly MIMIC and eICU datasets.
- On average, AUROC decreased by -0.037 in external data, with nearly half of studies showing a performance drop greater than 0.05.
Conclusions:
- External validation of ML risk scores in ICUs is infrequent, despite its increasing trend.
- The observed performance decline in external datasets raises concerns about the generalizability and reliability of many proposed ML scores.
- Over-reliance on limited datasets and exclusive use of AUROC hinder robust interpretation and validation.
More Related Videos
04:54Author Spotlight: IntelliSleepScorer — A High-Accuracy, Accessible GUI Software for Automated Sleep Stage Scoring in Mice and its Application in Psychiatric Research
Published on: November 8, 2024
10:38Observational Study Protocol for Repeated Clinical Examination and Critical Care Ultrasonography Within the Simple Intensive Care Studies
Published on: January 16, 2019
Related Concept Videos
Data Validation
Nursing assessment guides are generally based on holistic models rather than medical...
Imaging Studies for Cardiovascular System VI: Calcium -Scoring CT
Pre-Procedural Guidelines for Assessing Blood Pressure