Related Experiment Video
Updated: Sep 18, 2026

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Prediction Models for In-Hospital Delirium Using Routinely Collected Electronic Health Record Data: Systematic Review
Hung-Min Huang1, Chun-Shun Lu2, Geng-Wei Chang3
1Institute of Health Informatics, University College London, Gower Street, London, England, WC1E 6BT, United Kingdom, 44 7570909461.
Background:
Delirium is a common and clinically important form of acute in-hospital mental status deterioration. Electronic health record (EHR)-based prediction models may support early identification and targeted prevention, but their methodological quality, validation rigor, and clinical readiness remain uncertain.
Objective:
This systematic review aimed to synthesize and critically evaluate prediction models for in-hospital delirium developed using routinely collected EHR data, focusing on model characteristics, validation strategies, performance, risk of bias, and clinical applicability.
Methods:
We searched PubMed, MEDLINE, Embase, PsycINFO, and Web of Science from inception to November 11, 2025. Eligible studies developed, validated, or evaluated multivariable prediction models using routinely collected EHR or administrative data to predict acute mental status deterioration during adult hospital admissions. Although eligibility criteria were broad, all included studies operationalized deterioration as delirium. Data extraction was informed by CHARMS (Checklist for Critical Appraisal and Data Extraction for Systematic Reviews of Prediction Modeling Studies) and TRIPOD (Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis) or TRIPOD-artificial intelligence guidance. Model performance, validation, calibration, and implementation features were synthesized narratively. Risk of bias and applicability were assessed using PROBAST (Prediction Model Risk of Bias Assessment Tool).
Results:
Twenty-nine studies met the inclusion criteria. The evidence clustered into 4 overlapping prediction tasks: admission or early-stay risk stratification, perioperative or postoperative prediction, dynamic intensive care unit prediction, and external validation or workflow evaluation of existing tools. Most studies were retrospective cohorts (20/29, 69%) and were conducted in general ward, mixed ward-intensive care unit, intensive care unit, or emergency department settings. Machine learning or hybrid approaches were common (18/29, 62%), but more complex models did not consistently outperform statistical or rule-based approaches. Of 29 studies, internal discrimination was reported in 24 (83%; area under the receiver operating characteristic curve range 0.77-0.97) studies, whereas external discrimination was reported in 12 studies and calibration in 15 studies. Decision curve analysis was reported in 3 studies, and prospective evaluation or workflow integration remained limited. Overall risk of bias was low in 8 studies, unclear in 10 studies, and high in 11 studies, mainly because of analysis-domain limitations.
Conclusions:
Routinely collected EHR data can support delirium risk prediction across hospital settings, and many models show moderate to high discrimination. However, no single algorithm is ready for routine adoption. The field remains limited by heterogeneous prediction tasks, inconsistent outcome ascertainment, weak calibration and decision-analytic reporting, and insufficient external or prospective evaluation. Future studies should define the intended clinical use case before model development, evaluate calibration and clinical usefulness alongside discrimination, and test models across institutions, time periods, and workflows before deployment.