Related Experiment Video
Updated: Jan 15, 2026

14:56
An Experimental Paradigm for the Prediction of Post-Operative Pain PPOP
Published on: January 27, 2010
21.9K
Machine Learning-Based prediction models for postoperative delirium: a systematic review and Meta-Analysis
Yingying Tu1, Haoyuan Zhu2, Xiaozhen Zhang1
1Department of Nursing, The First Affiliated Hospital of Wenzhou Medical University, WenZhou, ZheJiang, China.
BMC Psychiatry
|October 7, 2025
Summary
Machine learning models show strong performance in predicting postoperative delirium (POD), demonstrating stable accuracy across diverse surgical and patient groups. This review evaluates existing POD risk models for clinical guidance.
Area of Science:
- Medical Informatics
- Clinical Prediction Models
- Machine Learning in Healthcare
Background:
- The proliferation of postoperative delirium (POD) risk prediction models necessitates an evaluation of their quality and clinical applicability.
- Existing models' performance and generalizability remain uncertain, impacting clinical practice and research.
- A systematic review is crucial to assess the current landscape of POD prediction models.
Purpose of the Study:
- To systematically review and evaluate published studies on POD risk prediction models.
- To assess the diagnostic accuracy, including sensitivity and specificity, of these models.
- To provide evidence-based guidance for the development and enhancement of future POD prediction models.
Main Methods:
- A comprehensive systematic search was conducted across major scientific databases (PubMed, Embase, Cochrane Library) up to February 15, 2025.
- Included studies provided essential data on the sensitivity and specificity of various predictive models for POD.
- Seventeen articles comprising 69 machine learning (ML) models were analyzed, covering 205,202 patients.
Main Results:
- ML models demonstrated a pooled AUROC of 0.83 (95% CI, 0.79-0.86) for POD prediction, with sensitivity of 0.73 and specificity of 0.79.
- The random forest model achieved the highest pooled AUROC (0.89). Models for orthopedic surgeries and younger patients (<60 years) showed superior performance.
- Models with external validation outperformed those with only internal validation. Asian population models showed better performance than European/American ones.
Conclusions:
- Machine learning models exhibit robust performance in predicting postoperative delirium.
- Model performance remains stable across different surgical types and demographic subgroups.
- Key predictors consistently identified include advanced age, cognitive impairment, comorbidities, anemia, and hypoalbuminemia.

