Related Experiment Videos
Artificial Intelligence for Predicting Length of Stay for General Surgery Patients: A Systematic Review
Joshua Blum1,2, Nomiki Glynatsis1, Brieanna Hill1
1Tasmanian School of Medicine, University of Tasmania, Hobart, Australia.
Background:
Length of stay (LoS) prediction is vital for optimizing surgical patient flow and resource allocation. Artificial intelligence (AI) and machine learning (ML) models may enhance LoS prediction, yet their validity and clinical utility in General Surgery remain unclear.
Methods:
Embase, Medline, Scopus, and Web of Science were systematically searched to June 2025. Studies employing AI or ML to predict LoS or discharge timing in adult General Surgery cohorts were included. Data on model development, predictor and target variables, performance, and validation were extracted. Risk of bias was appraised using PROBAST + AI. A checklist was developed to guide future LoS-prediction research.
Results:
Nine of 4115 screened records met inclusion criteria. Classifier models with low overall risk of bias achieved moderate discrimination (AUROC 0.71-0.83), whereas regression models performed poorly (RMSE 4.6-12.7 days). Methodological weaknesses were frequent, including data leakage, temporally invalid predictors, biased patient selection, limited external validation, and incomplete performance reporting. Important predictors commonly included ASA score, age and operative duration; psychosocial and home-support features were rarely considered. Outcome definitions and intended applications varied substantially, limiting direct comparison of performance. Models predicting discharge within short, dynamically reassessed time horizons may achieve greater accuracy, but offer limited capacity for longer-term bed forecasting. No included study benchmarked ML or AI predictions against usual clinician LoS estimates. Low-bias models were retrospective and generally internally validated, so performance in prospective or external cohorts remains unknown. More complex neural network models did not consistently outperform simpler methods. This review identifies recurring barriers to prospective clinical application for routine real-world use.
Conclusion:
AI and ML models can moderately predict LoS for General Surgery patients, but generalizability is limited by methodological bias, incomplete reporting, and non-standard outcome definitions. Current models appear more suited to hospital-wide forecasting than bedside decision support. Future development should use prospectively available predictors, clinically meaningful outcomes, transparent performance reporting, and external validation. The proposed framework highlights strategies to avoid common methodological flaws in models published to date.