Related Experiment Video
Updated: Sep 16, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Artificial Intelligence and Machine Learning for Emergency Department Overcrowding: A Systematic Review with Large
Zekai Wang1, Ahmed Qasem2, Lin Lu1
1Department of Analytics, Charles F. Dolan School of Business, Fairfield University, 1073 North Benson Road, Fairfield, CT 06824, USA.
Abstract:
Background/Objectives: Emergency department (ED) overcrowding contributes to delayed care, prolonged length of stay (LOS), resource strain, and adverse patient outcomes. This systematic review aimed to examine how artificial intelligence (AI) and machine learning (ML) have been used to address ED crowding and patient flow, with emphasis on modeling approaches, validation practices, and real-world implementation. Methods: Following PRISMA 2020 guidelines, Scopus, Embase, Ovid MEDLINE, and CENTRAL were searched for relevant studies published from 2020 onward. After deduplication, 1888 records underwent title and abstract screening using two locally deployed LLaMA models with human adjudication. Screening performance was assessed against 150 manually annotated records. Full-text eligibility assessment and structured data extraction were conducted independently by multiple reviewers, with disagreements resolved by consensus. Results: Thirty-two studies were included. Most were retrospective, single-site investigations using electronic health record, administrative, or operational data. Common outcomes included ED LOS, waiting time, occupancy, boarding, disposition, and crowding indices. Tree-based and boosting models frequently performed well, although no approach was consistently superior across tasks and settings. Most studies relied on same-site validation, while external and temporal validation were uncommon. Prospective implementation, workflow integration, model maintenance, and direct operational, clinical, economic, or equity impacts were rarely evaluated. For LLM-assisted screening, LLaMA 4 Scout achieved 84.0% accuracy, 80.0% recall, 88.9% precision, and an F1 score of 84.2%, compared with 78.0%, 67.5%, 88.5%, and 76.6%, respectively, for LLaMA 3.3 on 150 randomly sampled papers. Conclusions: AI and ML show promise for addressing ED overcrowding, but the literature remains concentrated at the model-development stage. Future research should prioritize standardized outcomes, multicenter validation, prospective implementation, and direct evaluation of operational and patient-care outcomes.