Related Experiment Video
Updated: Aug 21, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Prospective evaluation of a large language model clinical decision support system in the emergency department
Liron Leibovitch1, Adi Ahituv1,2, Alon Gorenshtein3
1Department of Neurology, Rambam Health Care Campus, Haifa, Israel.
Abstract:
Prospective evidence for artificial intelligence (AI)-based clinical decision support in emergency departments remains limited. Here we conducted a DECIDE-AI stage 1 evaluation of SHAKED, a clinical decision support system built on multiple large language models, in a tertiary emergency department. Over 4 weeks, 1,138 patients were analyzed across two parallel units-one using SHAKED and one following routine rotations. Clinical adoption of SHAKED declined from 68% to 30%, owing to workload-sensitive disengagement (OR = 0.72 per shift hour, 95% CI 0.62 to 0.83). Physicians preferred the use of SHAKED for radiology consultations (OR = 2.98, 95% CI 1.58 to 5.63). No adverse events were detected, and expert review rated 99 of 100 sampled outputs as clinically appropriate. Emergency department length of stay did not differ between wings (4.9 h in both, P = 0.99). Intention-to-treat analysis showed a non-significant trend toward shorter consultation cycle time (-9.4 min, P = 0.077). These findings suggest that sustained clinician engagement, rather than algorithmic accuracy, may be the key barrier to effective clinical AI use in emergency departments. They inform randomized trial design but do not justify clinical deployment of AI clinical decision support at this stage. ClinicalTrials.gov identifier: NCT06902675 .