Related Experiment Video
Updated: Aug 13, 2026

Translational Brain Mapping at the University of Rochester Medical Center: Preserving the Mind Through Personalized Brain Mapping
Published on: August 12, 2019
Efficacy of large language models in detecting postoperative delirium from unstructured clinical notes: A
Da-In Eun1,2,3, Hyung-Chul Lee1,2,3, Gang Heo2
1Interdisciplinary Program in Medical Informatics, Seoul National University, Seoul, Republic of Korea.
Abstract:
Early identification of postoperative delirium (POD) remains challenging. This retrospective observational study compared the performance of large language models (LLMs), Llama-3-70B and GPT-4o, and physicians in predicting clinically significant POD, defined as either requiring antipsychotics or diagnosis of delirium by neurologists following consultation for delirium-related symptoms. The c-statistics of Llama-3-70B and GPT-4o were 0.74 and 0.76, respectively. LLMs showed higher sensitivity (Llama-3-70B, 0.900; GPT-4o, 0.868; physicians, 0.723) and lower specificity (0.463, 0.547, and 0.814, respectively) than physicians. Inter-rater agreement was almost perfect for both Llama-3-70B and GPT-4o (Fleiss' kappa = 0.852 and 0.854, respectively) but fair for physicians (0.219). Both LLMs detected clinically significant POD approximately one day earlier than physicians (Kaplan-Meier analysis, median time to diagnosis: Llama-3-70B, 34.5 h; GPT-4o, 37.5 h; physicians, 62.9 h; log-rank P < 0.001). The integration of LLMs as a complementary screening tool under physician supervision may improve the early, reproducible diagnosis of clinically significant POD.
Related Concept Videos
Language and Cognition
Automatic Processing and Automatic Social Behavior

