Related Experiment Videos
Clinical Utility of LLM-assisted Chart Review for the Detection of Bleeding Events
Merijn C Reuland1, Oscar M van der Meer2, Alberto Testoni3
1Department of Intensive Care Medicine, Amsterdam UMC, Meibergdreef 9, Amsterdam, 1105 AZ, The Netherlands.
Introduction:
Red blood cell (RBC) transfusions are frequently administered in the intensive care unit (ICU) and are independently associated with increased mortality. While most studies have focused on transfusion events, bleeding events that do not result in transfusion remain understudied. The validated HEME scoring system enables structured bleeding assessment but requires labor-intensive chart review, limiting scalability. We hypothesized that large language models (LLMs) could function as screening instruments to facilitate scalable chart review and support the development of curated bleeding event datasets.
Methods:
This retrospective cohort study included critically ill patients at high risk of bleeding. Reference data on in-ICU bleeding events were derived from prior manual review using the Hemorrhage Measurement (HEME) scoring system. For the LLM-based analysis, unstructured clinical notes were extracted from the electronic health record (EHR) for the same patient cohort. Notes were processed via structured JSON-formatted prompts using the OpenAI gpt-4o-mini model in a secure environment. Next, these bleeding events were adjudicated to come to a verified reference set.
Results:
In 149 patients 36 490 notes were analyzed. A total of 654 adjudicated true bleeding events were identified in the reference set. Manual review detected 66 events, while the LLM detected 647 events. Manual review identified 7 true events missed by the LLM 1.1% (7/654). The LLM identified 588 true events not captured by manual review, corresponding to an incremental detection yield of 90.1% (588/654). Among the 739 events identified by the LLM, 85 were not confirmed upon adjudication, resulting in a false-positive rate of 11.5% (85/739). The LLM correctly classified bleeding site in 42% of cases (272/654) and bleeding severity in 65% of cases (424/654).
Conclusion:
Use of an LLM substantially increased detection of ICU bleeding events compared with manual review. However, an 11.5% false-positive rate and limited accuracy in site and severity classification indicate that expert adjudication remains necessary.