Related Experiment Videos
Comparison of physician-authored and artificial intelligence-generated after-visit summaries: A blinded comparative
Milla Kviatkovsky1, Caden Stewart2, Annalise McDonald1
1Department of Medicine, Division of Hospital Medicine, University of California San Diego, La Jolla, California, USA.
Background:
After-visit summaries (AVSs) are essential for a safe hospital discharge, yet are often written above literacy levels, omit key information, and are produced under substantial clinical time pressure. Large language models (LLMs) offer a potential solution, while performance and safety in real clinical workflows remain uncertain.
Objective:
To compare the quality, understandability, actionability, and safety of LLM-generated versus physician-authored after-visit summaries for hospitalized patients.
Methods:
We conducted a retrospective, blinded comparison of physician-authored AVSs and LLM-generated AVSs for 50 adults discharged from the University of California San Diego Health hospital medicine service in 2023. The physician-authored hospital course served as source text for generating AVSs using GPT-4 and Gemma 3n 2. Both were prompted to produce sixth-grade level, patient-centered AVSs. Five attending physicians independently evaluated each AVS using the Patient Education Materials Assessment Tool (PEMAT) for understandability and actionability and the AVSrubric, an instrument assessing accuracy, comprehensiveness, clarity, consistency with the medical record, tone and empathy, and potential for harm.
Results:
LLM-generated AVSs had higher PEMAT scores than physician-authored AVSs (understandability: 85.5% (GPT-4), 87.5% (Gemma), and 66.1% (physician-authored); actionability: 70.9% (GPT-4), 74.1% (Gemma) vs. 56.7% (physician-authored); all p < .001). On the AVSrubric, LLM-generated AVSs received higher ratings than physician-authored AVSs across all domains with GPT-4 demonstrating significantly higher in all five, while Gemma showed significant improvements in clarity, readability, tone, and empathy.
Conclusions:
In this blinded evaluation, LLM-generated AVSs were clearer, more comprehensive, and more empathetic than physician-authored AVSs and were associated with lower physician-rated potential for harm. A physician-in-the-loop LLM workflow may improve discharge communication while reducing clinician burden and warrants prospective evaluation.
Related Concept Videos
Blind Procedures
Blinding