Related Experiment Video
Updated: Apr 17, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Leveraging a Large Language Model to Generate Quality Improvement Feedback for Clinical Notes
Christopher J Kim1,2, Joseph Gelfinbein2, Nihan Gencerliler3
1Division of Hospital Medicine, Department of Medicine, NYU Langone Health, New York, New York, United States.
Background:
Poor documentation quality can significantly affect health care operations, but the feedback process for clinicians to improve clinical notes is time-consuming and often insufficient. Large language models (LLMs) such as Generative Pre-trained Transformer 4 (GPT-4) have the potential to streamline this process.
Objectives:
This study aimed to determine whether an LLM can generate feedback to improve the medical contingency and discharge planning (MCDP) component of clinical documentation that is non-inferior to feedback by physicians.
Methods:
A cross-sectional study of GPT-4 feedback and physician feedback on inpatient progress notes was conducted. A random sample of 64 inpatient progress notes identified by the validated artificial intelligence (AI) Audit Tool as having a low likelihood of containing MCDP was included from adult general medicine patients hospitalized at New York University Langone Health (NYULH) in December 2023. Both the GPT-4 model and attending physicians generated feedback on these inpatient progress notes. A/B testing was then conducted on the measures of understandability, usefulness, acceptability, and impartiality. Evaluations employed 5-point Likert scales that were converted to 10-point bidirectional interval scales for interpretability, ranging from -10 (human suggestions significantly better) to +10 (GPT-4 suggestions significantly better), with a non-inferiority threshold set to -1 for the primary endpoint.
Results:
Sixty-four inpatient progress notes were included, representing 55% female patients with a median age of 73 years. GPT-4 feedback was non-inferior to physician feedback in all measures: Understandability (mean, 1.27, 95% CI: 0.73-1.8, p < 0.001), usefulness (mean, 2.09, 95% CI: 1.27-2.91, p < 0.001), acceptability (mean, 2.07, 95% CI: 1.33-2.81, p < 0.001), and impartiality (mean, -0.20, 95% CI: -0.52 to 0.12, p < 0.001).
Conclusion:
This study shows that an LLM can be leveraged to generate note quality feedback that is non-inferior to expert clinician feedback.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
