Related Experiment Video
Updated: Sep 13, 2026

Fine-Tuning Large Language Models Using Entity Hallucination Index for Text Summarization
Published on: January 9, 2026
Optimizing generative artificial intelligence for clinical summarization: a blinded comparison study of automated
David A Dorr1, Nicole G Weiskopf1, Jean M Sabile1
1Division of Informatics, Clinical Epidemiology, and Translational Data Science, Oregon Health & Science University, Portland, OR 97239, United States.
Background:
Data, information, and knowledge in health care has expanded exponentially over the past 50 years, leading to significant challenges with information overload and complex, fragmented care plans. Generative artificial intelligence (AI) has the potential to facilitate summarization and integration of knowledge to enable efficient care planning and reduce cognitive load of clinicians.
Objective:
To determine the value of AI generated summarization through optimization of short synopsis creation at care transitions.
Design:
Blinded, randomized comparison study comparing human- vs AI-generated synopses using the data-information-knowledge-wisdom framework.
Participants:
De-identified records of 64 patients with multiple chronic conditions from the Medical Information Mart for Intensive Care III database.
Main Measures:
Accuracy, succinctness, synthesis, and usefulness of synopses using a standardized scale with scores > 80% indicating success.
Key Results:
Artificial intelligence and clinicians summarized 64 patients with 12% overlap. In blinded trials, AI synopses were rated as useful 75% of the time vs 76% for human synopses. AI had lower succinctness ratings for the data synopsis task (55%-67%). For accuracy and synthesis, AI had near equal or better scores in other domains (AI: 72%-79%, humans: 68%-84%), with best scores from AI in wisdom. Interrater agreement was variable.
Conclusions:
Artificial intelligence-created synopses that were nearly equivalent to human-created ones; they were slightly longer and did not always synthesize individual data elements compared to humans. We showed that expert-guided optimization of an open-source large language model can lead to synopses that match those produced by humans, reducing cognitive load.