Related Experiment Video
Updated: Aug 16, 2026

04:16
Fine-Tuning Large Language Models Using Entity Hallucination Index for Text Summarization
Published on: January 9, 2026
Leveraging Large Language Models to Summarize Data and Free-Text Comments From Resident Assessments
Carisa M Cooney1, Rebecca Slattery2, Damon S Cooney3
1is an Associate Professor and Director of Education Innovation, Department of Plastic and Reconstructive Surgery, Johns Hopkins University School of Medicine, Baltimore, Maryland, USA.
Journal of Graduate Medical Education
|August 15, 2026
Summary
Large language models (LLMs) can feasibly label resident assessment comments, but show higher recall for "Strengths" than "Areas for Improvement." This technology aids in summarizing feedback for competency committees.
Area of Science:
- Medical Education Technology
- Artificial Intelligence in Healthcare
- Clinical Competency Assessment
Background:
- Summarizing resident assessments for Clinical Competency Committee or mentoring meetings is challenging.
- Free-text comments in assessments offer valuable, yet difficult-to-process, feedback.
Purpose of the Study:
- To assess the feasibility and accuracy of using large language models (LLMs) for labeling free-text comments.
- To categorize comments as "Strengths" or "Areas for Improvement" from resident end-of-rotation assessments.
Main Methods:
- A mixed-methods study utilized de-identified resident assessments (July 2024-June 2025).
- A paid LLM (ChatGPT-4o) labeled free-text comments; confirmatory trials used free LLMs (MS CoPilot, ChatGPT-4).
- Feasibility assessed by data de-identification time, LLM processing time, and LLM recall; accuracy determined by expert comparison.
Main Results:
- LLM processing and labeling took <7 minutes per LLM, with data de-identification around 60 minutes.
- LLMs achieved 67.8% recall for "Strengths" and 23.9% for "Areas for Improvement."
- Human analysis identified 81.5% "Strengths" and 18.5% "Areas for Improvement" in 476 comments.
Conclusions:
- Prompts were designed to feasibly label end-of-rotation assessment free-text comments across three LLMs.
- LLMs demonstrated a 2.5 times higher recall rate for "Strengths" compared to "Areas for Improvement."
