Related Experiment Video
Updated: Oct 7, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large language model evaluation and feedback of transcribed simulated physician-patient verbal interactions
1Division of Hospital Medicine, Department of Internal Medicine, College of Medicine, The Ohio State University Wexner Medical Center, Columbus, OH, United States.
Introduction:
Physician-patient verbal interactions are fundamental to clinical care and medical education, yet their evaluation is time-intensive and subject to inter-rater variability. We evaluated whether large language models (LLMs) can accurately assess transcribed standardized patient-physician verbal interactions and generate useful narrative feedback.
Methods:
In this single-center study, 12 standardized patient-medical student encounters were independently evaluated by ChatGPT-4o, Gemini 3 Pro, and a physician reviewer using a gold standard rubric composed of 55 history-taking elements and 19 presentation elements assessing diagnostic reasoning, management planning, and patient counseling. Agreement was assessed using Cohen's kappa and intraclass correlation coefficients. Narrative feedback generated by the LLMs was independently evaluated by ChatGPT, Gemini, Claude, and a physician reviewer.
Results:
Across raters, mean scores were similar, with 35-36 of 55 history-taking elements and 13-14 of 19 presentation elements identified as present. Agreement was high for history-taking elements (κ = 0.84-0.89; intraclass correlation coefficients = 0.93-0.97) and lower for presentation elements (κ = 0.62-0.72; intraclass correlation coefficients = 0.61-0.83). Both LLMs generated detailed, accurate, and clinically useful narrative feedback.
Discussion:
These findings demonstrate that LLMs can reliably evaluate simulated physician-patient interactions and may support standardized assessment of clinical communication for research, quality improvement, and medical education.