Related Experiment Video
Updated: Jan 18, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Designing Patient-Centered Communication Aids in Pediatric Surgery Using Large Language Models
Arya S Rao1, Aneesh Mazumder2, Elizabeth Roux1
1Harvard Medical School, Boston, MA, United States; Medically Engineered Solutions in Healthcare Incubator, Innovation in Operations Research Center, Mass General Brigham, Boston, MA, United States.
Introduction:
Large language models (LLMs) have been shown to translate information from highly specific domains into lay-digestible terms. Pediatric surgery remains an area in which it is difficult to communicate clinical information in an age-appropriate manner, given the vast diversity in language comprehension levels across patient populations and the complexity of procedures performed. This study evaluates LLMs as tools for generating explanations of common pediatric surgeries to increase efficiency and quality of communication.
Methods:
Two generalist LLMs (GPT-4-turbo [OpenAI] and Gemini 1.0 Pro [Google]; accessed March 2024) were provided the following prompt: "Act as a pediatric surgeon and explain a [PROCEDURE] to a [AGE] old [GENDER] in age-appropriate language. Discuss indications for the procedure, steps of the procedure, possible complications, and post-operative recovery." Responses were generated for 4 common pediatric surgeries (appendectomy, umbilical hernia repair, cholecystectomy, and gastrostomy tube placement) for male and female children of ages 5, 8, 10, 13, and 16 years. Forty responses from each LLM were rated for accuracy, completeness, age-appropriateness, possibility of demographic bias, and overall quality by two pediatricians and two general surgeons using a five-point Likert scale. Numeric ratings were summarized as means and 95 % confidence intervals. An ordinal mixed-effects model with rater as a random effect was used to account for clustering by rater. P < 0.05 was considered statistically significant.
Results:
Responses from GPT-4-turbo and Gemini 1.0 Pro models were both rated with moderately high overall quality (GPT4: 3.97 [3.82, 4.12]; Gemini 1.0 Pro: 3.39 [3.20, 3.57]) and moderately low possibility of demographic bias (GPT4: 2.49 [2.38, 2.60]; Gemini 1.0 Pro: 2.93 [2.79, 3.07]). GPT-4-turbo responses were rated as highly accurate (4.18 [4.05, 4.32]), highly complete (4.21 [4.10, 4.33]), and highly age-appropriate (4.10 [3.96, 4.24]), while Gemini 1.0 Pro responses were rated as moderately accurate (3.83 [3.70, 3.96]), moderately complete (3.95 [3.83, 4.07]) and moderately age-appropriate (3.63 [3.47, 3.79]). With GPT-4-turbo, ratings on most measures tend to improve as patient age increases, whereas with Gemini 1.0 Pro, they tend to worsen as patient age increases. Ratings on all measures, with the exception of age-appropriateness, were slightly higher for responses generated for male patients as compared to female patients with GPT-4-turbo, while the gender differences were less pronounced with Gemini 1.0 Pro.
Discussion:
This study demonstrates that off-the-shelf LLMs have the potential to produce accurate, complete, and age-appropriate explanations of common pediatric surgeries with low possibility of demographic bias. Inter-model variability in areas such as quality of response, age-appropriateness and gender differences were also observed, signaling the need for additional validation and fine-tuning based on the clinical content. Such tools could be implemented at the point of care or in other patient education settings and personalized to ensure effective, equitable communication of pertinent medical information with demonstration of clinician-rated content quality.
Study Type:
This is a pilot study evaluating the performance of large language models (LLMs) as patient-centered communication aids in pediatric surgery.
Level Of Evidence:
Level IV (pilot study).
More Related Videos
Related Concept Videos
Barriers to Effective Communication II
Cultural barriers:
Differences in values, beliefs, religion, knowledge, and tradition can significantly impact communication. Awareness of nonverbal cues is critical, especially when conversing with a patient from a different culture. What appears appropriate in one culture may be inappropriate in another.
Semantic barriers:
As a result of their tendency to use...
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Role of Communication in the Nursing Process II: Planning and Implementation

