Related Experiment Videos
Evaluation of an Artificial Intelligence-Driven Language Model for Age-Specific Patient Education in Pediatric
Cole T Payne1, Andrew K Morse1, Brendan T O'Reilly1
1Department of Orthopaedic Surgery, University of Texas Health Science Center at Houston, Houston, TX, USA.
Background:
Describing complex orthopaedic surgical procedures and diagnoses to pediatric patients across a large age range and at varying levels of health literacy can be quite challenging. Generative artificial intelligence (GenAI) offers a potential solution by providing real-time age specific explanations for pediatric patients. The use of GenAI large language models (LLMs) could potentially allow providers to more effectively communicate complex medical problems in an age-appropriate and comprehensible manner. This study aims to validate the readability, accuracy, and comprehensiveness of responses produced by the ChatGPT-4 LLM (OpenAI, San Francisco, CA) when prompted to explain a specific pediatric orthopaedic procedure indicated for a given diagnosis to patients of 3 different age groups.
Methods:
ChatGPT-4 explanations were analyzed for 43 pediatric orthopaedic procedures and tests across prompts for 3 different age groups: 6 years old, 10 years old, and 14 years old. Responses were assessed for readability using 6 quantitative readability indices. The accuracy and comprehensiveness of the responses were rated by 7 board-certified pediatric orthopaedic surgeons, who also compared them to their own explanations used in clinical practice.
Results:
Each of the 6 indices showed a significant difference in readability between the 6-, 10-, and 14-year-old age group prompts (P < .001). The 6-year-old prompts were more readable than the 10- and 14-year-old prompts (P < .001), and the 10-year-old prompts were more readable than the 14-year-old prompts (P < .001) . Out of a 5-point scale, the mean (standard deviation) attending-rated accuracy for the artificial intelligence responses was 4.15 (0.81), with no significant differences in rated accuracy between procedures (P = .051) or age groups (P = .27). The mean attending-rated comfort was 3.95, with inter-rater reliability coefficients for accuracy and comfort of 0.92 and 0.94, respectively.
Conclusions:
GenAI LLMs can generate age-specific explanations for common surgical procedures, treatments, and diagnostic markers in pediatric orthopaedics. Physicians validated the accuracy of its responses and demonstrated a high level of comfort using the GenAI explanations in clinical settings.
Key Concepts:
(1)An artificial intelligence (AI) language model adjusted its explanations based on the child's age, providing simpler language for younger children and more advanced explanations for older children.(2)Across 6 established readability measures, explanations written for younger children were consistently easier to understand than those created for older age groups.(3)Seven board-certified pediatric orthopaedic surgeons found the AI-generated explanations to be highly accurate, with an average rating of 4.15 out of 5.(4)Physicians were generally comfortable with the use of AI-generated educational explanations in practice, and their ratings were highly consistent with one another.(5)When reviewed by a physician, generative AI could be a helpful tool for improving patient education, health literacy, and shared decision-making in pediatric orthopaedic care.
Level Of Evidence:
III/IV.