Related Experiment Video
Updated: Jun 20, 2026

Clinical-oriented Three-dimensional Gait Analysis Method for Evaluating Gait Disorder
Published on: March 4, 2018
Evaluating Chat Generative Pre-trained Transformer Responses to Common Pediatric In-toeing Questions
Jason Zarahi Amaral1, Rebecca J Schultz, Benjamin M Martin
1Department of Orthopaedic Surgery, Texas Children's Hospital and Baylor College of Medicine, Houston, TX.
Objective:
Chat generative pre-trained transformer (ChatGPT) has garnered attention in health care for its potential to reshape patient interactions. As patients increasingly rely on artificial intelligence platforms, concerns about information accuracy arise. In-toeing, a common lower extremity variation, often leads to pediatric orthopaedic referrals despite observation being the primary treatment. Our study aims to assess ChatGPT's responses to pediatric in-toeing questions, contributing to discussions on health care innovation and technology in patient education.
Methods:
We compiled a list of 34 common in-toeing questions from the "Frequently Asked Questions" sections of 9 health care-affiliated websites, identifying 25 as the most encountered. On January 17, 2024, we queried ChatGPT 3.5 in separate sessions and recorded the responses. These 25 questions were posed again on January 21, 2024, to assess its reproducibility. Two pediatric orthopaedic surgeons evaluated responses using a scale of "excellent (no clarification)" to "unsatisfactory (substantial clarification)." Average ratings were used when evaluators' grades were within one level of each other. In discordant cases, the senior author provided a decisive rating.
Results:
We found 46% of ChatGPT responses were "excellent" and 44% "satisfactory (minimal clarification)." In addition, 8% of cases were "satisfactory (moderate clarification)" and 2% were "unsatisfactory." Questions had appropriate readability, with an average Flesch-Kincaid Grade Level of 4.9 (±2.1). However, ChatGPT's responses were at a collegiate level, averaging 12.7 (±1.4). No significant differences in ratings were observed between question topics. Furthermore, ChatGPT exhibited moderate consistency after repeated queries, evidenced by a Spearman rho coefficient of 0.55 ( P = 0.005). The chatbot appropriately described in-toeing as normal or spontaneously resolving in 62% of responses and consistently recommended evaluation by a health care provider in 100%.
Conclusion:
The chatbot presented a serviceable, though not perfect, representation of the diagnosis and management of pediatric in-toeing while demonstrating a moderate level of reproducibility in its responses. ChatGPT's utility could be enhanced by improving readability and consistency and incorporating evidence-based guidelines.
Level Of Evidence:
Level IV-diagnostic.
Related Concept Videos
Assessment of apical pulse
Assessing the apical pulse is a critical nursing procedure, particularly indicated for:
Assessment of Respiration
Subjective Assessment: Nurses interview the patient to gather information directly during the subjective assessment. It includes questions about the individual's medical history, medications, and symptoms, focusing on past respiratory conditions like asthma or COPD,...
Physical Assessment of the Respiratory Tract IV: Auscultation
Breath Sounds
Breath sounds are categorized into vesicular, bronchovesicular, and bronchial.

