Related Experiment Video
Updated: Jan 17, 2026

Treatment Model for Young Patients with Psychogenic Erectile Dysfunction and Resultant Infertility
Published on: May 30, 2025
Evaluation of ChatGPT's performance on answering pediatric urology questions based on association guidelines
Wyatt MacNevin1, Nicholas Dawe2, Laura Harkness2
1Department of Urology, Dalhousie University, Halifax, NS, Canada.
Introduction:
ChatGPT has been shown to provide accurate and complete responses to clinically focused questions, although its ability to successfully answer common pediatric urology-based questions remains unexplored. Furthermore, the concordance of ChatGPT's answers with association recommendations has yet to be analyzed.
Methods:
A list of common pediatric urology questions of varying difficulty was developed in association with publicly available guidelines and resources from the Canadian Urological Association (CUA), American Urological Association (AUA), and the European Association of Urology (EAU). Questions were administered individually using three separate functions, and responses were evaluated for comprehensiveness and accuracy using a Likert scale. Descriptive statistics and analysis of variance were used for statistical analysis.
Results:
ChatGPT performed best in the domain of phimosis (mean ± standard deviation: 2.32/3.00±0.57) and VUR (2.11/3.00±0.63), and worst in acute scrotal pathology (1.90/3.00±0.58) and cryptorchidism (1.92/3.00±0.56) (p=0.031). "Easy" questions (2.31/3.00±0.09) had greater comprehensiveness scores compared to "medium" (1.92/3.00±0.07, p=0.003) and "difficult" questions (1.86/3.00±0.101, p=0.003). Definition-based questions had greater comprehensiveness scores across all guidelines. ChatGPT was more accurate and in concordance with EAU-based information (2.10±0.41) compared to AUA (1.95±0.41, p=0.04).
Conclusions:
ChatGPT answered questions with high levels of appropriateness and comprehensiveness. ChatGPT performed best in the areas of phimosis and VUR and worst in acute scrotal pathology. While ChatGPT performed well across all question domains, it performed best when referenced to EAU and CUA compared to AUA.
More Related Videos
13:44Project-Based Learning Guidelines for Health Sciences Students: An Analysis with Data Mining and Qualitative Techniques
Published on: December 9, 2022
05:04Author Spotlight: Evaluating Clinicians' Adoption of Ultrasound-Guided Vascular Cannulation Through Simulation Training
Published on: August 9, 2024
Related Concept Videos
Nursing Assessment of the Genitourinary System I: Health History
Imaging Studies V: Intravenous Urography and Retrograde Pyelography
Nursing Assessment of the Genitourinary System II: Inspection and Palpation
Chronic Kidney Disease III: Interprofessional Care
Pharmacokinetics in Pediatric Patients: Drug Excretion
Nursing Assessment of the Genitourinary System III: Percussion and Auscultation