Related Experiment Video
Updated: Jul 27, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
New Artificial Intelligence ChatGPT Performs Poorly on the 2022 Self-assessment Study Program for Urology
Linda My Huynh1, Benjamin T Bonebrake2, Kaitlyn Schultis2
1MD/PhD Scholars Program, University of Nebraska Medical Center, Omaha, Nebraska.
Large language models like ChatGPT showed limited accuracy on the American Urological Association exam. Despite improvements with multiple-choice questions, incorrect answers were consistently justified, posing a risk of medical misinformation.
Area of Science:
- Medical Education
- Artificial Intelligence in Medicine
- Urology
Background:
- Large language models (LLMs) show promise in various fields, but their utility in medical education requires thorough evaluation.
- The American Urological Association (AUA) Self-assessment Study Program serves as a key educational resource for urology professionals.
Purpose of the Study:
- To assess the effectiveness of ChatGPT as an educational tool for urology trainees and practicing physicians using the AUA Self-assessment Study Program.
- To evaluate ChatGPT's accuracy and consistency in answering urological assessment questions.
Main Methods:
- 135 questions from the 2022 AUA Self-assessment Study Program (excluding visual items) were used.
- Questions were categorized as open-ended or multiple-choice.
- ChatGPT responses were evaluated for correctness, with regeneration up to two times for indeterminate answers.
- Accuracy, concordance, and quality were assessed by independent researchers and physician adjudicators.
Main Results:
- ChatGPT achieved low accuracy: 26.7% for open-ended and 28.2% for multiple-choice questions.
- Indeterminate responses were frequent (29.6% open-ended, 3.0% multiple-choice).
- Regeneration did not significantly improve the proportion of correct answers, and ChatGPT consistently provided justifications for incorrect responses.
Conclusions:
- ChatGPT's performance on the AUA Self-assessment Study Program did not meet expectations, unlike its performance on general medical licensing exams.
- Accuracy was higher for multiple-choice questions compared to open-ended ones.
- The consistent generation of incorrect justifications highlights a significant risk of propagating medical misinformation if unchecked.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
05:04Author Spotlight: Evaluating Clinicians' Adoption of Ultrasound-Guided Vascular Cannulation Through Simulation Training
Published on: August 9, 2024
Related Concept Videos
Urologic Endoscopic Procedure: Cystoscopic Examination
Anatomy of the Genitourinary System II: Bladder and Urethra
Urodynamic Studies: Uroflowmetry
Imaging Studies V: Intravenous Urography and Retrograde Pyelography
Imaging Studies I: Kidney, Ureter, and Bladder Studies