Related Experiment Video For ChatGPT
Updated: Jan 16, 2026

Technical Modification of the Terminal Ureter During Total Transperitoneal Laparoscopic Nephroureterectomy for Upper Urinary Tract Urothelial Carcinoma
Published on: November 22, 2019
Expert Evaluation of ChatGPT-4 Responses to Upper Tract Urothelial Carcinoma Questions: A Prospective Comparative
Murat Beyatlı1, Hasan Samet Güngör1, Abdurrahman İnkaya1
1Department of Urology, Umraniye Training and Research Hospital, 34764 Istanbul, Turkey.
Abstract:
Background/Objectives: This study aimed to systematically evaluate the accuracy and clinical relevance of ChatGPT-4's answers to a set of questions on upper tract urothelial carcinoma (UTUC) as adjudged by expert urologists. To juxtapose performance, one question set consisted of queries derived from the 2025 EAU (European Association of Urology) guidelines, and the other was a miscellaneous set of frequently asked patient-oriented questions. Methods: Seventy-seven questions were posed to ChatGPT-4 in English: 60 were systematically selected from the 2025 EAU UTUC guidelines, and 17 were compiled as frequently asked questions (FAQs) from reputable urology sources. Two board-certified urologists independently scored the given responses, employing (a) binary scoring (correct = 1, incorrect = 0) and (b) detailed accuracy scoring (1 = completely accurate, 2 = accurate but inadequate, 3 = mixed accurate-misleading, 4 = completely inaccurate). Comparative analyses used the Mann-Whitney U test with effect size estimation. Results: Overall, 71 of 77 responses (92.2%) were correct. Accuracy rates were 90.0% (54/60) for EAU guideline questions and 100.0% (17/17) for the FAQs. The mean accuracy score for the guideline-based questions was 1.28 ± 0.74, compared with 1.00 ± 0.00 for the FAQs. Differences between the groups were not statistically significant (p = 0.094, r = 0.191). A subgroup analysis showed perfect accuracy (100%) in four EAU categories-Classification and Staging Systems, Diagnosis, Disease Management, and Metastatic Disease Management-while the Follow-up category had the lowest accuracy (25% correct, mean score = 2.75), indicating domain-specific limitations. Conclusions: ChatGPT-4 demonstrated high overall accuracy, particularly for patient-oriented UTUC questions, but showed reduced reliability in complex, guideline-specific areas, especially follow-up protocols. The model shows promise as an educational tool for patients but cannot replace expert clinical judgment for decision-making. These findings have important implications for the integration of AI tools in urological practice and highlight the need for domain-specific optimization.
More Related Videos
Related Concept Videos
Urologic Endoscopic Procedure: Cystoscopic Examination
Urinary Tract Infection III: Diagnostic Studies and Interprofessional Care
Imaging Studies V: Intravenous Urography and Retrograde Pyelography
Urinary Tract Calculi III: Medical Management

