Related Experiment Video
Updated: Aug 9, 2026

Treatment Model for Young Patients with Psychogenic Erectile Dysfunction and Resultant Infertility
Published on: May 30, 2025
Quality assessment of artificial intelligence responses in erectile dysfunction: a comparative study based on EAU
Gökhan Çeker1,2, Afonso Morgado3, Giorgio Ivan Russo4
1Department of Urology, Basaksehir Cam and Sakura City Hospital, Istanbul, Turkey. drgokhanceker@gmail.com.
Abstract:
Artificial intelligence (AI)-based language models are increasingly explored as tools for interpreting and applying clinical guideline recommendations. In urology, the European Association of Urology (EAU) recently introduced a guideline-specific chatbot; however, its comparative performance relative to contemporary general-purpose large language models (LLMs) remains unclear. In this structured comparative study, five AI systems-the EAU Guidelines Bot, ChatGPT-5, Gemini 2.5 Pro, Copilot - Smart GPT-5, and Perplexity Pro-were evaluated using 13 clinical questions derived directly from strongly recommended statements in the EAU erectile dysfunction (ED) guidelines. Responses were independently assessed by three senior reviewers across five predefined domains: relevance, clarity, structure, clinical utility, and factual accuracy, using a 5-point Likert scale. The primary outcome of the study was the composite performance score, which was calculated as the mean of the five domain scores. Inter-rater reliability was calculated using ICC(2,k), and differences among models were analyzed with the Friedman test followed by Holm-adjusted Wilcoxon post-hoc comparisons. Significant performance differences were observed across all domains (all p < 0.001). The highest composite scores were observed for Gemini 2.5 Pro [4.60 (4.40-4.73)] and the EAU Guidelines Bot [4.53 (4.47-4.80)], followed by ChatGPT-5 [4.27 (4.07-4.47)]. Lower composite scores were observed for Copilot - Smart GPT-5 [3.73 (3.40-3.87)] and Perplexity Pro [3.60 (3.47-3.80)]. Domain-level analysis showed consistently high median scores (≥ 4) for factual accuracy among top-performing models, whereas variability was more pronounced in clarity, structure, and clinical utility. These findings suggest that both guideline-specific systems and advanced general-purpose LLMs may generate responses broadly consistent with guideline-based recommendations in structured ED scenarios. However, variability across domains-particularly in structure and clinical utility-and modest differences in composite performance suggest that these models should be interpreted as supportive tools rather than definitive clinical decision-making systems, requiring further validation in real-world settings.