Related Experiment Video
Updated: Aug 9, 2026

04:22
Treatment Model for Young Patients with Psychogenic Erectile Dysfunction and Resultant Infertility
Published on: May 30, 2025
Quality assessment of artificial intelligence responses in erectile dysfunction: a comparative study based on EAU
Gökhan Çeker1,2, Afonso Morgado3, Giorgio Ivan Russo4
1Department of Urology, Basaksehir Cam and Sakura City Hospital, Istanbul, Turkey. drgokhanceker@gmail.com.
International Journal of Impotence Research
|August 7, 2026
Summary
This study compared AI language models for interpreting urology guidelines. Gemini 2.5 Pro and the EAU Guidelines Bot showed the highest performance, suggesting AI tools can support clinical decisions but need real-world validation.
Area of Science:
- Urology
- Artificial Intelligence
- Clinical Decision Support
Background:
- Artificial intelligence (AI) language models are increasingly used for clinical guideline interpretation.
- The European Association of Urology (EAU) introduced a guideline-specific chatbot, but its performance against general large language models (LLMs) is unknown.
Purpose of the Study:
- To comparatively evaluate the performance of five AI systems, including the EAU Guidelines Bot and general LLMs, in interpreting EAU erectile dysfunction (ED) guidelines.
- To assess AI model responses based on relevance, clarity, structure, clinical utility, and factual accuracy.
Main Methods:
- Five AI systems (EAU Guidelines Bot, ChatGPT-5, Gemini 2.5 Pro, Copilot - Smart GPT-5, Perplexity Pro) were tested with 13 clinical questions from EAU ED guidelines.
- Responses were evaluated by three senior reviewers using a 5-point Likert scale across five domains.
- Statistical analysis included inter-rater reliability (ICC) and comparative tests (Friedman, Wilcoxon).
Main Results:
- Significant performance differences were found across all domains (p < 0.001).
- Gemini 2.5 Pro (4.60) and the EAU Guidelines Bot (4.53) achieved the highest composite scores, followed by ChatGPT-5 (4.27).
- Top models demonstrated high factual accuracy, but clarity, structure, and clinical utility varied.
Conclusions:
- Both guideline-specific AI and advanced general-purpose LLMs can provide responses consistent with guideline recommendations for structured erectile dysfunction scenarios.
- Variability in performance across domains suggests AI models should be viewed as supportive tools, not definitive decision-makers.
- Further validation in real-world clinical settings is necessary.