Related Experiment Video
Updated: Jun 21, 2026

Treatment Model for Young Patients with Psychogenic Erectile Dysfunction and Resultant Infertility
Published on: May 30, 2025
Evaluation of ChatGPT, Gemini, and OpenEvidence in Obstetric and Gynecologic Clinical Decision Scenarios
Arif Onur Atay1, Feride Atay2, Samican Ozmen1
1Torbali State Hospital, Department of Obstetrics and Gynecology, Türkiye, Izmir.
Background:
Clinicians frequently face questions that require rapid, evidence-based answers. Artificial intelligence (AI) tools are increasingly used for this purpose, yet their reliability for clinical decision-making remains uncertain.
Objectives:
This study compared two generative large language model systems (ChatGPT and Gemini) and a retrieval-supported clinical platform (OpenEvidence) to determine which provides the most reliable, clear, and clinically applicable information in obstetrics, gynecology, and urogynecology.
Methods:
A cross-sectional comparative design was used to evaluate ChatGPT (GPT-5), Gemini (Gemini 2.5), and the retrieval-supported platform OpenEvidence. Twenty-four clinical questions across three subspecialties were independently assessed by two blinded specialists using the Expert-Adapted DISCERN (EA-DISCERN) tool, which rates 12 quality domains on a 5-point scale. Mean ± standard deviation scores were compared across systems and clinical domains using repeated-measures analysis.
Results:
OpenEvidence achieved the highest mean total score (54.0 ± 2.3), outperforming Gemini (50.3 ± 2.4) and ChatGPT (48.7 ± 2.4; p < 0.001). OpenEvidence scored significantly higher in evidence-based domains; clinical accuracy, guideline consistency, completeness, transparency, and reliability across all fields. As of this writing, Gemini ranked between the two, showing a modest advantage over ChatGPT in rationale explanation and evidence transparency, whereas both generative models scored higher in language fluency and readability. Overall, total EA-DISCERN scores ranked OpenEvidence highest, followed by Gemini, then ChatGPT. Interrater reliability for the total score was intraclass correlation coefficient [2,1] (absolute agreement = 0.391).
Conclusion:
OpenEvidence provided more guideline-aligned and transparent responses, whereas ChatGPT and Gemini were generally more fluent and readable. For obstetrics and gynecology clinicians, retrieval-supported platforms may be more suitable for point-of-care verification, while generative models should be used more cautiously and with clinician oversight.
Related Concept Videos
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast, controlled...
Clinical Trials: Overview
Gonadal and Placental Hormones
In males, testosterone is the primary gonadal androgen. It plays a central role in the maturation of male reproductive organs — the penis and testes. Additionally, testosterone is instrumental in the development of secondary sexual characteristics — a deep voice as well as facial and pubic hair growth — and...
Standards of Care II
Teratogenicity
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic illness...