Comparing ChatGPT and physicians' answers to endometriosis questions on Reddit: A blind expert evaluation
Clémence Beaulieu1, Aubert Agostini2, Patrice Crochet3
1Department of Gynecology and Obstetrics, AP-HM, Assistance Publique-Hôpitaux de Marseille, Marseille, France.
Objectives:
To compare the perceived quality, safety, and relevance of ChatGPT responses to those provided by verified physicians on Reddit, a large online discussion platform, in response to questions related to endometriosis.
Methods:
We selected 30 endometriosis-related questions posted on Reddit's r/AskDocs forum, each answered by a verified physician. Using the same question prompts, ChatGPT (GPT-3.5) generated matched-length responses. Responses were anonymized, randomized (A/B format), and assessed blindly by three university-affiliated physicians using a 11-item Likert-scale questionnaire covering medical accuracy, safety, clarity, empathy, and alignment with best practices. Evaluators also indicated which response they considered most pertinent and whether they suspected AI authorship.
Results:
ChatGPT responses were rated significantly higher than physicians' responses on most items, including medical coherence (mean 3.89 ± 0.89 vs. 3.08 ± 0.92), clarity (3.93 ± 0.95 vs. 3.04 ± 0.99), and empathy (3.91 ± 0.93 vs. 2.76 ± 1.09), all with p-values < 0.001. Experts selected ChatGPT as the most pertinent response in 63.3 % of cases. A substantial proportion of responses from both sources were considered potentially dangerous by at least one expert: 26.7 % for ChatGPT and 60.0 % for physicians (p = 0.019).
Conclusion:
ChatGPT outperformed Reddit physicians on multiple expert-rated criteria, particularly in terms of clarity, empathy, and adherence to clinical recommendations. However, the non-negligible proportion of responses considered potentially dangerous by experts underscores the need for cautious use and appropriate supervision of such technologies.


