Related Experiment Video
Updated: Jun 27, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Evaluating the quality and readability of AI-generated information on adenomyosis: a comparative analysis of ChatGPT
Ling Tian1, Yan Wang1,2, Ming-Tao Yang1
1Department of Obstetrics and Gynecology, Beijing Anzhen Nanchong Hospital, Capital Medical University and Nanchong Central Hospital, Nanchong, Sichuan, China.
Purpose:
To systematically evaluate and compare the quality, readability, and query-model consistency of adenomyosis-related content generated by two large language models, ChatGPT (GPT-5) and DeepSeek (R1).
Materials And Methods:
In total, 25 high-frequency patient queries were obtained based on Google Trends. Each query was processed using two interaction modes, namely, three consecutive repetitions and three independent cycles, on both large language models (ChatGPT GPT-5.0-web, released December 2025; DeepSeek R1-web, released November 2025). The generated texts (n = 300) were subsequently assessed for their readability [evaluated by Automated Readability Index (ARI), Flesch Reading Ease Score (FRES), and Gunning Fog Index (GFI)] and quality [assessed by DISCERN score, and Ensuring Quality Information for Patients (EQIP) tool]. Statistical comparisons were performed using non-parametric tests and t-tests.
Results:
In the cyclic mode, both ChatGPT and DeepSeek maintained stable output text readability and quality. DeepSeek-generated text demonstrated significantly superior readability across both interaction modes (lower ARI: 11.32 vs. 14.56, p < 0.001; higher FRES: 46 vs. 27, p < 0.001; lower GFI: 12.47 vs. 14.16, p < 0.001) and higher information quality (higher DISCERN: 62 vs. 43, p < 0.001; higher EQIP: 75 vs. 70, p < 0.001). Under the repetition mode, DeepSeek's output exhibited significant fluctuations across multiple metrics (ARI: p = 0.021; FRES: p = 0.015; GFI: p = 0.004; DISCERN: p = 0.013; EQIP: p < 0.001), while ChatGPT's output remained stable (all p > 0.05). Notably, the readability scores for both models indicated reading levels equivalent to undergraduate education, which is above the recommended level for general public health information.
Conclusion:
The findings of this study demonstrate that when generating information on adenomyosis, DeepSeek outperforms ChatGPT in terms of readability and several information quality metrics, whereas ChatGPT exhibits greater consistency in its outputs. However, the reading difficulty of texts generated by both models exceeds the level suitable for the general public, representing a key practical constraint limiting direct public use. Based on these results, AI chatbots may serve as complementary tools in patient education; however, their outputs should undergo expert review and be optimized for comprehensibility before broader clinical application. For clinicians and patients, these findings emphasize the importance of critically appraising AI-generated information and using it as a supplement to, rather than a substitute for, professional medical consultation.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy

