Related Experiment Video
Updated: Jan 6, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Evaluating Artificial Intelligence-Generated Patient Education Materials for Bariatric Surgery: Comparative Analysis
Shuai Guo1, Cheng-Li Yang2, Xiang-Ping Lin2
1Department of Pharmacy, The Affiliated Hospital of Guizhou Medical University, Guiyang, Guizhou, 550004, China. guoshuai@gmc.edu.cn.
Background:
Artificial intelligence (AI) models such as ChatGPT and DeepSeek have gained increasing attention for their potential to enhance patient education by delivering accessible and evidence-based health information. We designed the following study to evaluate the AI models-ChatGPT and DeepSeek-in generating patient education materials for bariatric surgery.
Methods:
Thirty commonly asked patient questions related to bariatric surgery were classified into four thematic domains: (1) surgical planning and technical considerations, (2) preoperative assessment and optimization, (3) postoperative care and complication management, and (4) long-term follow-up and disease management. Responses generated by ChatGPT and DeepSeek were evaluated using three key metrics: (1) response quality, assessed by the Global Quality Score, rated on a 5-point scale from 1 (poor) to 5 (excellent); (2) reliability, measured using modified DISCERN criteria, which assess adherence to clinical guidelines and evidence-based standards, with scores ranging from 5 (low) to 25 (high); and (3) readability, evaluated using two validated formulas: the Flesch-Kincaid Grade Level and the Flesch Reading Ease Score.
Results:
ChatGPT significantly outperformed DeepSeek in response quality, with a median (IQR) Global Quality Score of 5.00 (4.00, 5.00) vs. 4.00 (4.00, 5.00) (P = 0.002). Higher reliability was also observed in ChatGPT, as reflected by mDISCERN scores across all four domains (median [IQR], 22.0 [21.0, 23.25] vs. 19.7 [19.0, 20.75]; P < 0.001). While no significant difference was found in the Flesch Reading Ease Score (mean [SD], 26.11 [12.84] vs. 20.87 [12.20]; P = 0.110), ChatGPT yielded significantly higher Flesch-Kincaid Grade Level Scores (meaning its text was more complex) (mean [SD], 16.40 [2.43] vs. 13.48 [2.35]; P < 0.001). Both models produced responses at a readability level corresponding to college education.
Conclusions:
ChatGPT provided higher-quality and more reliable responses, while DeepSeek's answers were slightly easier to read. However, both models' answers lacked attention to psychosocial and cultural aspects of patient care, highlighting the need for more empathetic, adaptive AI to support inclusive patient education.
More Related Videos
Related Concept Videos
Assessment of the Gastrointestinal System II: Health Perception Pattern
Health Perception Patterns
Health perception patterns offer valuable insights into a patient's lifestyle habits and how they may impact their GI health. These patterns include:
Patient-centered Care

