Related Experiment Video
Updated: Jul 1, 2026

A Study on an Intelligent Diagnosis and Treatment Assistant System for Acupuncture in Diminished Ovarian Reserve Based on a Knowledge Graph
Published on: May 29, 2026
Assessment of the Accuracy and Readability of Artificial Generative Intelligence for Patient Questions on Heavy
Olivia Carere1, Lauren Clarfield2, M Alix Murphy3
1Temerty School of Medicine, University of Toronto, Toronto, ON.
Objectives:
To evaluate the quality, patient-centredness, clinician endorsement, and readability of generative artificial intelligence (AI) responses to patient questions about heavy menstrual bleeding (HMB) compared with high-quality patient-facing websites.
Methods:
A cross-sectional study compared responses generated by ChatGPT and Google Gemini to excerpts from the 5 highest-quality HMB patient resources identified through a Google Trends-informed search and the QUality Evaluation Scoring Tool (QUEST). Five layperson-style questions representing HMB subtopics (definition, causes, investigations, management, and safety) were submitted to each model. Responses were de-identified and independently evaluated by 5 gynecologists, blinded to source, using 5-point Likert scales for accuracy/comprehensiveness (quality), empathetic and validating communication style (patient-centredness), and expert comfort with patient use of the resource to guide understanding (clinician endorsement). Readability was assessed using the Flesch-Kincaid Grade Level. Inter-rater reliability was measured using intraclass correlation coefficients.
Results:
Thirty-six responses were reviewed (16 AI-generated and 20 web-based). Compared with web-based excerpts, AI responses showed similar quality (3.94/5 ± 0.70 vs. 3.53/5 ± 1.21; P = 0.21), lower patient-centredness (2.81/5 ± 0.82 vs. 3.52/5 ± 1.13; P = 0.04), and comparable clinician endorsement (3.56/5 ± 0.92 vs. 3.44/5 ± 1.35; P = 0.52). Without impacting content quality, AI responses were written at a lower grade level (7.6 ± 2.4) than web-based excerpts (11.1 ± 5.2; P = 0.042). Inter-rater reliability was high (intraclass correlation coefficient = 0.82-0.84).
Conclusions:
Compared with high-quality web-based resources, AI-generated responses to HMB questions were more inclusive for varying health literacy levels, comparable in quality and clinician endorsement, but performed worse in patient-centredness. In an evolving digital landscape for health information acquisition, AI-generated responses represent a valuable and safe supplementary or first-line resource for patients.
Related Concept Videos
Assessment of the Gastrointestinal System I: Subjective Data
Health History
The initial step in assessing the GI system is obtaining a comprehensive health history. This includes inquiring about the patient's history or presence of problems related to...
Assessment of the Cardiovascular System I: Subjective Data
Initial Enquiry
Ask the patient about their primary concern and thoroughly explore all reported symptoms.
Medical History
Investigate past illnesses affecting the cardiovascular system, such as angina, anemia, rheumatic fever, congenital heart disease, stroke, thrombophlebitis, dysrhythmias, varicosities
Inquire about symptoms...
Assessment of the Gastrointestinal System II: Health Perception Pattern
Health Perception Patterns
Health perception patterns offer valuable insights into a patient's lifestyle habits and how they may impact their GI health. These patterns include:
Pre-Procedural Guidelines for Assessing Blood Pressure
Nursing Assessment of the Genitourinary System I: Health History
