Related Experiment Video
Updated: Sep 14, 2025

Hemodynamic Precision in the Neonatal Intensive Care Unit using Targeted Neonatal Echocardiography
Published on: January 27, 2023
Assessing Information Provided by ChatGPT: Heart Failure Versus Patent Ductus Arteriosus
Meghana Bhupathi1, Jaza Mehweish Kareem2, Anjali Mediboina3
1Pediatric Medicine, Alluri Sitarama Raju Academy of Medical Sciences, Eluru, IND.
Insights
ChatGPT provides accurate and complete information on heart failure (HF) and patent ductus arteriosus (PDA), with better understandability for PDA. Further refinement of AI language models for medical information is needed.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Patient Education
Background:
- Assessing the accuracy, completeness, and understandability of AI-generated medical information is crucial for patient education.
- Heart failure (HF) and patent ductus arteriosus (PDA) represent conditions with varying prevalence and online information availability.
Purpose of the Study:
- To evaluate the effectiveness of ChatGPT (version 3.5) in delivering accurate, complete, and understandable information on heart failure (HF) versus patent ductus arteriosus (PDA).
- To determine if the commonality of a medical condition influences the quality of AI-generated responses.
Main Methods:
- Utilized 10 open-ended questions for HF and 11 for PDA, assessing ChatGPT 3.5 responses.
- Evaluated accuracy using a 6-point Likert scale and completeness using a 3-point Likert scale.
- Assessed understandability with the Patient Education Materials Assessment Tool for Printable Materials (PEMAT-P) and analyzed data using R Software.
Main Results:
- All questions generated relevant answers on the first attempt.
- HF responses: mean accuracy 5.4, completeness 2.3, understandability 83.75%.
- PDA responses: mean accuracy 5.09, completeness 2.27, understandability 94.5%. No significant differences in accuracy or completeness between HF and PDA were found.
Conclusions:
- ChatGPT demonstrated high accuracy and completeness for both HF and PDA, with superior understandability for PDA information.
- While AI shows promise in medical information dissemination, further refinement is necessary to ensure consistent quality and patient comprehension.
Abstract:
Introduction The study aims to provide insights regarding the effectiveness of ChatGPT in providing accurate, complete, and understandable information related to heart failure (HF) and patent ductus arteriosus (PDA). The study aimed to identify whether the abundance of online resources and patient inquiries leads to ChatGPT providing more accurate, complete, and understandable responses for HF compared to the less common PDA. Methodology Ten open-ended questions related to HF and 11 questions related to PDA were formed, and answers generated in ChatGPT 3.5 were assessed for accuracy via a 6-point Likert scale and for completeness via a 3-point Likert scale. Understandability was assessed using the Patient Education Materials Assessment Tool for Printable Materials (PEMAT-P). Data analysis and visualization was conducted using R Software version 4.3.1. Statistical significance between accuracy and completeness for both the conditions was determined using Spearman's R and Mann-Whitney U tests. Results Relevant answers were generated for all questions on the first attempt. For HF-related answers, the mean accuracy score was 5.4 (SD=0.7) and the mean completeness score was 2.3 (SD=0.7), with a median understandability of 83.75% (IQR=77.8%-94.5%). Spearman's correlation coefficient (rs) for HF was 0.44854 (p=0.19353). For PDA, the mean accuracy was 5.09 (SD=0.7), completeness was 2.27 (SD=0.5), and median understandability was 94.5% (IQR=88.9%-100%). Spearman's correlation coefficient for PDA was 0.21409 (p=0.52731). Furthermore, the Mann-Whitney U-test revealed no significant differences in accuracy (p=0.35758) or completeness (p=0.88866) between the two conditions. Conclusions While the accuracy and completeness of answers related to HF were slightly higher compared to those for PDA, the understandability of PDA responses was notably better. There is a need for refining artificial intelligence-driven language models for medical information dissemination.
Related Concept Videos
Heart Failure IV: Classification and Diagnostic Evaluation
Mitral Stenosis IV: Nursing Management
Cardiomyopathy II: Dilated Cardiomyopathy
Cardiac Catheterization II: Right Heart Catheterization
Aortic Regurgitation II: Clinical Features and Diagnostic Tests
Rheumatic Heart Disease II: Clinical Manifestations and Diagnostic Studies

