Related Experiment Video
Updated: Mar 20, 2026

Hydra, a Computer-Based Platform for Aiding Clinicians in Cardiovascular Analysis and Diagnosis
Published on: September 26, 2018
Assessment of ChatGPT-4.0 versus ChatGPT-Mini in Generating Guideline-Based Hypertension Content
Rômullo José Costa Ataídes1, Marcos Adriano Garcia Campos2, João Vítor Perez de Souza2
1Faculdade de Medicina da Universidade de São Paulo, São Paulo, SP - Brasil.
ChatGPT-4.0 showed minor improvements over ChatGPT-Mini in generating hypertension education content, but differences were not statistically significant. ChatGPT-Mini provided more consistent responses, highlighting the need for AI validation in medical education.
Area of Science:
- Artificial Intelligence in Healthcare
- Medical Education Technology
- Natural Language Processing
Background:
- AI language models are increasingly used for patient education materials.
- Concerns exist regarding the accuracy, completeness, and guideline adherence of AI-generated content.
- Evaluating AI performance in medical contexts is crucial.
Purpose of the Study:
- Compare ChatGPT-4.0 and ChatGPT-Mini for generating hypertension education content.
- Assess accuracy, completeness, structural quality (EQIP), response consistency, and guideline alignment.
- Determine the clinical utility of advanced AI models in patient education.
Main Methods:
- 31 standardized hypertension questions posed to ChatGPT-4.0 and ChatGPT-Mini.
- Independent blinded clinician evaluation using modified EQIP scores, accuracy, and completeness scales.
- BERTScore for response consistency and Wilcoxon rank-sum tests for comparisons.
Main Results:
- ChatGPT-4.0 showed slight advantages in accuracy, completeness, and EQIP scores, but differences were small and not statistically significant.
- Effect sizes (Hodges-Lehmann, Cliff's delta) were modest, with 95% CIs crossing zero.
- ChatGPT-Mini demonstrated higher response consistency (BERTScore F1).
Conclusions:
- ChatGPT-4.0 offers marginal improvements over ChatGPT-Mini for hypertension education content.
- The modest effect sizes and non-significant differences emphasize the need for careful AI evaluation.
- Consistent response generation by ChatGPT-Mini and the importance of reporting effect sizes are key takeaways.
Related Concept Videos
Pre-Procedural Guidelines for Assessing Blood Pressure
Chronic Kidney Disease III: Interprofessional Care
Peripheral Artery Disease III: Interprofessional Care
Hypertension III: Clinical Manifestations and Diagnostic Studies
Hypertension and Regulation of Blood Pressure
Hypertension IV: Drug Therapy and Lifestyle Modifications
