Related Experiment Video
Updated: Jun 27, 2026

E-Patient Counseling Trial (E-PACO): Computer Based Education versus Nurse Counseling for Patients to Prepare for Colonoscopy
Published on: August 1, 2019
Comparative Evaluation of ChatGPT-5.4 Mini and Gemini 3 Flash in Scar Patient Education: Quality, Reliability, and
Tianying Zang1, Xiujuan Yan, Yong Dong
1Department of Aesthetic Plastic Surgery and Laser Medicine, Beijing Anzhen Hospital Affiliated to Capital Medical University, Beijing, China.
Background:
Scar is an inevitable pathological product of tissue injury repair, and pathological scars often occur in exposed areas, bringing severe psychological burden and economic losses to patients. With the popularization of digital healthcare, patients increasingly rely on artificial intelligence (AI) for self-consultation, but the core capabilities of free generative AI in scar management have not been systematically evaluated.
Objective:
This study compared and evaluated the comprehensive performance of ChatGPT-5.4 mini and Gemini 3 Flash in answering clinical and psychological questions of scar patients, investigated multi-dimensional differences, and provided support for the application of AI in patient education.
Methods:
Fifteen core questions from scar patients were extracted and input into ChatGPT-5.4 mini and Gemini 3 Flash, respectively. The DISCERN-AI scale and Global Quality Scale (GQS) were used for evaluation, while multiple standardized tools were applied to quantify text readability and complexity. All data were subjected to a normality test and difference analysis using SPSS software.
Results:
Both models demonstrated high clinical reliability, with no significant difference in target topic clarity (P=0.806). ChatGPT had better overall quality, with a GQS score of 4.8 (4.5, 4.9), which was significantly higher than Gemini's 4.6 (4.4, 4.7) (P=0.033). ChatGPT was also more rigorous in stating medical limitations and uncertain treatment options (5.0 versus 4.5, P<0.05). In contrast, Gemini performed better in patient demand relevance and empathy (4.5 versus 4.0, P=0.026). Both models achieved moderate scores in shared decision-making support. Readability analysis showed that the reading thresholds of both models were excessively high, far exceeding the internationally recommended 6th- to 8th-grade standard for patient education materials.
Conclusion:
ChatGPT-5.4 mini and Gemini 3 Flash have complementary advantages and potential as auxiliary tools for digital health education in scar patients, but both have a serious readability gap. For future large-scale applications, readability prompt intervention should be introduced, and it should be clearly stated that AI cannot replace professional diagnosis and treatment to ensure the inclusiveness and safety of digital medical information.