Evaluating GPT-4 Responses on Scars or Keloids for Patient Education: Large Language Model Evaluation Study
Mingjun Rao1, Tang Xiujun1, Wang Haoyu1
1Department of Plastic Surgery, Guizhou Provincial People's Hospital, 83 Zhongshan East Road, Nanming District, Guiyang, 550002, China, 86 15343315902.
JMIR Medical Informatics
|March 3, 2026
Summary
Large language models like GPT-4 show promise for educating patients about scars and keloids, providing accurate and reliable information. However, improving readability and reducing AI-generated reference errors are crucial for patient trust and accessibility.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Patient Education
Background:
- Scars and keloids cause significant physical and psychological distress, impacting patients' quality of life.
- Patients seek online information, but existing resources are often unreliable or difficult to comprehend.
- Large language models (LLMs) offer potential for medical information delivery but require validation for accuracy and patient education.
Purpose of the Study:
- To systematically evaluate GPT-4's performance in providing patient education on scars and keloids.
- Focus areas included accuracy, reliability, readability, and the quality of references generated by GPT-4.
- Assess GPT-4's utility as a patient education tool for scar and keloid management.
Main Methods:
- Collected 354 questions on scars and keloids from Reddit communities and medical websites.
- Input questions into GPT-4 to simulate real-world patient interactions.
- Evaluated GPT-4 responses using AI-powered tools (PEMAT-AI, DISCERN-AI, GQS, REA) and expert surgeon review for accuracy, safety, and clinical appropriateness.
Main Results:
- GPT-4 demonstrated high accuracy (average 3.9/5) and reliability, with good understandability (75.5%) and information quality (4.28/5).
- Readability was moderate (12th-grade level), indicating a need for simplification.
- 11.8% of references were hallucinated, though 95.1% of real references were from authoritative sources.
Conclusions:
- GPT-4 shows significant potential as a patient education tool for scars and keloids, delivering accurate and reliable information.
- Enhancements in readability (targeting 6th-8th grade level) and reduction of reference hallucinations are necessary.
- Optimizing LLMs for simplified language and robust reference validation will maximize their clinical utility for patient education.


