Implementing GPT-4 Learning Models in Dermatology: An Assessment of Medical Quality and Utility
Aryan Naik1, Peter Vien1, Thanh-Nga Tran2
1Larner College of Medicine, The University of Vermont, Burlington, Vermont, USA.
Summary
Large language models (LLMs) like GPT-4 show potential in dermatology but require validation. While generally safe, GPT-4 outputs scored poorly on medical quality assessments, indicating a need for further optimization.
Area of Science:
- Artificial Intelligence in Medicine
- Dermatology AI Applications
Background:
- Large language models (LLMs) like GPT-4 can generate clinical information.
- Assessing the medical quality and accuracy of GPT-4 dermatology outputs is lacking.
Purpose of the Study:
- Evaluate the medical quality and treatment recommendation accuracy of GPT-4 models in dermatology.
- Compare GPT-4 outputs (ChatGPT-4, Copilot) against UpToDate (UTD) for 33 dermatologic conditions.
Main Methods:
- GPT-4 models generated summaries and treatments for 33 conditions.
- DISCERN scores assessed medical quality; concordance with UTD treatments was evaluated by a dermatologist.
- Statistical analyses included paired t-tests and ANOVA.
Main Results:
- UTD content quality was rated "fair," while ChatGPT-4 and Copilot were rated "poor" by DISCERN.
- ChatGPT-4 treatment recommendations showed significantly higher concordance (33.5%) with UTD compared to Copilot.
- GPT-4 models produced few harmful recommendations.
Conclusions:
- GPT-4 models show promise as dermatologic tools but require validation.
- LLM parameters and query structures may be optimized for dermatology.
- Future LLMs, used with dermatologist judgment, could enhance patient care and save time.

