Related Experiment Video
Updated: Jun 23, 2026

11:39
The Goeckerman Regimen for the Treatment of Moderate to Severe Psoriasis
Published on: July 11, 2013
39.1K
A Meta-Analysis of ChatGPT's Performance on Dermatology Specialty-Level (Board-Style) Certification Questions
1The School of Medicine, The University of Auckland, New Zealand.
Indian Dermatology Online Journal
|July 25, 2025
Summary
This meta-analysis shows ChatGPT achieved 69.97% accuracy on dermatology board-style questions, with GPT-4.0 outperforming GPT-3.5. It suggests ChatGPT as a valuable educational tool for dermatology certification preparation.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Education Technology
- Dermatology Specialty Training
Background:
- Artificial intelligence (AI) chatbots, like ChatGPT, demonstrate potential in medical applications.
- Dermatology is an area where AI tools are being explored for educational support.
Purpose of the Study:
- To evaluate the performance of ChatGPT on dermatology specialty certification (board-style) questions.
- To synthesize evidence from multiple studies on ChatGPT's accuracy in this domain.
Main Methods:
- A systematic meta-analysis searched PubMed/MEDLINE and EMBASE up to December 10, 2024.
- A random-effects model assessed overall ChatGPT performance, with subgroup analyses for GPT-3.5 and GPT-4.0.
- Heterogeneity was evaluated using the I2 statistic.
Main Results:
- Thirteen studies were included, with an overall pooled accuracy of 69.97% for ChatGPT.
- ChatGPT-4.0 achieved higher accuracy (79.46%) compared to ChatGPT-3.5 (61.07%).
- Significant heterogeneity was observed across studies (I2 = 93.28%).
Conclusions:
- ChatGPT shows potential as an supplementary educational resource for dermatology board certification candidates.
- Limitations include English-language focus and exam-specific performance, restricting generalizability to clinical practice.
- Further research is needed to explore heterogeneity factors like question difficulty and dermatology subdomains.

