Related Experiment Video
Updated: Jul 15, 2026

Cutaneous Leishmaniasis in the Dorsal Skin of Hamsters: a Useful Model for the Screening of Antileishmanial Drugs
Published on: April 21, 2012
A Comparative Study: Can Large Language Models be a Supportive Tool in the Diagnosis and Treatment of Scabies?
İstem Ş Çamur1, Eren Çamur2, Duru T Onan1
1From the Department of Dermatology and Venerology, Baskent University Ankara Hospital, Ankara, Türkiye.
Background:
The advent of large language models (LLMs) has endowed artificial intelligence with remarkable capabilities to comprehend and generate human-like text by leveraging extensive datasets. Numerous studies in medicine increasingly focus on LLMs, particularly highlighting their substantial potential in diagnostic and therapeutic management. Scabies is a contagious skin disease that continues to affect over 200 million people globally each year and is particularly prevalent in developing countries. In recent years, there has been a marked increase in scabies cases both in Turkey and worldwide, elevating scabies to the status of a public health concern. Therefore, the Clinical Practice Guideline for the Diagnosis and Treatment of Scabies (CPGDT) was published in 2024.
Aims And Objectives:
This study evaluates the performance of LLMs on scabies-related questions created from CPGDT and compares it to a dermatologist.
Materials And Methods:
Sixty-four multiple-choice questions in English were administered to nine LLMs, including Claude 3 Opus, ChatGPT (versions 3.5, 4, and 4o), Google Gemini (1.0 and 1.5 Pro), Microsoft Copilot, Meta LLaMA 3 70B, and Perplexity, as well as to a senior dermatologist with 22 years of experience in dermatology. Responses were scored as correct or incorrect. Statistical analyses, including the McNemar test, were conducted using IBM SPSS Statistics 23.0.
Results:
Claude 3 Opus achieved the highest accuracy (89%), significantly outperforming other LLMs and dermatologists (75%) (P < 0.05). ChatGPT 4o (79.6%) and ChatGPT 4 (78.1%) followed, with no significant differences observed among other LLMs or between LLMs and dermatologists (P > 0.05).
Conclusion:
This is the first study evaluating LLMs' performance on scabies-related knowledge. Claude 3 Opus' superior performance highlights its potential in dermatology, surpassing widely recognized LLMs like ChatGPT. Claude 3 Opus demonstrated exceptional accuracy in scabies-related queries, underscoring its value as adjunctive tool in dermatology. Enhancing LLMs' visual diagnostic capabilities is vital for their integration into clinical practice.