Related Experiment Video
Updated: May 1, 2026

05:56
Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
3.2K
Publicly Available Large Language Models for Trichoscopy: A Head-to-Head Comparison with Dermatologists
Basil Signer1, Ali Mokhtari2, Simone Cazzaniga1
1Department of Dermatology, Inselspital, Bern University Hospital, 3010 Bern, Switzerland.
Diagnostics (Basel, Switzerland)
|January 10, 2026
Summary
Large language models (LLMs) show limited diagnostic accuracy in trichoscopy, significantly underperforming human experts. Further development is needed for AI to assist reliably in hair and scalp disorder diagnoses.
Area of Science:
- Dermatology
- Artificial Intelligence
- Medical Imaging
Background:
- Trichoscopy is crucial for diagnosing hair and scalp disorders, demanding significant expertise.
- The diagnostic utility of publicly available large language models (LLMs) in trichology remains unexplored.
- This study assesses LLMs' accuracy in interpreting trichoscopic images against human dermatologists.
Purpose of the Study:
- To evaluate the diagnostic accuracy of four LLMs in trichoscopic image interpretation.
- To compare LLM performance against dermatology residents, board-certified dermatologists, and trichology experts.
- To determine the potential of AI in assisting trichological diagnoses.
Main Methods:
- A prospective comparative study using a preprocessed set of structurally transformed trichoscopic images.
- Fifteen dermatologists (residents, board-certified, experts) provided suspected and differential diagnoses.
- Four LLMs (ChatGPT-4o, Claude Sonnet 4, Gemini 2.5 Flash, Grok-3) evaluated images under identical conditions.
Main Results:
- Human dermatologists achieved an overall diagnostic accuracy of 58.1% for suspected diagnoses and 68.3% for suspected + differential diagnoses.
- AI models achieved lower accuracy: 18.2% for suspected diagnoses and 44.4% for suspected + differential diagnoses.
- Gemini 2.5 Flash showed the highest AI accuracy (62.5% for suspected + differential diagnoses), yet all AI models significantly underperformed human experts (p < 0.001).
Conclusions:
- Publicly available LLMs currently underperform human experts in trichoscopic diagnosis, particularly for single correct diagnoses.
- AI models demonstrate moderate to good agreement among themselves but show only slight to fair agreement with dermatologists.
- Specialized training and further development are essential for LLMs to be reliable tools in routine trichological care.

