Related Experiment Videos
Performance of multimodal Large Language Models in misdiagnosed dermatologic cases: a pilot study on diagnostic
1Department of Dermatology, Kayseri City Education and Research Hospital, Kayseri, Turkey.
Cutaneous and Ocular Toxicology
|June 24, 2026
Summary
Multimodal large language models show potential in diagnosing complex skin conditions, with Gemini 3 demonstrating improved accuracy when integrating visual data. However, current AI models struggle with detecting malignant lesions without specialized imaging.
Area of Science:
- Artificial Intelligence in Medicine
- Dermatology Diagnostics
- Machine Learning Applications
Background:
- Multimodal Large Language Models (LLMs) are emerging as diagnostic aids in dermatology.
- Existing research often overlooks performance in ambiguous "gray zone" cases.
- The impact of visual data integration on LLM diagnostic accuracy and cognitive bias correction is not well understood.
Purpose of the Study:
- To assess the diagnostic accuracy of three leading multimodal LLMs.
- To evaluate the effect of visual data integration on diagnostic performance.
- To determine the rate at which LLMs replicate initial human misdiagnoses.
Main Methods:
- A cross-sectional analysis of 30 histopathology-confirmed dermatologic cases.
- Models were queried using text-only and multimodal (image + text) prompts.
- Key outcomes included Top-1 accuracy and error replication rates.
Main Results:
- Gemini 3 achieved the highest Top-1 accuracy (60.0%) in multimodal mode.
- Gemini 3 showed improved accuracy in inflammatory dermatoses with visual data (45.5% to 72.7%).
- All models demonstrated limited accuracy for malignant lesions using standard macro-images.
Conclusions:
- Gemini 3 shows promise as a de-biasing tool for complex inflammatory dermatoses.
- Current multimodal LLMs lack precision for malignancy detection without dermoscopic data.
- Cautious integration of AI is recommended for high-stakes dermatologic diagnostics.