Related Experiment Video
Updated: Sep 19, 2025

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
Image-Based Diagnostic Performance of LLMs vs CNNs for Oral Lichen Planus: Example-Guided and Differential Diagnosis
Paak Rewthamrongsris1, Jirayu Burapacheep2, Ekarat Phattarataratip3
1Center of Artificial Intelligence and Innovation (CAII) and Center of Excellence for Dental Stem Cell Biology, Faculty of Dentistry, Chulalongkorn University, Bangkok, Thailand; Department of Conservative Dentistry and Periodontology, LMU University Hospital, LMU Munich, Germany.
Large language models (LLMs) show potential for diagnosing oral lichen planus (OLP) from images, but current models are not yet clinically viable. Convolutional neural networks (CNNs) demonstrated superior performance in OLP detection compared to LLMs.
Area of Science:
- Artificial Intelligence in Medicine
- Oral Pathology
- Medical Imaging Analysis
Background:
- Oral lichen planus (OLP) presents diagnostic challenges due to overlapping features with other oral lesions.
- Large language models (LLMs) with computer-vision capabilities offer a potential alternative diagnostic tool.
- Evaluating LLMs for detecting OLP and generating differential diagnoses is crucial for advancing diagnostic accuracy.
Purpose of the Study:
- To assess the diagnostic performance of seven LLMs (proprietary and open-source) in detecting oral lichen planus (OLP) from intraoral images.
- To compare the efficacy of LLMs in zero-shot recognition, example-guided recognition, and differential diagnosis generation for OLP.
- To benchmark LLM performance against established convolutional neural network (CNN) models for OLP detection.
Main Methods:
- A dataset of 1,142 clinical photographs of OLP, non-OLP lesions, and normal mucosa was utilized.
- LLMs were evaluated using zero-shot, example-guided, and differential diagnosis experimental designs.
- Performance metrics included accuracy, precision, recall, F1-score, and discounted cumulative gain (DCG); LLMs were compared to CNN models on a subset of 110 images.
Main Results:
- Gemini 1.5 Pro/Flash achieved highest accuracy (69.69%) in zero-shot, while GPT-4o led in F1-score (76.10%).
- Gemini 1.5 Flash showed highest accuracy (80.53%) and F1-score (84.54%) with example-guided prompts; Claude 3.5 Sonnet had the highest DCG (0.63).
- All LLMs were outperformed by CNN models; open-source Llama showed strengths in diagnosis ranking.
Conclusions:
- The seven evaluated LLMs currently lack sufficient diagnostic performance for clinical application in OLP detection.
- CNNs specifically trained for OLP detection demonstrated superior performance compared to the tested LLMs.
- Further development is needed to enhance LLM capabilities for reliable clinical use in oral mucosal lesion diagnosis.

