Related Experiment Video
Updated: Sep 18, 2025

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
Diagnostic Performance of Multimodal Large Language Models in the Analysis of Oral Pathology
Ana Suárez1, Yolanda Freire1, María Suárez1
1Department of Pre-Clinic Dentistry II, Faculty of Biomedical and Health Sciences, Universidad Europea de Madrid, Madrid, Spain.
Objective:
This study evaluated the accuracy and repeatability of ChatGPT-4o, a multimodal AI model, in interpreting photographs of oral mucosal lesions, and explored its potential as a diagnostic support tool for specialists and non-specialists.
Methods:
Thirty clinical photographs of oral and labial mucosal lesions were analysed using ChatGPT-4o. For each image, 30 responses were generated across 20 days. The model was asked to identify the anatomical location, suggest a diagnosis, and recommend diagnostic tests and treatments. Two oral pathology experts assessed 3600 responses using a three-point scale (0 = incorrect, 1 = partially correct, 2 = correct). Accuracy and repeatability were analysed using accuracy rates, Gwet's AC and percent agreement.
Results:
ChatGPT-4o achieved 71.4% accuracy in identifying lesion location and 58.2% in diagnosis. In cases with correct diagnoses, the model reached 90.7% and 95.8% accuracy in suggesting diagnostic tests and treatments, respectively. Repeated responses showed substantial to almost perfect agreement across all evaluated aspects.
Conclusions:
ChatGPT-4o showed potential as a reliable and accessible tool to support the initial assessment of oral lesions. Although not a substitute for clinical judgment, it may enhance diagnostic efficiency, particularly in resource-limited settings. Further validation is needed before clinical use.

