Related Experiment Video
Updated: Sep 22, 2026

Digital Hybrid Model Preparation for Virtual Planning of Reconstructive Dentoalveolar Surgical Procedures
Published on: August 5, 2021
Image-Based Diagnosis of Oral Lesions: Performance of a Vision-Language Model versus Human Clinicians
Luigi A Vaira1, Jerome R Lechien2,3, Fabiola Giudici4
1Maxillofacial Surgery Operative Unit, Department of Medicine Surgery and Pharmacy, University of Sassari, Sassari, Italy.
Objective:
To evaluate the real-world diagnostic performance of a multimodal large language model (LLM) for image-based assessment of oral mucosal lesions compared with clinicians of varying expertise.
Study Design:
Prospective international multicenter diagnostic accuracy study.
Setting:
Twenty university and tertiary head and neck centers in Italy, Belgium, France, Spain, and Israel.
Methods:
We enrolled 350 consecutive patients (320 with oral lesions, 30 with normal mucosa). Clinical photographs and basic epidemiologic data were analyzed using Gemini 2.5 Advanced with a standardized prompt. Model outputs for lesion detection, malignancy versus benign versus normal, precise histologic diagnosis, and urgency class were compared with histopathology and with 4 clinicians. Sensitivity, specificity, accuracy, and agreement were calculated.
Results:
AI-Gemini achieved 97.1% accuracy for lesion detection and malignancy classification, with sensitivity 98.5% and specificity 96.2% for malignancy, and 88.0% accuracy for precise histologic diagnosis. The head and neck surgeon achieved the highest accuracy for precise diagnosis (97.7%). Three-class diagnostic accuracy was 94.2% for AI-Gemini and 67.0% to 86.5% for nonsurgeon clinicians. Urgency assignment was correct in 70% of cases (κ = 0.716), with a conservative tendency to overestimate risk. Agreement with histologic diagnosis was almost perfect (κ = 0.929).
Conclusion:
In this exploratory study, a multimodal LLM showed encouraging performance in image-based evaluation of oral mucosal lesions. However, given the exploratory single-reader design, these findings should not be interpreted as evidence of equivalence or superiority relative to clinicians and require further prospective external validation before any potential clinical deployment in telemedicine or primary care settings.

