Related Experiment Video
Updated: Aug 28, 2026

Digital Hybrid Model Preparation for Virtual Planning of Reconstructive Dentoalveolar Surgical Procedures
Published on: August 5, 2021
Artificial Intelligence for Diagnosing Normal Anatomical Variants and Pathological Oral Mucosal Lesions: A
Ana Glavina1,2, Marija Galešić2, Bojan Poposki3
1Department of Dental Medicine, University Hospital of Split, 21000 Split, Croatia.
Abstract:
Background and Objectives: Artificial intelligence (AI) is increasingly used in clinical dentistry, but its diagnostic accuracy for oral mucosal lesions based on clinical photographs remains insufficiently validated. This prospective observational study compared the Top-1 diagnostic accuracy of ChatGPT-4o and ChatGPT-5 in identifying normal anatomical variants and pathological oral mucosal lesions and evaluated their performance across anatomical sites. Materials and Methods: Seventy adults with either normal anatomical variants (n = 21) or pathological oral mucosal lesions (n = 49) were consecutively recruited at the Department of Dental Medicine, University Hospital of Split, Croatia. One standardized clinical photograph per patient was analyzed by ChatGPT-4o and ChatGPT-5 under image-only and image-plus-text conditions using identical prompts. The reference diagnosis was established by an oral medicine specialist, with histopathological examination (HPE) performed when clinically indicated. Diagnostic performance was assessed using Top-1 accuracy and McNemar's test. Results: Both models showed low accuracy with image-only input, but performance improved significantly after clinical information was added (p < 0.001). Overall Top-1 accuracy increased from 19.0% to 69.0% for ChatGPT-4o and from 9.0% to 51.0% for ChatGPT-5. For normal anatomical variants, accuracy increased from 14.3% to 81.0% and from 14.3% to 76.2%, respectively. For pathological oral mucosal lesions, accuracy increased from 20.4% to 63.3% and from 6.1% to 40.8%, respectively. ChatGPT-4o showed numerically higher accuracy than ChatGPT-5, particularly for pathological oral mucosal lesions, but no statistically significant difference was found between the models in the corresponding paired comparisons. Conclusions: Diagnostic performance was limited with image-only input but improved substantially when standardized clinical information accompanied the images. The numerical differences between models, particularly for pathological oral mucosal lesions, may be clinically relevant but do not establish superiority or equivalence. Neither model can currently replace conventional clinical diagnosis, and AI should be regarded as a clinical decision-support tool for evaluating oral mucosal lesions.