Related Experiment Video
Updated: Sep 27, 2026

Reliability of Artificial Intelligence-Based Cone Beam Computed Tomography Integration with Digital Dental Images
Published on: February 23, 2024
Comparison of Multimodal Large Language Models and Oral and Maxillofacial Radiologists in the Detection of Incidental
İsmail Çapar1, Utku Cem Hasırcı1, Didem Dumanlı Kusay1
1Department of Oral and Maxillofacial Radiology, Faculty of Dentistry, Zonguldak Bülent Ecevit University, Zonguldak 67100, Turkey.
Abstract:
Background/Objectives: This study aimed to compare the diagnostic performance of multimodal (image-capable) large language models (LLMs) and oral and maxillofacial radiologists in detecting nine predefined incidental findings on panoramic radiographs, using a cone-beam computed tomography (CBCT)-derived reference standard. Methods: This retrospective diagnostic performance study included 500 purposively assembled, finding-enriched panoramic radiographs paired with CBCT images. CBCT images were evaluated by three radiologists to establish the reference standard. Two experts and three LLMs, the latter accessed through their consumer web interfaces, independently assessed the presence or absence of the nine findings; 4500 finding-level decisions were analyzed for each reader. Sensitivity, specificity, accuracy, and error rates were calculated. Within-patient clustering was accounted for using generalized estimating equations and a cluster bootstrap procedure with 5000 resamples. Finding-specific comparisons used Cochran's Q test with Benjamini-Hochberg correction. Results: Agreement between the two experts was very good (κ = 0.86). Expert sensitivity, specificity, and accuracy ranged from 89.5 to 91.6%, 92.0-93.0%, and 91.8-92.9%, respectively, compared with 75.1-83.1%, 88.0-90.0%, and 86.8-89.0% for the LLMs. The overall reader effect was significant for all three performance metrics (all p < 0.001). In the exploratory high-risk group, expert sensitivity ranged from 90.9 to 93.2% versus 62.9-78.0% for the LLMs. After correction for multiple comparisons, the difference between readers remained significant only for carotid artery calcification (q < 0.001) and extensive maxillary sinus pathology (q = 0.005). Conclusions: Under the consumer-interface conditions and access period tested, the LLMs performed below the experts, particularly for high-risk findings, and should not be used independently to evaluate panoramic radiographs. Because the dataset was finding-enriched and single-center, the absolute estimates cannot be transferred directly to routine clinical populations, and any future role for these models, more plausibly as an expert-supervised adjunct or screening aid than as a replacement for expert interpretation, remains to be tested prospectively.
