Related Experiment Video
Updated: Sep 27, 2026

Reliability of Artificial Intelligence-Based Cone Beam Computed Tomography Integration with Digital Dental Images
Published on: February 23, 2024
Classification Performance of General-Purpose Multimodal AI Chatbots Versus Expert Radiologists in the Winter
Ahmet Faruk Erturk1, Merve Yelken Kendirci1, Oyun Erdene Batgerel2
1Department of Oral and Maxillofacial Radiology, Faculty of Dentistry, Biruni University, Istanbul 34010, Turkey.
Abstract:
Background/Objectives: Cone-beam computed tomography (CBCT) provides three-dimensional anatomical detail that is used when classifying impacted mandibular third molars, yet the classification performance of general-purpose multimodal artificial intelligence (AI) chatbot services in this task remains unexplored. This study compared the Winter classification performance of five general-purpose multimodal AI chatbot services with that of expert oral and maxillofacial radiologists using identical standardized CBCT-derived image composites. Methods: A total of 114 impacted mandibular third molars were retrospectively evaluated. For each tooth, a standardized composite of axial, coronal and sagittal multiplanar reconstructions together with one three-dimensional surface rendering was exported as a single JPEG image and submitted to five multimodal AI chatbot services (ChatGPT, Gemini, Copilot, Perplexity and Grok) through their web interfaces using an identical structured prompt that named the target tooth by its FDI number (38 or 48). Each impacted tooth was submitted as a separate case in a separate cleared session; a composite was never used to query two teeth simultaneously. Two board-certified oral and maxillofacial radiologists independently classified the same JPEG composites to establish an adjudicated reference standard. Accuracy, balanced accuracy, macro-averaged recall, precision and F1 score, category-level sensitivity and specificity with exact 95% confidence intervals (CIs) and Cohen's kappa were calculated, and the accuracy of each service was compared with the majority-class (no-information) baseline. Results: The adjudicated reference standard comprised mesioangular (n = 28), distoangular (n = 9), vertical (n = 25), horizontal (n = 46) and inverted (n = 6) impactions; no transverse case occurred, so the six-category spectrum was incomplete. Inter-observer agreement between the two radiologists was 80.7% (92/114) with Cohen's kappa 0.688 (95% CI 0.571-0.805). Accuracy was 41.2% (95% CI 32.6-50.4) for ChatGPT, 30.7% (23.0-39.7) for Copilot and 24.6% (17.6-33.2) for Gemini, Perplexity and Grok. No service exceeded the 40.4% majority-class baseline (ChatGPT versus baseline, p = 0.85). Balanced accuracy was 23.8% or lower for every service, and agreement with the reference standard was slight to none (kappa -0.003 to 0.097; all 95% CIs included zero). No service produced a correct distoangular (0/9; 95% CI 0.0-33.6) or inverted (0/6; 95% CI 0.0-45.9) classification, although these categories were represented by very few cases and the corresponding estimates are therefore highly imprecise. Conclusions: Within the specific web interfaces, prompt, image composites and testing period evaluated, general-purpose multimodal AI chatbot services did not classify impacted mandibular third molars according to the Winter system more accurately than a trivial majority-class classifier, and their agreement with an expert-adjudicated reference standard was negligible. These findings apply to the tested consumer chatbot services and to static CBCT-derived composites and should not be generalized to all multimodal AI models or to task-specific dental AI systems.
