Related Experiment Video
Updated: Sep 5, 2026

A Postoperative Evaluation Guideline for Computer-Assisted Reconstruction of the Mandible
Published on: January 28, 2020
Accuracy of General-Use Multimodal AI Platforms for Pell and Gregory Classification of Impacted Mandibular Third
John A Fregene1, Aditya Tadinada2
1Dentistry, University of Connecticut Health, Farmington, USA.
Purpose:
The purpose of this study was to evaluate the performance of two general-use artificial intelligence models, ChatGPT and Grok, in classifying impacted mandibular third molars using the Pell and Gregory system on panoramic radiographs, compared with a resident consensus reference standard.
Materials And Methods:
One hundred panoramic radiographic images of impacted mandibular third molars were independently classified by two blinded resident reviewers using the Pell and Gregory classification system. Resident consensus was defined as exact agreement between both reviewers; the 94 concordant classifications were confirmed by a board-certified oral and maxillofacial radiologist (the second author), blinded to the AI outputs, and served as the reference standard. Cases without consensus were excluded from the AI accuracy analysis. Each image was then submitted individually to ChatGPT-5.2 and Grok 4.1 through their paid-tier chat interfaces, with each case initiated as a new conversation session. Inter-rater reliability was assessed using percent agreement and Cohen's kappa. AI accuracy was calculated as the proportion of correct classifications among consensus cases, and the two models were compared using McNemar's test with continuity correction.
Results:
The two resident reviewers reached consensus on 94 of 100 cases, with an observed agreement of 94.0% and a Cohen's kappa of 0.925 (95% CI: 0.867-0.983). ChatGPT correctly classified 19 of 94 consensus cases (20.2%; 95% CI: 12.1-28.3%), while Grok correctly classified 17 of 94 cases (18.1%; 95% CI: 10.3-25.9%). Both models performed only marginally above the 11.1% accuracy expected from random chance across nine classification categories. McNemar's test showed no statistically significant difference between the two models (χ²(1) = 0.045, p = 0.832).
Conclusions:
Both ChatGPT and Grok demonstrated poor accuracy in the Pell and Gregory classification of impacted mandibular third molars when compared with resident consensus. These findings suggest that currently available general-use AI models are not reliable for the independent radiographic classification of impacted mandibular third molars and should not replace expert clinical judgment for this task.
