Related Experiment Video
Updated: Jul 3, 2026

10:42
A Postoperative Evaluation Guideline for Computer-Assisted Reconstruction of the Mandible
Published on: January 28, 2020
6.9K
Can chatbots replace experts? Diagnostic accuracy of AI models in classifying impacted mandibular third molars
Müfide Bengü Erden1, Mehmet Gümüş Kanmaz2, Genta Agani Sabah3
1Department of Oral and Maxillofacial Surgery, Faculty of Dentistry, Izmir Tinaztepe University, Izmir, Turkey. benguerden@gmail.com.
Odontology
|September 25, 2025
Summary
AI chatbots show potential for interpreting dental X-rays but are not yet expert-level for classifying impacted third molars. ChatGPT-4o performed best, though no AI model achieved consistent diagnostic accuracy.
Area of Science:
- Dentistry
- Artificial Intelligence
- Radiology
Background:
- AI chatbots are increasingly used in dentistry.
- Their ability to interpret panoramic radiographs and classify impacted mandibular third molars is unevaluated.
Purpose of the Study:
- To assess the diagnostic performance of four leading AI chatbots in classifying impacted mandibular third molars using panoramic radiographs.
- To compare the accuracy of ChatGPT-4o, Gemini 2.5 Pro, Claude Sonnet 4.0, and Copilot (GPT-4) against expert evaluations.
Main Methods:
- 93 impacted mandibular third molars were analyzed from panoramic radiographs.
- Four AI chatbots classified molars using Pell and Gregory, Winter, and Rood and Shehab systems.
- Three dental experts rated chatbot responses using a Global Quality Score (GQS).
Main Results:
- No significant difference in GQS ratings among chatbots, with ChatGPT-4o scoring highest (2.41 ± 1.03).
- ChatGPT-4o showed superior performance in Winter classification (κ = 0.171).
- Gemini 2.5 Pro demonstrated moderate agreement in root findings; Copilot (GPT-4) showed consistency in canal parameters. No chatbot achieved acceptable Pell and Gregory classification agreement.
Conclusions:
- AI chatbots show promise for panoramic image interpretation but are currently suboptimal for third molar classification.
- ChatGPT-4o exhibited the best performance among the tested models, yet none reached expert-level accuracy.
- Further advancements in multimodal AI and large, labeled datasets are crucial for clinical integration.
Related Concept Videos
Stereotype Content Model
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence categorization, a person will feel...
Non-equilibrium in the Cell
An important concept in studying metabolism and energy is that of chemical equilibrium. Most chemical reactions are reversible. They can proceed in both directions, releasing energy into their environment in one direction, and absorbing it from the environment in the other direction. The same is true for the chemical reactions involved in cell metabolism, such as the breaking down and building up of proteins into and from individual amino acids, respectively. Reactants within a closed system...

