Related Experiment Video
Updated: Aug 23, 2026

Reliability of Artificial Intelligence-Based Cone Beam Computed Tomography Integration with Digital Dental Images
Published on: February 23, 2024
Multidisciplinary dental treatment planning by artificial intelligence: A comparative evaluation
Sohil A Kazim1, Bashaer A Alnoman2, Razan M Hashim3
1Department of Restorative Dentistry, Advanced Specialty Education Program in Prosthodontics, Rutgers School of Dental Medicine, Rutgers University, Newark, New Jersey, USA.
Purpose:
The purpose of this study was to evaluate and compare four AI software programs-ChatGPT-5, Microsoft Copilot, Google Gemini (V 2.5), and OpenEvidence-in generating comprehensive dental treatment plans for minimally destructed and severely mutilated teeth using identical clinical inputs.
Material And Methods:
Ten anonymized clinical cases, each consisting of 1 intraoral photograph and 1 corresponding periapical radiograph, were independently submitted to each AI software program using a standardized prompt. A reference standard was established through consensus among four calibrated specialists (one surgically trained prosthodontist, one restorative dentist, one prosthodontist, and one endodontist). AI-generated responses were evaluated using a structured scoring rubric across five domains: diagnostic accuracy, restorability assessment, multidisciplinary integration, treatment sequencing, and extraction appropriateness (score range: 0-10 per case). Agreement among evaluators was assessed using pairwise Cohen's kappa (κ) analysis based on initial independent scoring before consensus discussions.
Results:
Substantial inter-evaluator agreement was observed (κ = 0.74; 95% CI: 0.68-0.80), indicating consistent application of the scoring rubric. Mean total scores were highest for OpenEvidence (7.7 ± 1.4) and ChatGPT-5 (7.6 ± 1.1), followed by Google Gemini (5.1 ± 1.2) and Microsoft Copilot (4.7 ± 3.1). OpenEvidence and ChatGPT-5 more frequently incorporated phased treatment sequencing, ferrule assessment, restorability analysis, and multidisciplinary treatment considerations. Microsoft Copilot failed to generate responses in three of the 10 evaluated cases because of content-filtering restrictions.
Conclusions:
AI software programs generated structured and clinically relevant dental treatment proposals; however, meaningful variability existed in diagnostic interpretation, restorability assessment, multidisciplinary integration, and extraction thresholds. OpenEvidence and ChatGPT-5 demonstrated greater agreement with the expert reference standard, although clinically significant diagnostic errors were observed across all evaluated systems. AI-generated treatment plans should therefore be regarded as adjunctive decision-support tools requiring specialist oversight before clinical implementation.

