Related Experiment Video
Updated: Aug 5, 2026

Reliability of Artificial Intelligence-Based Cone Beam Computed Tomography Integration with Digital Dental Images
Published on: February 23, 2024
Accuracy and Completeness of Four Artificial Intelligence Chatbots in Assisting Diagnosis and Treatment of
Ana Paula Portes Zeno1, Breno Pereira Caetano1, Giullie Anne de Souza Giffoni da Conceição1
1Department of Pediatric Dentistry and Orthodontics, Faculty of Dentistry, Universidade Federal do Rio de Janeiro, Rio de Janeiro, Brazil.
Introduction:
This study aimed to evaluate the accuracy and completeness of the answers generated by four freely available artificial intelligence (AI) chatbots regarding the management of endodontic sequelae in traumatized permanent teeth and the regenerative endodontic procedure, analyzing their performance by question type and over time.
Methods:
Four AI chatbots (ChatGPT 3.5, Copilot, Perplexity, and DeepSeek) were evaluated with 14 structured questions developed by endodontists, applied at baseline (day 1) and repeated after 7 days (day 8). Nine questions involved clinical cases of traumatized immature and mature permanent teeth (diagnosis, treatment, and treatment with prior diagnosis), and five were conceptual about regenerative endodontic procedures (REPs). Two blinded endodontists independently evaluated response accuracy (1-6 points) and completeness (1-3 points) based on agreement with the expert benchmark, with results reported as median ± interquartile range. Reproducibility was analyzed with weighted Kappa, while Kruskal-Wallis with Dunn's post hoc and Wilcoxon signed-rank tests were used for comparisons.
Results:
On baseline, DeepSeek exceeded in diagnosis while Perplexity consistently outperformed others in treatment-related domains. Regarding response accuracy, Perplexity showed higher score values compared to ChatGPT and Microsoft Copilot, and similar scores to DeepSeek. In terms of completeness, no significant differences were observed among the AI chatbots (P > .05). Over time, only DeepSeek demonstrated a significant increase in completeness scores on day 8 (mean ± standard deviation: 1.56 ± 0.8) compared to baseline (mean ± standard deviation: 1.34 ± 0.55), with P < .05. Conceptual REP questions achieved high accuracy across chatbots.
Conclusions:
DeepSeek showed superior performance in diagnosis, while Perplexity outperformed others in treatment domains. All chatbots performed well on conceptual REP questions. Although chatbots demonstrated moderate-to-high performance in selected domains, none consistently matched the expert benchmark across all evaluated scenarios. Therefore, chatbot-generated information should be interpreted with caution and verified against current evidence-based guidelines before clinical application.
