Related Experiment Video
Updated: Sep 26, 2026

Reliability of Artificial Intelligence-Based Cone Beam Computed Tomography Integration with Digital Dental Images
Published on: February 23, 2024
Accuracy and Short-Term Consistency of Generative AI Chatbots in Guideline-Based Endodontic Decision Support: A
Ilgın İlgenli1, Ezgi Avcı1, Timur Köse2,3
1Department of Endodontics, Faculty of Dentistry, Ege University, 35040 Izmir, Turkey.
Abstract:
Background/Objectives: Large language model (LLM) chatbots are increasingly used for dental information and decision support, yet their accuracy and short-term reproducibility in endodontics remain insufficiently established. This study compared five chatbots using open-ended questions derived from established AAE and ESE endodontic guidelines. Methods: Twenty-six guideline-based questions were content-validated by five endodontists using Lawshe's Content Validity Ratio. ChatGPT-4o, Gemini 2.5 Pro, DeepSeek-V3-0324, ScholarGPT (academic version built on OpenAI's GPT-4 architecture), and MedGebra GPT-4 answered each question across three days and three sessions per day, yielding 1170 responses. Two blinded endodontists scored responses on a 5-point guideline-concordance scale. Brunner-Langer LD-F2 analyses assessed model and temporal effects, while weighted kappa evaluated response consistency. Results: The overall model effect was significant (p < 0.001). ChatGPT-4o, Gemini 2.5 Pro, DeepSeek-V3-0324, and ScholarGPT showed comparable performance, whereas MedGebra GPT-4 performed significantly lower after Bonferroni correction. The day effect was not significant (p = 0.054), while the session effect (p = 0.048) and model × session interaction (p = 0.030) were significant. Weighted kappa values varied across models and assessment days, ranging from 0.689-0.730 for Gemini 2.5 Pro, 0.606-0.662 for ChatGPT-4o, 0.520-0.645 for DeepSeek-V3-0324, 0.458-0.592 for ScholarGPT, and 0.240-0.739 for MedGebra GPT-4. Conclusions: Guideline-aligned performance and short-term reproducibility differed across the evaluated chatbots, showing that accuracy and consistency represent distinct aspects of performance. Repeated assessment captured variation missed by single-session testing. These findings support guideline-based evaluation and clinician verification when generative AI chatbots are used to provide endodontic information relevant to decision support or dental education.
