Related Experiment Video
Updated: May 31, 2026

Reliability of Artificial Intelligence-Based Cone Beam Computed Tomography Integration with Digital Dental Images
Published on: February 23, 2024
Reliability of Artificial Intelligence Chatbots in Answering Patient-Oriented Questions About Endodontic Apical
Faraj Alotaiby1, Waleed Almutairi2
1Department of Oral and Maxillofacial Diagnostic Sciences, College of Dentistry, Qassim University, Buraydah, SAU.
Objective:
This study aimed to evaluate the reliability and clinical appropriateness of responses generated by publicly accessible AI-based chatbot platforms (ChatGPT (OpenAI, San Francisco, CA, USA); Grok (xAI, Palo Alto, CA, USA); and DeepSeek (High-Flyer, Hangzhou, ZJ, CHN)) when addressing patient-oriented questions related to endodontic apical lesions.
Methods:
A total of 15 standardized, non-technical questions were developed to simulate typical patient inquiries. The questions covered four domains: identification of apical lesions, differentiation between odontogenic and non-odontogenic causes, management and recommended next steps, and potential risks and follow-up considerations. Each question was independently submitted once to ChatGPT, Grok, and DeepSeek. Two expert evaluators (a board-certified endodontist and a board-certified oral and maxillofacial pathologist) assessed the responses using a five-point Likert scale based on reliability, clarity, and clinical appropriateness for patient education. Inter-rater agreement was assessed using Cohen's kappa coefficient. Descriptive statistics were used to summarize response ratings, and differences among platforms were analyzed using the Kruskal-Wallis H test.
Results:
All three AI platforms demonstrated high levels of expert agreement, with a Cohen's kappa value of 0.85 indicating almost perfect inter-rater reliability. ChatGPT achieved the highest proportion of strongly agreeable responses, followed by Grok and DeepSeek. No statistically significant differences were observed among the platforms in agreement distributions (p = 0.36). Based on expert evaluation, ChatGPT responses tended to provide more detailed clinical explanations, whereas Grok and DeepSeek responses were perceived to use simpler and more accessible language.
Conclusion:
Publicly accessible AI chatbots can provide generally reliable and clinically appropriate responses to patient-oriented questions concerning endodontic apical lesions. ChatGPT demonstrated higher informational accuracy, whereas Grok and DeepSeek offered clearer patient-centered communication. These findings support the cautious integration of AI chatbots as adjunct tools for patient education, while emphasizing the continued necessity of professional clinical judgment.
