Related Experiment Video
Updated: Apr 1, 2026

Author Spotlight: Advancing CBCT and Digital Dental Image Integration with AI-Assisted Digitization
Published on: February 23, 2024
Are AI chatbots ready for endodontics? Evaluating their validity, consistency, and readability in patient-oriented
G Karakaya1, A Sirman2, I Öreroğlu2
1Department of Endodontics, Bahçeşehir University School of Dental Medicine, Gayrettepe Mahallesi, Barbaros Bulvarı, No: 153, Beşiktaş, Istanbul, Turkey. karakayagunes@gmail.com.
AI chatbots like ChatGPT-4o, Google Gemini, and Microsoft Copilot were evaluated for endodontic FAQs. While all performed well on basic validity, Google Gemini showed higher accuracy, though no single AI model excelled in all areas like readability and completeness.
Area of Science:
- Artificial Intelligence in Dentistry
- Natural Language Processing in Healthcare
- Clinical Decision Support Systems
Background:
- The integration of artificial intelligence (AI) chatbots into healthcare offers potential for patient education and information dissemination.
- Evaluating the accuracy, validity, and readability of AI-generated responses is crucial for their safe and effective use in specialized fields like endodontics.
Purpose of the Study:
- To compare the performance of ChatGPT-4o, Google Gemini (2.0 Flash), and Microsoft Copilot in answering endodontic frequently asked questions (FAQs).
- To assess the validity, consistency, and readability of responses from these AI models in an endodontic context.
Main Methods:
- Fifty patient-oriented, open-ended endodontic FAQs were developed by endodontic specialists.
- Each FAQ was posed to ChatGPT-4o, Google Gemini, and Microsoft Copilot three times, generating 450 responses.
- Responses were independently evaluated by two endodontists using a modified Global Quality Score (GQS) and analyzed for validity (low and high thresholds), consistency (Cronbach's alpha), and readability (Flesch Reading Ease Score, Flesch-Kincaid Grade Level).
Main Results:
- All chatbots performed adequately under a low-validity threshold, but performance decreased with stricter high-validity criteria.
- Google Gemini demonstrated significantly higher validity at the high-threshold compared to ChatGPT-4o (p=0.022).
- ChatGPT-4o produced more readable outputs (higher FRES, lower FKGL) but had a significantly lower mean overall score and did not achieve high-threshold validity as effectively as other models, suggesting a potential trade-off between readability and depth/accuracy.
Conclusions:
- No single AI chatbot currently provides optimal performance across all evaluated dimensions (readability, accuracy, completeness) for endodontic FAQs.
- While Google Gemini showed superior high-threshold validity, ChatGPT-4o offered better readability, indicating distinct strengths and weaknesses among current AI models.
- Further research is needed to refine AI models for clinical applications, ensuring both accuracy and accessibility of information in specialized dental fields.
More Related Videos
Related Concept Videos
Non-equilibrium in the Cell
Data Validation
Nursing assessment guides are generally based on holistic models rather than medical...

