Related Experiment Video
Updated: Jul 8, 2025

Author Spotlight: Advancing CBCT and Digital Dental Image Integration with AI-Assisted Digitization
Published on: February 23, 2024
Validity and reliability of artificial intelligence chatbots as public sources of information on endodontics
Hossein Mohammad-Rahimi1, Seyed AmirHossein Ourang2, Mohamad Amin Pourhoseingholi3
1Topic Group Dental Diagnostics and Digital Dentistry, ITU/WHO Focus Group AI on Health, Berlin, Germany.
Aim:
This study aimed to evaluate and compare the validity and reliability of responses provided by GPT-3.5, Google Bard, and Bing to frequently asked questions (FAQs) in the field of endodontics.
Methodology:
FAQs were formulated by expert endodontists (n = 10) and collected through GPT-3.5 queries (n = 10), with every question posed to each chatbot three times. Responses (N = 180) were independently evaluated by two board-certified endodontists using a modified Global Quality Score (GQS) on a 5-point Likert scale (5: strongly agree; 4: agree; 3: neutral; 2: disagree; 1: strongly disagree). Disagreements on scoring were resolved through evidence-based discussions. The validity of responses was analysed by categorizing scores into valid or invalid at two thresholds: The low threshold was set at score ≥4 for all three responses whilst the high threshold was set at score 5 for all three responses. Fisher's exact test was conducted to compare the validity of responses between chatbots. Cronbach's alpha was calculated to assess the reliability by assessing the consistency of repeated responses for each chatbot.
Results:
All three chatbots provided answers to all questions. Using the low-threshold validity test (GPT-3.5: 95%; Google Bard: 85%; Bing: 75%), there was no significant difference between the platforms (p > .05). When using the high-threshold validity test, the chatbot scores were substantially lower (GPT-3.5: 60%; Google Bard: 15%; Bing: 15%). The validity of GPT-3.5 responses was significantly higher than Google Bard and Bing (p = .008). All three chatbots achieved an acceptable level of reliability (Cronbach's alpha >0.7).
Conclusions:
GPT-3.5 provided more credible information on topics related to endodontics compared to Google Bard and Bing.
Related Concept Videos
The Availability Heuristic
Non-equilibrium in the Cell
Current Trends in Nursing II
Issues And Trends In Healthcare Delivery System
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
Tooth Anatomy
The Crown, Neck, and Root
The visible part of the tooth is referred to as the crown. It's covered by enamel, the hardest substance in the human body. The crown is uniquely shaped for each type of tooth, allowing for different functions such as cutting, tearing, or...
Teeth
In the bud stage, the tooth germ (an aggregation of cells) starts to form in the developing jawbone. During the cap stage, the tooth germ differentiates into enamel organ, dental papilla, and dental sac, which will later develop into the tooth's enamel, dentin...

