Related Experiment Videos
Evaluating AI Chatbots in Prosthodontics Education: A Quantitative MCQ-Based Assessment
Shankargouda Patil1, Samuel Bybee1, Venkata Suresh Venkataiah2
1College of Dental Medicine, Roseman University of Health Sciences, South Jordan 84095, Utah, USA, roseman.edu.
Purpose:
Prosthodontics is a specialized branch of dentistry that encompasses a wide range of dental procedures, requiring advanced knowledge and expertise for both healthcare professionals. In this study, we assessed the performance of four widely accessible chatbots, ChatGPT-5, Claude 4, Microsoft Copilot, and DeepSeek-V3, using 150 clinical scenario-based prosthodontics multiple-choice questions (MCQs). Evaluating their performance is crucial to ensure reliability.
Materials And Methods:
A total of 150 questions compiled from Bootcamp in English were presented to the four chatbots. We then employed multiple evaluation criteria, such as bilingual evaluation understudy (BLEU), word error rate (WER), cosine semantic similarity, and Flesch reading ease (FRE) score, to provide a more comprehensive assessment of lexical similarity, error rate, semantic consistency, and readability.
Results:
Among the four chatbots, Microsoft Copilot (accuracy: 74%; confidence interval [CI]: 72.2%-76.2%) and ChatGPT-5 (accuracy: 73%; CI: 71.5%-75.8%) demonstrated the best overall performance, achieving high accuracy at 95% CI. Particularly, ChatGPT-5 showed the fewest word errors (WER: 7.86 [95% CI: 6.61-9.10]) and FRE readability of FRE: 20.22 (95% CI: 16.26-24.19) compared to Microsoft Copilot. In terms of alignment with reference outputs, based on the BLEU score, Claude 4 (0.0106 [95% CI: 0.0058-0.0154]) maintained higher scores compared to other analyzed chatbots. Collectively, ChatGPT-5 produced a high level of observed accuracy with relatively fewer errors, moderate alignment with the reference answers, and favorable readability.
Conclusion:
This study suggests that ChatGPT-5 demonstrated a favorable overall performance in prosthodontic clinical scenarios, with high observed accuracy, good textual and semantic alignment, relatively fewer errors, and favorable readability.