Related Experiment Video
Updated: Jan 9, 2026

08:42
Assessment of Social Cognition in Non-human Primates Using a Network of Computerized Automated Learning Device ALDM Test Systems
Published on: May 5, 2015
12.6K
Assessing the accuracy of multiple-choice questions using different artificial intelligence-driven tools-an
1Department of Oral and Maxillofacial Surgery and Diagnostic Sciences, Division of Diagnostic Sciences, College of Dentistry, Jazan University, Jazan 82848, KSA.
Dento Maxillo Facial Radiology
|December 6, 2025
Summary
Microsoft Co-pilot and ChatGPT-4o demonstrated the highest accuracy in answering oral radiology multiple-choice questions (MCQs). This study reveals significant AI performance variations, offering insights into AI chatbots for dental education.
Area of Science:
- Artificial Intelligence in Dentistry
- Oral Radiology Education
- Medical Education Technology
Background:
- Artificial intelligence (AI) tools are increasingly integrated into medical education.
- Assessing AI's accuracy and efficiency in domain-specific knowledge testing is crucial.
- Oral radiology education requires precise and reliable assessment tools.
Purpose of the Study:
- To evaluate the accuracy and response times of various AI chatbots.
- To compare AI performance on multiple-choice questions (MCQs) in oral radiology.
- To determine the suitability of AI tools for dental education.
Main Methods:
- 80 oral radiology MCQs across 5 domains were used.
- Accuracy and response times of ChatGPT, ChatGPT-4o, Microsoft Co-pilot, DeepSeek, Gemini, and Meta AI were assessed.
- Chi-Square test and One-way ANOVA were employed for statistical analysis.
Main Results:
- Microsoft Co-pilot and ChatGPT-4o exhibited the highest overall accuracy.
- ChatGPT provided the fastest response times.
- Microsoft Co-pilot excelled in knowledge-based questions and radiographic safety (100% accuracy); DeepSeek was superior for radiographic diagnosis.
Conclusions:
- Microsoft Co-pilot demonstrated superior overall accuracy and performance in knowledge-based and radiographic safety questions.
- ChatGPT-4o was the second most accurate, while DeepSeek showed strength in radiographic diagnosis.
- AI chatbots exhibit performance variability, impacting their utility in dental education.

