Related Experiment Video
Updated: Sep 21, 2026

Digital Hybrid Model Preparation for Virtual Planning of Reconstructive Dentoalveolar Surgical Procedures
Published on: August 5, 2021
Comparative Performance of GPT-5.1, Claude Sonnet 4.5, Gemini 2.0 Flash, and DeepSeek-R1 on Oral and Maxillofacial
Gaye Keser1, Burcu Celen1, Filiz Namdar Peki̇ner2
1Department of Oral Medicine and Radiology, Marmara University, Istanbul, TUR.
Introduction:
Rapid advances in artificial intelligence (AI) and large language models (LLMs) have increased interest in their use in dental education and assessment. This study evaluated the accuracy of GPT-5.1, Claude Sonnet 4.5, Gemini 2.0 Flash, and DeepSeek-R1 on Turkish Dental Specialisation Exam (DUS) questions in oral and maxillofacial radiology.
Methods:
A total of 132 Turkish DUS questions from 2012 to 2021 were obtained from a publicly accessible question bank, including 126 theoretical and six image-based items. Each question had five options and one correct answer. Questions covered radiologic physics, radiation safety, dental anatomy, pathology classification, imaging instrumentation, and visual interpretation of panoramic, periapical, cone beam computed tomography (CBCT), and clinical images. All models received identical prompts, and image-based items were submitted through the image-upload functionality available in the respective web interfaces. Responses were scored against the official answer key by two oral and maxillofacial radiologists. Statistical analyses included Fisher's exact test, Cochran's Q test, and Bonferroni-corrected McNemar tests. An exploratory secondary analysis evaluated session-to-session variation using a model-based performance index.
Results:
GPT-5.1 answered all 132 questions correctly (100%), compared with Claude Sonnet 4.5 (119/132; 90.2%), Gemini 2.0 Flash (111/132; 84.1%), and DeepSeek-R1 (108/132; 81.8%). Overall performance differed significantly among the models (p < 0.001), and GPT-5.1 significantly outperformed each of the other three models after Bonferroni correction. On the six image-based questions, accuracy was 100% for GPT-5.1, 83.3% for Claude Sonnet 4.5, 50.0% for Gemini 2.0 Flash, and 66.7% for DeepSeek-R1. Because only six image-based items were available, these subgroup findings were considered exploratory and do not permit firm conclusions regarding radiographic interpretation. No significant temporal trend was identified in the exploratory model-based performance index.
Conclusions:
GPT-5.1 achieved the highest accuracy on this publicly accessible DUS benchmark and significantly outperformed Claude Sonnet 4.5, Gemini 2.0 Flash, and DeepSeek-R1. However, the public availability of the questions and answer keys introduces a risk of benchmark contamination, and the very small image-based subgroup precludes generalisation regarding radiographic interpretation. These findings should therefore be interpreted as comparative performance on a public examination benchmark rather than evidence of clinical diagnostic competence.