Related Experiment Video
Updated: Sep 4, 2026

Reliability of Artificial Intelligence-Based Cone Beam Computed Tomography Integration with Digital Dental Images
Published on: February 23, 2024
Comparison of the performances of different updated generative artificial intelligence models on the Japanese
Shuma Hamaguchi1, Masakazu Hamada2, Shunya Ikeda1
1Department of Pediatric Dentistry, Graduate School of Biomedical and Health Sciences, Hiroshima University, Hiroshima, Japan.
Background/Purpose:
Artificial intelligence (AI) has become widely used and applied in various fields. Although several studies have been conducted using generative AI for various qualification exams, to the best of our knowledge, none has focused on performance changes over time.
Materials And Methods:
In August 2025, ChatGPT 5, Gemini 2.5, Microsoft Copilot, and Medi-Search were asked to answer compulsory questions from five years of the Japanese National Dental Examination. In 2024, we also conducted similar tests on other ChatGPT series and Gemini, and the scores were compared.
Results:
In 2025, Copilot, Gemini, and MediSearch scored 80 % or higher, which was the passing standard, for all five years. Although ChatGPT 3.5 did not meet the passing standard for any of the five years, ChatGPT 4o mini and ChatGPT 5 exceeded it for two and three years, respectively. In addition, both Chat GPT's and Gemini's scores substantially improved over time and with each update.
Conclusion:
This report suggests that generative AI is improving annually and adapting to the National Dental Examination. Although each AI model is suited to different fields, the trends may change over time. It is necessary to continue comparing and analyzing AI models and provide users with the latest information.