Beyond exam accuracy: Tracking a persistent-failure set reveals visual dental reasoning gaps in multimodal LLMs
Yuichi Mine1, Tsuyoshi Taji2, Shota Okazaki3
1Project Research Center for Integrating Digital Dentistry, Hiroshima University, 1-2-3 Kasumi Minami-ku, Hiroshima 734-8553, Japan; Department of Medical Systems Engineering, Graduate School of Biomedical and Health Sciences, Hiroshima University, 1-2-3 Kasumi Minami-ku, Hiroshima 734-8553, Japan.
Objectives:
To benchmark late-2025 general-purpose multimodal large language models (LLMs) on the Japanese National Dental Examination (JNDE) and to reassess a previously identified persistent-failure set.
Methods:
GPT-5.2T, Claude 4.5, and Gemini 3 were tested in January 2026 using a zero-shot protocol (no prompt engineering) on 350 publicly released JNDE-2025 questions (202 text-only; 148 visually-based) and on 33 questions in a persistent-failure set that had been answered incorrectly by all models in a prior JNDE-2024 evaluation. Responses were scored against the official answer key; paired accuracies were compared using Cochran's Q test followed by post hoc pairwise comparisons with Bonferroni-adjusted p values.
Results:
Overall accuracy was high across all three models, but each model performed worse on visually-based than on text-only questions (71.0-79.7 % vs 91.6-95.5 %). Gemini 3 achieved the highest overall accuracy (88.9 %, 311/350), followed by GPT-5.2T (84.3 %, 295/350) and Claude 4.5 (84.0 %, 294/350). Twenty-one questions (6.0 %) were missed by all models and were predominantly visually-based (16/21), clustering in pediatric dentistry and orthodontics. In the persistent-failure set, 5/33 questions were answered correctly by all models, whereas 9/33 remained incorrect across all models.
Conclusions:
Licensing-examination benchmarking yields near-ceiling performance on text-only questions, but substantial gaps remain in visual dental reasoning. Whole-examination benchmarking should therefore be complemented by modality-stratified reporting and dentistry-specific challenge sets to track meaningful progress beyond aggregate examination accuracy.
Clinical Significance:
High aggregate examination accuracy should not be interpreted as sufficient evidence to judge the reliability of image-conditioned dental support; more clinically grounded evaluations are required.
More Related Videos
09:27Using Eye Movements Recorded in the Visual World Paradigm to Explore the Online Processing of Spoken Language
Published on: October 13, 2018
07:42Quasistatic Mechanical Testing for Computer-Aided Design and Manufacturing Occlusal Veneers Cemented to Milled Dentin Analog Material
Published on: December 20, 2024
Related Concept Videos
Visual Agnosia
Assessment of the Mouth
Mouth Inspection
The inspection begins with visually examining the mouth for symmetry, color, and size.
