Beyond exam accuracy: Tracking a persistent-failure set reveals visual dental reasoning gaps in multimodal LLMs.
Yuichi Mine1, Tsuyoshi Taji2, Shota Okazaki3
1Project Research Center for Integrating Digital Dentistry, Hiroshima University, 1-2-3 Kasumi Minami-ku, Hiroshima 734-8553, Japan; Department of Medical Systems Engineering, Graduate School of Biomedical and Health Sciences, Hiroshima University, 1-2-3 Kasumi Minami-ku, Hiroshima 734-8553, Japan.
General-purpose large language models (LLMs) show high accuracy on text-based dental exam questions but struggle with visual reasoning. Further evaluation is needed for reliable AI in dentistry.
Area of Science:
- Artificial Intelligence in Dentistry
- Medical Education Technology
- Large Language Models
Background:
- General-purpose multimodal large language models (LLMs) are increasingly capable of processing complex information.
- Benchmarking LLMs on professional licensing examinations like the Japanese National Dental Examination (JNDE) is crucial for assessing their real-world applicability.
- Previous evaluations identified persistent failure points for LLMs on the JNDE, particularly in visually-based questions.
Purpose of the Study:
- To benchmark late-2025 general-purpose multimodal large language models (LLMs) on the Japanese National Dental Examination (JNDE).
- To reassess a previously identified persistent-failure set of JNDE questions.
- To evaluate the performance differences between text-only and visually-based questions.
Main Methods:
- Three leading LLMs (GPT-5.2T, Claude 4.5, Gemini 3) were tested in a zero-shot protocol on 350 JNDE-2025 questions.
- The question set included 202 text-only and 148 visually-based items.
- Performance was also assessed on 33 persistent-failure questions from a prior JNDE evaluation.
Main Results:
- All models achieved high overall accuracy (84.0-88.9%) on the JNDE.
- Performance was significantly lower on visually-based questions (71.0-79.7%) compared to text-only questions (91.6-95.5%).
- Six percent of questions were missed by all models, predominantly visual items in pediatric dentistry and orthodontics; 9/33 persistent-failure questions remained unsolved.
Conclusions:
- While LLMs demonstrate near-ceiling performance on text-based dental knowledge, significant limitations exist in visual dental reasoning.
- Aggregate accuracy on licensing exams may overestimate AI reliability for image-conditioned dental support.
- Modality-stratified reporting and dentistry-specific challenge sets are essential for meaningful progress evaluation.
More Related Videos
09:27Using Eye Movements Recorded in the Visual World Paradigm to Explore the Online Processing of Spoken Language
Published on: October 13, 2018
07:42Quasistatic Mechanical Testing for Computer-Aided Design and Manufacturing Occlusal Veneers Cemented to Milled Dentin Analog Material
Published on: December 20, 2024
Related Concept Videos
Visual Agnosia
Assessment of the Mouth
Mouth Inspection
The inspection begins with visually examining the mouth for symmetry, color, and size.
