Related Experiment Video
Updated: Sep 6, 2026

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
Multimodal Large Language Models for Dental Chart Image Interpretation: Cross-Sectional Benchmarking Study With
Ah-Young Cho1, Soo-Heang Eo2,3, Mi-Jeong Jeon4
1Department of Conservative Dentistry and Dental Research Institute, School of Dentistry, Seoul National University, 101 Daehakno, Jongno-Gu, Seoul, 03080, Republic of Korea, 82 2-6256-3182.
Background:
Entering clinical training, dental students must learn to read tooth-centered electronic dental records, but limited teaching time in patient-centered clinics can leave gaps in chart-reading literacy. Multimodal large language models (LLMs) that process dental record images may offer scalable educational support, yet their performance has not been benchmarked.
Objective:
This study evaluated multimodal ChatGPT models on dental chart image interpretation and compared the best-performing model with dental students and residents.
Methods:
A retrospective, cross-sectional benchmark study used deidentified dental chart images from 15 patients at Seoul National University Dental Hospital (2017-2025). Charts containing Korean and English text were captured as sequential screenshots (154 images). For each patient, 24 Korean-language questions (360 total) spanned four categories: Type A, general factual retrieval; Type B, tooth- or procedure-specific retrieval; Type C, interpretation requiring multientry synthesis; and Type D, absent-information questions assessing abstention. Nine multimodal ChatGPT models (available August 2, 2025-August 9, 2025) were evaluated under standardized conditions. Outputs were scored against a gold standard using 7 metrics, with sentence bidirectional encoder representations from transformers (SBERT) similarity prespecified as the primary semantic measure. Human baselines included 2 third-year students and 2 first-year residents. Groups were compared with Kruskal-Wallis tests and Dunn post hoc analyses. Gold-standard reliability was assessed by independent senior-expert review and chance-corrected agreement (Gwet AC1) among clinical reference raters, and the model ranking was confirmed by content-based clinical accuracy analysis.
Results:
Across 360 items, GPT-5 Thinking achieved the highest median SBERT similarity (0.900, IQR 0.525-1.000), followed by GPT-5 Pro (0.861, IQR 0.501-1.000) and OpenAI o3 (0.831, IQR 0.489-1.000), with significant overall group differences (P<.001). Student 1 did not differ significantly from GPT-5 Thinking across Types A-D (all P≥.16), whereas Student 2 differed only on Type D (P=.002; Cliff δ=-0.27). Resident 1 scored higher than GPT-5 Thinking on Types A (P=.01) and C (P=.04) but lower on Type D (P<.001; δ=-0.35), whereas Resident 2 scored higher on Types A (P=.03), B (P=.02), and C (P=.04) and did not differ on Type D (P=.23). Type C tasks showed compressed SBERT distributions and low exact-match rates, indicating persistent difficulty in synthesis. The clinical reference rater agreement was high (Gwet AC1=0.99), and the content-based clinical-accuracy ranking was closely aligned with the SBERT ranking (Spearman ρ=0.90; P=.001).
Conclusions:
Statistically nonsignificant differences were observed between GPT-5 Thinking and dental students for most question-type contrasts, with Student 2 differing only on absent-information items. Compared with first-year residents, GPT-5 Thinking remained lower on several Type A-C contrasts, particularly interpretive Type C and one tooth- or procedure-specific Type B comparison; Type D contrasts require cautious interpretation because all groups had ceiling medians. Within this single-center benchmark, multimodal LLMs may have potential as supervised educational tools for chart-reading practice and verification, rather than as replacements for clinical expertise, pending external validation across institutions, specialties, and record systems.