Related Experiment Video
Updated: Jul 8, 2026

Measuring Maxillary Posterior Tooth Movement: A Model Assessment using Palatal and Dental Superimposition
Published on: February 23, 2024
Assessment of Large Language Models and Expert Clinicians in Grading Mandibular Third Molar Extraction Difficulty
Suk Min Yoon1, Jeong Hun Yoo1, Taejun Kim1
1Department of Oral and Maxillofacial Surgery, Daejeon Dental Hospital, Wonkwang University College of Dentistry, Daejeon, Republic of Korea.
Background:
The Pederson difficulty score (PDS) is widely used to assess mandibular third molar extraction difficulty, but is subject to interobserver variability because no universally accepted reference standard exists. This study investigates whether multimodal large language models (LLMs) demonstrate systematic grading tendencies comparable to those of human clinicians in the absence of a definitive standard.
Methods:
This retrospective cross-sectional study included 100 panoramic radiographs. Two LLMs (GPT and Gemini) and two oral and maxillofacial surgeons independently graded extraction difficulty using the PDS. No reference standard or gold-standard rater was designated; all raters were treated as independent assessors. Agreement was assessed using quadratic-weighted kappa and intraclass correlation coefficient values. Systematic bias was analysed using Bland-Altman plots and ordinal logistic regression. The effects of prompt language and session protocol were also evaluated.
Results:
The interexpert agreement was substantial (κ = 0.754). GPT demonstrated moderate agreement with experts (κ = 0.564-0.590), whereas Gemini exhibited fair-to-moderate agreement (κ = 0.356-0.461). Overall, the four-rater reliability was moderate (intraclass correlation coefficient = 0.553). Both LLMs displayed minimal systematic bias (GPT: -0.13; Gemini: -0.08), with no significant differences in grading tendencies compared with human experts (P > .05). However, sequential-session evaluation significantly reduced the scoring consistency for both models (GPT: P = .001; Gemini: P < .001), whereas prompt language had no significant effect on total PDS scores (GPT: P = .543; Gemini: P = .386). A component-level linguistic variation was observed only in GPT's assessment of the ramus relationship (P = .003).
Conclusion:
The examined LLMs showed grading tendencies partly comparable to those of human clinicians, without evidence of systematic directional bias; however, agreement remained below interexpert levels and was sensitive to conversational context. These findings should not be interpreted as evidence of diagnostic accuracy or clinical reliability, as no reference standard was established.
Clinical Relevance:
LLMs may provide supplementary grading information if used with independent-session protocols; however, their role in clinical decision-making requires further validation against defined outcome standards.
