Related Experiment Video
Updated: Jul 8, 2026

07:32
Measuring Maxillary Posterior Tooth Movement: A Model Assessment using Palatal and Dental Superimposition
Published on: February 23, 2024
Assessment of Large Language Models and Expert Clinicians in Grading Mandibular Third Molar Extraction Difficulty
Suk Min Yoon1, Jeong Hun Yoo1, Taejun Kim1
1Department of Oral and Maxillofacial Surgery, Daejeon Dental Hospital, Wonkwang University College of Dentistry, Daejeon, Republic of Korea.
International Dental Journal
|July 6, 2026
Summary
Large language models (LLMs) show grading tendencies similar to clinicians for assessing mandibular third molar extraction difficulty. However, their agreement is lower than experts and sensitive to conversational context.
Area of Science:
- Oral and Maxillofacial Surgery
- Artificial Intelligence in Medicine
- Radiographic Assessment
Background:
- The Pederson difficulty score (PDS) is a standard for evaluating mandibular third molar extraction difficulty.
- Interobserver variability in PDS grading exists due to the lack of a universal reference standard.
- This study explores the grading tendencies of large language models (LLMs) compared to human clinicians.
Purpose of the Study:
- To investigate if multimodal LLMs exhibit systematic grading tendencies comparable to oral and maxillofacial surgeons.
- To assess the agreement and bias of LLM-generated PDS scores against human expert assessments.
- To evaluate the impact of prompt language and session protocols on LLM grading.
Main Methods:
- Retrospective analysis of 100 panoramic radiographs.
- Independent grading of extraction difficulty using PDS by two LLMs (GPT, Gemini) and two surgeons.
- Agreement assessed via quadratic-weighted kappa and intraclass correlation coefficient.
- Systematic bias analyzed using Bland-Altman plots and ordinal logistic regression.
Main Results:
- Substantial interexpert agreement (κ = 0.754) was observed among human raters.
- GPT showed moderate agreement (κ = 0.564-0.590), while Gemini showed fair-to-moderate agreement (κ = 0.356-0.461).
- Both LLMs exhibited minimal systematic bias and no significant difference in grading tendencies compared to experts; however, sequential sessions reduced consistency.
Conclusions:
- LLMs demonstrate grading tendencies partly comparable to clinicians, without systematic directional bias.
- LLM agreement is below interexpert levels and sensitive to conversational context, limiting immediate clinical reliability.
- LLMs may offer supplementary grading data with independent-session protocols, pending further validation against outcome standards.
