Related Experiment Video
Updated: Aug 6, 2026

Reliability of Artificial Intelligence-Based Cone Beam Computed Tomography Integration with Digital Dental Images
Published on: February 23, 2024
Comparative Analysis of Agreement and Scoring Patterns Between Human Evaluators and Artificial Intelligence: The Case
Ayşegül Hazır1, Tansu Merve Beşparmak2, Eray Ceylanoğlu1
1Prosthetic Dental Technology Program, Vocational School of Health Services, Kırıkkale University, Kırıkkale, Turkey.
Objective:
This study aimed to compare agreement and scoring patterns between human evaluators and artificial intelligence (AI) models in the rubric-based evaluation of complete denture design applications.
Method:
Thirty complete denture fabrication assignments prepared by students in the Dental Prosthetics Technology Programme were photographed from five different angles according to a standard protocol and evaluated using a 20-item rubric. The assignments were independently assessed by three human evaluators with at least 5 years of experience in complete dentures, as well as by the AI models Gemini 3 Pro, Claude Sonnet 4.6, and Grok 4. Inter-group agreement was analysed using Fleiss' Kappa, while scoring differences were analysed using the Friedman and Wilcoxon Signed-Rank Tests.
Results:
While significant agreement was observed among human evaluators (κ = 0.63; p < 0.001), agreement among the AI models remained low (κ = 0.09; p = 0.012). AI models assigned significantly higher scores than human evaluators (1.72 ± 0.50 vs. 1.09 ± 0.85; p < 0.001).
Conclusion:
Current AI models do not yet demonstrate sufficient reliability to function as independent evaluators of laboratory-based tasks requiring precise visual and spatial assessment. Consequently, final assessment evaluation should therefore remain under the responsibility of experienced dental educators, with AI tools used only as a complementary feedback tool.
