Related Experiment Video
Updated: May 22, 2026

Roughness Impact of Piezoelectric Dental Scaler on Two Distinct Flowable Composite Filling Materials
Published on: January 10, 2025
A pilot study on the feasibility of using GPT-5.0 for objective esthetic assessment of single-tooth restorations
Guanqi Liu1, Yuanxiang Liu1, Xiaoyan Chen1
1Hospital of Stomatology, Guanghua School of Stomatology, Sun Yat-sen University, and Guangdong Provincial Clinical Research Center of Oral Diseases, Guangzhou, China.
Purpose:
Large language models (LLMs) are increasingly used in dental esthetic assessment, yet evidence of their feasibility and reliability for objective scoring of single-tooth restorations remains limited. This preliminary study explored the potential of GPT-5.0 for esthetic evaluation using established indices and compared its performance with experienced clinicians and senior dental students.
Methods:
A retrospective dataset of 46 maxillary central incisor restorations with natural adjacent teeth and standardized, high-quality photographs was analyzed. Each case was independently scored using the Pink Esthetic Score (PES) and White Esthetic Score (WES) by three expert prosthodontists, three calibrated senior dental students, and GPT-5.0 guided by a detailed prompt and expert examples. Intra- and inter-group reliability were assessed using intraclass correlation coefficients (ICCs), and group differences were evaluated with repeated-measures analysis of variance.
Results:
GPT-5.0 showed high internal consistency and strong agreement with human raters, particularly for restorative (WES) parameters. ICCs between GPT-5.0 and expert consensus were 0.944 (total), 0.868 (PES), and 0.943 (WES). Minor but significant differences occurred in certain soft tissue (PES) subitems (P = 0.035), with slightly greater variability than that of the human raters. Reliability was lower for subjective soft tissue parameters, especially the mesial papilla and soft tissue convexity/color/texture subitems; generalizability beyond idealized conditions remains to be established.
Conclusions:
With detailed guidance, GPT-5.0 can approach expert-level performance for objective scoring of single-tooth restorations in a controlled setting, but reliability is lower for subjective soft tissue parameters, and broader clinical validation is needed before routine implementation.
