Related Experiment Video
Updated: Sep 28, 2026

Midface Hypoplasia and Cranial Base Morphology in Syndromic Craniosynostosis: A Comparative Analysis Study Using a Predictive Regression Model
Published on: November 4, 2025
Evaluating ChatGPT's Reliability in Analyzing Facial Asymmetry: A Comparative Study With Human Evaluators Using Cross
Ali B Jafar1, Malak Motair2, Shahah Al Hajeri2
1Department of Surgery-Head and Neck Surgery, College of Medicine, Health Sciences Center Kuwait University Kuwait City Kuwait.
Objective:
Facial asymmetry assessment is often subjective and time intensive, thus we aim to evaluate reliability of ChatGPT in analyzing facial asymmetry compared to human raters.
Methods:
Thirty patients with unilateral facial paralysis who underwent facial reanimation surgery were included in this study. Sixty static 2D frontal images (pre and postoperative) were obtained from our database. Facial asymmetry was assessed using the Sunnybrook resting symmetry scale and a 0-4 global asymmetry rating scale. Two human raters evaluated all images independently. ChatGPT Pro 5.0 accessed from September 2025 to October 2025, was prompted through standardized instructions to evaluate the same set. Agreement was assessed using intraclass correlation coefficient (ICC), Cohen's kappa, the Wilcoxon signed-rank test, and Bland-Altman plots.
Results:
ChatGPT Pro 5.0 showed no statistically significant difference compared with human raters in preoperative facial assessment (p = 0.701), indicating high reliability in detecting pronounced asymmetry. However, a significant difference emerged in the postoperative assessment (p = 0.0001), where the ChatGPT was less sensitive to subtle facial asymmetry in postoperative stage. Cluster analysis confirmed agreement in high asymmetry cases and divergence in mild cases.
Conclusion:
ChatGPT reliably assesses pronounced facial asymmetry but is less accurate with subtle facial asymmetry features.
