Related Experiment Video
Updated: Sep 13, 2026

A Postoperative Evaluation Guideline for Computer-Assisted Reconstruction of the Mandible
Published on: January 28, 2020
Comparative Evaluation of Large Language Models in Frontal Sinus Fracture Management: A Multidisciplinary
Mehmet Sefik Oruc1, Ali C Gunenc1, Ovunc Akdemir1
1Department of Plastic, Reconstructive and Aesthetic Surgery, Faculty of Medicine, Istanbul Aydin University.
Abstract:
Frontal sinus fracture management requires integration of aesthetic, sinonasal, and intracranial risks, making it a demanding test of artificial intelligence-generated clinical advice. This study compared ChatGPT, Gemini, and DeepSeek across 30 systematically developed frontal sinus fracture vignettes. Ninety model-vignette responses were independently evaluated by 9 clinicians-3 plastic surgeons, 3 otorhinolaryngology/head and neck surgeons, and 3 neurosurgeons-using 9 one-to-five clinical domains and 5 binary safety outcomes. Three cyclic masking sequences counterbalanced model position; evaluators worked without a fixed time limit and could pause across sessions. The design generated 810 expert assessments. Mean total scores were 30.63±2.28 for ChatGPT, 30.50±2.37 for Gemini, and 30.72±2.13 for DeepSeek (maximum, 45), with no detectable overall difference (P=0.404; Kendall W=0.003). No domain-level difference remained significant after Holm adjustment. At least one safety concern was recorded in 45.2%, 41.1%, and 48.1% of individual assessments, respectively (P=0.255); majority consensus occurred for 35 of 90 responses and did not differ by model (P=0.056). Agreement on the composite safety outcome was limited. Under this evaluation framework, no model advantage was demonstrated, while safety concerns remained frequent and evaluator perspective materially influenced judgments. Large language model recommendations in frontal sinus trauma require expert review and multidisciplinary oversight.
