Related Experiment Video
Updated: Mar 19, 2026

Author Spotlight: Three-Dimensional Cephalometric Landmark Annotation Demonstration on Human Cone Beam Computed Tomography Scans
Published on: September 8, 2023
Evaluating the Performance of ChatGPT and DeepSeek in Bilingual Responses to Questions Regarding Craniofacial
Lei Li1, Shanbaga Zhao, Shi Feng
1Department of Cranio-Maxillofacial Surgery, Plastic Surgery Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College, Beijing, China.
Background:
Craniofacial microsomia (CFM) is the second most common congenital craniofacial anomaly. As patients increasingly seek health information online, large language models (LLMs) like ChatGPT and DeepSeek have emerged as potential sources of medical information. This study evaluates the performance of ChatGPT-5 and DeepSeek-V3.2 in providing bilingual responses to CFM-related questions.
Methods:
Twenty-two questions covering CFM definition, etiology, diagnosis, treatment, and prognosis were developed. Each question was submitted in English and Chinese to both LLMs using a zero-prompt approach. Responses were evaluated for accuracy using a predefined 4-point scale, with readability assessed using the Flesch Reading Ease score for English and the Chinese Readability Platform for Chinese. Safety statement frequency was also recorded.
Results:
DeepSeek demonstrated significantly higher accuracy than ChatGPT in both English (score 1: 86.4% versus 45.5%, P =0.004) and Chinese (77.3% versus 40.9%, P =0.014). However, only DeepSeek produced responses with inaccurate or misleading content (score 3). For English readability, DeepSeek scored significantly higher (39.4±5.5 versus 35.1±8.4, P =0.031), while Chinese readability was comparable. DeepSeek also included safety statements more frequently (54.5%-72.7% versus 4.5%-18.2%).
Conclusions:
Both LLMs show potential for CFM patient education, with DeepSeek offering superior accuracy and readability in English, though it occasionally produced misleading information. ChatGPT provided safer but less detailed responses. These findings highlight the need for model-specific optimization and clinician oversight when integrating LLMs into patient education for complex craniofacial conditions.
More Related Videos
08:03Midface Hypoplasia and Cranial Base Morphology in Syndromic Craniosynostosis: A Comparative Analysis Study Using a Predictive Regression Model
Published on: November 4, 2025
02:42Analysis of Craniomaxillofacial Malformations in Mice Using Three-dimensional Microcomputed Tomography
Published on: January 17, 2025