Related Experiment Video
Updated: May 25, 2026

Translationally-Relevant Tumor Resection Model for Murine Preclinical Models of Oral Squamous Cell Carcinoma
Published on: April 3, 2026
A comparative analysis of large language models for providing oral cavity cancer information
Eda İzgi1, Turan Canmurat İzgi2, Ceren Mordağ Çiçek3
1Department of Oral and Maxillofacial Surgery, Gülsüm Güral Faculty of Dentistry, Kütahya Health Science University, Inköy District, Eskisehir Highway Blvd. No: 65, Center, Kütahya, Turkey. eda.izgi@ksbu.edu.tr.
Abstract:
This study aimed to comparatively evaluate the medical information delivery capacity and content quality of current large language models (LLMs), specifically ChatGPT (GPT-5.2), Gemini (3.1), and DeepSeek (V4), regarding oral cavity cancer (OCC) based on expert opinions. 20 open-ended questions addressing the risk factors, diagnosis, and treatment of OCC were directed to the three models. The responses were evaluated using a blinded method by 31 expert physicians from Oral and Maxillofacial Surgery, Otorhinolaryngology (ENT), and Medical Oncology. The Modified Global Quality Scale (1-5 points) was utilised for evaluation. Statistical analyses were performed using Kruskal-Wallis, ANOVA, and Bonferroni post-hoc tests, while Fleiss' Kappa coefficient determined inter-expert consistency. The general performance scores of the models were high (3.57-4.15). In the overall assessment, Gemini received statistically significantly higher scores than the DeepSeek model (p = 0.036). Significant performance differences were identified across 15 of 20 questions (p < 0.05); ChatGPT excelled on clinical and treatment-oriented questions, while Gemini stood out on comprehensive informational items. While no statistically significant difference was found among the specialist groups for the overall evaluation and 19 out of 20 questions (p > 0.05), a significant difference was observed solely for Q6 (p = 0.042). Although LLMs have the potential to generate high-quality information about OCC, their performance varies by content type and model architecture. While Gemini demonstrated more consistent performance overall, expert supervision remains essential before these tools can be used as reliable sources of clinical information. Clinicians must be aware of the specific strengths and limitations of different LLMs in OCC to better guide patients who increasingly use such tools for medical information.
More Related Videos
07:13Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Related Concept Videos
Cancer Survival Analysis
Assessment of the Mouth
Mouth Inspection
The inspection begins with visually examining the mouth for symmetry, color, and size.
Mouse Models of Cancer Study
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...