Related Experiment Video
Updated: Sep 4, 2026

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
Performance of large language models on image-based oral pathology questions from the Japanese National Dental
Hikaru Watanabe1, Osamu Uehara2, Tetsuro Morikawa3
1Department of Oral and Maxillofacial Surgery, Ohu University School of Dentistry, Fukushima, Japan.
Background /Purpose:
Large language models (LLMs), such as Chat Generative Pre-trained Transformer (ChatGPT) and Gemini, have demonstrated promising capabilities for medical question-answering tasks. However, the diagnostic performance of LLMs in image-based oral pathologies remains largely unexplored. This study aimed to evaluate these capabilities using histopathological images obtained from the Japanese National Dental Examination.
Materials And Methods:
This study aimed to evaluate and compare the diagnostic accuracy and agreement of three LLMs (ChatGPT-4o [ChatGPT], Gemini 1.5 Pro [Gemini], and Claude 3.5 Sonnet [Claude]) on pathology image-based questions from the Japanese National Dental Examination.
Results:
Gemini achieved the highest accuracy (61.4 %), followed by Claude (52.3 %), and ChatGPT (45.4 %). Gemini and ChatGPT exhibited significant differences (P = 0.00054). Cohen's kappa values indicated moderate agreement for all models, with Gemini showing the highest agreement (κ = 0.599). Accuracy varied across disease categories: Gemini excelled in squamous cell carcinoma (92.0 %) and salivary gland tumors, whereas Claude performed best on soft tissue lesions. The confusion matrix analysis revealed distinct misclassification patterns in each model, particularly between odontogenic tumors and cystic lesions.
Conclusion:
LLMs demonstrated moderate diagnostic performance for image-based dental pathology questions, with Gemini demonstrating superior accuracy and consistency. Although promising decision support tools in education and clinical settings, LLMs still exhibit domain-specific limitations and require careful oversight. The integration of explainable artificial intelligence and real-world clinical validation is recommended for its safe and effective use.

