Related Experiment Video
Updated: Jul 14, 2026

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
Diagnostic accuracy of large language models in ICOP-based orofacial pain diagnosis: A comparative study
Mehmet Sait Şimşek1, Enis Esen1, Mahmut Koparal1
1Department of Oral and Maxillofacial Surgery, Faculty of Dentistry, Adıyaman University, Adıyaman, Türkiye.
Objective:
To compare the diagnostic performance of ChatGPT 5.5, Claude Opus 4.1, Gemini 3 Flash, and Grok 4 in International Classification of Orofacial Pain (ICOP)-based clinical scenarios.
Methods:
Thirty ICOP diagnoses were randomly selected, and corresponding clinical scenarios were manually developed. Each scenario was submitted to all models using standardized prompts in independent sessions. Two blinded evaluators assessed primary diagnosis accuracy, subclassification accuracy, clinical interpretation, and management recommendations.
Results:
Overall performance differed significantly among models (p < .001). Grok 4 achieved the highest total score and outperformed the other models. No significant differences were found among ChatGPT 5.5, Gemini 3 Flash, and Claude Opus 4.1. Subclassification accuracy was consistently lower than primary diagnosis accuracy, while management recommendations did not differ significantly.
Conclusion:
LLM performance varied across ICOP-based scenarios. Although Grok 4 showed the highest diagnostic concordance, current LLMs should support, not replace, clinician judgment.
