Related Experiment Video
Updated: May 11, 2026

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
Comparative analysis of diagnostic prediction and clinical reasoning of large language models in complex endodontic
Hrudi Sundar Sahoo1, A R Pradeep Kumar1, Archana Durvasulu1
1Department of Conservative Dentistry and Endodontics, Thai Moogambigai Dental College and Hospital, Dr. M.G.R. Educational and Research Institute, Golden George Nagar, Mogappair, Chennai, Tamil Nadu, India.
Objective:
To compare the diagnostic prediction accuracy and clinical reasoning of four state-of-the-art large language models (LLMs) - GPT-4o, Claude 3.7 Sonnet, DeepSeek R1, and Gemini 2.0 Flash - in complex endodontic case analysis using structured prompts.
Methods:
Nine complex endodontic case reports from peer-reviewed journals were converted into standardized narratives with concealed diagnoses. Each LLM generated three prioritized (Bayesian-ranked) diagnostic possibilities with corresponding justifications using structured prompts. Two endodontic specialists independently abstracted the gold-standard diagnoses into a standardized reference format and evaluated the LLM outputs using a rubric assessing clinical plausibility, radiographic justification, and terminological accuracy. Statistical analysis included Kruskal-Wallis testing, Bonferroni-adjusted pairwise comparisons, and the Diagnostic Agreement Index.
Results:
Gemini 2.0 Flash and Claude 3.7 Sonnet generally achieved higher interpretive reasoning quality than DeepSeek R1 across all diagnostic ranks. However, post-hoc Bonferroni-adjusted comparisons revealed that their differences relative to GPT-4o were not statistically significant (adjusted p = 1.000), whereas DeepSeek R1 differed significantly from both Claude and Gemini (adjusted p < 0.05) and approached significance compared with GPT-4o (adjusted p = 0.059).
Conclusions:
This study demonstrates significant variation in the interpretive reasoning performance of contemporary LLMs. Claude 3.7 Sonnet and Gemini 2.0 Flash outperformed DeepSeek R1, although their performance did not differ significantly from GPT-4o.
Clinical Significance:
This study demonstrates that model selection and structured prompting critically influence the reliability of AI-assisted endodontic diagnoses, with Gemini 2.0 and Claude 3.7 showing superior reasoning quality, highlighting the need for specialized training and human oversight for safe clinical integration.

