Related Experiment Video
Updated: Jun 21, 2026

Digital Hybrid Model Preparation for Virtual Planning of Reconstructive Dentoalveolar Surgical Procedures
Published on: August 5, 2021
Performance comparison of large language models in treatment planning for the restoration of endodontically treated
Mohammadjavad Shirani1, Maryam Emami2
1Department of Restorative Dentistry, Maurice H. Kornberg School of Dentistry, Temple University, Philadelphia, PA, USA.
Objectives:
This study aimed to compare the performance of five large language models (LLMs), including ChatGPT 4.5 (Deep Research), DeepSeek R1 (Deep Think), Gemini 2.5 Pro, Claude 3.7 Sonnet, and Microsoft Copilot (Think Deeper), in treatment planning for the restoration of endodontically treated teeth over time.
Methods:
Twenty-five case-based scenarios requiring restorative treatment were constructed using a 39-item checklist of clinically relevant factors. Each case was presented to the LLMs via three independent user accounts over three consecutive weeks. After each round, LLMs were shown example responses representing expert-developed answers to assess response variability over time. Two blinded evaluators independently assessed responses for accuracy (1-5 scale) and completeness (1-3 scale). Statistical analyses included Kendall's W, Generalized Estimating Equations, and Bonferroni-corrected pairwise comparisons (α = 0.05).
Results:
Gemini demonstrated the highest performance, significantly outperforming both DeepSeek and Microsoft Copilot across all sessions (P≤.005). Claude exhibited a significant improvement in accuracy during the third week (P<.01), and completeness scores significantly increased in Gemini and Claude following exposure to the exemplary correct responses (P<.0167). Nevertheless, none of the LLMs achieved perfect repeatability, and a substantial proportion of responses remained incomplete or partially accurate even by the third week.
Conclusions:
Among the five models assessed, Gemini consistently produced the most complete and accurate restorative treatment planning responses for endodontically treated teeth, while DeepSeek showed the lowest performance. Given the limitations in consistency and completeness across all models, current LLMs should be considered adjunctive tools requiring human supervision, rather than autonomous clinical decision-makers.

