Related Experiment Video
Updated: Aug 11, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluation of large language models in root resorption scenarios: an ESE-aligned comparative performance assessment
Öznur Küçük Keleş1, Zeynep Betül Arslan2
1Department of Endodontics, Faculty of Dentistry, Ankara Yıldırım Beyazıt University, 06220, Ankara, Turkey. oznurkucukkeles@aybu.edu.tr.
None:
This study aims to compare the diagnostic accuracy, appropriateness of treatment planning, and source citation performance of five large language models ChatGPT-4o (Free), ChatGPT-5.1 Plus, Microsoft Copilot, Google Gemini, and DeepSeek-R1 in root resorption scenarios. In December 2025, twelve clinical scenarios were created based on the classification of the European Society of Endodontology and each scenario was presented to all chatbots over four consecutive days. All responses were evaluated using a blinded assessment protocol and a binary scoring system. A total of 720 observations (12 cases × 4 repetitions × 3 criteria per model) were analyzed. The collected data were analyzed using chi-square, Fisher's exact, and Cochran Q tests. In terms of diagnostic accuracy, Microsoft Copilot (79.2%), ChatGPT-5.1 (77.1%), and ChatGPT-4o (Free) (75%) showed the highest performance. Google Gemini (68.8%) demonstrated a moderate level of accuracy, while DeepSeek (39.6%) showed markedly low performance. All models exhibited high accuracy in treatment plan recommendations, and no statistically significant differences were detected. Regarding citation accuracy, Copilot ranked first with 100% accuracy. Although large language models present potential as supportive decision-making tools in the evaluation of root resorption, diagnostic inconsistencies, limitations in source accuracy, and variability in responses restrict their independent use in clinical applications. Therefore, the outputs generated by these models should be interpreted cautiously within clinical decision-making processes.

