Related Experiment Video
Updated: Jul 5, 2026

Digital Hybrid Model Preparation for Virtual Planning of Reconstructive Dentoalveolar Surgical Procedures
Published on: August 5, 2021
Clinical Relevance of Large Language Models in Endodontics: Diagnostic Appropriateness Based on 50 Simulated Case
Büşra Karaca1, Yunus Emre Çakmak2, Damla Erkal3
1Department of Oral and Maxillofacial Surgery, Faculty of Dentistry, Burdur Mehmet Akif Ersoy University, Burdur, Turkiye.
None:
Large language models (LLMs) are increasingly used in healthcare, but their performance in endodontic decision-making remains unclear. This study aimed to compare six LLMs in terms of diagnostic appropriateness for endodontic treatment planning. Fifty clinical scenarios were developed and entered into six LLMs (ChatGPT-4o, ChatGPT-3.5, Claude 4, Copilot, DeepSeek-V3, Gemini 2.5). Two specialists scored responses as appropriate or inappropriate. Repeated measures ANOVA and chi-square tests were used for analysis. Claude showed the highest accuracy (76%), followed by DeepSeek and Gemini. ChatGPT-3.5 had the lowest (40%). Significant differences were found between models (p < 0.05). Performance was better on straightforward cases than on complex scenarios. LLMs vary widely in diagnostic accuracy for endodontic cases. While some models show promise, others may provide confidently incorrect recommendations. Caution and human oversight remain essential until domain-specific, fine-tuned models are developed.

