Related Experiment Video
Updated: Aug 6, 2026

07:14
Guided Endodontics: Three-Dimensional Planning and Template-Aided Preparation of Endodontic Access Cavities
Published on: May 24, 2022
Large Language Models for Endodontic Symptom Assessment and Treatment Planning Using Image-Free Clinical Records:
Dahyun Seo1, Jieun Cheong1, Yiseul Choi1,2
1Department of Advanced General Dentistry, College of Dentistry, Yonsei University, Seoul, Republic of Korea.
JMIR Medical Informatics
|July 24, 2026
Summary
Large language models (LLMs) show promise in aiding endodontic diagnosis, with ChatGPT demonstrating comparable performance to residents in screening and treatment planning. However, LLM accuracy requires further improvement and clinical oversight for safe application.
Area of Science:
- Dentistry
- Artificial Intelligence
- Endodontics
Background:
- Accurate pulpal status assessment is crucial for endodontic success but challenging due to calcified tissue.
- Current diagnostic methods rely on clinical and radiographic exams, demanding expertise and time, risking diagnostic errors.
- Large language models (LLMs) offer potential to enhance clinical reasoning and diagnostic decision-making in endodontics.
Purpose of the Study:
- To evaluate the clinical applicability of LLMs in endodontics.
- To compare LLM text-based screening performance and treatment plan validity against human evaluators.
- To assess LLM diagnostic capabilities using clinical records without radiographic images.
Main Methods:
- 100 endodontic cases (Jan 2011-Dec 2022) were selected from patient records.
- Four LLMs and human evaluators (dental specialists and residents) assessed cases using text-based records.
- Screening performance was measured by concordance score; treatment plan validity by Likert scale.
Main Results:
- ChatGPT achieved the highest mean concordance score (0.98) among LLMs on Korean-doctor prompts, but did not reach the 'partially correct' criterion.
- ChatGPT's diagnostic accuracy for pulpal (0.65) and periapical (0.57) disease was comparable to AGD and endodontic residents.
- AGD specialists showed the highest diagnostic accuracy (pulpal: 0.70; periapical: 0.65).
Conclusions:
- ChatGPT 4.0 demonstrated relatively high and consistent performance in screening and treatment planning among evaluated LLMs.
- LLM performance did not meet the 'partially correct' criterion, highlighting the need for improvement.
- Addressing LLM hallucinations and human interpretation biases requires continuous clinical supervision and user training for safe application.

