Related Experiment Video
Updated: Aug 16, 2026

Accuracy in Dental Medicine, A New Way to Measure Trueness and Precision
Published on: April 29, 2014
Diagnostic accuracy and repeatability of ChatGPT using textual and radiographic data in reversible pulpitis: a
María Llorente de Pedro1, Yolanda Freire1, Natalia Moneo1
1Department of Dentistry, Faculty of Biomedical and Health Sciences, Universidad Europea de Madrid, Villaviciosa de Odón, Spain.
Objectives:
Large language models have recently evolved into multimodal systems capable of interpreting images alongside text. The aim of this study was to assess the diagnostic accuracy and repeatability of ChatGPT-5 for reversible pulpitis, based on structured clinical information and periapical radiographs (PRs).
Methods:
Thirty expert-validated cases of reversible pulpitis were retrospectively selected. For each case, clinical data and PRs were compiled to create a reference dataset, with expert consensus serving as the reference standard. A structured prompt incorporating clinical information and PR was designed and included six predefined questions: four addressing diagnostic reasoning, one assessing whether PR information was used to establish the diagnosis, and one requiring radiographic findings description. Each case was entered thirty times into ChatGPT-5, generating a total of 900 answers. Answers were assessed by two independent experts, with disagreements resolved by a third. Radiographic descriptions were graded using a three-point Likert scale. Diagnostic accuracy and repeatability were analysed using binomial and agreement methods.
Results:
Diagnostic questions related to tooth identification, pulpal diagnosis, periapical diagnosis and treatment indication showed accuracies above 96%. Overall accuracy for radiographic description was low (4.67%). Coronal radiographic interpretation showed poor performance (5.15%), whereas accuracy for pulpal and periapical structures exceeded 96%. Repeatability was almost perfect for diagnostic questions and pulpal/periapical interpretation, but moderate for coronal descriptions.
Discussion:
ChatGPT-5 demonstrated high accuracy and repeatability in text-based diagnostic reasoning for reversible pulpitis. However, limitations in radiographic interpretation, particularly of coronal structures, indicate that its use should be limited to a supervised decision-support role.

