Related Experiment Video
Updated: Aug 30, 2026

Guided Endodontics: Three-Dimensional Planning and Template-Aided Preparation of Endodontic Access Cavities
Published on: May 24, 2022
Accuracy and Error Patterns of References Generated by Large Language Models in Endodontics: The Role of Prompt
Mehmet Adıgüzel1, Alparslan Mustafa Çeler2
1Department of Endodontics, Faculty of Dentistry, Hatay Mustafa Kemal University, Hatay, Turkey.
Abstract:
BACKGROUND Large language models (LLMs) are increasingly used in healthcare; concerns persist regarding the accuracy of generated bibliographic references. The effect of prompt design on reference reliability has not been clearly established. This comparative experimental study evaluated the impact of prompt specificity on LLM-generated reference accuracy in endodontics and compared model performance. MATERIAL AND METHODS We used ChatGPT 5 and Claude Sonnet 4.6. Ten predefined endodontic queries were combined with 3 prompt types of increasing specificity. Each model generated 5 references per query-prompt combination (total: 300 references). References were verified using PubMed, Google Scholar, and CrossRef. Accuracy was classified as fabricated (0), partially accurate (1; existing references containing ≥1 bibliographic inaccuracy), or fully accurate (2). Digital object identifier (DOI) accuracy was assessed separately. Statistical analyses were performed using mixed-effects models and Pearson's chi-square test or Fisher's exact test. RESULTS Accuracy scores tended to increase with greater prompt specificity (P=0.249). Claude demonstrated significantly higher accuracy than ChatGPT (mean score: 1.79 vs 1.25; P<0.001). DOI accuracy did not differ among prompt groups (P=0.338); it was significantly higher for Claude than for ChatGPT (90.0% vs 35.3%; P<0.001). ChatGPT produced significantly more title, journal, and DOI errors (P<0.001); author and year errors were similar between models. CONCLUSIONS Prompt specificity had limited effects on reference accuracy; model selection played a greater role. DOI accuracy was strongly model-dependent and largely unaffected by prompt design under the test conditions, highlighting the need for external verification of LLM-generated references.
