Related Experiment Video
Updated: Aug 30, 2026

Fine-Tuning Large Language Models Using Entity Hallucination Index for Text Summarization
Published on: January 9, 2026
Evaluating Large Language Models in Endodontic Irrigants: A Structured Assessment of Accuracy, Readability,
Damla Erkal1, Yunus Emre Çakmak2, Kürşat Er2
1Department of Endodontics, Faculty of Dentistry, Burdur Mehmet Akif Ersoy University, Burdur, Turkey.
Background:
Large language models (LLMs) are increasingly used by clinicians and learners for endodontic information, yet their reliability for irrigation-related knowledge remains unclear. This study evaluated the accuracy, readability, hallucination profile, clinical risk, and temporal stability of 4 general-purpose LLMs on endodontic irrigation questions, including false-premise prompts.
Methods:
A literature-based reference set was developed for sodium hypochlorite, calcium hypochlorite, ethylenediaminetetraacetic acid, and chlorhexidine. ChatGPT-5.2, Claude Sonnet 4.5, Gemini 3 Pro, and DeepSeek 3.2 were each asked 100 questions comprising 80 factual items and 20 contradiction-seeking items. Responses were independently scored by 2 blinded endodontists for accuracy, hallucination subtype, and clinical risk. Five readability indices were calculated, and model stability was reassessed 10 days later.
Results:
Inter-rater agreement was almost perfect (weighted κ = 0.92 for accuracy; κ = 0.97 for hallucination). Claude Sonnet 4.5 and Gemini 3 Pro showed the highest accuracy (1.75 and 1.74) and the lowest hallucination rates (7% and 9%). DeepSeek 3.2 showed the lowest accuracy (1.08), the highest hallucination rate (29%), and critical-risk outputs in 16% of responses. Gemini showed the highest test-retest stability (weighted κ = 0.95). Hallucination strongly correlated with clinical risk (ρ = 0.93; P < 0.001). Readability analysis showed a 2-tier pattern: Gemini and DeepSeek produced more accessible text, whereas ChatGPT and Claude generated denser outputs.
Conclusions:
LLM performance in endodontic irrigation is model- and irrigant-dependent. Hallucination profiling is clinically relevant, and high initial accuracy does not guarantee temporal stability. LLM outputs should be used only as clinician-verified adjuncts.
