Related Experiment Video
Updated: Jul 13, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance of Advanced Large Language Models in Caries Risk Assessment and Preventive Decision-Making: A
Berkant Sezer1, Tuğba Aydoğdu2
1Department of Pediatric Dentistry, School of Dentistry, Çanakkale Onsekiz Mart University, Çanakkale, Turkey, dt.berkantsezer@gmail.com; berkant.sezer@comu.edu.tr.
Introduction:
Large language models (LLMs) have recently been integrated into dental practice to support clinical reasoning and preventive decision-making. This study compared the performance of 5 advanced chatbots - ChatGPT-5, Claude 4.5 Sonnet, Gemini 2.5 Pro, LLaMA 3.1, and Mistral 7B - in providing evidence-based responses for caries risk assessment and preventive management in pediatric cases.
Methods:
Twenty-five validated, case-based questions were developed in accordance with internationally recognized pediatric and preventive dentistry guidelines. Responses were evaluated by 6 pediatric dentistry experts for accuracy, completeness, relevance, clarity, and usefulness using Likert-type scales. Response time, word count, and linguistic readability characteristics (Flesch Reading Ease Score and Flesch-Kincaid Grade Level) were additionally analyzed to compare textual complexity across chatbot-generated responses. Data normality was assessed using the Shapiro-Wilk test; parametric tests (ANOVA with Bonferroni correction) or nonparametric tests (Kruskal-Wallis with Dunn's post hoc) were applied as appropriate.
Results:
Statistically significant differences were observed across all qualitative criteria, including accuracy, completeness, relevance, clarity, and usefulness (p < 0.001). ChatGPT-5 consistently ranked among the top-performing models, showing balanced and high-quality responses across domains, while Claude 4.5 Sonnet achieved the highest accuracy and completeness scores. Gemini 2.5 Pro produced the fastest responses (p < 0.001), whereas Claude 4.5 Sonnet generated the longest and most linguistically complex outputs. Readability metrics also differed significantly among models (p < 0.001), with Mistral 7B and LLaMA 3.1 showing the highest readability.
Conclusions:
All evaluated chatbots generated generally relevant responses for caries risk assessment and preventive counseling; however, substantial inter-model differences were observed in qualitative performance, linguistic complexity, and response characteristics. Occasional inconsistencies and outdated content highlight the need for cautious interpretation and further externally validated evaluation before broader clinical implementation.
