Related Experiment Video
Updated: Jul 13, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance of Advanced Large Language Models in Caries Risk Assessment and Preventive Decision-Making: A
Caries Research
|July 11, 2026
Summary
This study evaluated five large language models (LLMs) for pediatric dental caries risk assessment. While all provided relevant information, significant differences in accuracy, completeness, and readability were found, necessitating careful clinical use.
Area of Science:
- Artificial Intelligence in Dentistry
- Clinical Decision Support Systems
- Natural Language Processing
Background:
- Large language models (LLMs) are increasingly used in dentistry for clinical reasoning and preventive care.
- This study assesses the performance of five advanced LLMs in pediatric dental applications.
Purpose of the Study:
- To compare the evidence-based response quality of ChatGPT-5, Claude 4.5 Sonnet, Gemini 2.5 Pro, LLaMA 3.1, and Mistral 7B for pediatric caries risk assessment and prevention.
- To analyze response characteristics including accuracy, completeness, relevance, clarity, usefulness, response time, word count, and readability.
Main Methods:
- Twenty-five validated pediatric dentistry case-based questions were used.
- Six pediatric dentistry experts evaluated chatbot responses using Likert scales.
- Quantitative analysis included response time, word count, and readability metrics (Flesch scores).
Main Results:
- Significant differences (p < 0.001) were found across all qualitative metrics.
- Claude 4.5 Sonnet showed highest accuracy and completeness; ChatGPT-5 offered balanced, high-quality responses.
- Gemini 2.5 Pro provided the fastest responses, while Mistral 7B and LLaMA 3.1 had the highest readability.
Conclusions:
- All LLMs generated relevant responses for caries risk assessment and prevention.
- Substantial inter-model variations exist in performance and linguistic complexity.
- Clinical implementation requires cautious interpretation due to potential inconsistencies and the need for external validation.
