Related Experiment Video
Updated: Jun 27, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Artificial Intelligence in Caries Risk Assessment: Evaluating the Current Status of Caries Management by Risk
Şevval Çakıcı1, Sümeyra Akkoç2
1Department of Pediatric Dentistry, Faculty of Dentistry, Kütahya Health Sciences University, Kütahya, Turkey, sevval.cakici@ksbu.edu.tr.
Introduction:
Caries risk assessment (CRA) is essential for individualized caries management in early childhood, but existing tools are difficult to standardize in routine practice. The ability of large language model (LLM)-based artificial intelligence (AI) systems to recognize and correctly apply validated CRA tools remains unclear. This study aimed to evaluate the ability of ChatGPT and Google Gemini, LLM-based AI systems, to apply and interpret CRA tools in early childhood under different prompting conditions.
Methods:
Thirty standardized clinical vignettes of children aged 0-5 years were evaluated using Caries Management by Risk Assessment (CAMBRA) and the Cariogram by ChatGPT-5.1 (OpenAI, San Francisco, CA, USA) and Google Gemini-3.0 Pro (Google LLC, Mountain View, CA, USA) under three prompting conditions: unguided, guideline-informed, and guideline-only. An expert panel provided reference classifications. AI outputs were assessed for categorical agreement, mean absolute error (MAE), quality, accuracy, and readability. MAE differences were analyzed using the Friedman test and mixed-design repeated-measures analysis of variance (ANOVA), while quality, accuracy, and readability were analyzed using repeated-measures ANOVA. Reliability was assessed using intraclass correlation coefficients (ICCs).
Results:
Inter- and intra-rater reliability were high (ICC = 0.89-0.93). Reference CAMBRA and Cariogram classifications showed a moderate correlation (ρ = 0.509; n = 30). MAE differed significantly across AI model-prompting condition combinations for both tools (p < 0.001). No significant difference was observed between AI models for CAMBRA, whereas ChatGPT showed lower MAE than Gemini for the Cariogram (p = 0.002). Guideline-informed and guideline-only prompting significantly reduced MAE and improved quality and accuracy compared with the unguided condition (p < 0.001).
Conclusion:
LLM-based AI systems can support CRA, particularly when guideline-based prompting is used. However, performance depends on the CRA tool and AI model, and numerical or algorithmic components, such as those in the Cariogram, remain challenging.
