Related Experiment Video
Updated: Sep 10, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large language models answering structured questions on external cervical resorption: a comparative study
Alexandre Barbosa1, Ana Cristina Braga2, Irene Pina-Vaz3
1Faculty of Dental Medicine, University of Porto, Rua Dr. Manuel Pereira da Silva, 4200-393, Porto, Portugal.
Abstract:
This study aimed to evaluate the accuracy, consistency, and temporal stability of six large language models (LLMs), when answering structured questions on external cervical resorption (ECR) across seven consecutive days, while assessing domain-specific error patterns. Four Base configurations (ChatGPT-5, Gemini, Claude, and Mistral) and two document-grounded (RAG) configurations (NotebookLM and Perplexity Pro) were evaluated using 46 validated dichotomous questions on ECR. Each model was queried using three independent user accounts over seven consecutive days. Accuracy was defined as agreement with predefined gold standard answers. Inter-account consistency and temporal stability were evaluated across accounts and days. Domain-specific error patterns were analyzed. Generalized estimating equation models were used to evaluate differences in response accuracy across LLM configurations, clinical domains, user accounts, and evaluation days. Overall accuracy was high (90.8%). Model configuration significantly influenced performance (p = 0.050), with NotebookLM and Gemini showing the lowest estimated error rates (3.0% and 4.0%, respectively), while Claude demonstrated the highest error rate (12.0%). Accuracy remained stable across user accounts and evaluation days, with no evidence of account-related variability (p = 0.654) or temporal drift (p = 0.875). Clinical domain significantly affected performance (p < 0.001); classification questions showed the highest error rate (37.0%). LLMs demonstrated high accuracy, reproducibility, and temporal stability. However, performance varied across models and clinical domains, with classification questions remaining challenging. These findings support the cautious integration of LLMs into clinical practice and endodontic education while underscoring the need for human oversight. Importantly, observed performance was restricted to standardized questions and should not be interpreted as evidence of comprehensive clinical decision-making capability.

