Related Experiment Video
Updated: Jul 3, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evidence-based evaluation of large language models in advanced clinical decision-making for removable prosthodontics
Savvas N Kamalakidis1, Damian J Lee2, Konstantinos Vazouras3
1Assistant Professor, Department of Prosthodontics, School of Dentistry, Aristotle University of Thessaloniki, Greece; and Adjunct Assistant Professor, Department of Prosthodontics, Tufts University School of Dental Medicine, Boston, MA.
Statement Of Problem:
Artificial intelligence is widely used to answer questions about prosthodontic treatments, but responses regarding removable prosthodontics remain limited with gaps.
Purpose:
Large language models (LLMs) have been increasingly used by clinicians and trainees to access clinical information. However, their ability to make accurate, evidence-based decisions in removable prosthodontics remains unclear. This study aimed to evaluate and compare the performance of 5 LLMs in responding to advanced clinical questions related to removable complete and partial dentures.
Material And Methods:
A curated set of 10 advanced, evidence-based clinical questions covering key domains in removable prosthodontics was developed based on current consensus statements, systematic reviews, and professional guidelines. Each question was posed independently to 5 LLMs using single-turn prompts and default settings. Two board-certified prosthodontists independently evaluated each response using a predefined 10-point rubric assessing scientific accuracy, comprehensiveness, clarity, and clinical relevance, with guideline-based reference answers serving as the standard. Intra-rater reliability was evaluated by repeat scoring after a 4-week interval. Statistical analyses included reliability testing, a linear mixed-effects model, and nonparametric comparisons among models (α=.05).
Results:
DeepSeek V3 recorded the highest mean scores, averaged across both evaluators and both scoring sessions (8.1/10), followed by ChatGPT-4o (7.2/10), Google Gemini Advanced (6.7/10), Microsoft Copilot (6.3/10), and ChatGPT-4.0 (6.0/10). Statistically significant differences were observed among the models (P<.001). The Cronbach α and intraclass correlation coefficients (ICC) indicated high internal consistency and reliability.
Conclusions:
Contemporary LLMs can provide clinically relevant responses to advanced questions in removable prosthodontics, extending into higher-level diagnostic and treatment-planning domains, with DeepSeek V3 demonstrating the highest performance. However, their output should be interpreted with caution and used only as adjunctive decision-support tools under expert clinician oversight.