Related Experiment Video
Updated: Sep 2, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Accuracy and response repeatability of three large language models on undergraduate operative dentistry
Jasmine Mary Antony1, Nikhil Harikrishnan1, Nishi Jayasheelan1
1Department of Conservative Dentistry and Endodontics, Yenepoya Dental College (Deemed to be University), Mangalore, Karnataka, India.
Introduction:
Artificial intelligence (AI) has emerged as a valuable tool in the field of dentistry to assist healthcare providers in diagnosis, optimize treatment outcomes, support research, and improve education. Furthermore, the generative aspects of AI have emerged as a powerful tool for dental professionals to tackle clinical and academic challenges. However, the reliability of AI in domains such as undergraduate training remains unverified.
Aim:
This study aimed to evaluate the accuracy and consistency of three large language models (LLMs) in answering multiple-choice questions on the undergraduate operative dentistry curriculum to determine their reliability as a supplementary learning resource.
Methods:
Sixty multiple-choice questions were formulated from the undergraduate operative dentistry curriculum. Each LLM (Claude, Gemini, and ChatGPT) was queried individually nine times over three days. The results were evaluated using a predetermined answer key. The accuracy of the LLMs was compared using a generalized estimating equation (GEE) with question-level clustering analysis. Consistency was assessed at the question level using response consistency, correctness consistency, and Fleiss' kappa. Statistical significance was set at P < 0.05.
Results:
Gemini demonstrated the highest overall accuracy (98.89%), followed by ChatGPT (98.33%) and Claude (97.78%). No statistically significant difference in accuracy was observed among the three LLMs (GEE, overall Wald χ 2(2) = 1.46, p = 0.481). All three models showed almost perfect intersession agreement (Fleiss' κ = 0.93-0.95), with response and correctness consistency ranging from 88.3% to 91.7% across models.
Conclusions:
All three LLMs demonstrated high accuracy on operative dentistry MCQs, suggesting their potential utility as a supplementary educational tool. However, the residual error rates in the answers warrant the cautious integration of LLMs into dental education.
