Related Experiment Video
Updated: Jun 14, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Exploring the Educational Utility of LLMs in Operative Dentistry Through MCQ-Based Accuracy, Semantic, and
Shilpa Bhandi1, Adam Lawlor1, Venkata Suresh Venkataiah2
1College of Dental Medicine, Roseman University of Health Sciences, South Jordan, Utah, USA.
Objectives:
The performance of five popular, widely available large language models (LLMs): ChatGPT-4o, Gemini 2.5 Flash, Llama 4, DeepSeek-V3, and Microsoft Copilot in operating dentistry education was evaluated by employing a multiple-choice question-based assessment system.
Material And Methods:
This was done using a set of 150 MCQs covering areas of endodontics, dental caries, paediatric, preventive, aesthetic and restorative dentistry, biomaterials, and periodontics. The LLM's performance was assessed using classification metrics (accuracy, sensitivity, predictive reliability), textual similarity metrics (BLEU score, cosine similarity, Word Error Rate), and readability metrics (Flesch Reading Ease score).
Results:
The highest classification accuracy was achieved by Gemini 2.5 Flash and ChatGPT-4o, showing their high sensitivity and high overall predictive reliability. The model with the most textual similarity to the reference answers was ChatGPT-4o with BLEU of 0.10 ± 0.0279, a high cosine similarity of 0.48 ± 0.0422, and a relatively low Word Error Rate (WER) of 5.57 ± 0.7301, and a Flesch Reading Ease score of 13.53 ± 4.9449.
Conclusion:
In medical education, ChatGPT-4o exhibited the highest accuracy, reference textual overlap, semantic alignment, lower number of errors, and readability among the five evaluated LLMs, making it a valuable assistant for dental healthcare professionals.
