Related Experiment Video
Updated: Jul 9, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance of large language models on undergraduate endodontic multiple-choice questions
Meltem Sümbüllü1, Oğuzhan Ünal2, İlke Menteş3
1Department of Endodontics, Faculty of Dentistry, Atatürk University, Erzurum, Türkiye. meltem_endo@hotmail.com.
BMC Oral Health
|July 7, 2026
Summary
ChatGPT-5.2 and Gemini-3 showed superior accuracy and consistency in answering endodontic questions compared to DeepSeek-V3.2. AI tools show promise for dental education but require careful evaluation.
Area of Science:
- Artificial Intelligence in Dental Education
- Large Language Models (LLMs) in Endodontics
Background:
- Undergraduate endodontic education relies on accurate knowledge assessment.
- Large Language Models (LLMs) are increasingly explored for educational support.
Purpose of the Study:
- To assess the accuracy and consistency of three LLMs (ChatGPT-5.2, Gemini-3, DeepSeek-V3.2) in responding to endodontic multiple-choice questions.
- To evaluate LLM performance across different endodontic topics and times of day.
Main Methods:
- 60 multiple-choice questions across six endodontic topics were used.
- Each LLM was queried at three daily intervals over four days.
- Accuracy and consistency were statistically analyzed.
Main Results:
- ChatGPT-5.2 and Gemini-3 significantly outperformed DeepSeek-V3.2 in accuracy and consistency (p < 0.05).
- Model performance varied by question topic, with ChatGPT-5.2 and Gemini-3 showing topic-dependent accuracy.
- LLM performance remained stable across different times of day.
Conclusions:
- Advanced LLMs show potential as supplementary tools for undergraduate endodontic education.
- Critical evaluation of AI-generated content is essential due to inter-model performance variations and topic-specific accuracy differences.