Related Experiment Video
Updated: May 13, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
463
Artificial intelligence performance in answering multiple-choice oral pathology questions: a comparative analysis.
Birkan Eyup Yilmaz1, Busra Nur Gokkurt Yilmaz2, Furkan Ozbey3
1Faculty of Dentistry, Department of Oral and Maxillofacial Surgery, Giresun University, Giresun, Türkiye. ylmzbirkan@gmail.com.
BMC Oral Health
|April 15, 2025
Summary
ChatGPT o1 demonstrated superior accuracy in answering oral pathology questions compared to other large language models (LLMs). This study highlights LLMs' potential as educational tools in dental training, though further validation is needed.
Area of Science:
- Artificial Intelligence in Dentistry
- Oral Pathology Education
- Large Language Models (LLMs)
Background:
- Artificial intelligence (AI) is increasingly integrated into healthcare and dental education.
- AI impacts diagnostic processes, treatment planning, and academic training.
- This study evaluates large language models (LLMs) in oral pathology.
Purpose of the Study:
- To assess and compare the performance of various LLMs.
- To analyze accuracy rates in answering multiple-choice oral pathology questions.
- To identify differences in LLM performance based on question type (case-based vs. knowledge-based).
Main Methods:
- Eight LLMs were evaluated using 100 multiple-choice oral pathology questions from the Turkish Dental Specialization Examination (2012-2021).
- Questions were categorized as case-based or knowledge-based.
- Responses were scored as correct or incorrect against official answer keys, with no feedback provided to prevent bias.
Main Results:
- Significant performance variations were observed among LLMs (p < 0.001).
- ChatGPT o1 achieved the highest accuracy (96%), followed by Claude 3.5 (84%), Gemini 2, and Deepseek (82% each).
- ChatGPT o1 and Claude excelled in case-based questions, while ChatGPT o1 and Deepseek led in knowledge-based questions.
Conclusions:
- LLMs exhibit variable proficiency in oral pathology, with ChatGPT o1 demonstrating superior accuracy.
- LLMs show promise as supplementary educational resources in dental training.
- Further research and validation are necessary to confirm the utility of LLMs in this field.

