Related Experiment Video
Updated: Jun 5, 2025

10:42
A Postoperative Evaluation Guideline for Computer-Assisted Reconstruction of the Mandible
Published on: January 28, 2020
6.5K
Evaluating Artificial Intelligence Chatbots in Oral and Maxillofacial Surgery Board Exams: Performance and Potential
Reema Mahmoud1, Amir Shuster2, Shlomi Kleinman3
1Resident, Department of Oral and Maxillofacial Surgery, Tel-Aviv Sourasky Medical Center, Tel Aviv, Israel.
Summary
Generative Pretrained Transformer 4o (GPT-4o) demonstrated superior accuracy and error correction on oral and maxillofacial surgery board questions compared to other large language models (LLMs). This highlights GPT-4o's potential for enhancing surgical education.
Area of Science:
- Artificial Intelligence in Medicine
- Oral and Maxillofacial Surgery (OMS) Education
Background:
- Large language models (LLMs) show promise in medicine but are under-explored in oral and maxillofacial surgery (OMS).
- Evaluating LLM performance on specialized medical examinations is crucial for understanding their educational utility.
Purpose of the Study:
- To assess and compare the accuracy of four leading LLMs on OMS board examination questions.
- To identify specific subject areas where LLMs require improvement for OMS education.
Main Methods:
- An in-silico study evaluated four LLMs (GPT-4o, GPT-3.5, Gemini, Copilot) using 714 OMS board questions.
- Accuracy was measured as the percentage of correct answers, with secondary outcomes including error correction ability across 11 OMS domains.
Main Results:
- GPT-4o achieved the highest accuracy (83.69%), significantly outperforming GPT-3.5 (64.83%), Gemini (66.85%), and Copilot (62.18%).
- GPT-4o also demonstrated superior error correction (98.2%), significantly better than Gemini (70.71%), GPT-3.5 (93.44%), and Copilot (29.26%).
Conclusions:
- GPT-4o shows significant potential as an educational tool in OMS due to its high accuracy and error correction capabilities.
- Variability in performance across domains necessitates ongoing refinement and evaluation for effective integration of LLMs in OMS.

