Related Experiment Video
Updated: Jun 13, 2026

Systematic Assessment of Mammalian Skull Specimens for Dental and Temporomandibular Joint Pathology
Published on: August 22, 2022
Benchmarking GPT-5, Gemini 2.5 Pro, Grok 4, and other LLMs on pediatric dentistry questions from a dental
Sukriye Turkoglu Kayaci1, Hamza Osman Ilhan2, Melek Tassoker3
1Department of Pediatric Dentistry, Faculty of Dentistry, University of Health Sciences, Istanbul, Turkey.
Abstract:
Artificial intelligence (AI), particularly large language models (LLMs), is an increasingly prominent tool in medical and dental education. Trained with deep learning and NLP techniques, these models interpret meaning, generate text, and manage complex information. They hold potential for practical educational applications, such as supporting exam preparation and personalized learning. Moreover, their performance in clinical case recognition suggests an emerging potential for use in diagnostic decision-support systems. This study aimed to evaluate the performance of state-of-the-art large language models (LLMs) on pediatric dentistry questions from the Dentistry Specialization Examination (DUS) in Türkiye, a high-stakes national exam for postgraduate training. A total of 119 pediatric dentistry questions from the past ten years of the DUS were compiled and presented to 11 recently developed LLMs (17 with reasoning mode activations), including GPT-5, Gemini 2.5 Pro, Grok-4, and DeepSeek R1. Each model's accuracy (%) and average response generation time (seconds) were calculated and compared. The Gemini 2.5 Pro (92.44%) demonstrated significantly higher mean scores compared to all other models except GPT-4 (78.15%), GPT-5 (90.76%), GPTOSS (78.99, 75.63%), and Grok-4 (88.24%). Similar patterns were also observed for GPT-5 and Grok-4. In contrast, Qwen-3 (49.58 - Reasoning, 54.62 - No Reasoning), and MedGemma (58.82%) exhibited notably lower accuracy rates across most comparisons. Overall, these findings highlight Gemini 2.5 Pro and GPT-5 as achieving the highest accuracy levels among the models, whereas Qwen-3, Qwen-3 (R), and MedGemma demonstrated the weakest performance. While models such as Gemma, LLaMA, and Mistral demonstrated faster response times (<1 s), they exhibited relatively low accuracy. In contrast, reasoning-intensive modes (e.g., DeepSeek R1) improved accuracy but required excessively long generation times (up to 68 s). LLM performance on pediatric dentistry questions was highly variable. Top models, notably Gemini 2.5 Pro (92.44%) and GPT-5 (90.76%), approached expert-level accuracy, while others (e.g., Qwen-3) performed poorly. A critical speed-accuracy trade-off was evident: reasoning modes improved scores but were impractically slow, whereas faster models had low accuracy. This variability necessitates careful validation before LLMs are used in high-stakes dental education or assessment.
Related Concept Videos
Teeth
In the bud stage, the tooth germ (an aggregation of cells) starts to form in the developing jawbone. During the cap stage, the tooth germ differentiates into enamel organ, dental papilla, and dental sac, which will later develop into the tooth's enamel, dentin and...
Assessment of the Mouth
Mouth Inspection
The inspection begins with visually examining the mouth for symmetry, color, and size.
Serum Laboratory Studies, Stool Test, Breath Test
