Related Experiment Video
Updated: Aug 6, 2026

Reliability of Artificial Intelligence-Based Cone Beam Computed Tomography Integration with Digital Dental Images
Published on: February 23, 2024
Large language model accuracy in dental radiology: effects of cognitive complexity and content domain
1Department of Dentomaxillofacial Radiology, Faculty of Dentistry, Afyonkarahisar Health Sciences University, 03030, Afyonkarahisar, Turkey. dtemresozen@gmail.com.
Abstract:
Large language models (LLMs) are increasingly used to answer medical questions; however, their performance may vary depending on task characteristics. This study evaluated the performance of multiple versions of two widely used LLM families on oral and maxillofacial radiology (OMFR) questions from the Turkish Dental Specialty Examination (DUS) across three assessment phases and examined the influence of cognitive complexity, model family, evaluation phase, and content domain on response accuracy. A comparative repeated-evaluation design was used. A total of 123 text-based OMFR questions from DUS examinations (2012-2021) were submitted to two widely used LLM families (ChatGPT and DeepSeek) across three evaluation phases (May 2025, August 2025, and February 2026). Questions were categorized by content domain and Bloom cognitive level (low vs. high). Model responses were evaluated using official answer keys, and generalized estimating equations (GEE) were applied to account for repeated measurements. A total of 1230 model responses were analyzed, yielding an overall accuracy of 83.7%. Agreement between repeated runs was substantial to almost perfect (κ = 0.689-0.912). Cognitive complexity emerged as the strongest determinant of performance, with low-level questions significantly more likely to be answered correctly than high-level questions (OR = 6.15, p = 0.003). Content domain was also associated with accuracy (p = 0.028), whereas no statistically significant associations were observed for model family or evaluation phase. LLMs demonstrated high accuracy in answering OMFR examination questions; however, performance was more strongly associated with cognitive complexity than with model family or evaluation phase.

