Related Experiment Video
Updated: Aug 7, 2026

Reliability of Artificial Intelligence-Based Cone Beam Computed Tomography Integration with Digital Dental Images
Published on: February 23, 2024
Evaluation of AI-Generated Multiple-Choice Questions for Periodontology Exams: A Quality Assessment Study
Bushra Ahmad1, Livia Valverde1, Shruti Jain1
1Department of Periodontology, Tufts University School of Dental Medicine, Boston, Massachusetts, USA.
Background:
This study evaluated the quality of multiple-choice questions (MCQs) generated by artificial intelligence (AI) using ChatGPT-4o compared with faculty-written items in periodontology using the Integrated National Board Dental Examination (INBDE) rubric.
Methods:
Thirty MCQs were assessed in a blinded cross-sectional comparison at Tufts University School of Dental Medicine. Sixteen questions were generated by ChatGPT-4o based on course objectives and INBDE guidelines, and 14 were randomly selected from the departmental exam bank. Fourteen periodontology faculty members rated each item on six INBDE criteria including clarity, content accuracy, distractor quality, fairness, curricular alignment and grammar using a five-point Likert scale ranging from poor (1) to excellent (5). Composite scores were analysed using a generalized linear mixed model.
Results:
AI-generated items achieved significantly higher composite scores than human-written items (20.7 ± 4.9 vs. 18.3 ± 5.1; p < 0.001). In descriptive comparisons, AI-generated items also received higher ratings across all six domains, particularly in clarity and grammar. Reviewers were unable to reliably identify the source of the items, and 84.1% of AI-generated items were judged suitable for exam use compared with 55.7% of faculty-written items.
Conclusions:
ChatGPT-4o produced high-quality and well-structured MCQs, and reviewers frequently reported difficulty distinguishing their origin in this blinded assessment. While these results highlight the potential value of AI-assisted assessment design, expert supervision remains essential to ensure accuracy, cognitive depth and alignment with educational standards. AI should be a supportive tool that complements rather than replaces faculty expertise in item development.