Related Experiment Videos
A Psychometric Comparison of Faculty-Authored and Large Language Model-Generated Multiple-Choice Questions in
Pradipkumar Damor1, Sarita Gill2, Wasim Wani3
1Department of Conservative Dentistry and Endodontics, School of Dental Sciences, Manav Rachna Dental College, Manav Rachna International Institute of Research and Studies, Faridabad, IND.
Abstract:
Background and objective Designing high-quality multiple-choice questions (MCQs) in dental education is highly resource-intensive. While generative artificial intelligence (AI) offers a rapid, scalable solution for automated item generation, there is limited student-tested psychometric evidence comparing uncurated large language model (LLM) outputs with traditional faculty-authored items in advanced dental specialties like endodontics. This study aimed to evaluate and compare the psychometric properties, specifically the difficulty index, discrimination index, distractor effectiveness, and internal consistency reliability, of faculty-authored and uncurated, LLM-generated endodontic MCQs. Materials and methods Sixty single-best-answer MCQs (30 authored by experienced endodontic faculty and 30 generated by the LLM Claude Sonnet 5 (Anthropic, San Francisco, CA) using a standardized single-prompt framework) were distributed evenly across 10 core endodontic domains. The unified 60-item examination was digitally administered to a convenience sample of 100 dental interns in a single, randomized, 60-minute session. Psychometric parameters were computed using classical test theory, and scale reliabilities were evaluated using Cronbach's alpha (α). Continuous variables were compared using independent-samples t-tests. Results The faculty-authored subscale demonstrated good internal consistency (α = 0.851), while the LLM subscale showed acceptable reliability (α = 0.798). Faculty-authored items had a significantly higher mean discrimination capacity (p = 0.00747) and a more balanced difficulty profile, with 27 (90.0%) of the items falling into the moderate difficulty range. Conversely, LLM-generated items were significantly easier (p = 0.000918), with 13 (43.3%) classified as easy, and had a high rate of non-functional distractors (NFDs). While five (16.7%) of the LLM questions had 100% NFDs (where all distractors failed to function), the faculty cohort had no such items. A concurrency analysis revealed that three (10.0%) of the faculty items simultaneously satisfied all ideal psychometric benchmarks, compared to only one (3.3%) of the LLM items. Conclusions Although next-generation LLMs can produce stylistically authentic clinical vignettes that closely mimic human-authored writing, their raw, uncurated outputs exhibit significant psychometric limitations, particularly regarding distractor plausibility and item discrimination. Fully autonomous AI item generation is not yet viable for high-stakes assessments. A hybrid "human-in-the-loop" model, where LLMs are leveraged for rapid preliminary drafting and experienced educators perform targeted distractor refinement, can safeguard assessment validity while substantially reducing the administrative burden on dental faculty.
Related Concept Videos
Multiple Comparison Tests
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Surveys