Related Experiment Videos
Psychometric performance and student perceptions of AI- versus student-generated multiple-choice questions: a
Dheyaa Al-Najafi1, Katherine D Krause1, Yundi Wang1
1UBC Faculty of Medicine, University of British Columbia, Vancouver, BC, Canada.
Background:
Developing high-quality multiple-choice examinations in medical education is time- and resource-intensive. Large language models (LLMs) offer a promising approach to accelerate question development; however, their utility for exam development remains underexplored.
Methods:
The trial was a participant-blinded, parallel-group randomized controlled trial conducted among first-year medical students. Students were randomized to complete a 112-item case-based, single-best-answer mock examination composed of either AI-generated or student-generated multiple-choice questions (MCQs). Questions were developed using identical curricular objectives. AI-generated items were produced via a dual-model workflow (ChatGPT for generation; Google Gemini for validation); student-generated items were authored by senior medical students. Outcomes were evaluated using Van der Vleuten's Assessment Utility Framework across feasibility, acceptability, item quality, internal consistency, validity evidence, and self-perceived educational impact. Primary analyses were conducted in the intention-to-treat (ITT) population using appropriate parametric or non-parametric tests, with effect sizes and 95% confidence intervals reported.
Results:
A total of 258 students were randomized, with 127 allocated to the AI-generated exam arm and 131 to the student-generated exam arm. LLM-assisted MCQ development achieved a 5.6-fold efficiency gain compared with student authorship (4.2 ± 1.9 vs. 19.6 ± 7.5 min per item; p < 0.0001). Student perceptions of exam acceptability-including clarity, difficulty, relevance, and educational value-were comparable between AI-generated and student-generated exams (all Bonferroni-adjusted p ≥ 0.12; all Cohen's |d| ≤ 0.31). Student-generated items demonstrated slightly higher discrimination indices than AI-generated items, though the effect size was small, and distractor efficiency did not differ between protocols. Student performance was marginally higher on the student-generated exam, though this difference was not significant in the ITT analysis. Exploratory analyses identified theme-specific performance variation between exam formats. Neither exam meaningfully changed students' perceived preparedness.
Conclusions:
In a formative, open-resource examination setting with student-generated comparators, LLMs can substantially accelerate MCQ development while producing assessments that are psychometrically comparable and acceptable to learners. These findings may not generalize to high-stakes summative or closed-book assessment settings. Although small differences persist, these findings support the integration of LLM-assisted item generation within a human-in-the-loop framework, combining AI efficiency with expert oversight to preserve psychometric quality.
Trial Registration:
This study was retrospectively registered on ClinicalTrials.gov (Identifier NCT07481162 registered March 18, 2026). Prospective registration was not performed as the study was conducted as an embedded educational intervention within a voluntary formative examination setting. The study protocol and statistical analysis plan were prespecified prior to data analysis. The trial is reported in accordance with CONSORT 2025 guidelines.