Related Experiment Videos
Psychometric performance and student perceptions of AI- versus student-generated multiple-choice questions: a
Dheyaa Al-Najafi1, Katherine D Krause1, Yundi Wang1
1UBC Faculty of Medicine, University of British Columbia, Vancouver, BC, Canada.
BMC Medical Education
|June 17, 2026
Summary
Large language models (LLMs) significantly accelerate medical education multiple-choice question (MCQ) development, offering comparable psychometric quality and learner acceptability in formative exams. This supports integrating LLM-assisted item generation with expert oversight.
Area of Science:
- Medical Education
- Artificial Intelligence in Assessment
- Psychometrics
Background:
- Developing high-quality medical education exams is resource-intensive.
- Large language models (LLMs) show potential for accelerating question development.
- LLM utility for medical exam development requires further exploration.
Purpose of the Study:
- To evaluate the efficiency, quality, and acceptability of LLM-generated multiple-choice questions (MCQs) compared to student-generated MCQs in medical education.
- To assess the psychometric properties and educational impact of AI-assisted versus traditional MCQ development methods.
Main Methods:
- A participant-blinded, randomized controlled trial involving first-year medical students.
- Comparison of a 112-item mock examination using either AI-generated (ChatGPT/Gemini) or senior student-generated MCQs.
- Evaluation using Van der Vleuten's Assessment Utility Framework, including feasibility, acceptability, item quality, and validity evidence.
Main Results:
- LLM-assisted MCQ development yielded a 5.6-fold efficiency gain over student authorship.
- Student acceptability and perceptions of exam quality were comparable between AI-generated and student-generated exams.
- Slightly higher discrimination indices for student-generated items, with no significant difference in overall student performance or perceived preparedness.
Conclusions:
- LLMs substantially accelerate MCQ development for formative assessments while maintaining comparable psychometric quality and learner acceptability.
- Findings suggest integrating LLM-assisted item generation within a human-in-the-loop framework is beneficial.
- Results may not generalize to high-stakes summative or closed-book examinations.