Related Experiment Video
Updated: Aug 12, 2026

07:08
Estimate the Cognitive Load Using Electrocardiographic Measure: A Human-AI Collaborative Task
Published on: December 5, 2025
Human-Edited Generative AI-Assisted Multiple-Choice Questions in Postgraduate Family Medicine: Blinded
Kai Ping Sze1,2, Jia Qing Lim2, Yng Miin Loke3
1LKC School of Medicine, Nanyang Technological University, Singapore, Singapore.
JMIR Medical Education
|August 10, 2026
Summary
Human-edited generative artificial intelligence (GenAI) multiple-choice questions (MCQs) showed plausible difficulty but lower psychometric quality than educator-crafted MCQs for postgraduate assessment. GenAI should augment, not replace, educator expertise in question development.
Area of Science:
- Medical Education
- Artificial Intelligence in Health Professions Education
- Psychometric Evaluation
Background:
- Generative artificial intelligence (GenAI) is increasingly used for drafting multiple-choice questions (MCQs) in health professions education.
- Evidence often focuses on raw model outputs or expert ratings, leaving the psychometric readiness of educator-edited GenAI items for postgraduate assessment unclear.
Purpose of the Study:
- To compare human-edited GenAI-assisted MCQs with educator-crafted MCQs for postgraduate Family Medicine assessment.
- To evaluate item difficulty, discrimination, reliability, distractor functioning, and participant perceptions.
Main Methods:
- A blinded, cross-sectional, within-participant comparative psychometric evaluation of 60 MCQs (30 GenAI-assisted, 30 educator-crafted) was conducted.
- Seventy-three postgraduate doctors preparing for the Family Medicine Applied Knowledge Test participated.
- Outcomes included item difficulty, discrimination (corrected point-biserial), reliability (KR-20), distractor functioning, and perceived item qualities.
Main Results:
- GenAI-assisted MCQs yielded lower scores (mean difference -1.97, P<.001) and lower reliability (0.38 vs 0.60) compared to educator-crafted items.
- Item difficulty was comparable (0.64 vs 0.70), but GenAI-assisted items had lower discrimination (0.09 vs 0.18, P=.04) and more negative discrimination.
- Distractor functioning was poorer for GenAI-assisted items, though not statistically significant; perceived qualities did not differ.
Conclusions:
- Human-edited GenAI-assisted MCQs can achieve acceptable difficulty but do not inherently ensure psychometric readiness for assessment.
- GenAI should serve as a drafting tool within educator-led workflows, emphasizing verification, distractor refinement, and empirical analysis.
- Rigorous piloting, analysis, and revision are crucial before incorporating GenAI-assisted items into item banks or summative assessments.