Related Experiment Video
Updated: Mar 7, 2026

Project-Based Learning Guidelines for Health Sciences Students: An Analysis with Data Mining and Qualitative Techniques
Published on: December 9, 2022
Assessing the Utility of AI Versus Human-Created MCQs in Pediatric Medical Education
James Knight1,2, Richard G McGee2,3,4, Bunmi S Malau-Aduli1,2
1Faculty of Medicine and Health, School of Rural Health, The University of New England, Armidale, NSW, Australia.
Background:
Multiple-choice questions (MCQs) remain central to assessment in medical education, but their development is resource intensive. Generative artificial intelligence (AI) offers a potential solution by automating MCQ creation. However, little is known about the psychometric quality of AI-generated MCQs compared with human-authored items, particularly in pediatric education.
Objective:
This study aimed to directly compare the quality of AI- and human-generated MCQs in pediatrics using item analysis grounded in classical test theory.
Methods:
A formative exam comprising both AI (Microsoft Copilot) and human-generated pediatric MCQs was administered to 4th-year medical students. Item analysis was performed to calculate difficulty indices, discrimination indices, item-total correlations, and distractor functioning. Reliability was assessed using KR-20. Descriptive and inferential statistics, including paired t-tests, compared performance between AI and human items.
Results:
Human-authored questions outperformed AI-generated questions across all quality indicators. AI questions showed lower discrimination (mean 0.19 vs. 0.29) and a higher proportion outside the acceptable difficulty range (56% vs. 32%). Distractor analysis also favored human questions, with fewer nonfunctioning distractors and more items containing fully functional distractors. While some AI items met ideal psychometric thresholds, overall consistency was lower.
Conclusion:
Generative AI in its current form cannot yet match human expertise in producing consistently high-quality MCQs for pediatrics. However, AI shows potential as a supplementary tool, particularly within hybrid human-AI workflows that combine efficiency with expert oversight. These findings highlight both the opportunities and limitations of AI in medical education assessment and underscore the importance of balancing reliability, validity, acceptability, and cost-effectiveness when integrating AI into assessment design.
Related Concept Videos
Methods of Documentation III: PIE
The Availability Heuristic