Related Experiment Video
Updated: Jan 17, 2026

Estimate the Cognitive Load Using Electrocardiographic Measure: A Human-AI Collaborative Task
Published on: December 5, 2025
Faculty versus artificial intelligence chatbot: a comparative analysis of multiple-choice question quality in
Anup Kumar D Dhanvijay1, Amita Kumari1, Mohammed Jaffer Pinjar1
1Department of Physiology, All India Institute of Medical Sciences, Deoghar, Jharkhand, India.
Abstract:
Multiple-choice questions (MCQs) are widely used for assessment in medical education. While human-generated MCQs benefit from pedagogical insight, creating high-quality items is time intensive. With the advent of artificial intelligence (AI), tools like DeepSeek R1 offer potential for automated MCQ generation, though their educational validity remains uncertain. With this background, this study compared the psychometric quality of Physiology MCQs generated by faculty and an AI chatbot. A total of 200 MCQs were developed following the standard syllabus and question design guidelines: 100 by the Physiology faculty and 100 by the AI chatbot DeepSeek R1. Fifty questions from each group were randomly selected and administered to undergraduate medical students in 2 hours. Item analysis was conducted postassessment using difficulty index (DIFI), discrimination index (DI), and nonfunctional distractors (NFDs). Statistical comparisons were made using t tests or nonparametric equivalents, with significance at P < 0.05. Chatbot-generated MCQs had a significantly higher DIFI (0.64 ± 0.22) than faculty MCQs (0.47 ± 0.19; P < 0.0001). No significant difference in DI was found between the groups (P = 0.17). Faculty MCQs had significantly fewer NFDs (median 0) compared to chatbot MCQs (median 1; P = 0.0063). AI-generated MCQs demonstrated comparable discrimination ability but were generally easier and contained more ineffective distractors. While chatbots show promise in MCQ generation, further refinement is needed to improve distractor quality and item difficulty. AI can complement but not yet replace human expertise in assessment design.NEW & NOTEWORTHY This study contributes to the growing research on artificial intelligence (AI)- versus faculty-generated multiple-choice questions in Physiology. Psychometric analysis showed that AI-generated items were generally easier but had comparable discrimination ability to faculty-authored questions, while containing more nonfunctional distractors. By focusing on Physiology, this work offers discipline-specific insights and underscores both the potential and current limitations of AI in assessment development.
Related Concept Videos
Non-equilibrium in the Cell
Multiple Comparison Tests
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Bioequivalence: Overview
Overview of Anatomy and Physiology