Related Experiment Video
Updated: Jan 17, 2026

Estimate the Cognitive Load Using Electrocardiographic Measure: A Human-AI Collaborative Task
Published on: December 5, 2025
Faculty versus artificial intelligence chatbot: a comparative analysis of multiple-choice question quality in
Anup Kumar D Dhanvijay1, Amita Kumari1, Mohammed Jaffer Pinjar1
1Department of Physiology, All India Institute of Medical Sciences, Deoghar, Jharkhand, India.
Artificial intelligence (AI) chatbots can generate multiple-choice questions (MCQs) for medical education, but faculty-created questions remain superior. AI-generated MCQs showed comparable discrimination but were easier and had more ineffective distractors than faculty-authored Physiology MCQs.
Area of Science:
- Medical Education
- Artificial Intelligence in Education
- Physiology Assessment
Background:
- Multiple-choice questions (MCQs) are crucial for medical education assessment.
- Human-generated MCQs require significant time and pedagogical expertise.
- AI offers automated MCQ generation, but its educational validity needs evaluation.
Purpose of the Study:
- To compare the psychometric quality of Physiology MCQs generated by faculty versus an AI chatbot (DeepSeek R1).
- To assess difficulty index, discrimination index, and nonfunctional distractors for both question types.
Main Methods:
- 200 MCQs developed: 100 by faculty, 100 by AI.
- 50 MCQs from each group administered to undergraduate medical students.
- Item analysis using difficulty index (DIFI), discrimination index (DI), and nonfunctional distractors (NFDs).
Main Results:
- AI-generated MCQs had a significantly higher DIFI (0.64) than faculty MCQs (0.47).
- No significant difference in DI between AI and faculty MCQs (P = 0.17).
- Faculty MCQs had significantly fewer NFDs (median 0) than AI MCQs (median 1).
Conclusions:
- AI-generated MCQs have comparable discrimination but are easier and contain more ineffective distractors.
- AI shows promise for MCQ generation but requires refinement in distractor quality and item difficulty.
- AI can supplement, but not replace, human expertise in assessment design.
Related Concept Videos
Non-equilibrium in the Cell
Multiple Comparison Tests
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Bioequivalence: Overview
Overview of Anatomy and Physiology