Related Experiment Video
Updated: Jul 6, 2025

Problem-Solving Before Instruction PS-I: A Protocol for Assessment and Intervention in Students with Different Abilities
Published on: September 11, 2021
ChatGPT 3.5 fails to write appropriate multiple choice practice exam questions
Alexander Ngo1, Saumya Gupta1, Oliver Perrine1
1Department of Pathology & Laboratory Medicine, Boston University Chobanian and Avidesian School of Medicine, Boston MA, USA.
Artificial intelligence (AI) can impact academic teaching. While ChatGPT shows potential for creating educational content like multiple-choice questions, it currently requires significant instructor editing due to accuracy issues.
Area of Science:
- Educational Technology
- Artificial Intelligence in Education
Background:
- Artificial intelligence (AI) presents both challenges and opportunities for traditional academic teaching.
- Concerns exist regarding AI tools like ChatGPT generating original essays.
- AI may serve as a valuable tool to augment existing teaching methodologies.
Purpose of the Study:
- To evaluate the efficacy of ChatGPT 3.5 in generating multiple-choice questions (MCQs) with explanations for academic purposes.
- To assess the accuracy and completeness of AI-generated MCQs and their explanations.
Main Methods:
- Utilized ChatGPT 3.5 to generate 60 MCQs based on uploaded author-written text.
- Instructed the AI to provide explanations for correct and incorrect answers for each MCQ.
- Analyzed the generated MCQs for accuracy of questions, correctness of answers, and quality of explanations.
Main Results:
- ChatGPT 3.5 successfully generated accurate questions and answers with explanations for only 32% (19 out of 60) of the MCQs.
- A significant portion of generated questions (25%) contained incorrect or misleading answers.
- In many cases, ChatGPT failed to provide explanations for incorrect answer choices.
Conclusions:
- Current AI models like ChatGPT 3.5 demonstrate limited reliability for autonomously generating accurate educational assessments.
- Extensive human review and editing are essential when using AI-generated content for practice exams or assessments.
- Despite limitations, AI tools may still offer supplementary benefits for instructors in creating draft assessment materials.
More Related Videos
08:13Development and Implementation of a Multi-Disciplinary Technology Enhanced Care Pathway for Youth and Adults with Concussion
Published on: January 20, 2019
00:08A Cross-Disciplinary and Multi-Modal Experimental Design for Studying Near-Real-Time Authentic Examination Experiences
Published on: September 4, 2019
Related Concept Videos
Multiple Comparison Tests
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Reliability and Validity
Testing a Claim about Population Proportion
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
Surveys
Detection of Gross Error: The Q Test
Cochran's Q Test