Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Multiple Comparison Tests01:13

Multiple Comparison Tests

3.9K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
3.9K
Reliability and Validity01:29

Reliability and Validity

12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
Testing a Claim about Population Proportion01:24

Testing a Claim about Population Proportion

3.3K
A complete procedure for testing a claim about a population proportion is provided here.
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
3.3K
Surveys02:16

Surveys

14.8K
Often, psychologists develop surveys as a means of gathering data. Surveys are lists of questions to be answered by research participants, and can be delivered as paper-and-pencil questionnaires, administered electronically, or conducted verbally. Generally, the survey itself can be completed in a short time, and the ease of administering a survey makes it easy to collect data from a large number of people.
14.8K
Detection of Gross Error: The Q Test01:00

Detection of Gross Error: The Q Test

6.1K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.1K
Cochran's Q Test01:17

Cochran's Q Test

347
Cochran's Q Test is a nonparametric statistical test used to determine if there are potential differences in the outcomes of three or more related groups on a binary (yes/no) or dichotomous outcome. It is essentially an extension of the McNemar Test, which is limited to two related samples - Cochran's Q test can handle three or more related samples, making it more versatile in scenarios where subjects are measured under multiple conditions. The test statistic follows a Chi-Square...
347

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Non-Pharmacologic Manual Therapies for Postoperative Bowel Dysfunction: A Systematic Review and Meta-Analysis.

Journal of clinical medicine·2026
Same author

Improving splice site usage prediction with SPLAIRE.

bioRxiv : the preprint server for biology·2026
Same author

Sex-related structural alterations across common epilepsies: a worldwide ENIGMA study.

bioRxiv : the preprint server for biology·2026
Same author

TOPODIFFUSIONNET: A TOPOLOGY-AWARE DIFFUSION MODEL.

... International Conference on Learning Representations·2026
Same author

Comparing the measurement properties, informativity and responsiveness of the DLQI, DLQI-Relevant (DLQI-R), and Skindex-29 in patients with chronic skin disease: a prospective cohort study.

Quality of life research : an international journal of quality of life aspects of treatment, care and rehabilitation·2026
Same author

Association of psychiatric comorbidity with worse quality of life despite clear or almost clear skin.

Journal of the European Academy of Dermatology and Venereology : JEADV·2026

Related Experiment Video

Updated: Jul 6, 2025

Problem-Solving Before Instruction PS-I: A Protocol for Assessment and Intervention in Students with Different Abilities
10:26

Problem-Solving Before Instruction PS-I: A Protocol for Assessment and Intervention in Students with Different Abilities

Published on: September 11, 2021

4.0K

ChatGPT 3.5 fails to write appropriate multiple choice practice exam questions.

Alexander Ngo1, Saumya Gupta1, Oliver Perrine1

  • 1Department of Pathology & Laboratory Medicine, Boston University Chobanian and Avidesian School of Medicine, Boston MA, USA.

Academic Pathology
|January 1, 2024
PubMed
Summary

Artificial intelligence (AI) can impact academic teaching. While ChatGPT shows potential for creating educational content like multiple-choice questions, it currently requires significant instructor editing due to accuracy issues.

Keywords:
Artificial intelligenceEducationImmunologyPathology

More Related Videos

Development and Implementation of a Multi-Disciplinary Technology Enhanced Care Pathway for Youth and Adults with Concussion
08:13

Development and Implementation of a Multi-Disciplinary Technology Enhanced Care Pathway for Youth and Adults with Concussion

Published on: January 20, 2019

6.6K
A Cross-Disciplinary and Multi-Modal Experimental Design for Studying Near-Real-Time Authentic Examination Experiences
00:08

A Cross-Disciplinary and Multi-Modal Experimental Design for Studying Near-Real-Time Authentic Examination Experiences

Published on: September 4, 2019

7.0K

Related Experiment Videos

Last Updated: Jul 6, 2025

Problem-Solving Before Instruction PS-I: A Protocol for Assessment and Intervention in Students with Different Abilities
10:26

Problem-Solving Before Instruction PS-I: A Protocol for Assessment and Intervention in Students with Different Abilities

Published on: September 11, 2021

4.0K
Development and Implementation of a Multi-Disciplinary Technology Enhanced Care Pathway for Youth and Adults with Concussion
08:13

Development and Implementation of a Multi-Disciplinary Technology Enhanced Care Pathway for Youth and Adults with Concussion

Published on: January 20, 2019

6.6K
A Cross-Disciplinary and Multi-Modal Experimental Design for Studying Near-Real-Time Authentic Examination Experiences
00:08

A Cross-Disciplinary and Multi-Modal Experimental Design for Studying Near-Real-Time Authentic Examination Experiences

Published on: September 4, 2019

7.0K

Area of Science:

  • Educational Technology
  • Artificial Intelligence in Education

Background:

  • Artificial intelligence (AI) presents both challenges and opportunities for traditional academic teaching.
  • Concerns exist regarding AI tools like ChatGPT generating original essays.
  • AI may serve as a valuable tool to augment existing teaching methodologies.

Purpose of the Study:

  • To evaluate the efficacy of ChatGPT 3.5 in generating multiple-choice questions (MCQs) with explanations for academic purposes.
  • To assess the accuracy and completeness of AI-generated MCQs and their explanations.

Main Methods:

  • Utilized ChatGPT 3.5 to generate 60 MCQs based on uploaded author-written text.
  • Instructed the AI to provide explanations for correct and incorrect answers for each MCQ.
  • Analyzed the generated MCQs for accuracy of questions, correctness of answers, and quality of explanations.

Main Results:

  • ChatGPT 3.5 successfully generated accurate questions and answers with explanations for only 32% (19 out of 60) of the MCQs.
  • A significant portion of generated questions (25%) contained incorrect or misleading answers.
  • In many cases, ChatGPT failed to provide explanations for incorrect answer choices.

Conclusions:

  • Current AI models like ChatGPT 3.5 demonstrate limited reliability for autonomously generating accurate educational assessments.
  • Extensive human review and editing are essential when using AI-generated content for practice exams or assessments.
  • Despite limitations, AI tools may still offer supplementary benefits for instructors in creating draft assessment materials.