Related Experiment Videos
Using randomization to compare AI and expert-generated formative assessment questions in medical education
Akshat Kumar1, Katie Stinson2, Litao Wang3
1Department of Internal Medicine, University of California San Diego Health, San Diego, United States.
Medical Education Online
|May 10, 2026
Summary
Artificial intelligence (AI) can generate medical education questions comparable to expert-created ones. Further evaluation methods are needed to assess AI-generated content effectively.
Area of Science:
- Medical Education
- Artificial Intelligence
- Educational Technology
Background:
- Artificial intelligence (AI) is increasingly used in education, including creating complex multiple-choice questions for medical training.
- Evaluating the quality and efficacy of AI-generated educational content presents significant challenges.
- Existing assessment methods may be insufficient for distinguishing AI-generated content from human expert content.
Purpose of the Study:
- To compare medical students' performance and subjective evaluations of AI-generated versus expert-generated questions.
- To test the hypothesis that there would be no significant difference in student outcomes or perceptions between AI and expert-generated questions.
Main Methods:
- A single-center, randomized study design was employed.
- Medical students received one AI- or expert-generated question daily for four weeks.
- Performance and subjective perceptions were analyzed using statistical comparisons.
Main Results:
- Participants reported similar perceptions of AI-generated and expert-generated questions (p=0.18).
- No significant difference was found in the proportion of correct responses between the two question types.
- AI-generated questions were rated as 'very easy' or 'easy' more frequently (53%) than expert-generated questions (31%).
Conclusions:
- Randomization demonstrated that AI-generated questions are nearly indistinguishable from expert-generated questions in terms of student performance and perception.
- The study highlights the necessity for developing advanced evaluation methodologies for AI-generated medical education materials.
- Current evaluation techniques may not adequately capture the subtle differences or potential impacts of AI-generated content.
Related Concept Videos
Randomized Experiments
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
Simple randomization
Simple...
Group Design
The most basic experimental design involves two groups: the experimental group and the control group. The two groups are designed to be the same except for one difference— experimental manipulation. The experimental group gets the experimental manipulation—that is, the treatment or variable being tested—and the control group does not. Since experimental manipulation is the only difference between the experimental and control groups, we can be sure that any differences between the two are due to...
Bioequivalence Experimental Study Designs: Completely Randomized and Randomized Block Designs
Bioequivalence experimental study designs are crucial methodologies used in evaluating and comparing the bioavailability of different drug products. These designs are categorized into various types: completely randomized, randomized block, repeated measures, cross and carry-over, and Latin square designs.Completely randomized designs involve randomly allocating treatments to all subjects participating in the experiment. This allocation is achieved by assigning unique random numbers to subjects...
Random Sampling Method
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. Data are the result of sampling from a population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest. Among the various sampling methods used by...
Bioequivalence Experimental Study Designs: Repeated Measures, Cross-Over, Carry-Over, and Latin Square Designs
Bioequivalence experimental study designs play a pivotal role in testing the effectiveness of various treatments. Key among these are the repeated measures, cross-over, carry-over, and Latin square designs. In the repeated measures design, each subject receives all treatments, allowing for temporal comparisons. This type of design is useful in reducing variability but requires careful planning to avoid bias.The cross-over design, an economical method, involves sequential administration of...
Random and Systematic Errors
Scientists always try their best to record measurements with the utmost accuracy and precision. However, sometimes errors do occur. These errors can be random or systematic. Random errors are observed due to the inconsistency or fluctuation in the measurement process, or variations in the quantity itself that is being measured. Such errors fluctuate from being greater than or less than the true value in repeated measurements. Consider a scientist measuring the length of an earthworm using a...