Related Experiment Video
Updated: May 13, 2026

08:33
A Cross-Disciplinary and Multi-Modal Experimental Design for Studying Near-Real-Time Authentic Examination Experiences
Published on: September 4, 2019
Within-session score gains for repeat examinees on a standardized patient examination
Alex K Chavez1, Kimberly A Swygert, Steven J Peitzman
1National Board of Medical Examiners, Philadelphia, Pennsylvania 19104, USA.
Summary
Examinees show score increases within standardized patient (SP) exams, a warm-up effect that resets between attempts. Across-attempt score gains in the United States Medical Licensing Examination Step 2 Clinical Skills reflect genuine improvement, not just this warm-up.
Area of Science:
- Medical Education
- Assessment and Evaluation
Background:
- Standardized patient (SP) exams are crucial for clinical skills assessment.
- Previous research indicates score improvements both within and across exam attempts.
Purpose of the Study:
- To quantify within-session score gains in repeated United States Medical Licensing Examination (USMLE) Step 2 Clinical Skills exams.
- To determine if within-session score patterns explain score increases observed across multiple exam attempts.
Main Methods:
- Analysis of encounter-level scores from 2,165 examinees who took the USMLE Step 2 Clinical Skills exam twice.
- Application of smoothing, regression, and statistical tests to model score patterns and compare them across attempts.
Main Results:
- Examinees demonstrated within-session score increases over the first 3-6 SP encounters, followed by performance stabilization.
- Across-session score gains could not be attributed to the within-session score trajectory observed in the first attempt.
Conclusions:
- A temporary "warm-up" effect influences performance within a single exam session, resetting between attempts.
- Score improvements across multiple attempts likely represent genuine enhancements in examinee competency.
Related Concept Videos
Reliability and Validity
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
Comparing Experimental Results: Student's t-Test
The t-test is a statistical method used to compare the sample mean with a population mean or compare two means from two data sets. The test statistic is calculated from the standard deviation, mean, and number of measurements in the data set at a selected confidence interval and then compared to a table of critical values at this confidence level. If the test statistic is smaller than the critical value, the null hypothesis is accepted. In this case, we state that the difference between the...
