Setting defensible standards in small cohort OSCEs: Understanding better when borderline regression can 'work'
Matt Homer1, Richard Fuller2, Jennifer Hallam1
1Leeds Institute of Medical Education, School of Medicine, University of Leeds, Leeds, UK.
Medical Teacher
|October 29, 2019
Summary
Borderline regression (BRM) is viable for small cohort Objective Structured Clinical Examinations (OSCEs). Careful scoring instrument design ensures defensible standards, though existing cut-scores are preferred if BRM issues arise.
Area of Science:
- Medical Education
- Assessment Methodology
- Psychometrics
Background:
- Borderline regression (BRM) is often deemed unsuitable for small cohort Objective Structured Clinical Examinations (OSCEs), leading institutions to use resource-intensive item-centered standard-setting methods.
- These traditional methods may lack defensibility, particularly in performance-based assessments.
Purpose of the Study:
- To investigate the applicability and robustness of BRM in small-cohort OSCE contexts.
- To challenge the assumption that BRM is problematic for smaller sample sizes in educational assessments.
Main Methods:
- Analysis of post-hoc station- and test-level metrics from three distinct small-cohort OSCEs.
- Contexts included international medical graduate exams, senior undergraduate exams, and Physician Associate exams.
Main Results:
- BRM yielded robust metrics and defensible cut scores in most stations across the studied contexts (5-14% problematic stations).
- Issues arose when the correlation between global grades and checklist scores was insufficient to support the BRM-determined standard.
Conclusions:
- BRM can provide defensible standards in small test cohorts if there is adequate ability spread and careful design of scoring instruments.
- Existing station cut-scores serve as a preferred alternative when BRM encounters standard-setting problems.
Related Concept Videos
Regression Toward the Mean
6.8K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.8K
Study Design in Statistics
9.9K
A study design is a set of techniques that allow a researcher to collect and analyze data from different variables defined for a specific research problem. Statistics is commonly for effective study design and more robust experiments,
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...
9.9K
Margin of Error
6.8K
The margin of error is also called the maximum error of an estimate. The margin of error is the maximum possible or expected difference between the observed sample parameter value and the actual population parameter value. For proportion, it is the maximum difference between the value of sample proportion obtained from the data and the true value of population proportion. As the true value of the population parameter is not known, the margin of error is calculated using the sample statistic.
6.8K
Confounding in Epidemiological Studies
544
Confounding in statistical epidemiology represents a pivotal challenge, referring to the distortion in the perceived relationship between an exposure and an outcome due to the presence of a third variable, known as a confounder. This variable is associated with both the exposure and the outcome but is not a direct link in their causal chain. Its presence can lead to erroneous interpretations of the exposure's effect, either exaggerating or underestimating the true association. This...
544
Comparing the Survival Analysis of Two or More Groups
519
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
519
Assumptions of Survival Analysis
361
Survival models analyze the time until one or more events occur, such as death in biological organisms or failure in mechanical systems. These models are widely used across fields like medicine, biology, engineering, and public health to study time-to-event phenomena. To ensure accurate results, survival analysis relies on key assumptions and careful study design.
361


