Related Experiment Video
Updated: Dec 10, 2025

10:39
Qualitative and Quantitative Validation of Tools with Rating Scales Aimed at Assessing the Quality of University Service-Learning
Published on: August 29, 2025
840
Are USMLE Scores Valid Measures for Chief Resident Selection?
Journal of Graduate Medical Education
|September 4, 2020
Summary
USMLE Step 1 and Step 2 scores do not predict chief resident selection. This study found no significant differences in scores between chief residents and non-chief residents, questioning their use in career decisions.
Area of Science:
- Medical Education
- Graduate Medical Education
- Physician Training
Background:
- The United States Medical Licensing Examination (USMLE) Steps 1 and 2 scores are frequently utilized for critical medical career decisions, including residency placement.
- However, the validity of these scores for predicting future performance, such as selection for chief residency, remains inadequately supported by evidence.
Purpose of the Study:
- To investigate the relationship between USMLE Step 1 and Step 2 Clinical Knowledge (CK) scores and the selection of chief residents (CRs) versus non-chief residents (non-CRs).
- To assess the predictive validity of USMLE scores for leadership roles within residency programs.
Main Methods:
- A retrospective cohort study analyzed USMLE Step 1 and Step 2 CK scores from 2015 to 2020.
- Data from 13 programs at a US academic medical center compared scores of 1334 non-CRs with 211 CRs.
Main Results:
- No statistically significant differences were observed in average USMLE Step 1 scores between non-CRs (239.81 ± 14.35) and CRs (240.86 ± 14.31; P = .32).
- Similarly, no significant differences were found in average USMLE Step 2 CK scores between non-CRs (251.06 ± 13.80) and CRs (252.51 ± 14.21; P = .16).
Conclusions:
- USMLE Step 1 and Step 2 CK scores showed no correlation with chief resident selection across various specialties over a six-year period.
- The findings suggest that relying on USMLE scores to identify potential chief residents is not advisable.
Related Concept Videos
Reliability and Validity
13.6K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
13.6K
Spearman's Rank Correlation Test
1.3K
Spearman's rank correlation test, also known as Spearman's rho, is a nonparametric method for assessing the strength and direction of association between two variables. This test is particularly valuable when the data distribution is unknown or when the assumption of normality does not hold. Named after the English psychologist and statistician Dr. Charles Edward Spearman, it serves as the nonparametric counterpart to Pearson's correlation coefficient.
Spearman's test calculates correlation by...
Spearman's test calculates correlation by...
1.3K
Wilcoxon Rank-Sum Test
533
The Wilcoxon rank-sum test, also known as the Mann-Whitney U test, is a nonparametric test used to determine if there is a significant difference between the distributions of two independent samples. This test is designed specifically for two independent populations and has the following key requirements:
533
Measures of Intelligence
8.1K
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
8.1K
Sensitivity, Specificity, and Predicted Value
1.1K
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
1.1K
Wilcoxon Signed-Ranks Test for Matched Pairs
357
The Wilcoxon signed-rank test for matched pairs evaluates the null hypothesis by combining the ranks of differences with their signs. It essentially tests whether the median of the differences in a population of matched pairs is zero. Since the test incorporates more information than the sign test, it generally yields more trustable conclusions. This test also does not require the data to follow a normal distribution, but two conditions must be met for it to be applicable: (1) the data must...
357

