Related Experiment Video
Updated: Jul 16, 2025

09:36
Assessment of Child Anthropometry in a Large Epidemiologic Study
Published on: February 2, 2017
27.1K
Is It All About the Form? Norm- vs Criterion-Referenced Ratings and Faculty Inter-Rater Reliability
Shannon A Scielzo1, Kareem Abdelfattah2, Hilary F Ryder3,4
1Department of Internal Medicine, University of Texas Southwestern Medical Center, Dallas, TX.
Ochsner Journal
|September 15, 2023
Summary
Criterion-referenced evaluations show higher inter-rater reliability than norm-referenced ones for resident performance assessments. This suggests criterion-referenced scaling offers more valid data for evaluating resident competence.
Area of Science:
- Medical Education Research
- Healthcare Quality Improvement
- Performance Assessment
Background:
- Limited research exists on the data quality of resident performance evaluations.
- This study addresses the need to compare different evaluation scaling approaches.
Purpose of the Study:
- To compare inter-rater reliability between norm-referenced and criterion-referenced evaluation scaling methods.
- To assess which scaling approach yields more reliable data for faculty performance evaluations of residents.
Main Methods:
- Examined resident performance evaluation data from 426 residents across 3 programs at 2 institutions.
- Calculated faculty inter-rater reliability using intraclass correlation coefficients (ICCs) for criterion-referenced and norm-referenced forms.
- Analyzed reliability across competency areas, evaluation forms, and scaling types.
Main Results:
- Criterion-referenced scaling demonstrated higher average inter-rater reliability across all competencies compared to norm-referenced scaling.
- Aggregate scores for criterion-referenced scaling showed significantly higher reliability (z=1.37) than norm-referenced scaling (z=0.88).
- Distributions of composite scores indicated criterion-referenced evaluations better represented the resident performance continuum.
Conclusions:
- Criterion-referenced evaluation approaches provide superior inter-rater reliability compared to norm-referenced approaches.
- Using criterion-referenced scaling may lead to more valid data in resident evaluations.
- Further research is needed to establish best practices for resident evaluation.
Related Concept Videos
Reliability and Validity
12.8K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.8K
Ratio Level of Measurement
18.4K
The way a set of data is measured is called its level of measurement. Correct statistical procedures depend on a researcher being familiar with levels of measurement. For analysis, data are classified into four levels of measurement—nominal, ordinal, interval, and ratio.
A set of data measured using the ratio scale takes care of the ratio problem and provides complete information. Ratio scale data are like interval scale data, except they have a zero point and ratios can be calculated....
A set of data measured using the ratio scale takes care of the ratio problem and provides complete information. Ratio scale data are like interval scale data, except they have a zero point and ratios can be calculated....
18.4K
Kendall's Coefficient of Concordance
402
Kendall's Coefficient of Concordance (W), also known as Kendall's W, is a non-parametric statistical measure used to assess the agreement or concordance between multiple raters or judges when they rank a set of items. It is often used when you have ordinal data (ranks) and you want to see if there is consistency or consensus among the raters. It is widely applied in research areas such as psychology, medicine, and social sciences, where multiple judges are asked to rank or rate subjects...
402
Review and Preview
7.6K
In statistics, several tools are used to interpret the data. Measures of central tendency represent the characteristics of the data, such as mean, median, and mode. Additionally, measures of variance like standard deviation and range are used to find the spread of data from the mean. Relative standing measures the distance between data locations. Commonly used measures of relative standings are percentile, z score, and quartiles.
Percentiles are a type of fractile that partition data into...
Percentiles are a type of fractile that partition data into...
7.6K
Measures of Intelligence
7.6K
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
7.6K
Surveys
14.8K
Often, psychologists develop surveys as a means of gathering data. Surveys are lists of questions to be answered by research participants, and can be delivered as paper-and-pencil questionnaires, administered electronically, or conducted verbally. Generally, the survey itself can be completed in a short time, and the ease of administering a survey makes it easy to collect data from a large number of people.
14.8K

