Related Experiment Videos
Standards and reliability in evaluation: when rules of thumb don't apply
1Institute for Clinical Evaluation, American Board of Internal Medicine, Philadelphia, Pennsylvania 19106-3699, USA. jnorcini@abim.org
Summary
Two common evaluation rules—absolute standards and high reliability—may not always apply. Relative standards can be better for selection and classroom tests, while other metrics may indicate reproducibility better than .80 reliability.
Area of Science:
- Educational Measurement
- Psychometrics
- Program Evaluation
Background:
- Established evaluation practices often rely on absolute standards and high reliability coefficients (e.g., .80).
- These rules of thumb are widely accepted but may not be universally applicable across all assessment contexts.
- Understanding the limitations of these guidelines is crucial for accurate and meaningful evaluation.
Purpose of the Study:
- To identify specific scenarios where the rules of thumb for absolute standards and high reliability in evaluation are not optimal.
- To explore alternative approaches to setting standards and assessing test reproducibility.
- To provide guidance on selecting appropriate evaluation metrics based on context.
Main Methods:
- Conceptual analysis of evaluation principles.
- Examination of decision-making contexts (e.g., selection, classroom testing).
- Discussion of alternative reproducibility indicators beyond traditional reliability coefficients.
Main Results:
- Absolute standards may be less suitable than relative standards in selection and classroom testing.
- Reliability coefficients of .80 or higher are not always the best measure of reproducibility.
- Standard error of measurement, pass/fail classification consistency, and domain-referenced reliability can be more informative indicators.
Conclusions:
- Evaluation standards should be context-dependent, with relative standards sometimes being more appropriate.
- The .80 reliability benchmark is not a universal requirement; other measures of reproducibility should be considered.
- Choosing the right evaluation metrics enhances the validity and utility of assessment results.