Related Experiment Video
Updated: Feb 17, 2026

A Protocol of Manual Tests to Measure Sensation and Pain in Humans
Published on: December 19, 2016
Establishing Inter- and Intrarater Reliability for High-Stakes Testing Using Simulation
Suzan Kardong-Edgren1, Marilyn H Oermann, Mary Anne Rizzolo
1About the Authors Suzan Kardong-Edgren, PhD, RN, CHSE, FAAN, ANEF, is a professor and director of the RISE Center, School of Nursing and Health Sciences, Robert Morris University, Moon Township, Pennsylvania. Marilyn H. Oermann, PhD, RN, FAAN, ANEF, is Thelma M. Ingles Professor of Nursing and director of evaluation and educational research, Duke University School of Nursing, Durham, North Carolina. Mary Anne Rizzolo, EdD, RN, FAAN, ANEF, is a consultant for the National League for Nursing. Tamara Odom-Maryon, PhD, is a professor of research, Washington State University College of Nursing, Spokane. For more information, contact Dr. Kardong-Edgren at kardongedgren@rmu.edu.
Developing standardized training for raters in high-stakes testing is crucial. This study found that not all faculty are expert evaluators, impacting reliability.
Area of Science:
- Medical Education
- Assessment and Evaluation
Background:
- Simulation is increasingly used in high-stakes testing.
- Research on developing inter- and intrarater reliability for raters in simulation-based assessments is lacking.
Purpose of the Study:
- To develop and report a standardized training method for raters.
- To establish inter- and intrarater reliability among a group of raters for high-stakes testing.
Main Methods:
- Eleven raters underwent standardized training.
- Raters scored 28 student videos over six weeks.
- Intrarater and interrater reliability were assessed by having raters rescore videos over two days.
Main Results:
- Kappa statistics indicated moderate to substantial agreement after excluding two outlier raters.
- One rater showed poor intrarater reliability; another failed all students.
- The standardized training method improved rater reliability.
Conclusions:
- Not all faculty possess the necessary expertise to be evaluators in high-stakes testing, despite content expertise.
- Standardized training is essential for ensuring reliable assessments.
- Careful selection of raters is necessary for valid high-stakes evaluations.
Related Concept Videos
Reliability and Validity
Modeling and Similitude
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
Uncertainty in Measurement: Accuracy and Precision
