Related Experiment Video
Updated: Feb 9, 2026

07:47
Measuring Delay Discounting in Humans Using an Adjusting Amount Task
Published on: January 9, 2016
16.0K
An Evaluation of Interrater Reliability Measures on Binary Tasks Using d-Prime
Malcolm J Grant1, Cathryn M Button1, Brent Snook1
1Memorial University of Newfoundland, St. John's, Newfoundland and Labrador, Canada.
Applied Psychological Measurement
|June 9, 2018
Summary
For binary tasks, Phi and Kappa coefficients are recommended for assessing interrater agreement. These indices reliably measure reliability across different prevalence, expertise, and bias conditions.
Area of Science:
- Psychometrics
- Statistical Modeling
- Research Methodology
Background:
- Assessing interrater agreement is crucial for reliability in binary tasks.
- Existing indices for interrater agreement have faced criticism regarding their performance.
- A need exists for robust agreement indices that are insensitive to variations in prevalence, rater expertise, and bias.
Purpose of the Study:
- To evaluate the performance of five common interrater agreement indices on binary tasks.
- To identify the most reliable indices under varying conditions of characteristic prevalence, rater expertise, and bias.
- To provide evidence-based recommendations for selecting appropriate interrater agreement measures.
Main Methods:
- Monte Carlo simulations were employed to evaluate five interrater agreement indices.
- The performance of each index was assessed against d-prime, an unbiased measure of rater ability.
- Simulations varied characteristic prevalence, rater expertise, and rater bias to test index robustness.
Main Results:
- Phi and Kappa coefficients demonstrated superior performance across various simulation conditions.
- Percentage Agreement, Scott's Pi, and Gwet's AC1 showed lower correlations with d-prime.
- Phi and Kappa were least affected by changes in prevalence, expertise, and bias.
Conclusions:
- Phi and Kappa are recommended for assessing interrater agreement in binary tasks.
- These indices offer greater reliability and are less influenced by environmental and decision-maker characteristics.
- Researchers should prioritize Phi or Kappa for robust interrater reliability assessments.
Related Concept Videos
Reliability and Validity
14.1K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
14.1K
Binary Fission
3.2K
Binary fission is the primary mode of asexual reproduction in prokaryotes, such as bacteria. It results in the production of two genetically identical daughter cells. This highly efficient process ensures the rapid propagation of bacterial populations under favorable conditions and involves coordinated cellular and molecular events.DNA Replication and SeparationThe process begins with the replication of the bacterial chromosome. The circular DNA molecule unwinds at a specific origin of...
3.2K
Binary Fission
64.0K
Fission is the division of a single entity into two or more parts, which regenerate into separate entities that resemble the original. Organisms in the Archaea and Bacteria domains reproduce using binary fission, in which a parent cell splits into two parts that can each grow to the size of the original parent cell. This asexual method of reproduction produces cells that are all genetically identical.
64.0K
Distribution Reliability and Automation
519
Distribution reliability in electrical power systems is critical for ensuring an uninterrupted power supply to consumers at minimal cost. According to IEEE Standard Terms, reliability is the probability that a device will function without failure over a specified time period or amount of usage. For electric power distribution, this translates to maintaining continuous power supply and addressing customer concerns over power outages. Several indices, as defined by IEEE Standard 1366-2012, are...
519
Self-Evaluation: Self-Enhancement and Self-Verification
5.8K
Social psychologists have documented that feeling good about ourselves and maintaining positive self-esteem is a powerful motivator of human behavior (Tavris & Aronson, 2008). In the United States, members of the predominant culture typically think very highly of themselves and view themselves as good people who are above average on many desirable traits (Ehrlinger, Gilovich, & Ross, 2005). Often, our behavior, attitudes, and beliefs are affected when we experience a threat to our...
5.8K
Nursing Evaluation
4.4K
The evaluation stage signals the end of the nursing process. The nurse gathers evaluative data to assess whether or not the patient has attained the expected results. Whereas the nurse collects data in the nursing assessment to identify the patient's health concerns, the evaluation stage data determines if the indicated health issues are resolved. Evaluative data collection includes two sections: the data acquired to evaluate patient outcomes and the time criteria for data collection.
4.4K

