Related Experiment Video
Updated: Feb 9, 2026

Measuring Delay Discounting in Humans Using an Adjusting Amount Task
Published on: January 9, 2016
An Evaluation of Interrater Reliability Measures on Binary Tasks Using d-Prime
Malcolm J Grant1, Cathryn M Button1, Brent Snook1
1Memorial University of Newfoundland, St. John's, Newfoundland and Labrador, Canada.
Abstract:
Many indices of interrater agreement on binary tasks have been proposed to assess reliability, but none has escaped criticism. In a series of Monte Carlo simulations, five such indices were evaluated using d-prime, an unbiased indicator of raters' ability to distinguish between the true presence or absence of the characteristic being judged. Phi and, to a lesser extent, Kappa coefficients performed best across variations in characteristic prevalence, and raters' expertise and bias. Correlations with d-prime for Percentage Agreement, Scott's Pi, and Gwet's AC1 were markedly lower. In situations where two raters make a series of binary judgments, the findings suggest that researchers should choose Phi or Kappa to assess interrater agreement as the superiority of these indices was least influenced by variations in the decision environment and characteristics of the decision makers.
Related Concept Videos
Reliability and Validity
Binary Fission
Binary Fission
Distribution Reliability and Automation
Self-Evaluation: Self-Enhancement and Self-Verification
Nursing Evaluation

