Related Experiment Video
Updated: Mar 27, 2026

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
The Effect of the Raters' Marginal Distributions on Their Matched Agreement: A Rescaling Framework for Interpreting
Tzur M Karelitz1, David V Budescu2
1a National Institute for Testing and Evaluation , Jerusalem , Israel.
Abstract:
Cohen's κ measures the improvement in classification above chance level and it is the most popular measure of interjudge agreement. Yet, there is considerable confusion about its interpretation. Specifically, researchers often ignore the fact that the observed level of matched agreement is bounded from above and below and the bounds are a function of the particular marginal distributions of the table. We propose that these bounds should be used to rescale the components of κ (observed and expected agreement). Rescaling κ in this manner results in κ', a measure that was originally proposed by Cohen (1960) and was largely ignored in both research and practice. This measure provides a common scale for agreement measures of tables with different marginal distributions. It reaches the maximal value of 1 when the judges show the highest level of agreement possible, given their marginal disagreements. We conclude that κ' should be used to measure the level of matched agreement contingent on a particular set of marginal distributions. The article provides a framework and a set of guidelines that facilitate comparisons between various types of agreement tables. We illustrate our points with simulations and real data from two studies-one involving judges' ratings of baseball players and one involving ratings of essays in high-stakes tests.
Related Concept Videos
Kendall's Coefficient of Concordance
Friedman Two-way Analysis of Variance by Ranks
Sign Test for Matched Pairs
To conduct the sign test, we first calculate the differences in...
Wilcoxon Signed-Ranks Test for Matched Pairs
McNemar's Test
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...

