Related Experiment Video
Updated: Mar 4, 2026

Measuring the Functional Abilities of Children Aged 3-6 Years Old with Observational Methods and Computer Tools
Published on: June 20, 2020
Kappa statistic to measure agreement beyond chance in free-response assessments
Marc Carpentier1, Christophe Combescure1, Laura Merlini2
1Division of Clinical Epidemiology, Geneva University Hospitals, and Faculty of Medicine, University of Geneva, Geneva, Switzerland.
Background:
The usual kappa statistic requires that all observations be enumerated. However, in free-response assessments, only positive (or abnormal) findings are notified, but negative (or normal) findings are not. This situation occurs frequently in imaging or other diagnostic studies. We propose here a kappa statistic that is suitable for free-response assessments.
Method:
We derived the equivalent of Cohen's kappa statistic for two raters under the assumption that the number of possible findings for any given patient is very large, as well as a formula for sampling variance that is applicable to independent observations (for clustered observations, a bootstrap procedure is proposed). The proposed statistic was applied to a real-life dataset, and compared with the common practice of collapsing observations within a finite number of regions of interest.
Results:
The free-response kappa is computed from the total numbers of discordant (b and c) and concordant positive (d) observations made in all patients, as 2d/(b + c + 2d). In 84 full-body magnetic resonance imaging procedures in children that were evaluated by 2 independent raters, the free-response kappa statistic was 0.820. Aggregation of results within regions of interest resulted in overestimation of agreement beyond chance.
Conclusions:
The free-response kappa provides an estimate of agreement beyond chance in situations where only positive findings are reported by raters.
Related Concept Videos
Kendall's Coefficient of Concordance
Kendall's Tau Test
A τ value of +1 indicates...
Comparing Experimental Results: Student's t-Test
Bonferroni Test
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
Critical Region, Critical Values and Significance Level
In hypothesis testing, a sample statistic is converted to a test statistic using z, t, or chi-square distribution. A critical region is an area under the curve in probability distributions demarcated by the critical value. When the test statistic falls in this region, it suggests that the null hypothesis must be rejected. As this region contains all those values of the...
Goodness-of-Fit Test

