Related Experiment Video
Updated: Jul 19, 2026

08:12
A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
Published on: March 1, 2022
Measuring agreement for ordered ratings in 3 x 3 tables
1Université Montpellier I, UFR Médecine, Département de l'Information Médicale, Montpellier, France. d-neveu@chu-montpellier.fr
Methods of Information in Medicine
|October 5, 2006
Summary
New metrics improve qualitative agreement assessment. The unweighted kappa index (kappa(c)) and a novel statistic, zeta, offer better insights into rater agreement and disagreement patterns than traditional methods.
Area of Science:
- Statistics
- Biostatistics
- Psychometrics
Background:
- Qualitative agreement is often measured using symmetrically weighted kappa statistics.
- Traditional kappa statistics can be insensitive to variations in complete agreement or disagreement.
- Paradoxical results can arise from current agreement assessment methods.
Purpose of the Study:
- To address limitations of traditional kappa statistics in assessing qualitative agreement.
- To develop and introduce a new statistic for evaluating qualitative agreements.
- To compare the performance of new and existing agreement metrics.
Main Methods:
- Calculation of symmetrically weighted kappa statistics.
- Development of a new statistic (zeta) to assess qualitative agreements.
- Analysis of agreement using relative amounts of complete agreements, partial, and maximal disagreements beyond chance.
- Illustration with data sets from existing literature.
Main Results:
- The unweighted kappa index (kappa(c)) and the new statistic zeta provide a better assessment of agreement.
- Zeta quantifies the excess of maximal disagreements over partial disagreements, independent of weighting.
- The (kappa(c), zeta) pair effectively identifies variations in agreement and disagreement.
- Increasing values of kappa(c) and zeta indicate improved qualitative agreement.
Conclusions:
- The (kappa(c), zeta) pair offers enhanced sensitivity to changes in agreement and disagreement.
- This new metric pair allows for precise localization of differences between qualitative agreements.
- The proposed method provides a more robust evaluation of qualitative agreement compared to traditional approaches.
Related Concept Videos
Kendall's Coefficient of Concordance
Kendall's Coefficient of Concordance (W), also known as Kendall's W, is a non-parametric statistical measure used to assess the agreement or concordance between multiple raters or judges when they rank a set of items. It is often used when you have ordinal data (ranks) and you want to see if there is consistency or consensus among the raters. It is widely applied in research areas such as psychology, medicine, and social sciences, where multiple judges are asked to rank or rate subjects or...
Friedman Two-way Analysis of Variance by Ranks
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures from...
Ratio Level of Measurement
The way a set of data is measured is called its level of measurement. Correct statistical procedures depend on a researcher being familiar with levels of measurement. For analysis, data are classified into four levels of measurement—nominal, ordinal, interval, and ratio.
A set of data measured using the ratio scale takes care of the ratio problem and provides complete information. Ratio scale data are like interval scale data, except they have a zero point and ratios can be calculated. For...
A set of data measured using the ratio scale takes care of the ratio problem and provides complete information. Ratio scale data are like interval scale data, except they have a zero point and ratios can be calculated. For...
Ordinal Level of Measurement
The way a set of data is measured is called its level of measurement. Correct statistical procedures depend on a researcher being familiar with levels of measurement. For analysis, data are classified into four levels of measurement—nominal, ordinal, interval, and ratio.
Data measured using an ordinal scale are similar to nominal scale data, but there is one major difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks in the...
Data measured using an ordinal scale are similar to nominal scale data, but there is one major difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks in the...
McNemar's Test
McNemar's Test is a nonparametric statistical test used to determine if there is a significant difference in proportions between two related groups when the outcome is binary (e.g., yes/no, success/failure). It is beneficial when we have paired data, such as pre-test/post-test designs, where the same subjects are measured under two different conditions. The test is named after the statistician Quinn McNemar, who introduced it in 1947. It is commonly used in situations where subjects are...
One-Way ANOVA: Equal Sample Sizes
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...