Assessing the inter-rater agreement for ordinal data through weighted indexes
Donata Marasini1, Piero Quatto1, Enrico Ripamonti2
1Statistical Section, Department of Economics, Management and Statistics, University of Milan-Bicocca, Milano, Italy.
Statistical Methods in Medical Research
|April 18, 2014
Summary
This study introduces a new statistic to assess agreement between multiple observers for ordinal data, overcoming limitations of existing methods like Fleiss' kappa. The modified statistic offers more reliable inter-rater agreement analysis in biomedical research.
Area of Science:
- Statistics
- Biomedical Research
- Ordinal Data Analysis
Background:
- Inter-rater agreement is crucial for statistical theory and biomedical applications, especially with ordinal variables.
- Existing methods like Cohen's and Fleiss' kappa can exhibit paradoxical behavior, complicating interpretation.
- There is a need for robust measures of agreement that avoid these paradoxes.
Purpose of the Study:
- To propose a novel modification of Fleiss' kappa for assessing inter-rater agreement with ordinal variables.
- To address the paradoxical behavior observed in traditional kappa statistics.
- To generalize the proposed statistic for broader applicability, including bivariate cases.
Main Methods:
- Development of a modified Fleiss' kappa statistic (s*) for ordinal variables.
- Utilizing Monte Carlo simulations for hypothesis testing and confidence interval calculation (percentile and bootstrap-t).
- Demonstration of the normal asymptotic distribution of the proposed statistic.
Main Results:
- The proposed statistic effectively assesses inter-rater agreement for ordinal data without paradoxical behavior.
- Monte Carlo simulations validated the statistical properties and reliability of the new measure.
- The method was successfully applied to a real-world dataset on cervical cancer classification.
Conclusions:
- The novel statistic provides a reliable and interpretable measure for inter-rater agreement with ordinal variables.
- This advancement improves the analysis of agreement in complex biomedical datasets.
- The generalization to bivariate cases expands its utility in statistical applications.
More Related Videos
Related Concept Videos
Ranks
591
Unlike parametric methods, nonparametric statistics are ideal for nominal and ordinal data, requiring fewer assumptions about the population's nature or distribution. This makes nonparametric methods easier to apply and interpret, as they do not depend on parameters like mean or standard deviation. One common approach in nonparametric analysis is to sort data according to a specific criterion. For instance, we might arrange weather data from hottest to coldest days in a month or rank cities...
591
Kendall's Coefficient of Concordance
1.3K
Kendall's Coefficient of Concordance (W), also known as Kendall's W, is a non-parametric statistical measure used to assess the agreement or concordance between multiple raters or judges when they rank a set of items. It is often used when you have ordinal data (ranks) and you want to see if there is consistency or consensus among the raters. It is widely applied in research areas such as psychology, medicine, and social sciences, where multiple judges are asked to rank or rate subjects...
1.3K
Weighted Mean
5.5K
While taking the arithmetic, geometric, or harmonic mean of a sample data set, equal importance is assigned to all the data points. However, all the values may not always be equally important in some data sets. An intrinsic bias might make it more important to give more weightage to specific values over others.
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
5.5K
Ordinal Level of Measurement
24.2K
The way a set of data is measured is called its level of measurement. Correct statistical procedures depend on a researcher being familiar with levels of measurement. For analysis, data are classified into four levels of measurement—nominal, ordinal, interval, and ratio.
Data measured using an ordinal scale are similar to nominal scale data, but there is one major difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks...
Data measured using an ordinal scale are similar to nominal scale data, but there is one major difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks...
24.2K
Friedman Two-way Analysis of Variance by Ranks
595
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
595
Ratio Level of Measurement
13.7K
The way a set of data is measured is called its level of measurement. Correct statistical procedures depend on a researcher being familiar with levels of measurement. For analysis, data are classified into four levels of measurement—nominal, ordinal, interval, and ratio.
A set of data measured using the ratio scale takes care of the ratio problem and provides complete information. Ratio scale data are like interval scale data, except they have a zero point and ratios can be calculated....
A set of data measured using the ratio scale takes care of the ratio problem and provides complete information. Ratio scale data are like interval scale data, except they have a zero point and ratios can be calculated....
13.7K


