Related Experiment Video
Updated: Jun 5, 2026

07:28
A Protocol of Manual Tests to Measure Sensation and Pain in Humans
Published on: December 19, 2016
Measures of interrater agreement
1Department of Health Sciences Research, Mayo Clinic, Rochester, Minnesota 55905, USA. mandrekar.jay@mayo.edu
Summary
Kappa statistics assess rater agreement for categorical data. This summary covers kappa features, prevalence impact, clinical utility, weighted kappa for ordinal data, and intraclass correlation for continuous data.
Area of Science:
- Statistics
- Biostatistics
- Clinical Research Methodology
Background:
- Inter-rater reliability is crucial for reproducible research.
- Categorical data analysis requires appropriate agreement metrics.
- Existing methods may not cover all data types (ordinal, continuous).
Purpose of the Study:
- To explain kappa statistics for assessing rater agreement.
- To discuss the influence of prevalence on kappa's interpretation.
- To introduce extensions for ordinal and continuous data.
Main Methods:
- Review and interpretation of kappa statistics.
- Discussion of prevalence's effect on kappa.
- Introduction of weighted kappa and intraclass correlation.
Main Results:
- Kappa statistics provide a robust measure for categorical agreement.
- Prevalence significantly impacts kappa's value and interpretation.
- Weighted kappa and intraclass correlation extend agreement assessment.
Conclusions:
- Kappa statistics are essential for evaluating inter-rater reliability in categorical studies.
- Understanding prevalence is key for accurate kappa interpretation.
- Weighted kappa and intraclass correlation offer versatile solutions for ordinal and continuous data agreement.
Related Concept Videos
Kendall's Coefficient of Concordance
Kendall's Coefficient of Concordance (W), also known as Kendall's W, is a non-parametric statistical measure used to assess the agreement or concordance between multiple raters or judges when they rank a set of items. It is often used when you have ordinal data (ranks) and you want to see if there is consistency or consensus among the raters. It is widely applied in research areas such as psychology, medicine, and social sciences, where multiple judges are asked to rank or rate subjects or...
Statistical Analysis: Overview
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
Calibration Curves: Correlation Coefficient
In a linear calibration curve, there is a value called the calibration coefficient, denoted by 'r,' which measures the strength and the direction of association between two variables. The correlation coefficient value ranges from −1 to +1. A value of +1 indicates a perfect positive linear correlation, −1 denotes a perfect negative correlation, and 0 implies no correlation between the two variables. A positive correlation value establishes that as one variable increases, the other increases, and...
Variability: Analysis
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
The range is a simple measure of variability, indicating the difference between the highest and...
Measures of Intelligence
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this; it...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this; it...
Comparing Experimental Results: Student's t-Test
The t-test is a statistical method used to compare the sample mean with a population mean or compare two means from two data sets. The test statistic is calculated from the standard deviation, mean, and number of measurements in the data set at a selected confidence interval and then compared to a table of critical values at this confidence level. If the test statistic is smaller than the critical value, the null hypothesis is accepted. In this case, we state that the difference between the...
