Pathologists should probably forget about kappa. Percent agreement, diagnostic specificity and related metrics

Alberto M Marchevsky1, Ann E Walts1, Birgit I Lissenberg-Witte2

  • 1Department of Pathology & Laboratory Medicine, Cedars-Sinai Medical Center, Los Angeles, CA, United States of America.

Insights

Kappa statistics in pathology show high diagnostic specificity but variable scores. Researchers propose using percent agreement and specificity for better interobserver variability evaluation.

Area of Science:

  • Pathology
  • Medical Statistics

Background:

  • Kappa statistics are commonly used to assess interobserver variability (IOV) in pathology.
  • Limited discussion exists regarding the clinical significance of kappa scores in pathology.

Purpose of the Study:

  • To evaluate the clinical applicability of kappa statistics for assessing IOV in pathology.
  • To propose alternative metrics for IOV evaluation.

Main Methods:

  • Reviewed five pathology papers for IOV evaluation methods and clinical applicability interpretation.
  • Recalculated kappa scores and diagnostic metrics (sensitivity, specificity, PPV, NPV) using a dataset of lung neuroendocrine neoplasms.
  • Compared kappa scores with percent agreement and diagnostic specificity against a gold standard.

Main Results:

  • Pathology papers commonly use kappa statistics (Cohen's, Fleiss') and the Landis and Koch scale for IOV.
  • No specific guidelines were found for interpreting the clinical applicability of kappa scores.
  • Diagnostic specificity among pathologists exceeded 90%, while kappa scores were more variable.
  • Kappa scores demonstrated limited clinical applicability in pathology.

Conclusions:

  • Positive and negative percent agreement are more clinically applicable for evaluating IOV between two raters.
  • Diagnostic specificity against a gold reference diagnosis is recommended for evaluating IOV among multiple raters.
  • Rethinking distinct diagnostic categories versus morphologic continua may be necessary.

Related Concept Videos

Sensitivity, Specificity, and Predicted Value01:13

Sensitivity, Specificity, and Predicted Value

In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
1.1K
Receiver Operating Characteristic Plot01:15

Receiver Operating Characteristic Plot

A ROC (Receiver Operating Characteristic) plot is a graphical tool used to assess the performance of a binary classification model by illustrating the trade-off between sensitivity (true positive rate) and specificity (false positive rate). By plotting sensitivity against 1 - specificity across various threshold settings, the ROC curve shows how well the model distinguishes between classes, with a curve closer to the top-left corner indicating a more accurate model. The area under the ROC curve...
421
Variability: Analysis01:11

Variability: Analysis

Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
364
Statistical Analysis: Overview01:11

Statistical Analysis: Overview

When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
13.4K
Review and Preview01:10

Review and Preview

In statistics, several tools are used to interpret the data. Measures of central tendency represent the characteristics of the data, such as mean, median, and mode. Additionally, measures of variance like standard deviation and range are used to find the spread of data from the mean. Relative standing measures the distance between data locations. Commonly used measures of relative standings are percentile, z score, and quartiles.
Percentiles are a type of fractile that partition data into...
8.2K