Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Kendall's Coefficient of Concordance01:20

Kendall's Coefficient of Concordance

972
Kendall's Coefficient of Concordance (W), also known as Kendall's W, is a non-parametric statistical measure used to assess the agreement or concordance between multiple raters or judges when they rank a set of items. It is often used when you have ordinal data (ranks) and you want to see if there is consistency or consensus among the raters. It is widely applied in research areas such as psychology, medicine, and social sciences, where multiple judges are asked to rank or rate subjects...
972
Kendall's Tau Test01:16

Kendall's Tau Test

1.1K
Kendall's tau test, also known as the Kendall rank coefficient test, is a nonparametric method for assessing association between two variables. This test is particularly useful for identifying significant correlations when the distributions of the sample and population are unknown. Developed in 1938 by the British statistician Sir Maurice George Kendall, the tau coefficient (denoted as τ) serves as a rank correlation coefficient, with values ranging from -1 to +1.
A τ value of +1 indicates...
1.1K
Spearman's Rank Correlation Test01:20

Spearman's Rank Correlation Test

1.4K
Spearman's rank correlation test, also known as Spearman's rho, is a nonparametric method for assessing the strength and direction of association between two variables. This test is particularly valuable when the data distribution is unknown or when the assumption of normality does not hold. Named after the English psychologist and statistician Dr. Charles Edward Spearman, it serves as the nonparametric counterpart to Pearson's correlation coefficient.
Spearman's test calculates correlation by...
1.4K
Theory of Attribution II: Kelley's Covariation Theory01:29

Theory of Attribution II: Kelley's Covariation Theory

505
Attribution theory plays a crucial role in social psychology, helping to explain how individuals interpret the causes of behavior. One prominent model within this field is Harold Kelley's covariation theory, which provides a systematic approach to determining whether internal traits or external circumstances drive a person's actions. The model posits that individuals rely on three key types of information—consensus, consistency, and distinctiveness—to make these judgments.Consensus:...
505
Cochran's Q Test01:17

Cochran's Q Test

962
Cochran's Q Test is a nonparametric statistical test used to determine if there are potential differences in the outcomes of three or more related groups on a binary (yes/no) or dichotomous outcome. It is essentially an extension of the McNemar Test, which is limited to two related samples - Cochran's Q test can handle three or more related samples, making it more versatile in scenarios where subjects are measured under multiple conditions. The test statistic follows a Chi-Square...
962
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation01:24

One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation

1.1K
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
1.1K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Predictive risk models for COVID-19 patients using the multi-thresholding meta-algorithm.

Scientific reports·2024
Same author

The effect of seasonality in predicting the level of crime. A spatial perspective.

PloS one·2023
Same author

Exploring the role of ecology and social organisation in agropastoral societies: A Bayesian network approach.

PloS one·2022
Same author

Detecting target species: with how many samples?

Royal Society open science·2022
Same author

Bayesian network-based over-sampling method (BOSME) with application to indirect cost-sensitive learning.

Scientific reports·2022
Same author

Misuse of Beer-Lambert Law and other calibration curves.

Royal Society open science·2022

Related Experiment Video

Updated: Jan 19, 2026

Measuring the Functional Abilities of Children Aged 3-6 Years Old with Observational Methods and Computer Tools
11:29

Measuring the Functional Abilities of Children Aged 3-6 Years Old with Observational Methods and Computer Tools

Published on: June 20, 2020

9.7K

Why Cohen's Kappa should be avoided as performance measure in classification.

Rosario Delgado1, Xavier-Andoni Tibau2

  • 1Department of Mathematics, Universitat Autònoma de Barcelona, Campus de la UAB, Cerdanyola del Vallès, Spain.

Plos One
|September 27, 2019
PubMed
Summary

Cohen's Kappa and Matthews Correlation Coefficient (MCC) often correlate but diverge with imbalanced data. A worse classifier can achieve a higher Kappa score, questioning its reliability for comparing classification performance.

More Related Videos

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.9K
Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
07:13

Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model

Published on: April 18, 2025

495

Related Experiment Videos

Last Updated: Jan 19, 2026

Measuring the Functional Abilities of Children Aged 3-6 Years Old with Observational Methods and Computer Tools
11:29

Measuring the Functional Abilities of Children Aged 3-6 Years Old with Observational Methods and Computer Tools

Published on: June 20, 2020

9.7K
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.9K
Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
07:13

Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model

Published on: April 18, 2025

495

Area of Science:

  • Machine Learning
  • Data Science
  • Statistical Modeling

Background:

  • Cohen's Kappa and Matthews Correlation Coefficient (MCC) are key metrics for multi-class classification performance evaluation.
  • Both metrics assess classifier performance but can yield different results, especially in imbalanced datasets.

Purpose of the Study:

  • To investigate the relationship between Cohen's Kappa and MCC under various data imbalances.
  • To identify specific conditions where Kappa exhibits anomalous behavior compared to MCC.

Main Methods:

  • Comparative analysis of Cohen's Kappa and MCC across different multi-class classification scenarios.
  • Experimental study focusing on confusion matrix properties, particularly the entropy of off-diagonal elements.

Main Results:

  • Cohen's Kappa and MCC are generally correlated but diverge significantly in specific unbalanced situations.
  • Anomalous Kappa behavior emerges when the entropy of off-diagonal confusion matrix elements decreases, indicating a potential pitfall.

Conclusions:

  • Cohen's Kappa may not be a reliable metric for comparing classifiers due to its potential for counterintuitive results.
  • The findings suggest limitations in using Kappa for performance evaluation, especially with imbalanced datasets.