Related Experiment Video
Updated: Jul 25, 2025

12:38
Three-dimensional Reconstruction of the Vascular Architecture of the Passive CLARITY-cleared Mouse Ovary
Published on: December 10, 2017
8.8K
A statistical comparison between Matthews correlation coefficient (MCC), prevalence threshold, and Fowlkes-Mallows
Davide Chicco1, Giuseppe Jurman2
1University of Toronto, Canada.
Journal of Biomedical Informatics
|June 23, 2023
Summary
The Matthews correlation coefficient (MCC) is often superior for evaluating binary classifications. This study shows MCC is more informative than prevalence threshold (PT) and Fowlkes-Mallows index when data elements are equally weighted.
Area of Science:
- Statistics
- Machine Learning
- Data Science
Background:
- Assessing binary classifications is crucial in scientific research, but a consensus on a single summary statistic for confusion matrices remains elusive.
- Previous work highlighted the advantages of the Matthews correlation coefficient (MCC) over various other metrics like accuracy, F1 score, and Cohen's kappa.
- The need for robust metrics is evident across diverse scientific disciplines.
Purpose of the Study:
- To compare the Matthews correlation coefficient (MCC) with prevalence threshold (PT) and Fowlkes-Mallows index.
- To investigate the interrelationships between these three statistical metrics.
- To determine the conditions under which MCC offers superior information for binary classification evaluation.
Main Methods:
- Comparative analysis of statistical metrics for binary classification.
- Investigation of mutual relations among Matthews correlation coefficient (MCC), prevalence threshold (PT), and Fowlkes-Mallows index.
- Evaluation using relevant use cases from scientific research.
Main Results:
- The Matthews correlation coefficient (MCC) was compared against prevalence threshold (PT) and Fowlkes-Mallows index.
- Mutual relations among the three metrics were analyzed.
- Results indicate MCC can be more informative than PT and Fowlkes-Mallows index when positive and negative data elements hold equal importance.
Conclusions:
- The Matthews correlation coefficient (MCC) demonstrates significant advantages in evaluating binary classifications, particularly when class balance is considered.
- MCC provides more comprehensive information compared to prevalence threshold (PT) and Fowlkes-Mallows index under specific conditions of equal data element importance.
- The findings support the broader adoption of MCC as a preferred metric for binary classification tasks in scientific research.
Related Concept Videos
Identifying Statistically Significant Differences: The F-Test
1.7K
The F-test is used to compare two sample variances to each other or compare the sample variance to the population variance. It is used to decide whether an indeterminate error can explain the difference in their values. The underlying assumptions that allow the use of the F-test include the data set or sets are normally distributed, and the data sets are independent of each other. The test statistic F is calculated by dividing one variance by another. In other words, the square of one standard...
1.7K
Multiple Comparison Tests
3.9K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
3.9K
Bonferroni Test
2.8K
The Bonferroni test is a statistical test named after Carlo Emilio Bonferroni, an Italian mathematician best known for Bonferroni inequalities. This statistical test is a type of multiple comparison test to determine which means are different than the rest. Bonferroni test can minimize the Type 1 error by reducing the significance level alpha, which otherwise increases with sample pairs.
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
2.8K
Coefficient of Correlation
6.2K
The correlation coefficient, r, developed by Karl Pearson in the early 1900s, is numerical and provides a measure of strength and direction of the linear association between the independent variable x and the dependent variable y.
If you suspect a linear relationship between x and y, then r can measure how strong the linear relationship is.
What the VALUE of r tells us:
The value of r is always between –1 and +1: –1 ≤ r ≤ 1.
The size of the correlation r indicates the...
If you suspect a linear relationship between x and y, then r can measure how strong the linear relationship is.
What the VALUE of r tells us:
The value of r is always between –1 and +1: –1 ≤ r ≤ 1.
The size of the correlation r indicates the...
6.2K
Spearman's Rank Correlation Test
877
Spearman's rank correlation test, also known as Spearman's rho, is a nonparametric method for assessing the strength and direction of association between two variables. This test is particularly valuable when the data distribution is unknown or when the assumption of normality does not hold. Named after the English psychologist and statistician Dr. Charles Edward Spearman, it serves as the nonparametric counterpart to Pearson's correlation coefficient.
Spearman's test calculates...
Spearman's test calculates...
877
Fisher's Exact Test
643
Fisher's exact test is a statistical significance test widely used to analyze 2x2 contingency tables, particularly in situations where sample sizes are small. Unlike the chi-squared test, which approximates P-values and assumes minimum expected frequencies of at least five in each cell, Fisher's exact test calculates the exact probability (P-value) of observing the data or more extreme results under the null hypothesis. This feature makes it especially valuable when the assumptions of...
643

