Related Experiment Video
Updated: Sep 29, 2025

How to Calculate and Validate Inter-brain Synchronization in a fNIRS Hyperscanning Study
Published on: September 8, 2021
Confidence interval for micro-averaged F 1 and macro-averaged F 1 scores
Kanae Takahashi1,2, Kouji Yamamoto3, Aya Kuchiba4,5
1Department of Medical Statistics, Osaka City University Graduate School of Medicine, Osaka, Japan.
This study introduces new statistical methods for estimating F1 scores in multi-class classification problems. These methods provide confidence intervals for improved classifier performance evaluation, especially when prevalence is low.
Area of Science:
- Statistics
- Computer Science
- Machine Learning
Background:
- Binary classification metrics like sensitivity and specificity are common in medicine.
- Precision and recall are standard for classifier evaluation in computer science.
- The F1 score, a harmonic mean of precision and recall, is crucial for imbalanced datasets but lacks multi-class extensions.
Purpose of the Study:
- To address the lack of statistical inference methods for F1 scores in multi-class classification.
- To propose novel methods for estimating F1 scores and their confidence intervals.
- To enhance the evaluation of classifier performance in complex scenarios.
Main Methods:
- Development of statistical methods based on the large sample multivariate central limit theorem.
- Application to estimating three types of F1 scores in multi-class classification.
- Focus on providing confidence intervals for F1 score estimates.
Main Results:
- Proposed methods enable statistical inference for multi-class F1 scores.
- Confidence intervals can be estimated for various F1 score types.
- The methods are grounded in established statistical theory.
Conclusions:
- The developed methods fill a critical gap in multi-class classification evaluation.
- Provides a robust framework for understanding F1 score uncertainty.
- Facilitates more reliable performance assessment of classifiers in diverse applications.
More Related Videos
Related Concept Videos
Identifying Statistically Significant Differences: The F-Test
F Distribution
Uncertainty: Confidence Intervals
Confidence Interval for Estimating Population Mean
A confidence interval for the mean is a range of values that provides an estimate of the population mean. As the...
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
One-Way ANOVA: Unequal Sample Sizes

