Related Experiment Video
Updated: Apr 18, 2026

16:23
Automated, Quantitative Cognitive/Behavioral Screening of Mice: For Genetics, Pharmacology, Animal Cognition and Undergraduate Instruction
Published on: February 26, 2014
15.0K
Facing Imbalanced Data Recommendations for the Use of Performance Metrics
László A Jeni1, Jeffrey F Cohn2, Fernando De La Torre1
1Carnegie Mellon University, Pittsburgh, PA.
Summary
Facial action unit (AU) detection is significantly impacted by imbalanced data, which can skew performance metrics. Researchers recommend reporting skew-normalized scores to ensure accurate evaluation of AU detection models.
Area of Science:
- Computer Vision
- Machine Learning
- Human-Computer Interaction
Background:
- Facial action unit (AU) recognition is crucial for automated video analysis.
- Existing research primarily focuses on face tracking, registration, and feature selection.
- The impact of imbalanced data on AU detection performance metrics remains underexplored.
Purpose of the Study:
- To investigate the influence of data skew on performance metrics for facial action unit detection.
- To evaluate how skewed data distributions affect various threshold and rank metrics.
- To propose methods for mitigating bias in performance evaluations.
Main Methods:
- Experiments were conducted using simulated classifiers and three diverse facial action unit databases.
- The study analyzed the effect of data skew on threshold metrics (Accuracy, F-score, Cohen's kappa, Krippendorf's alpha).
- Rank metrics, including area under the ROC curve and precision-recall curve, were also evaluated.
Main Results:
- Skewed data distributions significantly attenuated most evaluated metrics, except for the area under the ROC curve.
- Precision-recall curves indicated that ROC might obscure poor performance in skewed datasets.
- The degree of performance attenuation varied across different metrics and datasets.
Conclusions:
- Data skew is a critical factor that can bias performance evaluations in facial action unit detection.
- Standard performance metrics can be misleading when dealing with imbalanced datasets.
- Reporting skew-normalized scores is recommended to obtain more reliable performance estimates.
Related Concept Videos
Weighted Mean
7.6K
While taking the arithmetic, geometric, or harmonic mean of a sample data set, equal importance is assigned to all the data points. However, all the values may not always be equally important in some data sets. An intrinsic bias might make it more important to give more weightage to specific values over others.
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
7.6K
Skewness
21.6K
The measures of central tendency calculated from a data set may not reveal much about its intrinsic distribution. If a plot is made of the data set’s values, the mean and the median may not only differ, but also the plot may have more values on one side of the central tendencies. Such a data set is said to be skewed towards that side.
The longer the tail of the plot on one side, the more skewed it is. The skewness of a data set’s values suggests that the measures of central tendency...
The longer the tail of the plot on one side, the more skewed it is. The skewness of a data set’s values suggests that the measures of central tendency...
21.6K
Data: Types and Distribution
2.3K
In biostatistics, data are the observations collected for analysis. There are two main types: parametric and non-parametric. Parametric data, which include continuous (e.g., weight) and discrete numerical data (e.g., number of tablets), assume a particular distribution pattern, often the normal distribution. Non-parametric data do not adhere to a specific distribution and typically comprise nominal (e.g., gender) and ordinal categorical data (e.g., pain scale ratings).
Distributions in...
Distributions in...
2.3K
Statistical Analysis: Overview
18.8K
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
18.8K
Trimmed Mean
3.6K
While measuring the mean of a data set, care needs to be taken when associating the mean to its central tendency. The same goes for the arithmetic mean, the geometric mean, or the harmonic mean. This is because the presence of a single outlier data value can significantly affect the mean. That is, the mean is sensitive to fluctuations in the data set.
Although certain measures of central tendency are not sensitive to outliers, there are alternative versions of the mean that get around the...
Although certain measures of central tendency are not sensitive to outliers, there are alternative versions of the mean that get around the...
3.6K
Ratio Level of Measurement
22.6K
The way a set of data is measured is called its level of measurement. Correct statistical procedures depend on a researcher being familiar with levels of measurement. For analysis, data are classified into four levels of measurement—nominal, ordinal, interval, and ratio.
A set of data measured using the ratio scale takes care of the ratio problem and provides complete information. Ratio scale data are like interval scale data, except they have a zero point and ratios can be calculated....
A set of data measured using the ratio scale takes care of the ratio problem and provides complete information. Ratio scale data are like interval scale data, except they have a zero point and ratios can be calculated....
22.6K

