Related Experiment Video
Updated: Sep 25, 2025

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
Published on: March 1, 2022
Assessing Fairness in the Presence of Missing Data
1University of Pennsylvania, Philadelphia, PA 19104, USA.
Analyzing incomplete data presents fairness challenges. This study introduces theoretical bounds for estimating fairness in complete data when using only complete cases, addressing a critical gap in machine learning research.
Area of Science:
- Machine Learning and Data Science
- Algorithmic Fairness
- Statistical Analysis
Background:
- Missing data is a common issue in real-world datasets, complicating data analysis and predictive modeling.
- Existing research on algorithmic fairness primarily focuses on fully observed data, leaving a gap in understanding fairness with incomplete data.
- A common approach to handle missing data is complete case analysis, which may lead to biased results due to distributional shifts.
Purpose of the Study:
- To theoretically investigate and estimate fairness in the complete data domain when analysis is performed using only complete cases from incomplete datasets.
- To address the potential for biased predictions towards marginalized groups when fairness is assessed solely on complete cases.
- To provide the first theoretical framework for fairness guarantees in the analysis of incomplete data.
Main Methods:
- Development of theoretical upper and lower bounds for estimating fairness error in the complete data domain.
- Evaluation of an arbitrary prediction model using only complete cases.
- Conducting numerical experiments to validate the theoretical findings.
Main Results:
- The study establishes theoretical bounds on the error of fairness estimation when using complete cases.
- Numerical experiments demonstrate the practical implications of the theoretical results.
- The findings quantify the potential discrepancy in fairness when evaluating models on complete cases versus the complete data.
Conclusions:
- This research provides the foundational theoretical results for ensuring fairness in machine learning models trained on incomplete data.
- The developed bounds offer a way to understand and potentially mitigate fairness estimation errors arising from complete case analysis.
- The work highlights the importance of considering missing data mechanisms when assessing algorithmic fairness.
More Related Videos
09:00Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
Published on: August 16, 2024
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Related Concept Videos
One-Way ANOVA: Unequal Sample Sizes
Detection of Gross Error: The Q Test
Test for Homogeneity
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Censoring Survival Data
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...