Related Experiment Video
Updated: Sep 17, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Exploring the inequitable impact of data missingness on fairness in machine learning.
Sitao Min1, Hafiz Asif2, Jaideep Vaidya1
1Rutgers University, Newark, NJ, 07102, USA.
Missing data in artificial intelligence (AI) and machine learning (ML) models can worsen fairness disparities, especially when missingness correlates with sensitive attributes. Addressing data gaps is crucial for equitable AI.
Area of Science:
- Computer Science
- Data Science
- Artificial Intelligence
Background:
- Data-driven models and AI/ML are integral to societal decision-making.
- Concerns regarding algorithmic fairness are significant.
- The impact of missing data on fairness remains understudied despite its prevalence.
Purpose of the Study:
- To systematically evaluate how missing data affects classifier fairness.
- To investigate the role of missing data correlated with protected classes and outcomes.
Main Methods:
- Analysis of 150 experimental dataset variants reflecting real-world scenarios.
- Utilizing a comprehensive framework covering missing data patterns, rates, and mitigation strategies.
Main Results:
- Missing data, particularly when correlated with sensitive attributes and outcomes, can exacerbate fairness disparities.
- Even small amounts of missingness can significantly impact fairness.
- Systemic missingness poses a substantial challenge to equitable AI.
Conclusions:
- Addressing missing data is critical for evaluating and ensuring algorithmic fairness.
- Fairness evaluations must account for the presence and patterns of missing data.
- Mitigation strategies for missing data are essential for developing trustworthy AI systems.
Related Concept Videos
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
Bias in Epidemiological Studies
Censoring Survival Data
Detection of Gross Error: The Q Test
Weighted Mean
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...

