Related Experiment Video
Updated: Sep 17, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Exploring the inequitable impact of data missingness on fairness in machine learning
Sitao Min1, Hafiz Asif2, Jaideep Vaidya1
1Rutgers University, Newark, NJ, 07102, USA.
Abstract:
Today, data-driven models and artificial intelligence / machine learning underlie decision making in almost all aspects of society. However, significant concerns have been raised over the fairness of such models. While various aspects of algorithmic fairness have been studied, the effect of missing data on fairness remains understudied. This is a significant problem since data in real-world settings is almost never complete, and may often suffer from systemic missingness. This article systematically evaluates how missing data, particularly when correlated with protected classes and outcome variables, affects the fairness of classifiers. Utilizing a comprehensive framework covering various missing data patterns, rates, and mitigation methods, we analyze 150 experimental dataset variants derived from real-world scenarios, and find that missing data correlated with sensitive attributes and outcomes can exacerbate disparities, even for little missingness, making it crucial to address missingness in fairness evaluations.
Related Concept Videos
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
Bias in Epidemiological Studies
Censoring Survival Data
Detection of Gross Error: The Q Test
Weighted Mean
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...

