Related Experiment Videos
Reliable fairness auditing with semi-supervised inference
Jianhui Gao1, Jessica Gronsbell1
1Department of Statistical Science, University of Toronto, Toronto, ON M7A 2S4, Canada.
Abstract:
Machine learning (ML) models often exhibit bias that can exacerbate inequities in biomedical applications. Fairness auditing, the process of evaluating a model's performance across subpopulations, is critical for identifying and mitigating these biases. However, audits typically rely on large volumes of labeled data, which are costly and labor-intensive to obtain. To address this challenge, we introduce Infairness, a unified framework for auditing a wide range of fairness criteria using semi-supervised inference. Our approach combines a small labeled dataset with a large unlabeled dataset by imputing missing outcomes via regression with carefully selected nonlinear basis functions. Through extensive theoretical and empirical analyses, we show that our proposed estimator is (1) robust to specification of the ML or imputation model and (2) substantially more efficient than supervised estimation based solely on the labeled data. In two real-world fairness audits using electronic health record and medical imaging data, Infairness reduces variance by 40 - 60% compared to supervised estimation, underscoring its value for reliable fairness auditing with limited labeled data.
Related Concept Videos
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
Equity Theory
Reliability and Validity