Related Experiment Videos
Data Representation Bias and Conditional Distribution Shift Drive Predictive Performance Disparities in
Sandeep Kumar1,2, Yan Cui1,2,3
1Department of Genetics, Genomics and Informatics, University of Tennessee Health Science Center, Memphis, TN 38163, USA.
Machine learning struggles with population-stratified data due to bias and distribution shifts. Conditional shifts and data bias significantly impact model performance across diverse groups, affecting transfer learning effectiveness.
Area of Science:
- Genetics and Bioinformatics
- Machine Learning and Artificial Intelligence
- Population Health
Background:
- Machine learning models often face performance and generalizability issues with population-stratified datasets.
- Data representation bias and distribution shifts are key challenges impacting model fairness across diverse ancestry groups.
- Understanding these challenges is crucial for developing equitable AI in various scientific domains.
Purpose of the Study:
- To systematically investigate the influence of data representation bias and distribution shifts on multi-population machine learning.
- To evaluate the effectiveness of mixture learning, independent learning, and transfer learning in mitigating disparities.
- To provide insights for building robust and equitable machine learning models for diverse populations.
Main Methods:
- Utilized synthetic genotype-phenotype datasets representing five continental populations.
- Evaluated three distinct machine learning approaches: mixture learning, independent learning, and transfer learning.
- Analyzed the impact of conditional and marginal distribution shifts alongside data representation bias.
Main Results:
- Conditional distribution shifts, coupled with data representation bias, significantly degrade machine learning performance across diverse populations.
- The effectiveness of transfer learning as a disparity mitigation strategy is notably influenced by these factors.
- Marginal distribution shifts demonstrated a limited impact compared to conditional shifts.
Conclusions:
- The interplay between data representation bias and distribution shifts critically affects multi-population machine learning outcomes.
- Conditional distribution shifts are a primary driver of performance disparities in population-stratified machine learning.
- Findings offer critical insights for developing equitable and high-performing machine learning models for diverse datasets.
Related Concept Videos
Analysis of Population Pharmacokinetic Data
Distributions to Estimate Population Parameter
Mechanistic Models: Compartment Models in Individual and Population Analysis
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance, comparing...
Testing a Claim about Population Proportion
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
Choosing Between z and t Distribution