Related Experiment Videos
Data Representation Bias and Conditional Distribution Shift Drive Predictive Performance Disparities in
Sandeep Kumar1,2, Yan Cui1,2,3
1Department of Genetics, Genomics and Informatics, University of Tennessee Health Science Center, Memphis, TN 38163, USA.
Abstract:
Machine learning frequently encounters challenges when applied to population-stratified datasets, where data representation bias and data distribution shifts substantially impact model performance and generalizability across different population groups. These challenges are well illustrated in the context of polygenic prediction for diverse ancestry groups, and the underlying mechanisms are broadly applicable to machine learning with population-stratified data across domains. Using synthetic genotype-phenotype datasets representing five continental populations, we evaluate three approaches for utilizing population-stratified data, mixture learning, independent learning, and transfer learning, to systematically investigate how data representation bias and distribution shifts influence multi-population machine learning. Our results show that conditional distribution shifts, in combination with data representation bias, significantly influence machine learning performance across diverse populations and the effectiveness of transfer learning as a disparity mitigation strategy, while the effect of marginal distribution shifts is limited. The joint effects of data representation bias and distribution shifts demonstrate distinct patterns under different multi-population machine learning approaches, providing critical insights for the development of effective and equitable machine learning models for population-stratified data.
Related Concept Videos
Analysis of Population Pharmacokinetic Data
Distributions to Estimate Population Parameter
Mechanistic Models: Compartment Models in Individual and Population Analysis
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance, comparing...
Testing a Claim about Population Proportion
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
Choosing Between z and t Distribution