Evaluating and Reducing Subgroup Disparity in AI Models: An Analysis of Pediatric COVID-19 Test Outcomes

Alexander Libin1, Jonah T Treitler2, Tadas Vasaitis3

  • 1AIM AHEAD Consortium, Georgetown-Howard Universities Center for Clinical and Translational Science (GHUCCTS), Medstar Research Health Institute, Georgetown University, Washington, D.C., USA.

Insights

Subgroup disparities in artificial intelligence (AI) healthcare models are common, impacting fairness. Synthetic data shows potential to reduce these disparities, improving AI model equity in pediatric COVID-19 testing.

Area of Science:

  • Healthcare AI
  • Machine Learning Fairness
  • Health Disparities

Background:

  • Artificial intelligence (AI) in healthcare raises concerns about perpetuating health disparities.
  • The frequency and extent of subgroup fairness in AI models require further investigation.
  • Understanding AI fairness is crucial for equitable healthcare delivery.

Purpose of the Study:

  • To assess the prevalence and extent of subgroup fairness disparities in AI models predicting pediatric COVID-19 test outcomes.
  • To evaluate the impact of synthetic data on mitigating identified subgroup disparities.

Main Methods:

  • Utilized a nationally representative pediatric dataset (ages 0-17, n=9,935) from the US National Health Interview Survey (NHIS) for COVID-19 test outcomes.
  • Trained 50 machine learning models using five algorithms to assess subgroup disparities.
  • Evaluated models' area under the curve (AUC) on 12 small subgroups defined by socioeconomic factors against the overall population.
  • Explored synthetic data generation techniques (resampling, generative adversarial networks) to mitigate disparities.

Main Results:

  • Subgroup disparities were prevalent, found in 50.7% of the models.
  • Subgroup AUCs were generally lower than overall AUCs, with a mean difference of 0.01 (range: -0.29 to +0.41).
  • Four out of 12 subgroups exhibited statistically significant disparities across models.
  • Synthetic data introduction exacerbated disparities in 57.7% of models, but reduced mean AUC disparities by 0.03 (resampling) and 0.04 (GANs).

Conclusions:

  • Significant subgroup disparities exist in AI models for pediatric COVID-19 testing.
  • Synthetic data shows promise in reducing AI fairness gaps, though careful implementation is needed.
  • Further research is essential to ensure equitable AI deployment in healthcare settings.