Related Experiment Video
Updated: Jun 28, 2025

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
Generative models improve fairness of medical classifiers under distribution shifts
Ira Ktena1, Olivia Wiles2, Isabela Albuquerque3
1Google DeepMind, London, UK. iraktena@google.com.
Generative AI, specifically diffusion models, can create synthetic medical data to improve machine learning model fairness and robustness. This approach enhances diagnostic accuracy for underrepresented groups, especially in real-world, out-of-distribution scenarios.
Area of Science:
- Medical Imaging
- Machine Learning
- Generative Artificial Intelligence
Background:
- Domain generalization is a key challenge in healthcare machine learning, where models underperform due to data discrepancies between development and deployment.
- Underrepresentation of specific groups or conditions in training data leads to reduced model performance and fairness.
- Acquiring and labeling extensive clinical data is often infeasible due to cost and rarity of conditions.
Purpose of the Study:
- To investigate the use of generative artificial intelligence, specifically diffusion models, to create synthetic data for improving machine learning model robustness and fairness.
- To address the unmet need for label-efficient data augmentation in medical machine learning.
- To evaluate the effectiveness of learned augmentations across diverse medical imaging tasks.
Main Methods:
- Utilized diffusion models to learn realistic data augmentations in a label-efficient manner.
- Enriched training datasets with synthetic examples to address underrepresented conditions and subgroups.
- Evaluated model performance and fairness on three distinct medical imaging datasets: histopathology, chest X-ray, and dermatology images.
Main Results:
- Learned augmentations generated by diffusion models significantly improved model robustness across all tested medical imaging tasks.
- The approach enhanced statistical fairness, particularly improving diagnostic accuracy for underrepresented groups, especially in out-of-distribution settings.
- Synthetic data augmentation proved effective in histopathology, chest X-ray, and dermatology image analysis.
Conclusions:
- Generative AI, through diffusion models, offers a steerable and label-efficient method to create synthetic data for medical machine learning.
- Synthetic data augmentation enhances model robustness and fairness, crucial for real-world clinical deployment.
- This approach shows promise in mitigating performance gaps caused by data underrepresentation in healthcare AI.
Related Concept Videos
Types of Skewness
For instance, in the middle of a pandemic, the geographical distribution of vaccine coverage may be positively skewed towards populations in the global north countries. However,...
Improving Translational Accuracy
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Bias in Epidemiological Studies
Regression Toward the Mean
Distributions to Estimate Population Parameter

