Related Experiment Video
Updated: Jun 15, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Evaluating Synthetic Data Augmentation to Correct for Data Imbalance in Realistic Clinical Prediction Settings
Nina Wahler1, Bayrem Kaabachi1, Bogdan Kulynych1
1Lausanne University Hospital (CHUV), Switzerland.
Synthetic data generation shows limited improvement for imbalanced clinical datasets. This study found that complex methods did not significantly outperform standard baselines in predictive modeling for class imbalance.
Area of Science:
- Medical Informatics
- Machine Learning
- Data Science
Background:
- Clinical decision-making relies on predictive modeling, often challenged by imbalanced datasets.
- Data augmentation is crucial for improving model performance on underrepresented classes.
Purpose of the Study:
- To evaluate the effectiveness of synthetic data generation for enhancing predictive models on small, imbalanced clinical datasets.
- To compare advanced synthetic data techniques against traditional methods for class imbalance correction.
Main Methods:
- Investigated Generative Adversarial Networks (GANs), Normalizing Flows, and Variational Autoencoders (VAEs).
- Compared these methods to standard baselines for class underrepresentation.
- Utilized four realistic clinical datasets for evaluation.
Main Results:
- Some synthetic data methods showed marginal improvements in F1 scores.
- No statistically significant evidence was found that synthetic data generation outperformed standard baselines.
- Results were consistent even after multiple repetitions.
Conclusions:
- The efficacy of synthetic data for augmenting imbalanced clinical data requires careful evaluation.
- Complex synthetic data generation methods may not consistently outperform simpler, established techniques.
- Emphasizes the need to benchmark novel methods against standard baselines in predictive modeling research.
More Related Videos
Related Concept Videos
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Improving Translational Accuracy
Statistical Software for Data Analysis and Clinical Trials
Data Collection by Experiments
An example of the experimental method is a public...
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Clinical Trials: Overview

