Related Experiment Video
Updated: Jul 4, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Identifying and handling data bias within primary healthcare data using synthetic data generators.
Barbara Draghi1,2, Zhenchen Wang1, Puja Myles1
1Medicines and Healthcare products Regulatory Agency, London, UK.
Advanced synthetic data generators can create realistic medical data while protecting privacy. This study introduces methods to detect and correct biases in synthetic data, improving AI model performance for better healthcare outcomes.
Area of Science:
- Medical Informatics
- Artificial Intelligence
- Data Science
Background:
- Advanced synthetic data generators can simulate sensitive patient data, reducing identification risks and enabling AI development in medicine.
- Massive datasets like UK-NHS records are available, but biases from under-represented cohorts can transfer to synthetic data.
- Machine learning models can perpetuate data biases, leading to inaccurate correlations and distributions in synthetic datasets.
Purpose of the Study:
- To enhance synthetic data generators by addressing bias and improving predictive model performance.
- To introduce probabilistic methods for detecting and boosting difficult-to-predict samples in ground truth data.
- To develop strategies for generating bias-reduced synthetic data that also boosts AI model accuracy.
Main Methods:
- Probabilistic approaches to identify challenging data samples within ground truth datasets.
- Techniques to "boost" these difficult samples during the synthetic data generation process.
- Exploration of bias reduction strategies integrated with performance enhancement for predictive models.
Main Results:
- Improved detection of under-represented or complex data points in original datasets.
- Enhanced synthetic data generation that mitigates bias propagation.
- Demonstrated potential for synthetic data to improve the accuracy and fairness of AI diagnostic and management tools.
Conclusions:
- Probabilistic methods can effectively identify and address biases in synthetic data generation.
- The proposed techniques improve the quality of synthetic data for medical AI applications.
- This approach facilitates the development of more robust, equitable, and accurate AI-driven healthcare solutions.
More Related Videos
07:31Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
09:33Visualizing Field Data Collection Procedures of Exposure and Biomarker Assessments for the Household Air Pollution Intervention Network Trial in India
Published on: December 23, 2022
Related Concept Videos
Bias in Epidemiological Studies
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
Data Collection I
Data Validation
Nursing assessment guides are generally based on holistic models rather than medical...
Primary Healthcare Services
In 1978, international leaders convened in Alma-Ata, Kazakhstan, for what would be a pivotal event in global health. The Alma-Ata Declaration was the first to call...
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...