Related Experiment Video
Updated: Aug 16, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
An Approach to Identifying and Quantifying Bias in Biomedical Data.
M Clara De Paolis Kaluza1, Shantanu Jain, Predrag Radivojac
1Northeastern University, Boston, MA 02115, USA.
This study introduces a method to detect and quantify bias in labeled data for machine learning. The approach helps ensure trustworthy AI in biomedical applications by analyzing data distributions.
Area of Science:
- Machine Learning
- Biomedical Data Analysis
- Statistical Modeling
Background:
- Data biases impede trustworthy machine learning in biomedical fields.
- Unrepresentative labeled data requires methods utilizing unlabeled data.
- Identifying and quantifying bias is crucial for reliable AI models.
Purpose of the Study:
- To develop methods for detecting and quantifying bias in labeled data within a semi-supervised learning framework.
- To address the challenge of unrepresentative data in machine learning for biomedical applications.
Main Methods:
- Utilized a binary semi-supervised setting assuming component distributions with varying proportions in labeled and unlabeled data.
- Modeled training data using nested mixtures of multivariate Gaussian distributions.
- Developed a multi-sample expectation-maximization algorithm to learn model parameters.
- Created a statistical test to detect bias and estimated bias level using distribution distances.
Main Results:
- Successfully applied a novel expectation-maximization algorithm to learn model parameters from combined labeled and unlabeled data.
- Developed and validated a statistical test for detecting general bias in labeled data.
- Quantified the extent of bias by measuring the distance between class-conditional distributions.
- Demonstrated the effectiveness of the bias estimation procedure on both synthetic and real-world biomedical data.
Conclusions:
- The developed methods can effectively identify and quantify bias in labeled data for machine learning.
- This approach is crucial for building trustworthy AI in biomedical applications by mitigating the impact of data representativeness issues.
- The bias estimation procedure is shown to be both feasible and effective in practice.
More Related Videos
08:53Strand-Specific Analysis of Proteins at Replicating DNA Strands by Enrichment and Sequencing of Protein-Associated Nascent DNA Method
Published on: May 2, 2025
07:41Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
Published on: May 17, 2019
Related Concept Videos
Bias in Epidemiological Studies
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Biostatistics: Overview
Discrete variables are...
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...