Related Experiment Videos
Detecting dataset bias in medical AI using a generalized and modality agnostic auditing approach
Nathan Drenkow1,2, Mitchell Pavlak3, Keith Harrigian3
1The Johns Hopkins University, Baltimore, MD, USA. nathan.drenkow@jhuapl.edu.
NPJ Digital Medicine
|May 29, 2026
Summary
Artificial Intelligence (AI) in healthcare can exhibit bias due to non-representative data. We introduce Generalized Attribute Utility and Detectability-Induced bias Testing (G-AUDIT) to automatically detect these AI bias issues in medical datasets.
Area of Science:
- Medical Informatics
- Artificial Intelligence
- Machine Learning
Background:
- Artificial Intelligence (AI) shows promise in healthcare but faces challenges with bias amplification from non-representative datasets.
- Latent bias in AI models can be introduced during training and remain undetected during testing, posing risks in clinical deployment.
Purpose of the Study:
- To develop and present a novel method for detecting AI bias in medical datasets.
- To introduce Generalized Attribute Utility and Detectability-Induced bias Testing (G-AUDIT) as a tool to quantify shortcut learning risks.
Main Methods:
- G-AUDIT is a data modality-agnostic approach for auditing datasets.
- It examines the relationship between task-level annotations, sensor-level measurements, and patient, environmental, and acquisition characteristics.
- The method automatically quantifies risks associated with shortcut learning in AI models.
Main Results:
- The applicability of G-AUDIT was demonstrated across diverse medical datasets, including images, text, and tabular data.
- The method successfully identified potential shortcuts in machine learning tasks that traditional qualitative assessments often miss.
- G-AUDIT provides a quantitative measure of bias, enhancing the reliability of AI in healthcare.
Conclusions:
- G-AUDIT offers a robust solution for identifying and mitigating AI bias in medical applications.
- This automated auditing approach can improve the safety and trustworthiness of AI tools deployed in healthcare.
- The method's modality-agnostic nature makes it broadly applicable for ensuring AI fairness and performance.
Related Concept Videos
Data Validation
Data validation is an essential part of a comprehensive assessment. Validation is confirming or verifying and opening the door to gathering more assessment data as it clarifies vague or unclear data. The process of checking and verifying the collected information is called data validation. The primary purpose of data validation is to ensure data is as free from error, bias, and misinterpretation as possible.
Nursing assessment guides are generally based on holistic models rather than medical...
Nursing assessment guides are generally based on holistic models rather than medical...
Bias in Epidemiological Studies
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
Bias
Bias refers to any tendency that prevents a question from being considered unprejudiced. In research, bias occurs when one outcome or answer is selected or encouraged over others in sampling or testing. Bias can occur during any research phase, including study design, data collection, analysis, and publication.
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...