Related Experiment Video
Updated: Jul 15, 2025

Design and Analysis for Fall Detection System Simplification
Published on: April 6, 2020
Confound-leakage: confound removal in machine learning leads to leakage
Sami Hamdan1,2, Bradley C Love3,4,5, Georg G von Polier1,6,7
1Institute of Neuroscience and Medicine, Brain and Behaviour (INM-7), Forschungszentrum Jülich, 52428 Jülich, Germany.
Linear confound regression (CR) can unexpectedly increase confounding in machine learning (ML) models. This confound-leakage pitfall can lead to biased predictions, highlighting the need for careful handling in ML pipelines.
Area of Science:
- Computational Biology
- Epidemiology
- Medical Informatics
Background:
- Machine learning (ML) is vital for data analysis in epidemiology and medicine, with nonlinear methods excelling at complex predictions.
- ML models can be biased by confounding information in features.
- Featurewise linear confound regression (CR) is a standard method to address confounding, but its pitfalls in ML are not fully understood.
Purpose of the Study:
- To investigate the potential risks of using linear confound regression (CR) within machine learning (ML) pipelines.
- To demonstrate how CR can inadvertently amplify confounding effects, leading to unreliable predictions.
Main Methods:
- Developed a simple framework using the target variable as a confound to analyze CR's impact.
- Employed feature shuffling to distinguish between confound-leakage and genuine information revelation.
- Evaluated the method in a real-world clinical prediction task for attention-deficit/hyperactivity disorder.
Main Results:
- Linear CR can increase confounding risk when combined with nonlinear ML, contrary to expectations.
- Information leakage via CR can inflate effect sizes, leading to overestimated predictive accuracy.
- Demonstrated overestimation in predicting ADHD using speech features when depression was used as a confound.
Conclusions:
- Improper use of CR can lead to untrustworthy, biased, and unfair ML predictions due to confound-leakage.
- Understanding and addressing the confound-leakage pitfall is crucial for developing robust and reliable ML models.
- Guidelines are provided to mitigate these risks in ML applications.
Related Concept Videos
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Confounding in Epidemiological Studies
Difference from Background: Limit of Detection
The LOD indicates the presence or absence...
Quantifying and Rejecting Outliers: The Grubbs Test
Random and Systematic Errors
Parameters Affecting Nonlinear Elimination: Zero-Order Input, First-Order Absorption and Two-Compartment Model
When a drug is administered through a constant intravenous infusion and eliminated via nonlinear pharmacokinetics, it follows zero-order input. For example, oral drugs undergo first-order absorption upon administration and are eliminated through nonlinear pharmacokinetics.
In the case of subcutaneously administered drugs,...

