Related Experiment Video
Updated: Jul 16, 2025

An Open Source Technology Platform to Manufacture Hydrogel-Based 3D Culture Models in an Automated and Standardized Fashion
Published on: March 31, 2022
Leakage and the reproducibility crisis in machine-learning-based science
Sayash Kapoor1, Arvind Narayanan1
1Department of Computer Science and Center for Information Technology Policy, Princeton University, Princeton, NJ 08540, USA.
Machine-learning (ML) methods often suffer from data leakage, leading to overoptimistic results. Correcting these errors reveals that complex ML models do not outperform traditional logistic regression (LR) in many scientific applications.
Area of Science:
- Quantitative sciences
- Computational social science
Background:
- Machine-learning (ML) methods are increasingly used in quantitative sciences.
- Methodological pitfalls, such as data leakage, can compromise the reliability of ML-based research.
- Reproducibility issues are a significant concern in the application of ML in science.
Purpose of the Study:
- To systematically investigate reproducibility issues in ML-based science, focusing on data leakage.
- To identify and categorize different types of data leakage in ML research.
- To propose a method for testing and mitigating data leakage.
Main Methods:
- Systematic literature survey across 17 scientific fields that utilize ML methods.
- Development of a detailed taxonomy of eight types of data leakage.
- Introduction of model info sheets for researchers to test for leakage.
- Reproducibility study comparing complex ML models with traditional logistic regression (LR) in civil war prediction.
Main Results:
- Data leakage was identified in 294 papers across 17 fields, often resulting in exaggerated conclusions.
- A comprehensive taxonomy of eight data leakage types was established.
- Model info sheets were proposed as a tool for assessing leakage.
- In a civil war prediction case study, corrected ML models showed no substantial performance advantage over LR models.
Conclusions:
- Data leakage is a widespread problem in ML-based science, undermining reproducibility and leading to inflated performance metrics.
- The proposed taxonomy and model info sheets offer a framework for identifying and addressing leakage.
- Complex ML models may not inherently outperform simpler, established statistical methods like LR when methodological rigor is applied.
More Related Videos
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Related Concept Videos
Systematic Error: Methodological and Sampling Errors
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Propagation of Uncertainty from Systematic Error
Random and Systematic Errors
Survival Tree
Building a Survival Tree
Constructing a...
Uncertainty in Measurement: Accuracy and Precision
Steps in Outbreak Investigation