Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches01:23

Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches

124
Biopharmaceutical studies constitute a vital field aiming to enhance drug delivery methods and refine therapeutic approaches, drawing upon diverse interdisciplinary knowledge. In research methodologies, the choice between controlled and non-controlled studies significantly influences the study's reliability and accuracy.
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
124
Bias in Epidemiological Studies01:29

Bias in Epidemiological Studies

190
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:  
190
Mouse Models of Cancer Study02:43

Mouse Models of Cancer Study

5.5K
Mice have long served as models for studying human biology and pathology because of their phylogenetic and physiological similarity with humans. They are also easy to maintain and breed in the laboratory, and hence, many inbred strains are now available for research. Studies on mice have contributed immeasurably to our understanding of cancer biology.
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...
5.5K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Inferring and evaluating network medicine-based disease modules with nextflow.

Bioinformatics (Oxford, England)·2026
Same author

EpiATLAS - a reference for human epigenomic research.

bioRxiv : the preprint server for biology·2026
Same author

SATB1 is a targetable modulator of JAK-STAT signaling and cytokines in human Treg and Tconv cells.

EMBO reports·2026
Same author

Aggregicyclins Shed Light on Type II Polyketide Biosynthesis in <i>Myxococcota</i>.

JACS Au·2026
Same author

Critical evaluation of drug response prediction models with DrEval.

Nature communications·2026
Same author

Drugst.One DREAM-Drug repurposing through expert annotation and modification.

British journal of pharmacology·2026

Related Experiment Video

Updated: Jun 17, 2025

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
07:15

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model

Published on: August 16, 2020

6.8K

Guiding questions to avoid data leakage in biological machine learning applications.

Judith Bernett1, David B Blumenthal2, Dominik G Grimm3,4,5

  • 1TUM School of Life Sciences, Technical University of Munich, Freising, Germany.

Nature Methods
|August 9, 2024
PubMed
Summary

Data leakage in machine learning can inflate biological model performance. This study introduces seven key questions to help researchers identify and prevent data leakage, ensuring more reliable biological data analysis and reproducible research.

More Related Videos

Author Spotlight: AI-Driven Trypanosome Species Detection from Microscopic Images
08:20

Author Spotlight: AI-Driven Trypanosome Species Detection from Microscopic Images

Published on: October 27, 2023

1.4K
A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
07:35

A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports

Published on: October 13, 2023

1.6K

Related Experiment Videos

Last Updated: Jun 17, 2025

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
07:15

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model

Published on: August 16, 2020

6.8K
Author Spotlight: AI-Driven Trypanosome Species Detection from Microscopic Images
08:20

Author Spotlight: AI-Driven Trypanosome Species Detection from Microscopic Images

Published on: October 27, 2023

1.4K
A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
07:35

A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports

Published on: October 13, 2023

1.6K

Area of Science:

  • Bioinformatics
  • Computational Biology
  • Machine Learning in Biology

Background:

  • Machine learning (ML) is crucial for pattern extraction in high-dimensional biological data.
  • Reported ML prediction performance in biology often fails in real-world applications.
  • Data leakage, or information sharing between training and test sets, is a primary cause of inflated performance estimates.

Purpose of the Study:

  • To address the challenge of data leakage in biological machine learning.
  • To provide a practical framework for preventing data leakage in biological datasets.
  • To promote robust and reproducible machine learning research in the life sciences.

Main Methods:

  • Development of seven critical questions to guide the prevention of data leakage.
  • Application of these questions to non-trivial biological dataset examples.
  • Illustrative analysis to demonstrate the utility of the proposed questions.

Main Results:

  • Data leakage is a significant and often undetected issue in biological ML.
  • The proposed seven questions effectively highlight potential data leakage scenarios.
  • The framework aids in identifying and mitigating information contamination in biological models.

Conclusions:

  • Awareness of potential data leakage is essential for biological ML.
  • Implementing the seven questions can improve the reliability of ML models in biology.
  • This approach supports the development of more robust and reproducible biological research using ML.