Related Experiment Video
Updated: Oct 11, 2025

Constructing and Visualizing Models using Mime-based Machine-learning Framework
Published on: July 22, 2025
Navigating the pitfalls of applying machine learning in genomics
Sean Whalen1, Jacob Schreiber2, William S Noble3
1Gladstone Institutes, San Francisco, CA, USA.
Abstract:
The scale of genetic, epigenomic, transcriptomic, cheminformatic and proteomic data available today, coupled with easy-to-use machine learning (ML) toolkits, has propelled the application of supervised learning in genomics research. However, the assumptions behind the statistical models and performance evaluations in ML software frequently are not met in biological systems. In this Review, we illustrate the impact of several common pitfalls encountered when applying supervised ML in genomics. We explore how the structure of genomics data can bias performance evaluations and predictions. To address the challenges associated with applying cutting-edge ML methods to genomics, we describe solutions and appropriate use cases where ML modelling shows great potential.
More Related Videos
Related Concept Videos
Genomics
Evolutionary Relationships through Genome Comparisons
Modern Molecular Taxonomy
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...

