Related Experiment Video
Updated: Jul 31, 2025

Droplet Barcoding-Based Single Cell Transcriptomics of Adult Mammalian Tissues
Published on: January 10, 2019
Genomic benchmarks: a collection of datasets for genomic sequence classification
Katarína Grešová1,2, Vlastimil Martinek1,2, David Čechák1,2
1Centre for Molecular Medicine, Central European Institute of Technology (CEITEC), Masaryk University, Brno, Czechia.
Researchers have developed new benchmark datasets for genomics to advance deep learning applications. This collection of genomic sequence datasets aims to improve the reproducibility and comparability of machine learning models in the field.
Area of Science:
- Genomics
- Bioinformatics
- Machine Learning
Background:
- Deep neural networks show promise in biological fields, exemplified by AlphaFold's success in protein structure prediction.
- However, the advancement of deep learning in genomics is hindered by the lack of curated benchmark datasets, unlike in protein folding.
- Genomics faces challenges in genome annotation and functional element identification, requiring standardized benchmarks.
Purpose of the Study:
- To create a collection of curated and accessible sequence classification datasets for genomics research.
- To provide a baseline model and framework for machine learning in genomics.
- To enhance the comparability, reproducibility, and accessibility of machine learning applications in genomics.
Main Methods:
- Compiled novel datasets by mining public databases and incorporating existing published datasets.
- Focused on regulatory elements, including promoters, enhancers, and open chromatin regions.
- Developed a Python package 'genomic-benchmarks' containing nine datasets for three model organisms (human, mouse, roundworm) and a baseline convolutional neural network model.
Main Results:
- A collection of nine curated genomic sequence classification datasets is now available.
- These datasets cover regulatory elements across human, mouse, and roundworm.
- A baseline convolutional neural network model and associated code are provided for immediate use.
Conclusions:
- The proposed benchmark datasets and baseline model facilitate machine learning in genomics.
- This effort aims to make machine learning for genomics more comparable and reproducible.
- The repository is intended to reduce researcher overhead and foster competition and discovery in the field.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Genomics
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Modern Molecular Taxonomy
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Genome Annotation and Assembly

