Related Experiment Video
Updated: Nov 28, 2025

Exploring Sequence Space to Identify Binding Sites for Regulatory RNA-Binding Proteins
Published on: August 9, 2019
The impact of different negative training data on regulatory sequence predictions
Louisa-Marie Krützfeldt1,2, Max Schubach1,2, Martin Kircher1,2
1Charité-Universitätsmedizin Berlin, Berlin, Germany.
Choosing the right negative sequences is crucial for training accurate deep learning models for DNA regulatory elements. Genomic background sequences generally yield better performance than shuffled sequences for predicting regulatory activity.
Area of Science:
- Genomics
- Computational Biology
- Bioinformatics
Background:
- Regulatory DNA elements (promoters, enhancers) are vital for gene regulation and human variation.
- Understanding the functional encoding of these regions is limited.
- Machine learning, including gapped k-mer support vector machines (gkm-SVMs) and convolutional neural networks (CNNs), shows promise for deciphering DNA encoding.
Purpose of the Study:
- To investigate the impact of negative sequence selection on the performance of machine learning models trained on regulatory DNA.
- To compare different strategies for generating negative training datasets.
Main Methods:
- Trained gapped k-mer support vector machines (gkm-SVMs) and convolutional neural networks (CNNs) on open chromatin data.
- Compared two negative training data approaches: genomic background sequences and sequence shuffles of positive sequences.
- Evaluated model performance on predicting cell-type activity, cell-type specificity, and quantitative activity.
Main Results:
- Genomic background sequences generally led to better model performance compared to shuffled sequences.
- Models trained on highly shuffled sequences underperformed on complex prediction tasks and learned artificial sequence features.
- CNNs outperformed gkm-SVMs on larger datasets, while gkm-SVMs provided robust results for typical dataset sizes without extensive tuning.
Conclusions:
- Negative sequence selection significantly impacts the performance of machine learning models for regulatory element prediction.
- Genomic background sequences are a more effective choice for negative training data.
- The choice of model and negative data strategy should consider dataset size and prediction task complexity.
More Related Videos
08:04Identification and Classification of Position-specific GABAA Receptor Subunit Missense Variants for Their Role In Hippocampal Pyramidal Neurons
Published on: June 6, 2025
03:37Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
Related Concept Videos
Cis-regulatory Sequences
Cis-regulatory Sequences
Negative Regulator Molecules
Gene Evolution - Fast or Slow?
In contrast, regions which code...
Cooperative Binding of Transcription Regulators
Survival Tree
Building a Survival Tree
Constructing a...