Related Experiment Video
Updated: Jul 18, 2025

Exploring Sequence Space to Identify Binding Sites for Regulatory RNA-Binding Proteins
Published on: August 9, 2019
Prediction accuracy of regulatory elements from sequence varies by functional sequencing technique.
Ronald J Nowling1, Kimani Njoya2, John G Peters1
1Electrical Engineering and Computer Science, Milwaukee School of Engineering, Milwaukee, WI, United States.
Machine learning models predict cis-regulatory elements using DNA sequence. DNase-seq and STARR-seq data accurately predict regulatory activity, unlike ChIP-seq methods, highlighting sequence-based prediction potential.
Area of Science:
- Genomics
- Computational Biology
- Molecular Biology
Background:
- * Identifying cis-regulatory elements (CREs) is crucial for understanding genome regulation.
- * Various sequencing techniques (ChIP-seq, ATAC-seq, DNase-seq, STARR-seq) are used, relying on direct sequence or indirect markers.
- * CRE activity depends on DNA sequence and secondary epigenetic processes.
Purpose of the Study:
- * To evaluate the predictive accuracy of CREs based solely on DNA sequence using machine learning.
- * To differentiate sequence-driven regulatory activity from that influenced by secondary processes.
- * To compare the suitability of different sequencing datasets for training sequence-based prediction models.
Main Methods:
- * Machine learning models were trained and evaluated on cis-regulatory element sequences from *D. melanogaster*.
- * Datasets included DNase-seq, STARR-seq, and various ChIP-seq (H3K4me1, H3K4me3, H3K27ac), FAIRE-seq, and ATAC-seq data.
- * Experimental validation using luciferase assays compared enhancer activity of selected sequences.
Main Results:
- * Models trained on DNase-seq and STARR-seq data showed significantly higher accuracy than those using ChIP-seq, FAIRE-seq, or ATAC-seq.
- * DNase-seq and STARR-seq activities are largely explained by DNA sequence, independent of secondary epigenetic marks.
- * Luciferase assays confirmed STARR-seq sequences are enriched for enhancer activity, while DNase-seq and H3K4me1 ChIP-seq sequences are not.
Conclusions:
- * DNase-seq identifies a broad range of regulatory elements, suitable for general sequence-based activity prediction.
- * STARR-seq data are optimal for training models to predict enhancer-specific sequences.
- * H3K4me1 ChIP-seq data are not well-suited for training sequence-based models for CRE prediction.
Related Concept Videos
Cis-regulatory Sequences
Evolutionary Relationships through Genome Comparisons
Cooperative Binding of Transcription Regulators
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Next-generation Sequencing
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....

