Related Experiment Video
Updated: Jun 18, 2025

09:40
Novel Sequence Discovery by Subtractive Genomics
Published on: January 25, 2019
8.6K
Genomic background sequences systematically outperform synthetic ones in de novo motif discovery for ChIP-seq data
Vladimir V Raditsa1, Anton V Tsukanov1, Anton G Bogomolov2
1Department of System Biology, Institute of Cytology and Genetics, Novosibirsk 630090, Russia.
NAR Genomics and Bioinformatics
|July 29, 2024
Summary
Choosing the right background sequences is crucial for accurate transcription factor motif discovery. A genomic background approach, implemented in the AntiNoise web service, offers more robust motif detection and better exclusion of non-specific repeats.
Area of Science:
- Genomics
- Bioinformatics
- Molecular Biology
Background:
- Accurate de novo motif discovery from ChIP-seq data relies heavily on appropriate background sequence selection.
- ChIP-seq peaks can contain both specific transcription factor binding motifs and non-specific motifs like simple sequence repeats, complicating analysis.
- Existing methods for generating background sequences, such as the synthetic approach, may not adequately account for genomic biases.
Purpose of the Study:
- To compare the effectiveness of synthetic versus genomic background sequence generation for de novo motif discovery.
- To evaluate the performance of these approaches across different species, including mammals and plants.
- To develop a user-friendly web service for implementing an improved background sequence generation method.
Main Methods:
- Compiled benchmark ChIP-seq datasets for mouse, human, and Arabidopsis.
- Performed de novo motif discovery using both synthetic (shuffled peak nucleotides) and genomic (random or promoter-based genome sequences) background approaches.
- Developed and validated the AntiNoise web service utilizing the genomic approach for background sequence extraction.
Main Results:
- The genomic background approach demonstrated more robust detection of known transcription factor motifs compared to the synthetic approach.
- The genomic approach showed more effective exclusion of simple sequence repeats, reducing false positive motif identification.
- The advantage of the genomic approach was more pronounced in plant datasets than in mammalian datasets.
- The AntiNoise web service was developed, supporting twelve eukaryotic genomes.
Conclusions:
- The genomic approach is superior to the synthetic approach for generating background sequences in de novo motif discovery from ChIP-seq data.
- The AntiNoise web service provides a valuable tool for researchers needing reliable background sequences for motif analysis.
- The findings highlight the importance of species-specific genomic characteristics when selecting background sequences for motif discovery.
Related Concept Videos
RNA-seq
9.9K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
9.9K
Maxam-Gilbert Sequencing
11.1K
In the same year as the discovery of the Sanger sequencing method, another group of scientists, Allan Maxam and Walter Gilbert, demonstrated their chemical-cleavage method for DNA sequencing. The Maxam-Gilbert method relies on using different chemicals that can cleave the DNA sequence at specific sites, the separation of resulting DNA fragments of variable size using electrophoresis, and deciphering the DNA sequence from the resulting gel bands.
Challenges of the Maxam-Gilbert Method
The...
Challenges of the Maxam-Gilbert Method
The...
11.1K
DNA Microarrays
17.3K
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
17.3K
Next-generation Sequencing
88.5K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
88.5K
Ribosome Profiling
3.5K
Ribosome profiling or ribo-sequencing is a deep sequencing technique that produces a snapshot of active translation in a cell. It selectively sequences the mRNAs protected by ribosomes to get an insight into a cell’s translation landscape at any given point in time.
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
3.5K
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K

