Related Experiment Video
Updated: May 9, 2026

08:23
De novo Identification of Actively Translated Open Reading Frames with Ribosome Profiling Data
Published on: February 18, 2022
Overview of repeat annotation and de novo repeat identification.
1Department of Horticulture, Michigan State University, East Lansing, MI, USA.
Methods in Molecular Biology (Clifton, N.J.)
|August 7, 2013
Summary
Automating transposable element (TE) identification is crucial for plant genome annotation. This review focuses on de novo repeat finders to overcome the upcoming bottleneck in genomic data analysis.
Area of Science:
- Genomics
- Bioinformatics
- Plant Science
Background:
- Large-scale genomic sequencing is rapidly advancing, particularly in crop plants.
- Transposable elements (TEs) constitute a significant portion of plant genomes.
- Genome annotation is becoming a bottleneck due to the increasing volume of sequence data.
Purpose of the Study:
- To review repeat-finding tools for automated transposable element (TE) identification.
- To focus on de novo repeat identification programs for plant genomes.
- To guide the processing of results and construction of repeat libraries for downstream analysis.
Main Methods:
- Review of functions and mechanisms of various repeat-finding tools.
- Emphasis on de novo repeat identification strategies.
- Discussion of post-identification processing and library construction.
Main Results:
- Identification and classification of TEs present challenges in plant genome annotation.
- De novo repeat identification programs are essential for efficient genome characterization.
- Automated methods are needed to handle the growing genomic data.
Conclusions:
- Automation of TE identification is critical for future plant genome annotation.
- Understanding TE composition and dynamics is facilitated by efficient identification tools.
- The development and application of de novo repeat finders are key to overcoming annotation bottlenecks.
Related Concept Videos
Genome Annotation and Assembly
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
RNA-seq
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Next-generation Sequencing
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.

