Related Experiment Video
Updated: Mar 21, 2026

08:23
De novo Identification of Actively Translated Open Reading Frames with Ribosome Profiling Data
Published on: February 18, 2022
4.3K
OrfM: a fast open reading frame predictor for metagenomic data
Ben J Woodcroft1, Joel A Boyd1, Gene W Tyson1
1Australian Centre for Ecogenomics, School of Chemistry and Molecular Biosciences, University of Queensland, Brisbane, QLD 4072, Australia.
Bioinformatics (Oxford, England)
|May 7, 2016
Summary
OrfM is a new bioinformatics tool that rapidly identifies open reading frames (ORFs) in DNA sequences. It is four times faster than existing methods, addressing a key bottleneck in genomic data analysis.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Identifying open reading frames (ORFs) is crucial for DNA sequence analysis.
- Current computational tools for ORF finding are slow, creating a bottleneck with increasing sequence data volumes.
- This bottleneck is particularly significant in metagenomics for analyzing unassembled reads and assembled contigs.
Purpose of the Study:
- To develop a faster computational tool for identifying open reading frames (ORFs).
- To address the performance limitations of existing ORF-finding software.
- To provide a solution for efficient gene discovery in large-scale sequencing datasets.
Main Methods:
- Developed OrfM, a novel tool for rapid ORF identification.
- Employed the Aho-Corasick algorithm to efficiently locate regions without stop codons.
- Benchmarked OrfM against established tools like GetOrf and Translate.
Main Results:
- OrfM identifies identical ORFs compared to existing tools.
- OrfM demonstrates a 4-5 times speed improvement over similar ORF-finding software.
- The tool is sequencing platform-agnostic, excelling with large, high-quality datasets (e.g., Illumina).
Conclusions:
- OrfM significantly accelerates the identification of open reading frames (ORFs).
- The tool effectively overcomes computational bottlenecks in sequence data analysis.
- OrfM is a valuable asset for genomic and metagenomic research, especially with high-throughput sequencing data.
More Related Videos
Related Concept Videos
RACE - Rapid Amplification of cDNA Ends
7.5K
Rapid Amplification of cDNA Ends, or RACE, is one of the most effective methods to obtain a full-length cDNA from an mRNA sequence between a known internal region to the unknown sequence at the 5’ or 3’ end. The unknown region is cloned in the cDNA by a gene-specific primer that binds the known end, and a hybrid primer that attaches a predefined anchor sequence to the unknown end of the cDNA. The sequence in between is amplified by PCR with an anchor primer and a gene-specific...
7.5K
Next-generation Sequencing
100.8K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
100.8K
Genomics
41.7K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
41.7K
Sanger Sequencing
777.7K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
777.7K
RNA-seq
12.4K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
12.4K
Point and Frameshift Mutations
1.6K
Point mutations are genetic alterations involving the change of a single nucleotide base pair in DNA. Depending on how the alteration affects protein synthesis, they can lead to various consequences.Point mutations fall into the following types:Silent mutations occur when a nucleotide change does not alter the amino acid sequence due to the redundancy of the genetic code. For instance, changing ACC to ACA still encodes threonine, leaving the protein function unaffected. This occurs because...
1.6K

