Related Experiment Video
Updated: Nov 21, 2025

10:41
Identifying Amino Acid Overproducers Using Rare-Codon-Rich Markers
Published on: June 24, 2019
8.6K
Utilizing Amino Acid Composition and Entropy of Potential Open Reading Frames to Identify Protein-Coding Genes.
Katelyn McNair1, Carol L Ecale Zhou2, Brian Souza3
1Computational Sciences Research Center, San Diego State University, 5500 Campanile Drive, San Diego, CA 92182, USA.
Microorganisms
|January 12, 2021
Summary
We developed GOODORFS, a novel method for identifying protein-coding genes in prokaryotes. GOODORFS outperforms existing tools in accuracy and F1-score for gene prediction.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Prokaryotic gene finding requires distinguishing protein-coding open reading frames (ORFs) from chance occurrences.
- Current homology-based methods are limited to identifying previously discovered genes.
- Existing gene prediction programs often rely on protein-coding training models, with some allowing *ab initio* model creation.
Purpose of the Study:
- To introduce GOODORFS, a new method for identifying protein-coding genes from all possible ORFs.
- To enhance the accuracy of prokaryotic gene prediction by improving training model creation.
- To evaluate GOODORFS against established gene prediction tools.
Main Methods:
- GOODORFS utilizes amino acid frequencies to calculate an entropy density profile (EDP) for each ORF.
- KMeans clustering is applied to EDPs to group potential coding regions.
- The cluster with the lowest variation is selected as the set of coding ORFs.
Main Results:
- GOODORFS achieved the highest accuracy (0.94) and F1-score (0.85) across 14,179 annotated phage genomes.
- Compared to Glimmer, MED2, PHANOTATE, and Prodigal, GOODORFS demonstrated superior overall performance.
- Glimmer showed the highest precision (0.92), while PHANOTATE had the highest recall (0.96).
Conclusions:
- GOODORFS provides a more accurate and effective approach for identifying protein-coding genes in prokaryotes.
- The entropy density profile and KMeans clustering method offers an improved strategy for training model creation in gene prediction.
- This method addresses limitations of homology-based approaches by enabling *de novo* gene discovery.
Related Concept Videos
Ribosome Profiling
3.9K
Ribosome profiling or ribo-sequencing is a deep sequencing technique that produces a snapshot of active translation in a cell. It selectively sequences the mRNAs protected by ribosomes to get an insight into a cell’s translation landscape at any given point in time.
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
3.9K
From DNA to Protein
21.0K
The flow of genetic information in cells from DNA to mRNA to protein is described by the central dogma, which states that genes specify the sequence of mRNAs, which in turn specify the sequence of amino acids making up all proteins. The decoding of one molecule to another is performed by specific proteins and RNAs. Because the information stored in DNA is so central to cellular function, it makes intuitive sense that the cell would make mRNA copies of this information for protein synthesis...
21.0K
Leaky Scanning
5.4K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.4K
Genome Annotation and Assembly
19.9K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.9K
The Central Dogma
30.9K
The central dogma explains the flow of genetic information from DNA nucleotides to the amino acid sequence of proteins.
RNA is the Missing Link Between DNA and Proteins
In the early 1900s, scientists discovered that DNA stores all the information needed for cellular functions and that proteins perform most of these functions. However, the mechanisms of converting genetic information into functional proteins remained unknown for many years. Initially, it was believed that a single gene is...
RNA is the Missing Link Between DNA and Proteins
In the early 1900s, scientists discovered that DNA stores all the information needed for cellular functions and that proteins perform most of these functions. However, the mechanisms of converting genetic information into functional proteins remained unknown for many years. Initially, it was believed that a single gene is...
30.9K
The Central Dogma
135.7K
Overview
135.7K

