Related Experiment Video
Updated: Oct 19, 2025

16:17
The ITS2 Database
Published on: March 12, 2012
31.2K
Accurate annotation of protein coding sequences with IDTAXA.
Nicholas P Cooley1, Erik S Wright1
1Department of Biomedical Informatics, University of Pittsburgh, Pittsburgh, PA 15206, USA.
NAR Genomics and Bioinformatics
|September 20, 2021
Summary
A new algorithm, IDTAXA, accurately classifies protein functions from sequences, outperforming BLAST and HMMER. This tool aids genome annotation and discovers novel antibiotic resistance associations.
Area of Science:
- Bioinformatics
- Genomics
- Computational Biology
Background:
- Protein sequence data grows faster than functional knowledge.
- Current protein annotation relies on homology searches (BLAST, HMMER).
- Existing methods may misclassify novel protein functions.
Purpose of the Study:
- Develop a novel protein classification algorithm.
- Improve accuracy and reliability in genome annotation.
- Identify novel protein-function associations, including antibiotic resistance.
Main Methods:
- Developed IDTAXA, a novel sequence-based protein classification algorithm.
- Compared IDTAXA performance against BLAST and HMMER for KEGG ortholog assignment.
- Applied IDTAXA for eukaryotic and prokaryotic genome annotation and contamination detection.
- Re-annotated microbial genomes to investigate antibiotic resistance phenotypes.
Main Results:
- IDTAXA demonstrated higher accuracy than BLAST and HMMER in assigning KEGG ortholog groups.
- IDTAXA effectively distinguished novel functions, avoiding common misclassification errors.
- Successfully applied to multi-level genome annotation and detection of eukaryotic genome contamination.
- Discovered two novel associations between proteins and antibiotic resistance in microbial genomes.
Conclusions:
- IDTAXA offers a more accurate and reliable method for protein function classification.
- The algorithm enhances genome annotation pipelines and aids in discovering biological insights.
- IDTAXA has practical applications in identifying genome contamination and exploring genotype-phenotype relationships.
Related Concept Videos
Genome Annotation and Assembly
19.6K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.6K
Modern Molecular Taxonomy
280
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...
280
Applications of Molecular Taxonomy
211
Molecular taxonomy has revolutionized the understanding and classification of bacteria, providing precise insights into their diversity, evolutionary relationships, and ecological roles. By utilizing molecular techniques such as DNA sequencing and fingerprinting, researchers have made significant strides in various fields related to bacterial studies.Resolving Taxonomic AmbiguitiesMolecular taxonomy has been instrumental in distinguishing closely related bacterial species initially thought to...
211
Peptide Identification Using Tandem Mass Spectrometry
7.3K
Tandem mass spectrometry, also known as MS/MS or MS2, is an analytical technique that employs two mass analyzers. Essentially it is a series of mass spectrometers that helps isolate a particular biomolecule and then helps study its chemical properties.
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
7.3K
Tagging and Fusion Proteins
7.4K
Proteins are involved in several cellular processes and biochemical reactions. Analyzing a specific protein of interest requires it to be isolated from the other proteins in the cell. This is achieved by overexpressing the specific gene in a suitable host to produce large quantities of the target protein. A tag or label is recombined with the gene to produce a fusion protein containing the target protein and the tag. The tags on these fusion proteins can then be used for easy detection and...
7.4K

