Related Experiment Video
Updated: May 26, 2026

09:37
An Integrated Approach for Microprotein Identification and Sequence Analysis
Published on: July 12, 2022
MetaDomain: a profile HMM-based protein domain classification tool for short sequences
1Department of Computer Science and Engineering, Michigan State University, East Lansing, MI 48824, USA. zhangy72@msu.edu
Pacific Symposium on Biocomputing. Pacific Symposium on Biocomputing
|December 17, 2011
Summary
MetaDomain improves protein domain classification for short sequencing reads. This new tool enhances functional profiling in metagenomic annotation by increasing sensitivity for identifying protein domains in short sequences.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Protein homology search is crucial for functional profiling in metagenomic annotation.
- Profile Hidden Markov Model (HMM)-based methods offer high sensitivity for remote homology detection but struggle with short reads.
- Short reads from next-generation sequencing (NGS) present a challenge for accurate protein domain classification.
Purpose of the Study:
- To introduce MetaDomain, a novel computational tool for protein domain classification specifically designed for short sequencing reads.
- To enhance the sensitivity and accuracy of identifying protein domains within short DNA sequences.
Main Methods:
- MetaDomain employs relaxed position-specific score thresholds for aligning reads to profile HMMs.
- It utilizes the distribution of alignment positions as an additional constraint to minimize false positive matches.
- The tool was applied to bacterial genome transcriptomic data and a soil metagenomic dataset.
Main Results:
- MetaDomain demonstrated superior sensitivity compared to state-of-the-art profile HMM alignment tools for short sequences.
- The tool effectively identified encoded protein domains from short reads in both tested datasets.
- Experimental results validated MetaDomain's improved performance in metagenomic annotation.
Conclusions:
- MetaDomain offers a significant advancement in protein domain classification for short reads.
- The tool enhances functional profiling capabilities in metagenomics, particularly for NGS data.
- MetaDomain provides a valuable solution for analyzing complex biological datasets with short sequence reads.
Related Concept Videos
Conservation of Protein Domains Over Different Proteins
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Conservation of Protein Domains
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Protein Families
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key locations, protein...
Protein Families
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key locations, protein...
Modern Molecular Taxonomy
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...
Membrane Domains
The membrane domains concentrate specific lipids and proteins at one place within the membrane, which helps in cell signaling, adhesion, and other critical cellular processes. These domains can differ in size, composition, function, and lifespan.
Protein Domains
The membrane comprises a group of distinct proteins responsible for carrying out a cell's specific function. For example, the plasma membrane of the human sperm, or a single germ cell, contains a unique set of proteins in the anterior...
Protein Domains
The membrane comprises a group of distinct proteins responsible for carrying out a cell's specific function. For example, the plasma membrane of the human sperm, or a single germ cell, contains a unique set of proteins in the anterior...

