Related Experiment Video
Updated: Jul 13, 2026

09:37
An Integrated Approach for Microprotein Identification and Sequence Analysis
Published on: July 12, 2022
Automated protein subfamily identification and classification
Duncan P Brown1, Nandini Krishnamurthy, Kimmen Sjölander
1Department of Bioengineering, University of California, Berkeley, California, United States of America.
Plos Computational Biology
|August 22, 2007
Summary
This study introduces SCI-PHY, a phylogenomic pipeline for accurate protein function prediction. It automates subfamily identification using Hidden Markov Models (HMMs), improving gene annotation and avoiding errors common in homology-based methods.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Homology-based gene function prediction is common but prone to errors.
- Phylogenomic analysis offers accurate predictions but is difficult to automate.
- Existing methods struggle with high-throughput annotation and error propagation.
Purpose of the Study:
- To develop a computationally efficient pipeline for automated phylogenomic protein classification.
- To improve the accuracy and reliability of gene function prediction.
- To address the limitations of homology-based annotation and manual phylogenomic analysis.
Main Methods:
- Utilized the SCI-PHY (Subfamily Classification in Phylogenomics) algorithm for automated subfamily identification.
- Employed subfamily Hidden Markov Models (HMMs) for sequence classification.
- Implemented logistic regression to differentiate novel subfamilies and an information-sharing protocol for HMM parameter estimation.
Main Results:
- SCI-PHY subfamilies closely align with expert-defined functional subtypes and conserved phylogenetic clades.
- Subfamily HMMs significantly enhance the discrimination between homologous and non-homologous proteins in database searches.
- Achieved extremely high specificity in classification and demonstrated potential for predicting novel subtypes.
Conclusions:
- The SCI-PHY pipeline provides an automated and efficient method for accurate protein subfamily classification.
- Subfamily HMMs represent a significant advancement over family HMMs for precise functional annotation.
- The SCI-PHY Web server and PhyloFacts resource offer valuable tools for the research community.
Related Concept Videos
Protein Families
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key locations, protein...
Protein Networks
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...

