Related Experiment Video
Updated: Jul 7, 2026

An Integrated Approach for Microprotein Identification and Sequence Analysis
Published on: July 12, 2022
Subfamily specific conservation profiles for proteins based on n-gram patterns.
1Department of Computational Biology, School of Medicine, University of Pittsburgh, Pittsburgh, PA 15213, USA. vries@ccbb.pitt.edu
A novel algorithm generates subfamily-specific conservation profiles using n-gram patterns, reflecting evolutionary history. This method accurately captures evolutionary signals even for small subfamilies, offering a powerful tool for sequence analysis.
Area of Science:
- Bioinformatics
- Computational Biology
- Evolutionary Biology
Background:
- A new algorithm has been developed for generating conservation profiles based on n-gram patterns (NP{n,m}).
- These profiles reflect the evolutionary history of a subfamily associated with a query sequence.
- The algorithm treats profile generation as a signal-to-noise problem, differentiating evolutionary signals from noise using singular value decomposition.
Purpose of the Study:
- To introduce and evaluate a novel algorithm for generating subfamily-specific conservation profiles.
- To assess the algorithm's ability to capture evolutionary history and its performance compared to existing methods.
- To demonstrate the utility of the algorithm for subfamilies with limited representation in databases.
Main Methods:
- Utilizes n-gram patterns (sets of residues and wildcards) to represent sequence features.
- Applies singular value decomposition to rank-ordered target sequences to separate evolutionary signal from noise.
- Compares generated profiles against those derived from multiple sequence alignments using a consensus approach.
Main Results:
- Successfully generated 4,248 conservation profiles from 120 Pfam-A families.
- New profiles closely matched consensus profiles when subfamilies were well-represented in multiple alignments.
- The algorithm effectively constructed subfamily-specific profiles for subfamilies with as few as five members, with comparable speed to multiple alignment methods.
Conclusions:
- Subfamily-specific conservation profiles can be generated without prior knowledge of family relationships or domain architecture.
- The algorithm is particularly useful for subfamilies with multiple domains or sparse representation in protein databases.
- This approach is applicable to small subfamily sample sizes that are insufficient for traditional multiple alignment methods.
More Related Videos
16:02Demonstration of the Sequence Alignment to Predict Across Species Susceptibility Tool for Rapid Assessment of Protein Conservation
Published on: February 10, 2023
07:49Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group
Published on: August 16, 2017
Related Concept Videos
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Conservation of Protein Domains
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...
Protein Families
Protein Families
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...