Related Experiment Videos
Structural similarity to link sequence space: new potential superfamilies and implications for structural genomics
Patrick Aloy1, Baldomero Oliva, Enrique Querol
1EMBL, Biocomputing, Meyerhofstrasse 1, D-69117 Heidelberg, Germany.
Protein Science : a Publication of the Protein Society
|April 23, 2002
Summary
This study links protein structure and sequence databases to predict distant protein relationships and infer function. It identifies new protein superfamilies, aiding structural genomics and sequence comparison methods.
Area of Science:
- Structural biology
- Bioinformatics
- Computational biology
Background:
- Protein structure determination now often precedes functional characterization.
- Structure-based homology assignment is crucial for inferring protein function.
- Sequence similarity after structure-based alignment is a strong indicator of homology.
Purpose of the Study:
- To develop a method for predicting distant homologous relationships by merging protein structure and sequence databases.
- To identify novel homologous protein families and superfamilies.
- To infer protein function for proteins with known structures but unknown functions.
Main Methods:
- Utilized the Structural Classification of Proteins (SCOP) database to link sequence alignments from SMART and Pfam.
- Developed new sequence alignments leveraging known three-dimensional protein structures.
- Extended Murzin's method to statistically assess sequence identities after structural alignment.
Main Results:
- Successfully linked several distantly related protein sequence families with confidence.
- Identified new potential protein superfamilies based on conserved structural features and substrate binding sites.
- Demonstrated a robust method for inferring homology and function from structural data.
Conclusions:
- The integrated approach effectively predicts homologous relationships and infers function for proteins with known structures.
- The findings have significant implications for Structural Genomics initiatives and improving sequence comparison algorithms.
- This method enhances our ability to assign function in the post-genomic era.