Related Experiment Videos
Assessing annotation transfer for genomics: quantifying the relations between protein sequence, structure and
C A Wilson1, J Kreychman, M Gerstein
1Department of Molecular Biophysics and Biochemistry, Yale University, New Haven, CT 06520, USA.
Journal of Molecular Biology
|March 8, 2000
Summary
This study quantifies how much protein sequence, structure, and function information transfers between related proteins. It reveals clear thresholds for functional conservation based on sequence identity, crucial for genome annotation.
Area of Science:
- * Bioinformatics
- * Structural Biology
- * Genomics
Background:
- * Robust genome annotation requires quantitative measures of information transfer between related protein sequences.
- * Understanding the relationship between protein sequence, structure, and function is key to predicting protein roles.
- * Previous studies established general relationships, but a larger-scale analysis with modern scoring was needed.
Purpose of the Study:
- * To statistically measure the transfer of structural and functional information across protein sequence similarities.
- * To refine the understanding of sequence-structure relationships using traditional and modern scoring methods.
- * To establish thresholds for functional conservation based on sequence and structural similarity.
Main Methods:
- * Performed pairwise comparisons of approximately 30,000 protein domains with known structures and functions, categorized by SCOP fold.
- * Analyzed sequence and structure similarity using traditional scores (e.g., percent identity, RMSD) and modern scores (e.g., Smith-Waterman, P-values).
- * Developed a functional classification scheme by integrating existing annotations to assess functional similarity levels.
Main Results:
- * Confirmed an exponential relationship between structural divergence (RMSD) and sequence divergence (percent identity).
- * Demonstrated that modern scoring methods provide more precise sequence-structure relationship quantification.
- * Identified sigmoidal relationships between functional and sequence similarity, with precise function conserved down to ~40% sequence identity and broad function to ~25%.
Conclusions:
- * The fundamental exponential sequence-structure relationship is general across different protein classes and scoring schemes.
- * Percent identity is a surprisingly effective metric for quantifying functional conservation compared to advanced scores.
- * Findings provide crucial quantitative data for improving genome annotation and understanding protein evolution.