Related Experiment Videos
Weighting aligned protein or nucleic acid sequences to correct for unequal representation
1European Molecular Biology Laboratory, Heidelberg, F.R.G.
Journal of Molecular Biology
|December 20, 1990
Summary
Sequence databases are often unrepresentative of entire protein families due to skewed organism representation and limited sequencing. This study introduces a novel algorithm for weighting sequences to improve database search accuracy for distantly related family members.
Area of Science:
- Bioinformatics
- Computational Biology
- Molecular Evolution
Background:
- Sequence databases exhibit significant bias, overrepresenting certain organisms and underrepresenting the full diversity of protein families.
- Existing sequence alignments and profiles may not accurately reflect the complete evolutionary landscape of protein families.
- This skewed representation poses challenges for identifying distantly related proteins using computational methods.
Purpose of the Study:
- To address the limitations of biased sequence databases in protein family analysis.
- To develop and present an algorithm for effectively weighting individual sequences within protein families.
- To enhance the accuracy of database searches for distantly related family members.
Main Methods:
- Development of a novel sequence weighting algorithm.
- Application of the algorithm to protein families, using haemoglobins as an example.
- Evaluation of the algorithm's efficacy in correcting for unequal sequence representation.
Main Results:
- The proposed algorithm effectively assigns weights to individual sequences, accounting for database biases.
- Weighted sequence sets provide a more representative view of protein families compared to unweighted sets.
- Demonstrated improvement in the identification of distantly related proteins through database searches.
Conclusions:
- Sequence weighting is crucial for accurate analysis of protein families from skewed databases.
- The developed algorithm offers a robust solution for correcting representation bias in sequence data.
- This approach enhances the utility of sequence alignments and profiles for broader evolutionary and functional studies.