Related Experiment Videos
Blocks+: a non-redundant database of protein alignment blocks derived from multiple compilations
S Henikoff1, J G Henikoff, S Pietrokovski
1Howard Hughes Medical Institute, Fred Hutchinson Cancer Research Center, 1100 Fairview Avenue North, Seattle, WA 98109-1024, USA. steveh@fhcrc.org
Bioinformatics (Oxford, England)
|June 26, 1999
Summary
Blocks+ unifies protein family data from multiple databases, doubling coverage for improved sequence classification. This enhanced database helps accurately identify protein functions and relationships.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Protein sequence classification and function prediction are vital as biological databases expand.
- The original Blocks Database, based on Prosite, offered valuable sequence classification but had incomplete coverage.
- Expanding the Blocks Database with data from other protein family resources is necessary for broader application.
Purpose of the Study:
- To create a unified and comprehensive protein family database by integrating data from multiple sources.
- To enhance sequence classification accuracy and functional prediction capabilities.
- To develop a non-redundant, hierarchical compilation of protein families.
Main Methods:
- Constructed Blocks+ by integrating data from five protein family databases (Prosite, Prints, Pfam-A, ProDom, Domo).
- Utilized the PROTOMAT/BLOSUM scoring model and a single search algorithm for consistent classification.
- Employed the LAMA blocks-versus-blocks searching program to identify and manage overlapping protein families for non-redundancy.
Main Results:
- Blocks+ comprises 8909 blocks representing 1995 protein families, effectively doubling the coverage of the original Blocks Database.
- The unified database successfully integrates blocks from various sources, including those not present in Prosite.
- Demonstrated the ability to retain related yet distinct families and minimize redundancy, using the SNF2 family as a case study.
Conclusions:
- Blocks+ provides a significantly expanded and unified resource for protein family classification and functional prediction.
- The integration strategy effectively addresses the challenge of non-redundancy in compiling diverse protein family data.
- This enhanced database facilitates more accurate and comprehensive analysis of protein sequences and functions.