Related Experiment Videos
Removing near-neighbour redundancy from large protein sequence collections
Bioinformatics (Oxford, England)
|July 31, 1998
Summary
A new non-redundant protein sequence database (nrdb90) reduces redundancy and computation time for biological discovery. This database accelerates homology searches and unifies annotations, saving significant computer resources.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Sequence databases are rapidly growing and contain redundant information.
- This redundancy strains computational resources and reduces the quality of annotations.
- Efficient homology searching requires up-to-date and non-redundant sequence collections.
Purpose of the Study:
- To address the challenges of large and redundant sequence databases.
- To develop a method for creating a representative subset of sequences with reduced redundancy.
- To improve the efficiency and accuracy of biological discovery through enhanced homology searching.
Main Methods:
- Clustering of highly similar sequences to form a representative set with >90% mutual sequence identity.
- Utilizing deca- and pentapeptide composition filters to reduce the need for explicit sequence alignment.
- Applying the algorithm to a comprehensive union of major sequence databases (e.g., Swissprot, Trembl, Genbank).
Main Results:
- Generated a non-redundant database (nrdb90) with a 46% size reduction (260,000 to 140,000 sequences).
- Achieved the all-against-all comparison for 90% sequence identity in 2 days of CPU time.
- Demonstrated faster homology searches and unified annotation for clustered sequences.
Conclusions:
- The developed method effectively reduces sequence redundancy while preserving biological information.
- The nrdb90 database and associated tools offer significant computational savings for biological research.
- This approach enhances the efficiency of large-scale biological discovery and data analysis.