Related Experiment Videos
Rapid automatic detection and alignment of repeats in protein sequences
Proteins
|August 31, 2000
Summary
We developed RADAR, an automatic algorithm to identify protein sequence repeats, crucial for protein function and structure. This method efficiently finds novel repeats in large databases, aiding in understanding protein evolution.
Area of Science:
- Bioinformatics
- Computational Biology
- Molecular Biology
Background:
- Large proteins often evolve through internal duplication, with sequence repeats correlating to functional and structural units.
- Identifying these repeats is essential for understanding protein evolution and function.
Purpose of the Study:
- To develop an automated algorithm for accurate segmentation of protein sequences into repeats.
- To identify novel and previously unannotated repeats within protein databases.
Main Methods:
- The RADAR algorithm determines repeat length via suboptimal self-alignment traces.
- It optimizes repeat borders for a maximal integer number of repeats.
- Iterative profile alignment is used to validate distant repeats.
Main Results:
- RADAR identifies various repeat types, including short, composition-biased, gapped, and complex architectures.
- The algorithm requires no manual intervention or prior assumptions on repeat number or length.
- Comparison with Pfam-A showed good coverage, accurate alignments, and reasonable repeat borders.
- Screening Swissprot identified 3,000 novel, unannotated repeats, many previously undescribed.
Conclusions:
- Automated sequence analysis, like that performed by RADAR, efficiently captures important protein repeat information.
- This approach complements curated databases, especially given increasing data backlogs.
- RADAR facilitates the discovery of novel protein repeats, advancing the understanding of protein structure and function.