Related Experiment Videos
Improved database searches for orthologous sequences by conditioning on outgroup sequences
Philip J Cotter1, Daniel R Caffrey, Denis C Shields
1Department of Clinical Pharmacology, Royal College of Surgeons in Ireland, 123 Stephen's Green, Dublin 2, Ireland.
Bioinformatics (Oxford, England)
|February 12, 2002
Summary
Distinguishing closely related gene sequences (orthologues) from distant ones (homologues) is challenging. Conditioning database searches with an outgroup sequence significantly improves the identification of true positives, enhancing biological sequence analysis.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Biological sequence database searches aim to identify significant matches.
- Distinguishing orthologues from homologues is a growing challenge due to abundant related sequences.
- Partial sequences and non-conserved regions complicate orthologue identification.
Purpose of the Study:
- To improve the accuracy of biological sequence database searches.
- To enhance the ability to distinguish orthologues from homologues.
- To develop a method for better identification of closely related sequences.
Main Methods:
- Conditioning search results by subtracting the score of an outgroup sequence.
- Testing the method on Caenorhabditis elegans kinase sequences against human EST sequences.
- Validating the approach using a dataset of 151 C.elegans proteins with automatically assigned outgroups.
Main Results:
- Outgroup conditioning identified 58% more true positives ahead of false positives in one test.
- A similar test dataset found 50% more true positives using the outgroup conditioning method.
- The method improves true positive identification with minimal increase in computational time.
Conclusions:
- Outgroup conditioning is an effective strategy for improving biological sequence database searches.
- This method enhances the distinction between orthologues and homologues.
- The approach offers a practical improvement for analyzing large sequence datasets.