Related Experiment Videos
PhyME: a probabilistic algorithm for finding motifs in sets of orthologous sequences
Saurabh Sinha1, Mathieu Blanchette, Martin Tompa
1Center for Studies in Physics and Biology, The Rockefeller University, New York, NY 10021, USA. saurabh@lonnrot.rockefeller.edu
BMC Bioinformatics
|October 30, 2004
Summary
This study introduces a new algorithm for finding transcription factor binding sites by combining sequence overrepresentation and cross-species conservation. The method effectively utilizes evolutionary relationships for improved motif discovery in diverse genomic data.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Identifying transcription factor binding sites (TFBS) is crucial for understanding gene regulation.
- Heterogeneous sequence data, including orthologous sequences across species, presents challenges for accurate TFBS discovery.
Purpose of the Study:
- To develop a novel algorithm for discovering transcription factor binding sites.
- To integrate sequence overrepresentation and cross-species conservation into a unified probabilistic score for motif significance.
Main Methods:
- The algorithm employs the Expectation-Maximization (EM) technique.
- It accommodates user-defined phylogenetic trees to model evolutionary relationships between orthologous sequences.
- The method is designed to scale efficiently with the number of species and sequence length.
Main Results:
- The proposed algorithm successfully integrates overrepresentation and cross-species conservation.
- Evaluations on synthetic data and real datasets from yeast, fly, and human demonstrate its efficacy.
- The approach shows improved performance in motif discovery compared to methods not utilizing cross-species information.
Conclusions:
- The developed algorithm enhances transcription factor binding site discovery.
- Exploiting information from multiple species significantly improves motif identification accuracy.
- This approach offers a robust tool for analyzing regulatory sequences in comparative genomics.