Related Experiment Videos
Information assessment on predicting protein-protein interactions.
Nan Lin1, Baolin Wu, Ronald Jansen
1Department of Mathematics, Washington University in St. Louis, St. Louis, MO 63130, USA. nlin@math.wustl.edu <nlin@math.wustl.edu>
BMC Bioinformatics
|October 20, 2004
Summary
Functional similarity data, particularly from MIPS and Gene Ontology (GO), is key for predicting protein-protein interactions in yeast. Machine learning models like random forests outperform Bayesian networks when using this complete genomic information.
Area of Science:
- Computational biology
- Genomics
- Bioinformatics
Background:
- Identifying protein-protein interactions is crucial for understanding cellular mechanisms.
- High-throughput experimental methods for detecting these interactions have limitations, including high false positive and negative rates.
- Integrating diverse genomic data can improve prediction accuracy and biological insights.
Purpose of the Study:
- To assess the contribution of different genomic features to predicting protein-protein interactions.
- To compare the effectiveness of various computational models for interaction prediction.
- To identify the most informative data types for robust protein-protein interaction prediction.
Main Methods:
- Utilized a Bayesian network approach to model protein-protein interactions genome-wide in yeast.
- Analyzed the predictive power of genomic features including mRNA expression, localization, essentiality, and functional annotation.
- Compared Bayesian networks with alternative models like logistic regression and random forests.
Main Results:
- Functional classification data (MIPS, Gene Ontology) significantly contributed to predicting protein-protein interactions.
- In a complete information subset, functional similarity was more informative than expression correlations or essentiality.
- Random forests and logistic regression showed potential to be more effective than Bayesian networks.
Conclusions:
- MIPS and Gene Ontology functional similarity datasets are dominant contributors for predicting protein-protein interactions.
- Random forests using only MIPS and GO data achieved high classification accuracy.
- Adding other genomic data provided minimal improvement in the complete information subset, and Bayesian discretization reduced performance.