Related Experiment Videos
Filtering high-throughput protein-protein interaction data using a combination of genomic features
Ashwini Patil1, Haruki Nakamura
1Institute for Protein Research, Osaka University, 3-2 Yamadaoka, Suita, Osaka 565-0871, Japan. ashwini@protein.osaka-u.ac.jp
BMC Bioinformatics
|April 19, 2005
Summary
High-throughput protein-protein interaction data can be noisy. Combining genomic features like sequence, structure, and annotations helps validate these interactions, improving molecular network predictions.
Area of Science:
- Systems Biology
- Bioinformatics
- Computational Biology
Background:
- High-throughput experiments generate vast protein-protein interaction data.
- This data often contains numerous spurious or incorrect interactions.
- Validation is crucial for accurate molecular network construction and prediction.
Purpose of the Study:
- To develop a method for validating protein-protein interactions from high-throughput experiments.
- To assess the reliability of these interactions using genomic features.
- To estimate the proportion of true interactions in various model organisms.
Main Methods:
- Utilized a combination of three genomic features: Pfam domains, Gene Ontology annotations, and sequence homology.
- Employed Bayesian network approaches to assign reliability scores to protein-protein interactions.
- Applied the method to high-throughput data from Saccharomyces cerevisiae, Caenorhabditis elegans, Drosophila melanogaster, and Homo sapiens.
Main Results:
- Protein-protein interactions supported by genomic features showed higher likelihood ratios, indicating greater reliability.
- The developed method achieved 90% sensitivity and 63% specificity.
- 56% of high-throughput interactions in Saccharomyces cerevisiae were deemed highly reliable.
- Estimated true interaction proportions were 27% (C. elegans), 18% (D. melanogaster), and 68% (H. sapiens).
Conclusions:
- A combination of sequence, structure, and annotation data effectively predicts true protein-protein interactions in noisy datasets.
- The method provides a reliable way to assign likelihood ratios to interactions, enhancing data quality.
- This approach significantly improves the accuracy of molecular network analysis.