Related Experiment Videos
A probabilistic model for mining implicit 'chemical compound-gene' relations from literature.
Shanfeng Zhu1, Yasushi Okuno, Gozoh Tsujimoto
1Bioinformatics Center, Institute for Chemical Research, Kyoto University Gokasho, Uji, Japan.
Bioinformatics (Oxford, England)
|October 6, 2005
Summary
This study introduces a probabilistic model to identify relationships between chemical compounds and genes from scientific literature. Incorporating compound-compound co-occurrences significantly improved predictive performance for drug-gene interactions.
Area of Science:
- Computational Biology
- Bioinformatics
- Cheminformatics
Background:
- Chemical genomics is a growing field in molecular biology, necessitating the identification of associations between chemical compounds (drugs) and genes.
- Co-occurrence analysis of biological entities in literature is a common method for inferring relationships.
Purpose of the Study:
- To develop a probabilistic model for mining implicit relationships between chemical compounds and genes from literature co-occurrence data.
- To evaluate the model's performance using diverse co-occurrence datasets and validate predicted drug-gene pairs.
Main Methods:
- Proposed a Mixture Aspect Model (MAM) and an associated parameter estimation algorithm.
- Utilized MEDLINE records and the ChEBI database for validation.
- Experimented with compound-gene, gene-gene, and compound-compound co-occurrence datasets.
Main Results:
- The MAM trained on all co-occurrence datasets significantly outperformed simpler models.
- Incorporating compound-compound co-occurrences proved most effective in enhancing predictive accuracy.
- The top 20 predicted drug-gene pairs were identified and validated from biological, medical, and pharmaceutical perspectives.
Conclusions:
- The proposed MAM effectively identifies chemical compound-gene relationships from literature.
- The model provides a robust framework for leveraging various co-occurrence data types, with compound-compound interactions being particularly valuable.
- This approach aids in discovering novel drug-gene associations relevant to multiple scientific disciplines.