Related Experiment Videos
Literature mining and database annotation of protein phosphorylation using a rule-based system.
Z Z Hu1, M Narayanaswamy, K E Ravikumar
1Department of Biochemistry and Molecular Biology, Georgetown University Medical Center, Washington, DC 20057, USA. zh9@georgetown.edu
Bioinformatics (Oxford, England)
|April 9, 2005
Summary
This study introduces RLIMS-P, a computational tool for automatically extracting protein phosphorylation data from scientific literature. This system significantly improves the efficiency of annotating protein phosphorylation information for databases.
Area of Science:
- Biochemistry
- Bioinformatics
- Computational Biology
Background:
- Vast amounts of protein phosphorylation data exist in scientific literature, but are difficult to access and curate for databases.
- Manual literature curation is time-consuming and limits the comprehensive collection of this valuable data.
Purpose of the Study:
- To develop and evaluate a computational system for automated literature mining of protein phosphorylation.
- To facilitate the efficient extraction and annotation of protein phosphorylation information for biological databases.
Main Methods:
- Utilized a rule-based system named RLIMS-P (Rule-based LIterature Mining System for Protein Phosphorylation).
- Applied RLIMS-P to MEDLINE abstracts to identify and extract phosphorylation-related entities (kinases, substrates, sites).
- Evaluated system performance using an established, annotation-tagged literature corpus.
Main Results:
- RLIMS-P achieved high precision (91.4%) and recall (96.4%) for retrieving relevant phosphorylation papers.
- The system demonstrated high precision (97.9%) and recall (88.0%) for extracting protein substrates and phosphorylation sites.
- Combined high recall in paper retrieval with high precision in information extraction.
Conclusions:
- RLIMS-P effectively automates the extraction of protein phosphorylation data from scientific literature.
- The system significantly enhances the process of literature mining and database annotation for protein phosphorylation.
- This approach promises to accelerate the curation of crucial biological information.