Related Experiment Videos
A sentence sliding window approach to extract protein annotations from biomedical articles
Martin Krallinger1, Maria Padron, Alfonso Valencia
1Protein Design Group, National Center of Biotechnology, CNB-CSIC, Cantoblanco, E-28049 Madrid, Spain. martink@cnb.uam.es
BMC Bioinformatics
|June 18, 2005
Summary
A novel sentence sliding window approach efficiently extracts protein annotations from biomedical texts. This method improves the accuracy of identifying proteins and GO terms compared to extracting complete protein-function relations.
Area of Science:
- Biomedical Informatics
- Computational Biology
- Natural Language Processing
Background:
- Text mining and statistical natural language processing (NLP) are rapidly advancing in biomedical research.
- A need exists for comparative assessments and standardized evaluation criteria for these methods.
- The Critical Assessment of Text Mining Methods in Molecular Biology (BioCreative) contest addresses these needs.
Purpose of the Study:
- To evaluate text mining system performance on biomedical texts.
- To assess tools for recognizing named entities like genes and proteins.
- To evaluate tools for automatic protein annotation extraction.
Main Methods:
- Development and application of a "sentence sliding window" approach.
- Extraction of text fragments from full-text articles for protein annotations.
- Analysis of the accuracy of extracted individual entities (proteins, GO terms) and complete annotations (protein-function relations).
Main Results:
- The sentence sliding window approach demonstrated high efficiency in extracting protein annotations.
- This method yielded the highest number of correctly predicted annotations.
- Correct extraction of individual entities (proteins, GO terms) significantly surpassed the extraction of complete protein-function relations.
Conclusions:
- Averaging sentence sliding windows show promise for information extraction, particularly without conventional training data.
- Combining this approach with advanced statistical estimators and machine learning can enhance future annotation extraction in biomedical text mining.