Related Experiment Videos
Automatic extraction of gene/protein biological functions from biomedical text
Asako Koike1, Yoshiki Niwa, Toshihisa Takagi
1Department of Computational Biology, Graduate School of Frontier Science, The University of Tokyo Kiban-3A1(CB01) 5-1-5, Kashiwanoha Kashiwa, Chiba 277-8561, Japan. akoike@hgc.jp
Bioinformatics (Oxford, England)
|October 29, 2004
Summary
This study introduces an automated method to extract gene biological functions from text using Gene Ontology (GO) IDs. The system achieves high precision in identifying gene-GO relationships, aiding biomedical data interpretation.
Area of Science:
- Biomedical Informatics
- Computational Biology
- Text Mining
Background:
- High-throughput analysis in biomedical science necessitates efficient information extraction.
- Automatic gene functional annotation is crucial for interpreting large datasets.
- Increasing demand exists for automated extraction of gene functions from scientific literature.
Purpose of the Study:
- To develop a method for automatically extracting biological process functions of genes, proteins, and families from text.
- To assign corresponding Gene Ontology (GO) IDs to gene/protein/family names based on their described functions.
- To enhance the recognition of gene/protein/family functions using GO and various text analysis techniques.
Main Methods:
- Utilized a shallow parser and sentence structure analysis for function extraction.
- Identified gene/protein/family names using specialized dictionaries.
- Employed co-occurrence, collocation similarities, and rule-based techniques for semi-automatic gathering of GO-based functional terms.
- Leveraged ACTOR-OBJECT relationships to link gene/protein/family names with their functions.
Main Results:
- Achieved an estimated recall of 54-64% and a precision of 91-94% for extracting described functions in abstracts.
- Successfully extracted over 190,000 gene-GO relationships.
- Extracted approximately 150,000 family-GO relationships for major eukaryotes from the PubMed database.
Conclusions:
- The developed method effectively automates the extraction of gene/protein/family functions and their associated GO IDs from biomedical text.
- The system demonstrates high precision, making it a valuable tool for large-scale biomedical data analysis.
- The large-scale extraction of gene-GO and family-GO relationships provides significant resources for biological research.