Related Experiment Videos
iProLINK: an integrated protein resource for literature mining
Zhang-Zhi Hu1, Inderjeet Mani, Vincent Hermoso
1Georgetown University Medical Center, 3900 Reservoir Road, NW, Washington, DC 20057, USA.
Computational Biology and Chemistry
|November 24, 2004
Summary
The Protein Information Resource developed iProLINK, a curated data resource to aid biological literature mining and protein annotation. This resource supports text mining research for improved genome and proteome annotation.
Area of Science:
- Bioinformatics
- Computational Biology
- Molecular Biology
Background:
- Exponential growth in molecular sequence data and scientific literature necessitates advanced biological literature mining.
- Existing text mining methodologies lack adequate curated data for training and benchmarking.
- The Protein Information Resource (PIR) curates the UniProt protein sequence database.
Purpose of the Study:
- To develop iProLINK, an integrated Protein Literature INformation and Knowledge resource.
- To provide curated data sources for text mining research in protein annotation.
- To support bibliography mapping, annotation extraction, protein named entity recognition, and protein ontology development.
Main Methods:
- Developed iProLINK, a resource for protein literature mining.
- Created mapped citations linking PubMed IDs to protein entries and feature lines.
- Compiled annotation-tagged literature corpora, including abstracts and full-text articles with experimentally validated post-translational modifications (PTMs).
- Assembled a protein name dictionary, word token dictionaries, and protein name-tagged literature corpora with tagging guidelines.
- Established a protein ontology based on PIRSF protein family names.
Main Results:
- iProLINK provides curated data sources for various text mining applications.
- The resource includes mapped citations and tagged literature corpora for bibliography mapping and annotation extraction.
- iProLINK offers resources for entity recognition and ontology development, including a protein name dictionary and a protein ontology.
Conclusions:
- iProLINK addresses the need for curated data in biological literature mining.
- The resource facilitates research in protein annotation, named entity recognition, and ontology development.
- iProLINK is freely accessible, promoting advancements in bioinformatics and computational biology.