Related Experiment Videos
Distributed modules for text annotation and IE applied to the biomedical domain
Harald Kirsch1, Sylvain Gaudan, Dietrich Rebholz-Schuhmann
1EMBL-EBI, Wellcome Trust Genome Campus, Hinxton, Cambridge CB10 1SD, UK. kirsch@ebi.ac.uk
International Journal of Medical Informatics
|August 9, 2005
Summary
This study introduces a modular server infrastructure to automate the extraction of biological information from scientific literature, aiding manual curation efforts for databases.
Area of Science:
- Bioinformatics
- Computational Biology
- Scientific Literature Analysis
Background:
- Manual curation of biological databases is essential for data quality but is labor-intensive.
- Automated information extraction methods can support and accelerate the curation process.
Purpose of the Study:
- To develop a flexible server software infrastructure for integrating information extraction modules.
- To facilitate the identification and presentation of biologically relevant text to curators.
Main Methods:
- A modular server architecture allowing easy integration of plug-in modules.
- Modules for identifying UniProt, UMLS, GO terminology, gene/protein names, mutations, and protein-protein interactions.
- Utilizes XML annotated text streams for inter-server communication and distributed processing.
Main Results:
- Modules successfully identify and link key biological entities such as UniProt, UMLS, and GO concepts.
- Specific modules employ syntax patterns for mutation identification and chunk parsing for protein-protein interactions.
- The system supports distributed processing and pipeline construction for flexible workflows.
Conclusions:
- The presented server infrastructure streamlines the extraction of biological information from literature.
- This tool enhances the efficiency of biological database curation by automating key tasks.
- The software and server are publicly available to the research community.