Related Experiment Videos
GENETAG: a tagged corpus for gene/protein named entity recognition
Lorraine Tanabe1, Natalie Xie, Lynne H Thom
1National Center for Biotechnology Information, National Library of Medicine, NIH, 8600 Rockville Pike, Bethesda, MD 20894, USA. tanabe@ncbi.nlm.nih.gov
BMC Bioinformatics
|June 18, 2005
Summary
GENETAG is a new corpus of 20,000 MEDLINE sentences for gene/protein named entity recognition (NER). It enables standardized evaluation of biomedical NER systems, crucial for text mining the scientific literature.
Area of Science:
- Biomedical Natural Language Processing
- Computational Biology
- Bioinformatics
Background:
- Named Entity Recognition (NER) is vital for biomedical text mining.
- Standardized test corpora are essential for evaluating NER system performance.
- Annotating gene/protein names for NER is challenging due to name complexity.
Purpose of the Study:
- To describe the construction and annotation of the GENETAG corpus.
- To provide a standardized test corpus for gene/protein NER.
- To facilitate the evaluation of biomedical NER systems.
Main Methods:
- Constructed GENETAG corpus using 20,000 MEDLINE sentences, selected for term similarity to known gene names.
- Manually annotated gene/protein names with a wide definition and specificity constraint.
- Included acceptable alternatives and applied semantic constraints for robust annotation.
Main Results:
- GENETAG corpus comprises 20,000 annotated MEDLINE sentences for gene/protein NER.
- 15,000 sentences were utilized for the BioCreAtIvE Task 1A Competition.
- Annotation incorporated acceptable alternatives and semantic constraints for improved evaluation.
Conclusions:
- Manual annotation of GENETAG introduced tagging inconsistencies.
- Word-based segmentation was used, though character-based indices may be more robust.
- GENETAG data and programs are publicly available; a newer version (GENETAG-05) is forthcoming.