Related Experiment Videos
Generation of a large gene/protein lexicon by morphological pattern analysis.
Lorraine Tanabe1, W John Wilbur
1National Center for Biotechnology Information, 8600 Rockville Pike, Bethesda, MD 20894, USA. tanabe@ncbi.nlm.nih.gov
Journal of Bioinformatics and Computational Biology
|August 4, 2004
Summary
This study developed a method using inductive logic programming to filter gene and protein names from millions of entries. The resulting lexicon contains over 1.1 million high-quality gene/protein names for research use.
Area of Science:
- Bioinformatics
- Computational Biology
- Natural Language Processing
Background:
- Accurate identification of gene and protein names is crucial for biological research.
- Previous methods yielded large datasets of potential gene/protein names with significant noise.
- A high-quality, curated lexicon is needed to advance biological text mining.
Purpose of the Study:
- To develop and validate a computational approach for purifying large gene/protein name datasets.
- To create a high-accuracy lexicon of gene and protein names from MEDLINE documents.
- To improve the efficiency and reliability of named entity recognition in biomedical literature.
Main Methods:
- Generated name classes based on common morphological features.
- Applied inductive logic programming (ILP) within each class to learn gene/protein name characteristics.
- Implemented a false positive filter to refine the curated name set.
Main Results:
- Identified 193 distinct name classes.
- ILP criteria selected 1,240,462 potential gene/protein names.
- A final lexicon of 1,145,913 names was generated with 82% accuracy.
- The lexicon demonstrated high precision, with 82% accurate gene/protein names.
Conclusions:
- The ILP-based approach effectively purifies large gene/protein name collections.
- The generated lexicon provides a valuable resource for biomedical text mining and analysis.
- This method enhances the quality and reliability of gene/protein name recognition in scientific literature.