Related Experiment Video
Updated: Jul 10, 2026

Navigating MARRVEL, a Web-Based Tool that Integrates Human Genomics and Model Organism Genetics Information
Published on: August 15, 2019
Normalization of gene/protein names in biological literatures using Vector-Space Model
Joon-Ho Lim1, Hyunchul Jang, Jaesoo Lim
1Lifeinformatics Team, Electronics and Telecommunication Research Institute, Gajeong-Dong, Yuseon-Gu, Daejeon, 305-700, Korea. joonho.lim@etri.re.kr
This study introduces a new text mining normalization method for gene and protein names using the Vector-Space Model. The approach effectively handles name variations, improving data integration from biological literature.
Area of Science:
- Bioinformatics
- Computational Biology
- Natural Language Processing
Background:
- The exponential growth of biological literature necessitates advanced text mining systems.
- Gene and protein name normalization is crucial for integrating information and building databases or ontologies.
- Existing normalization methods struggle with highly variable gene/protein names found in literature.
Purpose of the Study:
- To propose a novel normalization method for gene and protein names in biological text mining.
- To address the limitations of direct comparison methods in handling name variations.
- To improve the accuracy and efficiency of mapping gene/protein names to databases.
Main Methods:
- Utilizing the Vector-Space Model for gene and protein name normalization.
- Ranking potential database identifiers based on similarity to extracted names.
- Implementing a method to identify the most similar identifier for each gene/protein name.
Main Results:
- The proposed Vector-Space Model method achieved a 70.7% f-measure in experimental results.
- Demonstrated improved performance compared to previous direct comparison methods.
- Successfully handled variational gene/protein names in text mining.
Conclusions:
- The Vector-Space Model offers a robust solution for gene and protein name normalization.
- This method enhances the ability to create comprehensive databases and ontologies from diverse literature.
- The findings contribute to more effective information extraction in bioinformatics.
Related Concept Videos
Organization of Genes
Organization of Genes
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Genetic Lingo
Structure of a Gene
However, only 1% of the DNA is composed of genes that encode proteins; the rest, 99% is non-coding DNA. This non-coding DNA performs...
Gene Families
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...

